Point cloud generation method, device and equipment
By using the verification simulator, binocular camera and lidar to generate training data sets, extracting and optimizing disparity features, the problem of point cloud sparsity and incompleteness is solved, and the generation quality of point clouds and the reliability of robot environment perception is improved.
Patent Information
- Application Number
- CN202510865114.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-06-25
AI Technical Summary
In the prior art, the original point cloud generation technology generates sparseness, incompleteness, and insufficient adaptability of dynamic scenes, which seriously affects the reliability and robustness of downstream perception algorithms.
The initial model is trained by using a verification simulator to generate a training data set, and data acquisition is collected for the calibration scene by combining a binocular camera and a lidar, multi-dimensional parallax features are extracted and optimized to generate depth map data to generate point clouds.
It improves the integrity and density of point clouds, enhances the generalization ability of point cloud generation models, and improves the robot's three-dimensional perception of the environment.
Smart Images

Figure CN120374701A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a method, apparatus, and device for generating point clouds. Background Art
[0002] In environmental perception tasks in fields such as autonomous driving, industrial robots, and intelligent service robots, high-precision and high-density point clouds are the basis for realizing core functions such as obstacle detection and semantic segmentation. Currently, point cloud generation technology mainly relies on direct collection by sensors such as lidar (LiDAR) and depth cameras. However, limited by hardware performance and environmental complexity, the generated original point clouds generally have problems such as sparsity, incompleteness, and insufficient adaptability to dynamic scenes, seriously restricting the reliability and robustness of downstream perception algorithms. Summary of the Invention
[0003] Embodiments of this application provide a method, apparatus, and device for generating point clouds, which can effectively improve the integrity and density of point clouds.
[0004] The technical solution of the embodiments of this application is implemented as follows: In a first aspect, embodiments of this application provide a method for generating point clouds. The method includes: Extracting feature data to be processed from the data to be processed by using a point cloud generation model, obtaining the feature data to be processed; wherein, the data to be processed includes at least image data; the point cloud generation model is obtained by training an initial model using a training data set generated by a verified simulator; the verified simulator is verified based on first data, first internal and external parameters, second data, and second internal and external parameters; the first data is obtained by using a binocular camera and a lidar to collect data from a calibration scene; the first internal and external parameters represent the internal and external parameters of the binocular camera and the lidar obtained by using the first data; the second data is obtained by using the simulator to be verified to simulate the data collection of the binocular camera and the lidar in the calibration scene; the second internal and external parameters represent the internal and external parameters of the binocular camera and the lidar obtained by using the second data; Extracting multi-dimensional parallax features from the feature data to be processed by using the point cloud generation model, obtaining parallax feature data; Performing optimization processing on the parallax feature data by using the point cloud generation model to obtain optimized parallax data; Generating depth map data according to the optimized parallax data, and generating point clouds according to the depth map data.
[0005] In a second aspect, embodiments of this application provide a point cloud generation apparatus, including a first generation unit and a second generation unit; The first generation unit is configured to extract features from the data to be processed by using a point cloud generation model to obtain the feature data to be processed; and extract multi-dimensional parallax features from the feature data to be processed by using the point cloud generation model to obtain parallax feature data; and perform parallax feature optimization processing on the parallax feature data by using the point cloud generation model to obtain the optimized parallax data; wherein, the data to be processed includes at least image data; the point cloud generation model is obtained by training an initial model by using a training data set generated by a verified simulator; the verified simulator is verified based on the first data, the first internal and external parameters, the second data, and the second internal and external parameters; the first data is obtained by using a binocular camera and a lidar to collect data for a calibration scene; the first internal and external parameters represent the internal and external parameters of the binocular camera and the lidar obtained by using the first data; the second data is obtained by simulating the data collection of the binocular camera and the lidar in the calibration scene by using the simulator to be verified; the second internal and external parameters represent the internal and external parameters of the binocular camera and the lidar obtained by using the second data; The second generation unit is configured to generate depth map data according to the optimized parallax data, and generate a point cloud according to the depth map data.
[0006] In a third aspect, an embodiment of the present application provides a point cloud generation device, including a processor and a memory storing processor-executable instructions; when the instructions are executed by the processor, the above-mentioned point cloud generation method is implemented.
[0007] In a fourth aspect, an embodiment of the present application provides an electronic device, including the point cloud generation device, the electronic device is configured to construct map data according to the point cloud generated by the point cloud generation device; and perform corresponding operations according to the map data.
[0008] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the above-mentioned point cloud generation method is implemented.
[0009] In a sixth aspect, an embodiment of the present application provides a computer program product, including a computer program or instruction, and when the computer program or instruction is executed by a processor, the above-mentioned point cloud generation method is implemented.
[0010] The embodiments of the present application provide a point cloud generation method, apparatus and device. The point cloud generation apparatus can extract features from the data to be processed by using a point cloud generation model to obtain the feature data to be processed. Among them, the data to be processed at least includes image data. The point cloud generation model is obtained by training an initial model with a training data set generated by a verified simulator. The verified simulator is verified based on the first data, the first internal and external parameters, the second data, and the second internal and external parameters. The first data is obtained by using a binocular camera and a lidar to collect data from a calibration scene. The first internal and external parameters represent the internal and external parameters of the binocular camera and the lidar obtained by using the first data. The second data is obtained by using the simulator to be verified to simulate the data collection of the binocular camera and the lidar in the calibration scene. The second internal and external parameters represent the internal and external parameters of the binocular camera and the lidar obtained by using the second data. The point cloud generation model is used to extract multi-dimensional disparity features from the feature data to be processed to obtain disparity feature data. The point cloud generation model is used to optimize the multi-scale disparity features of the disparity feature data to obtain the optimized disparity data. The depth map data is generated according to the optimized disparity data, and the point cloud is generated according to the depth map data.
[0011] It can be seen that the present application can pre-verify the simulator to be verified, including completing the verification by using the first data obtained by using a binocular camera and a lidar to collect data from a real calibration scene, the first internal and external parameters of the binocular camera and the lidar, the second data obtained by using the simulator to be verified to simulate the data collection scene of the binocular camera and the lidar, and the internal and external parameters of the binocular camera and the lidar calculated by using the second data, so as to determine whether the simulator to be verified has a good simulation effect, that is, whether it can obtain data that is basically no different from the real calibration scene by using the simulator to be verified. Thus, the initial model is trained by using the data generated by the verified simulator to obtain the point cloud generation model, which can improve the generalization ability of the point cloud generation model. Furthermore, when the point cloud generation apparatus generates a point cloud, it can first extract features from the data to be processed by using the constructed point cloud generation model, and then extract multi-dimensional disparity features from the feature data to be processed, so as to obtain disparity features in different dimensions. Then, the disparity features in different dimensions will be optimized with different scales of disparity features to further optimize the disparity features. Thus, the depth map data is generated according to the optimized disparity data, and the point cloud is generated based on the depth map data, which can effectively improve the integrity and density of the generated point cloud. Description of the Drawings
[0012] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present application.
[0013] Figure 1 Schematic diagram of the implementation process of the point cloud generation method proposed in the embodiment of the present application; Figure 2 Schematic of the implementation of the point cloud generation method proposed in the embodiment of the present application Figure 1 ; Figure 3 Schematic of the implementation of the point cloud generation method proposed in the embodiment of the present application Figure 2 ; Figure 4 Schematic of the implementation of the point cloud generation method proposed in the embodiment of the present application Figure 3 ; Figure 5 Schematic of the implementation of the point cloud generation method proposed in the embodiment of the present application Figure 4 ; Figure 6 Schematic of the implementation of the point cloud generation method proposed in the embodiment of the present application Figure 5 ; Figure 7 Schematic of the implementation of the point cloud generation method proposed in the embodiment of the present application Figure 6 ; Figure 8 Schematic of the implementation of the point cloud generation method proposed in the embodiment of the present application Figure 7 ; Figure 9 Schematic of the implementation of the point cloud generation method proposed in the embodiment of the present application Figure 8 ; Figure 10 Schematic of the implementation of the point cloud generation method proposed in the embodiment of the present application Figure 9 ; Figure 11 Schematic of the implementation of the point cloud generation method proposed in the embodiment of the present application Figure 10 ; Figure 12 Schematic of the implementation of the point cloud generation method proposed in the embodiment of the present application Figure 10 One; Figure 13 Schematic of the implementation of the point cloud generation method proposed in the embodiment of the present application Figure 10 Two; Figure 14 Schematic diagram of the implementation process of the emulator verification method proposed in the embodiment of the present application; Figure 15 Schematic of the composition structure of the point cloud generation device proposed in the embodiment of the present application Figure 1 ; Figure 16 Schematic of the composition structure of the point cloud generation device proposed in the embodiment of the present application Figure 2 . Detailed implementation manners
[0014] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. It can be understood that the specific embodiments described herein are only used to explain the relevant application, rather than limiting the application. In addition, it should be noted that for the convenience of description, only the parts related to the relevant application are shown in the drawings.
[0015] Currently, with the continuous development of artificial intelligence technology, robots are evolving from the relatively single forms of patrol inspection, delivery, and floor cleaning machines in the past to the direction of general humanoid robots represented by embodied intelligence, enabling robots to play a productive role in more human scenarios; such as including commercial retail, industrial, household and other fields. Embodied intelligence endows the "body" of the robot with a more powerful "brain", enabling the robot to perceive, learn and dynamically interact with the physical environment like a human, which poses very high requirements for the robot to accurately perceive the surrounding environment, and requires artificial intelligence algorithms to accurately analyze the pose information of the objects around the robot from sensor data.
[0016] Currently, point cloud generation technology mainly relies on direct acquisition by sensors such as lidar (LiDAR) and depth cameras. However, due to hardware performance and environmental complexity limitations, the generated raw point clouds generally have problems such as sparsity, incompleteness, and insufficient adaptability to dynamic scenes, seriously restricting the reliability and robustness of downstream perception algorithms.
[0017] The stereo matching technology calculates depth through disparity, and its principle can be expressed as , where Z is the depth, is the baseline distance, is the focal length, Parallax; binocular depth estimation is a commonly used technique in robot perception. It mainly uses binocular images to calculate depth information for three-dimensional reconstruction or restoration of the environment. Traditional algorithms often manually select or design image features according to task requirements, and then use the selected or designed image features for depth estimation. They are more vulnerable to high-texture, uniform regions, and object uprightness. In recent years, due to the rapid development of artificial intelligence technology, more and more methods use deep neural networks to solve the binocular depth estimation problem. However, most methods are often trained only on relatively fixed small datasets, which are usually collected from real scenes or synthesized on a small scale. At the same time, although deep neural networks have made great progress, the past depth estimation networks are still relatively small. The models trained based on this type of method often only have a certain generalization ability in a single scene domain or dataset, and the cross-domain generalization ability is poor. Once the test set or scene domain changes, the robustness of the algorithm will be greatly reduced, resulting in the robot being unable to accurately perceive the three-dimensional environment. At the same time, although lidar directly provides the three-dimensional point cloud information of the scene, the data often contains noise, the ranging accuracy is greatly affected by the environment, and the texture information of the target object is lacking. Therefore, in the field of three-dimensional perception, the processing of real-scene datasets is often very difficult, generally taking a long time, resulting in a small dataset scale and average quality, which further limits the performance improvement of depth estimation methods.
[0018] To solve the problems of sparsity, incompleteness, and insufficient adaptability to dynamic scenes in the current point cloud, the embodiments of the present application provide a point cloud generation method, device, and equipment. The point cloud generation device can extract features from the data to be processed using a point cloud generation model to obtain the feature data to be processed. Among them, the data to be processed at least includes image data. The point cloud generation model is obtained by training an initial model using a training dataset generated by a verified simulator. The verified simulator is verified based on the first data, the first internal and external parameters, the second data, and the second internal and external parameters. The first data is obtained by collecting data from a calibration scene using a binocular camera and a lidar. The first internal and external parameters represent the internal and external parameters of the binocular camera and the lidar obtained using the first data. The second data is obtained by simulating the data collection of the binocular camera and the lidar in the calibration scene using the simulator to be verified. The second internal and external parameters represent the internal and external parameters of the binocular camera and the lidar obtained using the second data. The point cloud generation model is used to extract multi-dimensional parallax features from the feature data to be processed to obtain parallax feature data. The point cloud generation model is used to optimize the multi-scale parallax features of the parallax feature data to obtain the optimized parallax data. Depth map data is generated based on the optimized parallax data, and a point cloud is generated based on the depth map data.
[0019] It can be seen that the present application can pre-check the emulator to be verified, including the first data obtained by collecting data on a real calibration scene using a binocular camera and a lidar, the first internal and external parameters of the binocular camera and the lidar, the second data obtained by simulating the data collection scene of the binocular camera and the lidar using the emulator to be verified, and the internal and external parameters of the binocular camera and the lidar calculated using the second data to complete the verification. It can verify whether the emulator to be verified has a good simulation effect, that is, verify whether it is possible to obtain data that is basically indistinguishable from the real calibration scene using the emulator to be verified. Thus, the initial model can be trained using the data generated by the verified emulator to obtain a point cloud generation model, so as to improve the generalization ability of the point cloud generation model. Furthermore, when generating a point cloud, the point cloud generation device can first extract features from the data to be processed using the constructed point cloud generation model, and then extract multi-dimensional disparity features from the feature data to be processed, capable of obtaining disparity features of different dimensions. Then, the disparity features of different dimensions will be optimized with disparity features of different scales to further optimize the disparity features. Thus, depth map data can be generated based on the optimized disparity data, and a point cloud can be generated based on the depth map data, effectively improving the integrity and density of the generated point cloud.
[0020] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application.
[0021] An embodiment of the present application provides a point cloud generation method, which is applied to a point cloud generation device. As Figure 1 shown, the point cloud generation method of the point cloud generation device may include the following steps: Step 101: Extract features from the data to be processed using a point cloud generation model to obtain feature data to be processed; wherein, the data to be processed includes at least image data; the point cloud generation model is obtained by training an initial model using a training data set generated by a verified emulator.
[0022] In an embodiment of the present application, the point cloud generation device may first obtain the data to be processed. The data to be processed may include image data, or may include image data and radar data; wherein, the image data may be obtained by a camera, and the radar data may be obtained by a lidar.
[0023] In some embodiments of the present application, the image data may be obtained by the camera collecting images of the environment where the point cloud generation device is located; the radar data may be obtained by the lidar scanning the environment where the point cloud generation device is located.
[0024] In an embodiment of the present application, the point cloud generation device may be a device in an electronic device for generating a point cloud. The electronic device may include a binocular camera and a lidar. The binocular camera may include a left camera and a right camera. For example, the electronic device may be a robot equipped with a binocular camera and a lidar. Thus, the point cloud generation device may generate a point cloud by obtaining the image data collected by the binocular camera mounted on the robot, or may generate a point cloud by obtaining the image data collected by the binocular camera and the radar data collected by the lidar mounted on the robot.
[0025] In an embodiment of the present application, the image data obtained by the binocular camera may include the image data of the left camera and the image data of the right camera. The type of the binocular camera is not limited in the present application. For example, it may be a binocular red, green, blue (RGB) camera or a binocular infrared (IR) camera.
[0026] In an embodiment of the present application, the emulator to be verified is verified based on the first data, the first internal and external parameters, the second data, and the second internal and external parameters. The first data is obtained by using the binocular camera and the lidar to collect data from a calibration scene. The first internal and external parameters represent the internal and external parameters of the binocular camera and the lidar calculated using the first data. The second data is obtained by simulating the data collection of the binocular camera and the lidar in the calibration scene using the emulator to be verified. The second internal and external parameters represent the internal and external parameters of the binocular camera and the lidar calculated using the second data.
[0027] In some embodiments of the present application, when verifying the emulator to be verified based on the first data, the first internal and external parameters, the second data, and the second internal and external parameters, the first data and the second data may be visualized, and then the visualized first data and the visualized second data may be compared to obtain a first comparison result. At the same time, an error calculation may be performed on the first internal and external parameters and the second internal and external parameters to obtain the error information between the first internal and external parameters and the second internal and external parameters. An error calculation may also be performed on the coordinate information of the feature points of the feature object in the first data and the second data respectively to obtain the error result of the feature point coordinates. Thus, the verification result of the emulator to be verified is determined according to the first comparison result, the error information, and the error result of the feature point coordinates. The verification result may be used to indicate whether the emulator to be verified passes the verification.
[0028] In some embodiments of the present application, during verification, the feature object is various objects in the calibration scene, and the feature points may be the key points of these objects. For example, the inner corner points of a calibration board, etc.
[0029] In some embodiments of the present application, if the verification result of the emulator to be verified fails the verification, the emulator can be adjusted, and then the adjusted emulator can be verified again to determine whether the adjusted emulator can pass the verification.
[0030] In some embodiments of the present application, if the verification result of the emulator to be verified passes the verification, it can be determined that the emulator is an emulator that has passed the verification.
[0031] In some embodiments of the present application, there can be multiple emulators to be verified; that is, different emulators can be verified separately to determine the verification result of each emulator. Then, among the verification results of each emulator, the emulators that have passed the verification can be selected for subsequent generation of the training dataset, the emulators that have not passed the verification can be adjusted to make the adjusted emulators pass the verification, or these emulators that have not passed the verification can be directly excluded.
[0032] In some embodiments of the present application, different types of objects can be set in the calibration scene, and these different types of objects can be used to represent different types of scenes. For example, for some scenes with objects having high-reflection surfaces, high-reflection objects can be placed in the calibration scene, and their surfaces can have high-reflection points and / or high-reflection strips, so as to detect the data synthesis effect of the emulator to be verified on such scenes with high-reflection surface objects. Thus, when the verification result is that the verification passes, it indicates that the emulator is applicable to the scene corresponding to the objects set in the calibration scene.
[0033] In some embodiments of the present application, a binocular camera and a lidar are used to collect data from the calibration scene, that is, the binocular camera collects image data of the calibration scene, and the lidar collects lidar data of the calibration scene; that is, the first data can include the image data of the calibration scene collected by the binocular camera and the point cloud of the calibration scene collected by the lidar.
[0034] In some embodiments of the present application, the first internal and external parameters can be calculated using the calibration relationship and the first data; the first internal and external parameters include the first internal parameter and the first external parameter of the binocular camera, and the first internal parameter and the first external parameter of the lidar.
[0035] In some embodiments of the present application, the second data can include the synthetic image and the synthetic point cloud obtained by simulating the data collection of the binocular camera and the lidar by the emulator to be verified in the calibration scene, where the synthetic image is the image obtained by simulating and emulating the data collection of the binocular camera by the emulator to be verified in the calibration scene, and the synthetic point cloud is the point cloud obtained by simulating and emulating the data collection of the lidar by the emulator to be verified in the calibration scene.
[0036] In some embodiments of the present application, the second internal and external parameters can be calculated based on the first data. The second internal and external parameters include the second internal parameters and the second external parameters of the binocular camera, as well as the second internal parameters and the second external parameters of the lidar.
[0037] In some embodiments of the present application, the first comparison result can be used to indicate the difference between the image synthesized by the emulator to be verified and the image of the real calibration scene collected by the binocular camera, and the difference between the point cloud synthesized by the emulator to be verified and the point cloud of the real calibration scene collected by the lidar; the error information can be used to indicate the error between the second internal and external parameters calculated based on the second data and the first internal and external parameters calculated based on the first data.
[0038] In some embodiments of the present application, the first comparison result can be understood as a qualitative evaluation, which can determine whether there are problems such as offset and stratification compared with the corresponding images and point clouds of the real scene through the visualized synthetic images and synthetic point clouds; the error information and the error results of the feature point coordinates can be understood as a quantitative evaluation. By the error value between the specific second internal and external parameters and the first internal and external parameters, and the numerical values of the feature point coordinates in the first data and the second data respectively, it is determined whether the deviation degree of these data meets the preset error conditions. If the preset error conditions are met, the error is considered small, otherwise the error is considered large.
[0039] In some embodiments of the present application, when the first comparison result is that the difference situation meets the preset difference conditions and the error information meets the preset error conditions, the emulator to be verified can be determined as the emulator that passes the verification; wherein, the present application does not make specific limitations on the preset difference conditions and the preset error conditions. For example, the preset difference conditions can be no offset and stratification, and the preset error conditions can include that the error value between the second internal and external parameters and the first internal and external parameters is less than the first error threshold, and the error result of the feature point coordinates is less than the second error threshold.
[0040] In some embodiments of the present application, the preset difference condition can be determined according to the deviation allowed in the actual application field, and the preset error condition can be determined according to the sensor errors of the binocular camera and the lidar; for example, if the actual application field of the point cloud generation method is the field of navigation devices, and the geometric deviation in this field allows an offset of less than 0.1 mm, then the preset difference condition can be an offset less than 0.1 mm. If the offset is less than 0.1 mm, it is considered to meet the preset difference condition; if the focal length error in the internal and external parameters of the sensor of the binocular camera is less than 0.1%, then the preset error condition can include a focal length error less than 0.1%. It can be understood that in the embodiments of the present application, when using the point cloud generation model to extract features from the data to be processed, it can be to extract image features from the image data, or to extract point cloud features from the radar data while extracting image features from the image data.
[0041] In some embodiments of the present application, the point cloud generation model may include a monocular depth estimation network, a multi-scale feature extraction network, and a point cloud feature extraction network; wherein, the monocular depth estimation network can be used to perform depth estimation on the image data corresponding to the left camera and the right camera respectively; the multi-scale feature extraction network can be used to perform multi-scale feature extraction on the image data corresponding to the left camera and the right camera respectively; the point cloud feature extraction network can be used to extract point cloud features from the radar data.
[0042] In the embodiments of the present application, the multi-scale feature extraction network performing multi-scale feature extraction on the image data means extracting the representations of the image data at different resolutions and different levels, so as to be able to understand the image data from different perspectives; that is to say, "scale" can be understood as resolution and level.
[0043] In some embodiments of the present application, when the point cloud generation device uses the point cloud generation model to extract features from the data to be processed and obtains the to-be-processed feature data, in the case where the data to be processed includes image data and radar data, the monocular depth estimation network can be used to extract the monocular depth feature information of the image data, and the multi-scale feature extraction network can be used to extract the multi-scale depth feature information of the image data; the point cloud feature extraction network can be used to extract the point cloud features of the radar data; thus, the monocular depth feature information, the multi-scale depth feature information, and the point cloud features are determined as the to-be-processed feature data.
[0044] In some embodiments of the present application, when the point cloud generation device uses the point cloud generation model to extract features from the data to be processed and obtains the to-be-processed feature data, in the case where the data to be processed includes image data, the monocular depth estimation network can be used to extract the monocular depth feature information of the image data, and the multi-scale feature extraction network can be used to extract the multi-scale depth feature information of the image data, and the monocular depth feature information and the multi-scale depth feature information are determined as the to-be-processed feature data.
[0045] It can be understood that in the embodiments of the present application, when extracting image features, it may include using a monocular depth estimation network to extract the monocular depth feature information of the left camera image data, and using a multi-scale feature extraction network to extract the multi-scale depth feature information of the left camera image data, so as to use the monocular depth feature information of the left camera image data and the multi-scale depth feature information of the left camera image data as the to-be-processed feature data of the left camera image data; at the same time, it may also include using a monocular depth estimation network to extract the monocular depth feature information of the right camera image data, and using a multi-scale feature extraction network to extract the multi-scale depth feature information of the right camera image data, so as to use the monocular depth feature information of the right camera image data and the multi-scale depth feature information of the right camera image data as the to-be-processed feature data of the right camera image data.
[0046] In the embodiments of the present application, since the to-be-processed feature data may include the depth features corresponding to the image data of the binocular camera and the point cloud features corresponding to the lidar data, the above feature extraction process realizes a cross-modal feature extraction, which can complement the advantages of the visually rich color, texture information and the accurate ranging ability of the point cloud, and make up for the weaknesses of a single-modal sensor. For example, if only the binocular camera is used to obtain images, there is a lack of ranging ability, and if only the lidar is used to obtain point clouds, there is a lack of target color, texture information, etc.
[0047] In the embodiments of the present application, the to-be-processed feature data may be a mixed cost volume feature obtained by fusing the depth feature information of the image data, that is, the monocular depth feature information and the multi-scale depth feature information, and the point cloud features; among them, the cost volume can be understood as a three-dimensional data structure for storing the pixel matching costs under different disparity hypotheses, and the pixel matching cost is a quantization value of the similarity between the pixels in the left camera image data and the pixels in the right camera image data. The three dimensions of the three-dimensional data structure include the disparity hypothesis dimension, the feature map height dimension, and the feature map width dimension.
[0048] In some embodiments of the present application, when fusing the depth feature information of the image data and the point cloud features to obtain the final to-be-processed feature data, the point cloud features may be projected onto the two-dimensional image coordinate system for fusion with the image features; the advantages of the lidar include active detection and accurate ranging, but it lacks semantic information, the point cloud is sparse in the distance, and the point cloud quality is greatly affected by the environment; the binocular camera performs poorly in weak texture areas and low-light conditions. By fusing the features of these two modalities, the stereo matching effect can be improved to generate a more complete and denser point cloud.
[0049] Exemplarily, the principle of projecting the point cloud features onto the two-dimensional image coordinate system can be expressed by the following formula: (1) Wherein, , , are the world coordinates of the point cloud, , represents the image feature coordinates obtained after transformation, and respectively represent the rotation matrix and displacement of the point cloud from the world coordinate system to the camera coordinate system, is the camera internal parameter, is the depth of the point cloud in the camera coordinate system, is the scaling factor between the transformed image feature and the original image feature.
[0050] Step 102: Use the point cloud generation model to extract multi-dimensional disparity features from the to-be-processed feature data to obtain disparity feature data.
[0051] In the embodiments of the present application, after the point cloud generation device uses the point cloud generation model to extract features from the to-be-processed data to obtain the to-be-processed feature data, it can use the point cloud generation model to extract multi-dimensional disparity features from the to-be-processed feature data to obtain disparity feature data.
[0052] In some embodiments of the present application, the point cloud generation model may include a global feature extractor and a hybrid convolutional feature extractor.
[0053] In the embodiments of the present application, the global feature extractor may be a feature extractor based on Transformer and including an attention mechanism; the global feature extractor may include a self-attention layer, a first summation normalization layer, a feed-forward neural network, and a second summation normalization layer.
[0054] In some embodiments of the present application, when using the point cloud generation model to extract multi-dimensional disparity features from the to-be-processed feature data to obtain disparity feature data, the global feature extractor can be used to perform convolutional downsampling processing on the to-be-processed feature data to obtain the downsampled disparity features, and perform global feature extraction processing based on the downsampled disparity features and the position encoding information corresponding to the to-be-processed data to obtain the first disparity feature; use the hybrid convolutional feature extractor to perform multi-scale local disparity feature extraction processing on the to-be-processed feature data to obtain multi-scale local disparity feature information, and perform hybrid convolutional feature extraction processing based on the multi-scale local disparity feature information to obtain the second disparity feature; determine the disparity feature data based on the first disparity feature and the second disparity feature.
[0055] In some embodiments of the present application, determining the disparity feature data based on the first disparity feature and the second disparity feature can be understood as using the first disparity feature and the second disparity feature as the disparity feature data; in the actual operation process of determining the disparity feature data, the first disparity feature and the second disparity feature can be concatenated to obtain the disparity feature data.
[0056] In the embodiments of the present application, the present application does not limit the acquisition method of the position encoding information corresponding to the data to be processed. For example, the position encoding information can be determined according to the position information of the binocular camera, or the position encoding information can be determined according to the lidar data.
[0057] In the embodiments of the present application, the global feature extraction process may include using a self-attention layer, a first sum normalization layer, a feed-forward neural network, and a second sum normalization layer to repeatedly perform the feature extraction process and the three-dimensional convolutional upsampling process on the input first disparity feature and position encoding information; for example, as Figure 2 shown, after performing the downsampling process 21 of three-dimensional convolution on the feature data to be processed, the downsampled disparity feature 22 can be obtained. The downsampled disparity feature can be understood as a kind of coarse-grained disparity feature, which can significantly reduce the computing power; then the downsampled disparity feature 22 and the three-dimensional position encoding information 23 corresponding to the data to be processed are sequentially input into the self-attention layer 24, the first sum normalization layer 25, the feed-forward neural network 26, and the second sum normalization layer 27, and the processing of these four network layers is repeatedly performed N times, and then after the upsampling 28 of three-dimensional convolution, the first disparity feature 29 is output; where N≥2.
[0058] In the embodiments of the present application, the specific value of N for repeatedly performing N times is not limited in the present application. In the actual application process, the value of N can be determined by comprehensively considering the requirements for computing efficiency and the model accuracy of the global feature extractor.
[0059] In the embodiments of the present application, the global feature extractor can use the global self-attention mechanism for global feature fusion. The advantage is that it can extract global context modeling information. Since the amount of calculation of global information is often large, the downsampling of three-dimensional convolution is used to obtain the coarse-grained disparity feature, and then the coarse-grained disparity feature is processed, which can significantly reduce the computing power requirement; at the same time, the hybrid convolutional feature extractor focuses more on the perception of local information. Based on the global feature extractor and the hybrid convolutional feature extractor, the feature extraction can be greatly enriched to improve the generation effect of the point cloud.
[0060] In some embodiments of the present application, when performing hybrid convolution feature extraction processing based on multi-scale local disparity feature information to obtain the second disparity feature, three-dimensional convolution processing, spatial convolution processing, and disparity convolution processing may be performed on the local disparity feature information at the target scale in the multi-scale local disparity feature information to obtain the third feature information; then, the multi-scale image features of the image data are fused using attention weights to obtain the fourth feature information; thereby, the second disparity feature is determined based on the third feature information and the fourth feature information.
[0061] In an embodiment of the present application, the hybrid convolution feature extractor may be constructed based on a convolutional neural network. For example, the hybrid convolution feature extractor may be constructed based on the network structure of Hourglass or based on the network structure of U-Net.
[0062] In an embodiment of the present application, the target scale may be any one of the scales in the multi-scale.
[0063] In an embodiment of the present application, when determining the second disparity feature based on the third feature information and the fourth feature information, the third feature information and the fourth feature information may be fused to obtain a fusion result, and the hybrid convolution feature extractor is used to continue performing feature extraction at subsequent scales on the fusion result, including the process of downsampling or upsampling, and finally the second disparity feature is obtained.
[0064] Exemplarily, as Figure 3 shown, the hybrid convolution feature extractor may extract multi-scale disparity features from the feature data 33 to be processed by means of downsampling 31 and upsampling 32, that is, multi-scale local disparity features are obtained to obtain the second disparity feature 34; wherein, in the above process of extracting multi-scale disparity features, the features at a certain scale may be processed as Figure 4 shown. This processing process includes sequentially performing three-dimensional convolution processing 36, spatial convolution processing 37, and disparity convolution processing 38 on the local disparity feature information 35 at the target scale, that is, this scale, to obtain the third feature information 39, and using the multi-scale attention weights 310 to process the multi-scale image features 311 to obtain the fourth feature information 312, and then fusing the third feature information and the fourth feature information 313. It can be understood that after the features at a certain scale are processed as Figure 4 shown, the hybrid convolution feature extractor may continue to perform subsequent downsampling and / or upsampling processes on the fused result to obtain the second disparity feature; through the hybrid convolution feature extractor, the computational efficiency can be significantly improved while effectively extracting disparity features.
[0065] In the embodiments of the present application, the multi-scale image features can be obtained by extracting multi-scale image features from the image data, and the present application does not limit the extraction method of the multi-scale image features.
[0066] In the embodiments of the present application, the three-dimensional convolution processing is a processing that performs convolution simultaneously in the spatial and temporal dimensions, where the space includes length and width; the spatial convolution processing represents the processing of convolving the disparity features in the spatial dimension; the disparity convolution processing represents the processing of convolving the disparity features in the disparity dimension.
[0067] Exemplarily, the three-dimensional convolution is represented as K_s×K_s×K_d, and the three-dimensional convolution can be decomposed into a spatial convolution K_s×K_s×1 and a disparity convolution 1×1×K_d to utilize the spatial convolution to implement the spatial convolution processing and utilize the disparity convolution to implement the disparity convolution processing. The execution order of the spatial convolution processing and the disparity convolution processing is not limited in the present application; for example, the way of first performing the spatial convolution processing and then performing the disparity convolution processing is as Figure 5 shown. After performing the spatial convolution processing 42 on the disparity feature 41, the disparity convolution processing 43 can be performed to obtain the disparity feature 44 after the spatial convolution and the disparity convolution, which can realize the extraction of the disparity features in different dimensions, greatly reduce the number of calculation parameters and computing power, and thus improve the calculation efficiency.
[0068] Step 103: Use the point cloud generation model to optimize the multi-scale disparity features of the disparity feature data to obtain the optimized disparity data.
[0069] In the embodiments of the present application, after the point cloud generation device extracts the multi-dimensional disparity features of the to-be-processed feature data by using the point cloud generation model to obtain the disparity feature data, the point cloud generation model can be used to optimize the multi-scale disparity features of the disparity feature data to obtain the optimized disparity data.
[0070] In some embodiments of the present application, the point cloud generation model may include an optimization iterator for optimizing the multi-scale disparity features of the disparity feature data.
[0071] In the embodiments of the present application, the optimization iterator can be a model based on a Convolutional Gated Recurrent Unit (ConvGRU) and including a multi-level hidden state update and an attention selection mechanism. Based on the optimization iterator, the disparity features can be further refined and optimized.
[0072] In some embodiments of the present application, when the point cloud generation device optimizes the multi-scale disparity features of the disparity feature data using the point cloud generation model to obtain the optimized disparity data, it can use an optimization iterator to fuse the image features of the image data and the disparity feature data in the data to be processed to obtain the fused features; then determine the disparity correction value according to the first feature information at different scales in the fused features; and further optimize the initial disparity value according to the disparity correction value to obtain the optimized disparity data.
[0073] In the embodiments of the present application, the fused features can be presented in the form of a correlation coefficient feature pyramid; the fused features can include the first feature information at multiple scales.
[0074] In the embodiments of the present application, the process of the optimization iterator performing the optimization process of the disparity features can be understood as determining the multi-scale correlation coefficient feature pyramid after fusing the image features and the disparity feature data, and then calculating the disparity correction value for the first feature information at each scale in the multi-scale correlation coefficient feature pyramid, so as to obtain the optimized disparity data according to the disparity correction value and the initial disparity value, and complete the optimization of the disparity features.
[0075] In some embodiments of the present application, when optimizing the initial disparity value according to the disparity correction value to obtain the optimized disparity data, the disparity correction value corresponding to each scale can be superimposed on the initial disparity value and then passed to the next scale for further correction until the disparity correction value corresponding to the last scale is obtained as the disparity correction result, which is superimposed on the initial disparity value to obtain the optimized disparity data.
[0076] In the embodiments of the present application, the initial disparity value can be calculated according to the disparity feature data, and the present application does not limit the acquisition method of the initial disparity value. For example, it can be obtained after feature extraction, cost calculation, and cost aggregation of the disparity feature data.
[0077] Exemplarily, as Figure 6 shown, after fusing the image features, including the image feature 51 of the left camera and the image feature 52 of the right camera, and the disparity feature data 53, the fused feature 54 presented in the form of a correlation coefficient feature pyramid is obtained, and then the disparity correction value at the current scale is calculated according to the first feature information at each scale in the fused feature, and based on the result of the superposition of the disparity correction value at the current scale and the initial disparity value 55 and the first feature information of the next scale, the disparity correction value of the next scale is continuously calculated, and so on, until the disparity correction value 56 of the last scale is obtained. The disparity correction value 56 of the last scale is superimposed on the initial disparity value 55 to obtain the optimized disparity data 57.
[0078] In some embodiments of the present application, when the point cloud generation device determines the disparity correction value based on the first feature information at different scales in the fused features, the device may downsample the first feature information at multiple resolutions to obtain multiple second feature information; then fuse the multiple second feature information to obtain fused feature information; and then determine the disparity correction value based on the fused feature information.
[0079] In some embodiments of the present application, the first feature information may be downsampled by 32 times, 16 times, and 8 times to obtain resolutions of the original resolution of the first feature information. , as well as That is, the multiple resolutions may include the original resolution of the first feature information. , as well as ; Then the second feature information obtained after the three downsampling is fused to obtain fused feature information.
[0080] For example, if the feature update is only performed at a fixed resolution, the image receptive field will be relatively small even if the update is repeated multiple times, which has a negative impact on weak textures and large-scale targets. Figure 7 As shown, when executing the calculation of the disparity correction value at a certain scale, the embodiment of the present application can first perform downsampling 62 of 32 times the resolution, downsampling 63 of 16 times the resolution, and downsampling 64 of 8 times the resolution on the first feature information 61 at the scale, and fuse the second feature information obtained after the above-mentioned downsampling at different magnifications to obtain fused feature information 65; thereby, the disparity correction value at the scale can be calculated based on the fused feature information at the scale.
[0081] That is to say, in the embodiments of the present application, the disparity features can be refined by optimizing the iterator, and the local details can be gradually corrected through multi-stage correction, and finally a precisely optimized disparity can be obtained in the last layer, which takes into account both computational efficiency and precision optimization.
[0082] Step 104: Generate depth map data according to the optimized disparity data, and generate a point cloud according to the depth map data.
[0083] In an embodiment of the present application, after the point cloud generation device optimizes the multi-scale disparity features of the disparity feature data using the point cloud generation model to obtain the optimized disparity data, it can generate depth map data based on the optimized disparity data and generate a point cloud based on the depth map data.
[0084] In the embodiments of the present application, the point cloud generated by the above method has good integrity and density, accurately represents the spatial environment, and is also richer in the representation of details. Therefore, when performing subsequent environment perception tasks based on such a point cloud, the reliability and robustness of environment perception can be greatly improved.
[0085] In some embodiments of the present application, after obtaining the point cloud, map data can be constructed based on the point cloud, and then related operations such as road planning, positioning, or navigation can be implemented based on the map data.
[0086] In some embodiments of the present application, before the point cloud generation device extracts features from the data to be processed using the point cloud generation model to obtain the feature data to be processed, the following steps 105 to 109 may further be included: Step 105: Obtain first scene data determined through three-dimensional modeling.
[0087] In the embodiments of the present application, the point cloud generation device may obtain first scene data determined through three-dimensional modeling.
[0088] In the embodiments of the present application, the first scene data is scene data generated through three-dimensional modeling; for example, three-dimensional scene modeling can be implemented through three-dimensional modeling software to obtain the first scene data.
[0089] Step 106: Obtain second scene data; wherein, the second scene data is scene data reconstructed from a simulation scene according to the first image information.
[0090] In the embodiments of the present application, the point cloud generation device may obtain second scene data; wherein, the second scene data is scene data reconstructed from a simulation scene according to the first image information.
[0091] In the embodiments of the present application, the specific manner of simulation scene reconstruction is not limited in the present application. For example, it can be realized through any method such as multi-view stereo reconstruction (MVS) reconstruction, neural radiance fields (NeRF), 3D Gaussian splatting (3DGS), or four-dimensional Gaussian reconstruction, simultaneous localization and mapping (SLAM) combined with visual reconstruction, and data twin. These methods can quickly convert the real-world scene into three-dimensional scene data, thereby realizing the construction of low-cost and large-scale simulation scene data.
[0092] In the embodiments of the present application, through three-dimensional modeling and simulation scene reconstruction, it is possible to generate scene data covering diverse scenarios such as retail, industry, home, and autonomous driving, including indoor and outdoor, dynamic and static, weak texture, transparent objects, and objects with highly reflective surfaces. Subsequently, when synthesizing sensor data based on this scene data, it is possible to generate tens of millions of high-fidelity stereo image data and point clouds covering diverse scenarios.
[0093] Step 107: Perform sensor data synthesis on the first scene data and the second scene data based on the verified emulator and pre-configured parameters to obtain synthesized image data and synthesized point clouds; wherein, the pre-configured parameters include simulation configuration parameters of the camera and the radar, as well as environmental parameters.
[0094] In the embodiments of the present application, after the point cloud generation device obtains the first scene data determined through three-dimensional modeling and obtains the second scene data; wherein, the second scene data is the scene data obtained through simulation scene reconstruction based on the first image information, it can perform sensor data synthesis on the first scene data and the second scene data based on the verified emulator and pre-configured parameters to obtain synthesized image data and synthesized point clouds; wherein, the pre-configured parameters include simulation configuration parameters of the camera and the radar, as well as environmental parameters.
[0095] In the embodiments of the present application, the verified emulator can be used to parse and present the scene data.
[0096] In the embodiments of the present application, the simulation configuration parameters of the camera can include parameters such as the internal parameters of the binocular camera, the baseline, and the focal length, etc.; the simulation configuration parameters of the radar can include the number of beams, angular resolution, point cloud resolution, and ranging accuracy, etc.; the environmental parameters can include lighting, material, object layout, and dynamic physical simulation parameters, etc.
[0097] In the embodiments of the present application, sensor data synthesis can be understood as a synthesis of imaging data, that is, synthesizing images and point clouds from different perspectives in the scene to obtain synthesized image data and synthesized point clouds.
[0098] Step 108: Determine the training data set based on the synthesized image data and the synthesized point clouds.
[0099] In the embodiments of the present application, after the point cloud generation device performs sensor data synthesis on the first scene data and the second scene data based on the verified emulator and pre-configured parameters to obtain synthesized image data and synthesized point clouds, it can determine the training data set based on the synthesized image data and the synthesized point clouds.
[0100] In some embodiments of the present application, when the point cloud generation device determines the training dataset based on the synthesized image data and the synthesized point cloud, it may use the data quality assessment model to classify the synthesized image data and the synthesized point cloud to obtain the classified image data and the classified point cloud; and determine the training dataset according to the image data with qualified classification results in the classified image data and the point cloud with qualified classification results in the classified point cloud.
[0101] In the embodiments of the present application, to ensure that the data in the training dataset can meet the requirements for generating a dense point cloud, the data quality assessment model can be used to classify the synthesized image data and the synthesized point cloud, so as to determine the training dataset according to the image data with qualified classification results and the point cloud with qualified classification results; that is to say, the data quality assessment model can be a classification model.
[0102] In some embodiments of the present application, in addition to determining the image data with qualified classification results and the point cloud with qualified classification results as the training dataset, it is also possible to obtain the image data and point cloud that have undergone manual verification. The image data and point cloud that have undergone manual verification refer to the image data and point cloud obtained by manually verifying the classified image data and the classified point cloud again. In this way, the fuzzy or unreliable samples can be identified and removed, and then the qualified data that has undergone secondary verification can be determined as the training dataset, which can further improve the quality of the dataset.
[0103] In some embodiments of the present application, the data quality assessment model can also be updated using the data with qualified classification results and the data with unqualified classification results to obtain an updated data quality assessment model that can more accurately classify qualified and unqualified data.
[0104] Exemplarily, as Figure 8 shown, after obtaining the synthesized data 71, where the synthesized data includes synthesized image data and synthesized point cloud, the data quality assessment model can be used to classify the synthesized data 72 to obtain qualified data 73 and unqualified data 74. For example, the qualified data can be data with the picture centered and the object shown completely, and the unqualified data can be data with an overly deviated display angle or unclear data, etc.; then the classified data can be verified, for example, by using the method of manual verification, so as to obtain the verified qualified data 75, and determine the training dataset 76 according to the verified qualified data; in addition, both the qualified data and the unqualified data in the classified data can be used to iteratively update the data quality assessment model.
[0105] Step 109: Train the initial model using the training dataset to obtain a point cloud generation model.
[0106] In an embodiment of the present application, after determining a training dataset based on the synthesized image data and the synthesized point cloud, the point cloud generation device can use the training dataset to train an initial model to obtain a point cloud generation model.
[0107] In an embodiment of the present application, the initial model represents an untrained model for point cloud generation. The initial model may include an untrained initial global feature extractor, an initial hybrid convolutional feature extractor, an initial optimization iterator, an initial monocular depth estimation network, an initial multi-scale feature extraction network, and an initial point cloud feature extraction network.
[0108] In an embodiment of the present application, based on the synthesis of multi-source sensor data for rich-source synthetic scenes, the sources of the synthetic scenes include the construction of a pure simulation environment through 3D modeling and the migration from real to simulation (Real2Sim) scene through simulation scene reconstruction; different scene construction methods in the above ways can greatly increase data richness; multi-source sensors can support the synthesis of binocular RGB cameras, infrared cameras, and lidar point clouds, so that image data and radar data with variable resolutions and field of view ranges can be generated based on rich sensor configurations, greatly improving data richness.
[0109] In an embodiment of the present application, training the initial model with the above-mentioned massive simulation synthesis data can effectively solve the problem of the degradation of cross-domain performance of traditional stereo matching methods and multi-source sensor data fusion in real complex scenes, so that the finally obtained point cloud generation model can maintain strong generalization ability in rich scenes of the real world, different sensor combinations or configuration methods.
[0110] An embodiment of the present application provides a point cloud generation method. A point cloud generation device can use a point cloud generation model to extract features from data to be processed, obtaining feature data to be processed. Among them, the data to be processed at least includes image data. The point cloud generation model is obtained by training an initial model using a training data set generated by a verified simulator. The verified simulator is verified based on first data, first internal and external parameters, second data, and second internal and external parameters. The first data is obtained by using a binocular camera and a lidar to collect data from a calibration scene. The first internal and external parameters represent the internal and external parameters of the binocular camera and the lidar obtained using the first data. The second data is obtained by simulating the data collection of the binocular camera and the lidar in the calibration scene using the simulator to be verified. The second internal and external parameters represent the internal and external parameters of the binocular camera and the lidar obtained using the second data. Use the point cloud generation model to extract multi-dimensional disparity features from the feature data to be processed, obtaining disparity feature data. Use the point cloud generation model to optimize the multi-scale disparity features of the disparity feature data, obtaining optimized disparity data. Generate depth map data according to the optimized disparity data, and generate a point cloud according to the depth map data.
[0111] Based on the above embodiment, in another embodiment of the present application, exemplarily, as Figure 9 shown, after the images of the binocular camera collected from the real environment, including the image 81 of the left camera and the image 82 of the right camera, are input into the point cloud generation model, the monocular depth estimation network 83 and the multi-scale feature extraction network 84 in the point cloud generation model can be used to extract depth features of the images of the binocular camera. Then, the extracted depth features, including the depth feature 85 of the left camera and the depth feature 86 of the right camera, are fused to obtain a mixed cost volume feature 87. Then, the mixed cost volume feature 87 is input into the global feature extractor 88 and the mixed convolutional feature extractor 89 to extract multi-dimensional disparity features. Then, the obtained disparity feature data 810 is input into the optimization iterator 811 to optimize the disparity features, obtaining optimized disparity data 812. Thus, depth map data can be generated according to the optimized disparity data, and a point cloud can be generated according to the depth map data.
[0112] Exemplarily, as Figure 10As shown in the figure, after collecting the images of the binocular camera from the real environment, including the image 91 of the left camera and the image 92 of the right camera, and the radar data 93 obtained by the lidar, the above data can be input into the point cloud generation model. The monocular depth estimation network 94 and the multi-scale feature extraction network 95 in the point cloud generation model are used to extract the depth features of the images, and at the same time, the point cloud feature extraction network 96 in the point cloud generation model is used to extract the point cloud features of the radar data. Then, the extracted depth features, including the depth feature 97 of the left camera and the depth feature 98 of the right camera, and the point cloud feature 99 are fused to obtain the mixed cost volume feature 910. Then, the mixed cost volume feature is input into the global feature extractor 911 and the mixed convolutional feature extractor 912 to extract multi-dimensional disparity features. Then, the obtained disparity feature data 913 is input into the optimization iterator 914 to optimize the disparity features and obtain the optimized disparity feature 915. Thus, the depth map data can be generated according to the optimized disparity data, and then the point cloud can be generated according to the depth map data.
[0113] In some embodiments of the present application, as Figure 11 shown, the first scene data 111 can be obtained through 3D modeling, and the second scene data 112 can be obtained through Real2Sim for simulation scene reconstruction. The specific method of implementing Real2Sim is not limited in this application. For example, MVS reconstruction, NeRF, 3D Gaussian splash, etc. Then, the first scene data and the second scene data are imported into the simulator 113, and the parameters of the binocular camera, lidar, and environmental parameters are configured 114. Based on the random pose, the synthetic data 115 is obtained. The synthetic data can include synthetic image data and synthetic point cloud.
[0114] Exemplarily, as Figure 12 shown, the images of the binocular camera collected from the real environment, including the image 121 of the left camera and the image 122 of the right camera, after being input into the point cloud generation model, the monocular depth estimation network and the multi-scale feature extraction network in the point cloud generation model can be used to extract the depth features of the images of the binocular camera 123. Then, the extracted depth features, including the depth feature of the left camera and the depth feature of the right camera, are fused to obtain the mixed cost volume feature 124. Then, multi-dimensional disparity features are extracted from the mixed cost volume feature 125, including inputting the mixed cost volume feature into the global feature extractor and the mixed convolutional feature extractor to obtain multi-dimensional disparity features. Then, multi-scale disparity feature optimization processing is performed on the disparity features 126, including inputting the obtained disparity features into the optimization iterator to optimize the multi-dimensional disparity features and obtain the optimized disparity data. Thus, the depth map data can be generated according to the optimized disparity data 127, and then the point cloud 128 can be generated according to the depth map data 127.
[0115] Exemplarily, as Figure 13 shown, after acquiring the images of the binocular camera from the real environment, including the image 131 of the left camera and the image 132 of the right camera, and the lidar data 133 obtained by the lidar, the above data can be input into the point cloud generation model. The monocular depth estimation network and the multi-scale feature extraction network in the point cloud generation model are used to extract depth features from the images 134, and at the same time, the point cloud feature extraction network in the point cloud generation model is used to extract point cloud features from the lidar data 135. Then, after fusing the extracted depth features, including the depth features of the left camera and the right camera, and the point cloud features, a mixed cost volume feature 136 is obtained, and then multi-dimensional disparity features are extracted from the mixed cost volume feature 137, including inputting the mixed cost volume feature into the global feature extractor and the mixed convolutional feature extractor to obtain multi-dimensional disparity features, and then performing optimization processing on the multi-scale disparity features of the disparity features 138, including inputting the disparity features into the optimization iterator to optimize the disparity features and obtain the optimized disparity features; thus, depth map data 139 can be generated according to the optimized disparity data, and then point cloud 1310 and point cloud 1311 can be generated according to the depth map data.
[0116] In some embodiments of the present application, when it is necessary to enhance the generalization ability of the point cloud generation model in a specific type of scene, a small-sample real data acquisition method can also be used for model fine-tuning. During fine-tuning, the images and the lidar point cloud are used as the original inputs of the model, and synthetic data is extracted from the high-precision three-dimensional scene reconstruction data according to the real-time pose of the sensor in space to obtain the model training ground truth. Only samples with high data accuracy and general point cloud thickening effect will be added to the model fine-tuning. In this way, the model performance will be further improved in a specific scene; that is to say, the point cloud generation model trained based on a large amount of synthetic data in the embodiments of the present application has high generalization in real scene data. When it is necessary to further improve the point cloud thickening performance in a certain type of scene, a small amount of real machine data fine-tuning method can be used to further improve the performance of the point cloud generation model.
[0117] In the embodiments of the present application, the point cloud generation model includes a monocular depth estimation network, a multi-scale feature extraction network, a point cloud feature extraction network, a global feature extractor, a mixed convolutional feature extractor, and an optimization iterator. By proposing the point cloud generation model with the above structure, the present application can maintain good generalization ability in any real complex scene and generate point clouds with good integrity and density.
[0118] In some embodiments of the present application, the global feature extractor may be a feature extractor based on Transformer and incorporating an attention mechanism; the global feature extractor may include a self-attention layer, a first sum normalization layer, a feed-forward neural network, and a second sum normalization layer; the global feature extractor can perform global feature fusion, and its advantage lies in being able to extract global context modeling information. However, since the global information calculation amount of Transformer is often large, the present application adopts a method of processing the coarsened disparity features obtained after downsampling, which can significantly reduce the computing power; but this part focuses on global feature extraction and has a relatively weak perception of local information. Therefore, in the overall architecture, it is necessary to combine a hybrid convolutional feature extractor and an optimization iterator to further enrich feature extraction.
[0119] In some embodiments of the present application, the hybrid convolutional feature extractor can automatically perform feature extraction, has the advantages of weight parameter sharing and local connection, retains the spatial structure relationship while performing feature extraction, and can efficiently process large-scale data. Since the disparity feature has one more dimension of disparity than general image features, the hybrid convolutional feature extractor of the present application fuses multi-scale features through downsampling and then upsampling, and when extracting features at a certain scale, it adopts a method of partial three-dimensional convolution plus two-dimensional spatial convolution plus one-dimensional disparity convolution for feature extraction, and performs feature extraction through an attention mechanism guided by multi-scale image features. Compared with the current related stereo matching networks that use three-dimensional convolution for feature extraction, it can significantly improve the calculation efficiency while effectively extracting disparity features.
[0120] In some embodiments of the present application, an optimization iterator can be used to optimize the disparity features. The process of the optimization iterator performing the optimization process of the disparity features can be understood as determining a multi-scale correlation coefficient feature pyramid after fusing the image features and the disparity feature data, and then calculating the disparity correction value for the first feature information at each scale in the multi-scale correlation coefficient feature pyramid, so as to obtain the optimized disparity data according to the disparity correction value and the initial disparity value, and complete the optimization of the disparity features; in addition, the optimization iterator has the characteristic of multi-level hidden state update, and can perform feature update simultaneously on the 8-fold, 16-fold, and 32-fold downsampled features, so that the optimization iterator can produce good responses to regions of different sizes and can adapt to large-scale, highly random, and diverse synthetic data.
[0121] In some embodiments of the present application, for robot scene interaction, refined perception is an important prerequisite for a robot to perform physical interaction and good manipulation. An accurate three-dimensional structure perception of the target scene by the robot is achieved through a point cloud generation model generalized based on a large amount of synthetic data. The multi-modal data fusion technology can utilize the different sensor advantages of cameras and lidar to identify the accurate position and attitude information of the target object, thus laying a solid perception foundation for a series of actions such as obstacle avoidance, grasping, placing, and assembling by the robot. The point cloud generation model can become a key support for the robot manipulation to move from the laboratory to practical applications by enhancing the environmental perception ability and operation planning accuracy. In the field of autonomous driving, the point cloud generation model can also improve the perception ability of autonomous driving vehicles. By realizing the generalization of the model trained with a large amount of simulation data in the real scene, it can help improve the use effect of stereo matching in the autonomous driving field, overcome the bottleneck problem of data scarcity, and enhance the three-dimensional scene perception ability.
[0122] In some embodiments of the present application, the effect of point cloud thickening depends on the point cloud generation model trained based on large-scale synthetic data, and the synthetic data depends on the generation of the simulator. Therefore, the alignment degree between the simulator synthetic data and the real machine collected data will affect the final point cloud thickening effect. Therefore, the present application proposes a method for calibrating the simulator to improve the generation effect of the synthetic data, thereby ensuring the point cloud thickening effect. Exemplarily, as Figure 14 shown, a calibration scene can be set up in the real environment (step 1301), and then the first data in the calibration scene is collected using a binocular camera and lidar (step 1302). Then, the first internal and external parameters of the binocular camera and lidar are calculated using the first data (step 1303). The calibration scene can also be simulated based on the first internal and external parameters and the simulator to be calibrated to synthesize images and point clouds to obtain the second data (step 1304). Furthermore, the second internal and external parameters of the binocular camera and lidar are calculated based on the second data (step 1305). Then, qualitative and quantitative evaluations are performed based on the first data, the second data, the first internal and external parameters, and the second internal and external parameters to complete the calibration of the simulator to be calibrated (step 1306).
[0123] Exemplarily, in the first step, to construct the same scenario for the real machine and the simulation, a calibration room, i.e., a calibration scenario, can be built in the real world first. There can be multiple calibration boards, QR codes, high-reflectivity objects for radar in the calibration room, such as high-reflectivity points or high-reflectivity strips, etc., as well as some objects with fixed shapes, such as cubes, cuboids, etc. The positions and postures of different targets in the real machine environment are calculated through calibration methods. Then, targets of the same size are constructed in the simulator, and the targets are placed with the same pose as in the real environment. In the second step, real machine data can be collected and the internal and external parameters of the sensors can be calculated, including collecting data of different objects in the real calibration room through sensors such as binocular cameras and lidar, i.e., the first data. Then, the internal parameters of the sensors and the external parameters of the cameras and lidar in different frames of data are calculated through the calibration relationship, i.e., the first internal and external parameters. In the third step, data synthesis of the sensors is performed through the simulator, including simulating the data of different objects collected by the binocular camera and lidar in the real calibration room in the simulator to render and synthesize images and lidar point clouds of different frames to obtain the second data, and calculating the internal and external parameters of the simulated binocular camera and lidar in the simulator again according to the second data, i.e., the second internal and external parameters. In the fourth step, the alignment degree between the simulation rendering result and the real machine collected data is evaluated, that is, qualitative evaluation and quantitative evaluation are carried out using the first data, the second data, the first internal and external parameters, and the second internal and external parameters. For example, in qualitative evaluation, the visualized first data and second data can be merged to evaluate whether the data is consistent. For example, it can be judged whether there are situations such as offsets and stratifications. Quantitative evaluation can evaluate the coordinate errors of the feature points of the feature objects in the first data and the second data, such as the coordinate errors of the inner corner points of the calibration board and the key points of other feature objects, and can also evaluate whether the second internal and external parameters calculated by the simulation are consistent with the first internal and external parameters of the real sensor, or whether the error magnitude is within a reasonable range. Thus, it can be judged whether the simulation data and the real machine data can be effectively aligned based on the results of comprehensive qualitative and quantitative evaluations, that is, it can be judged whether the simulator can effectively simulate real data.
[0124] In the embodiment of the present application, the average distance alignment error of the calibrated simulator is only 0.2 mm, and the average angle alignment error is only 0.01°. Therefore, the calibrated simulator is used to generate massive scene data and is used to train the initial model to obtain a point cloud generation model, which can effectively improve the generalization ability of the point cloud generation model, so that in any scenario, a point cloud with good integrity and density can be generated.
[0125] Exemplarily, the method of the embodiments of the present application can be applied to a variety of actual scenarios. For example, in the scenario of robot interaction, refined perception is an important prerequisite for robots to perform physical interactions and good manipulations. Through the point cloud thickening technology based on the generalization of massive synthetic data, the robot can achieve accurate three-dimensional structure perception of the target scene. The multi-modal data fusion technology can utilize the different sensor advantages of cameras and lidar to identify the accurate position and attitude information of the target object, thus laying a solid perception foundation for a series of actions such as obstacle avoidance, grasping, placing, and assembling of the robot. The point cloud thickening technology can become a key support for the transformation of robot manipulation from the laboratory to practical applications by enhancing the environmental perception ability and operation planning accuracy. It can also be applied to the field of autonomous driving. The multi-modal point cloud thickening technology can also improve the perception ability of autonomous driving vehicles. By realizing the generalization of the training model of massive simulation data in the real scene, it can help improve the use effect of stereo matching in the field of autonomous driving, overcome the bottleneck problem of data scarcity, and improve the three-dimensional scene perception ability. It can also be applied to scene three-dimensional reconstruction. The method proposed in the present application can be used for three-dimensional reconstruction. For example, by modifying the input of the laser point cloud into thickened point cloud and color data as new inputs, it can simultaneously utilize three-dimensional data and target color texture information, thereby improving the three-dimensional reconstruction effect.
[0126] In summary, through the alignment verification of simulation-real machine data, massive data synthesis, and a network architecture with easy-to-learn features, the point cloud generation model trained by the embodiments of the present application can have good simulation-real machine transfer ability, and can effectively generate point clouds when using real machine data as input, enhancing the integrity and density of the point clouds.
[0127] An embodiment of the present application provides a point cloud generation method. A point cloud generation device can obtain data to be processed; use a point cloud generation model to extract features from the data to be processed to obtain feature data to be processed; wherein, the data to be processed at least includes image data; the point cloud generation model is obtained by training an initial model using a training data set generated by a verified simulator; the verified simulator is verified based on first data, first internal and external parameters, second data, and second internal and external parameters; the first data is obtained by using a binocular camera and a lidar to collect data from a calibration scene; the first internal and external parameters represent the internal and external parameters of the binocular camera and the lidar calculated using the first data; the second data is obtained by using the simulator to be verified to simulate the data collection of the binocular camera and the lidar in the calibration scene; the second internal and external parameters represent the internal and external parameters of the binocular camera and the lidar calculated using the second data; use the point cloud generation model to extract multi-dimensional disparity features from the feature data to be processed to obtain disparity feature data; use the point cloud generation model to optimize the multi-scale disparity features of the disparity feature data to obtain optimized disparity data; generate depth map data according to the optimized disparity data, and generate a point cloud according to the depth map data.
[0128] Based on the above embodiment, in another embodiment of the present application, a point cloud generation device is provided, as Figure 15 shown, the point cloud generation device 1 may include a first generation unit 11, a second generation unit 12, and a training unit 13.
[0129] The first generation unit 11 can be used to extract features from the data to be processed using a point cloud generation model to obtain feature data to be processed; wherein, the data to be processed at least includes image data; the point cloud generation model is obtained by training an initial model using a training data set generated by a verified simulator; the verified simulator is verified based on first data, first internal and external parameters, second data, and second internal and external parameters; the first data is obtained by using a binocular camera and a lidar to collect data from a calibration scene; the first internal and external parameters represent the internal and external parameters of the binocular camera and the lidar obtained using the first data; the second data is obtained by using the simulator to be verified to simulate the data collection of the binocular camera and the lidar in the calibration scene; the second internal and external parameters represent the internal and external parameters of the binocular camera and the lidar obtained using the second data; and use the point cloud generation model to extract multi-dimensional disparity features from the feature data to be processed to obtain disparity feature data; and use the point cloud generation model to optimize the disparity features of the disparity feature data to obtain optimized disparity data.
[0130] The second generation unit 12 can be used to generate depth map data according to the optimized disparity data, and generate a point cloud according to the depth map data.
[0131] In some embodiments of the present application, the point cloud generation model may include a global feature extractor and a hybrid convolutional feature extractor; the first generation unit 11 may also be configured to perform convolutional downsampling on the to-be-processed feature data by using the global feature extractor to obtain the downsampled disparity feature, and perform global feature extraction on the downsampled disparity feature and the position encoding information corresponding to the to-be-processed data to obtain the first disparity feature; and perform multi-scale local disparity feature extraction on the to-be-processed feature data by using the hybrid convolutional feature extractor to obtain multi-scale local disparity feature information, and perform hybrid convolutional feature extraction on the multi-scale local disparity feature information to obtain the second disparity feature; and determine the disparity feature data according to the first disparity feature and the second disparity feature.
[0132] In some embodiments of the present application, the point cloud generation model may include an optimization iterator; the first generation unit 11 may also be configured to fuse the image feature and the disparity feature data of the image data in the to-be-processed data by using the optimization iterator to obtain the fused feature; and determine the disparity correction value according to the first feature information at different scales in the fused feature; and optimize the initial disparity value according to the disparity correction value to obtain the optimized disparity data.
[0133] In some embodiments of the present application, the first generation unit 11 may also be configured to perform three-dimensional convolution processing, spatial convolution processing, and disparity convolution processing on the local disparity feature information at the target scale in the multi-scale local disparity feature information to obtain the third feature information; and fuse the multi-scale image features of the image data by using the attention weight to obtain the fourth feature information; and determine the second disparity feature based on the third feature information and the fourth feature information.
[0134] In some embodiments of the present application, the first generation unit 11 may also be configured to perform feature downsampling at multiple resolutions on the first feature information to obtain a plurality of second feature information; and fuse the plurality of second feature information to obtain the fused feature information; and determine the disparity correction value according to the fused feature information.
[0135] In some embodiments of the present application, the point cloud generation model may include a monocular depth estimation network, a multi-scale feature extraction network, and a point cloud feature extraction network; the first generation unit 11 may also be configured to, when the data to be processed includes image data and radar data, extract the monocular depth feature information of the image data by using the monocular depth estimation network, extract the multi-scale depth feature information of the image data by using the multi-scale feature extraction network; and extract the point cloud features of the radar data by using the point cloud feature extraction network; and determine the monocular depth feature information, the multi-scale depth feature information, and the point cloud features as the feature data to be processed; and when the data to be processed includes image data, extract the monocular depth feature information of the image data by using the monocular depth estimation network, extract the multi-scale depth feature information of the image data by using the multi-scale feature extraction network, and determine the monocular depth feature information and the multi-scale depth feature information as the feature data to be processed.
[0136] The training unit 13 may be configured to obtain the first scene data determined by 3D modeling; and obtain the second scene data; wherein, the second scene data is the scene data reconstructed by simulating a scene according to the first image information; and perform sensor data synthesis on the first scene data and the second scene data based on the verified simulator and pre-configured parameters to obtain the synthesized image data and the synthesized point cloud; wherein the pre-configured parameters include the simulation configuration parameters of the camera and the radar, and the environmental parameters; and determine the training data set based on the synthesized image data and the synthesized point cloud; and use the training data set to train the initial model to obtain the point cloud generation model.
[0137] In the embodiments of the present application, further, Figure 16 is a schematic diagram of the composition structure of the point cloud generation device proposed in the embodiments of the present application Figure 2 , as Figure 16 shown, the point cloud generation device 1 proposed in the embodiments of the present application may further include a processor 14, a memory 15 storing instructions executable by the processor 14; further, the point cloud generation device 1 may further include a communication interface 16, and a bus 17 for connecting the processor 14, the memory 15, and the communication interface 16.
[0138] In an embodiment of the present application, the above-mentioned processor 14 may be at least one of an Application Specific Integrated Circuit (ASIC), a Digital Signal Processor (DSP), a Digital Signal Processing Device (DSPD), a Programmable Logic Device (PLD), a Field Programmable Gate Array (FPGA), a Central Processing Unit (CPU), a controller, a microcontroller, and a microprocessor. It can be understood that for different devices, the electronic devices for implementing the functions of the above-mentioned processor may also be others, and the embodiments of the present application do not make specific limitations. The point cloud generation device 1 may further include a memory 15, and the memory 15 may be connected to the processor 14. Among them, the memory 15 is used to store executable program codes, and the program codes include computer operation instructions. The memory 15 may include a high-speed RAM memory and may also include a non-volatile memory, for example, at least two disk memories.
[0139] In an embodiment of the present application, the bus 17 is used to connect the communication interface 16, the processor 14, and the memory 15 and for mutual communication between these devices.
[0140] In an embodiment of the present application, the memory 15 is used to store instructions and data.
[0141] Further, in the embodiments of the present application, the above-mentioned processor 14 is configured to extract features from the data to be processed by using a point cloud generation model to obtain the feature data to be processed; wherein, the data to be processed includes at least image data; the point cloud generation model is obtained by training an initial model by using a training data set generated by a verified simulator; the verified simulator is verified based on the first data, the first internal and external parameters, the second data, and the second internal and external parameters; the first data is obtained by using a binocular camera and a lidar to collect data from a calibration scene; the first internal and external parameters represent the internal and external parameters of the binocular camera and the lidar obtained by using the first data; the second data is obtained by using the simulator to be verified to simulate the data collection of the binocular camera and the lidar in the calibration scene; the second internal and external parameters represent the internal and external parameters of the binocular camera and the lidar obtained by using the second data; the point cloud generation model is used to extract multi-dimensional disparity features from the feature data to be processed to obtain disparity feature data; the point cloud generation model is used to optimize the multi-scale disparity features of the disparity feature data to obtain the optimized disparity data; a depth map data is generated according to the optimized disparity data, and a point cloud is generated according to the depth map data.
[0142] In practical applications, the above-mentioned memory 15 may be a volatile memory, such as a Random-Access Memory (RAM); or a non-volatile memory, such as a Read-Only Memory (ROM), a flash memory, a Hard Disk Drive (HDD), or a Solid-State Drive (SSD); or a combination of the above types of memories, and provides instructions and data to the processor 14.
[0143] In addition, each functional module in this embodiment may be integrated into a processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional module.
[0144] When the integrated unit is implemented in the form of a software functional module and is not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the method of this embodiment. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.
[0145] An embodiment of this application provides a point cloud generation device, which is used to extract features from data to be processed by using a point cloud generation model to obtain feature data to be processed; wherein, the data to be processed at least includes image data; the point cloud generation model is obtained by training an initial model by using a training data set generated by a verified emulator; the verified emulator is verified based on first data, first internal and external parameters, second data, and second internal and external parameters; the first data is obtained by using a binocular camera and a lidar to collect data for a calibration scene; the first internal and external parameters represent the internal and external parameters of the binocular camera and the lidar obtained by using the first data; the second data is obtained by using the emulator to be verified to simulate the data collection of the binocular camera and the lidar in the calibration scene; the second internal and external parameters represent the internal and external parameters of the binocular camera and the lidar obtained by using the second data; the point cloud generation model is used to extract multi-dimensional disparity features from the feature data to be processed to obtain disparity feature data; the point cloud generation model is used to optimize the multi-scale disparity features of the disparity feature data to obtain optimized disparity data; depth map data is generated according to the optimized disparity data, and a point cloud is generated according to the depth map data.
[0146] It can be seen that the present application can pre-check the emulator to be checked, including completing the check by using the first data obtained by collecting data on a real calibration scene using a binocular camera and a lidar, the first internal and external parameters of the binocular camera and the lidar, the second data obtained by simulating the data collection scene of the binocular camera and the lidar using the emulator to be checked, and the internal and external parameters of the binocular camera and the lidar calculated using the second data, so as to determine whether the emulator to be checked has a good simulation effect, that is, whether it is possible to obtain data that is basically indistinguishable from the real calibration scene using the emulator to be checked. Thus, the initial model can be trained using the data generated by the emulator that passes the check to obtain a point cloud generation model, which can improve the generalization ability of the point cloud generation model. Furthermore, when generating a point cloud, the point cloud generation device can first use the constructed point cloud generation model to extract features from the data to be processed, and then extract multi-dimensional disparity features from the data to be processed feature data, so as to obtain disparity features of different dimensions. Then, the disparity features of different dimensions will be optimized with different scales of disparity features to further optimize the disparity features. Thus, depth map data can be generated based on the optimized disparity data, and a point cloud can be generated based on the depth map data, which can effectively improve the integrity and density of the generated point cloud.
[0147] Specifically, the program instructions corresponding to a point cloud generation method in this embodiment can be stored on storage media such as optical discs, hard disks, and USB flash drives. When the program instructions corresponding to a point cloud generation method in the storage media are read or executed by a point cloud generation device, the following steps are included: Using the point cloud generation model to extract features from the data to be processed to obtain the data to be processed feature data; wherein, the data to be processed at least includes image data; the point cloud generation model is obtained by training the initial model using a training data set generated by the emulator that passes the check; the emulator that passes the check is obtained by checking the emulator to be checked based on the first data, the first internal and external parameters, the second data, and the second internal and external parameters; the first data is obtained by collecting data on the calibration scene using a binocular camera and a lidar; the first internal and external parameters represent the internal and external parameters of the binocular camera and the lidar obtained using the first data; the second data is obtained by simulating the data collection of the binocular camera and the lidar in the calibration scene using the emulator to be checked; the second internal and external parameters represent the internal and external parameters of the binocular camera and the lidar obtained using the second data; Using the point cloud generation model to extract multi-dimensional disparity features from the data to be processed feature data to obtain disparity feature data; Using the point cloud generation model to optimize the multi-scale disparity features of the disparity feature data to obtain the optimized disparity data; Generating depth map data based on the optimized disparity data and generating a point cloud based on the depth map data.
[0148] In an embodiment of the present application, an electronic device is provided. The electronic device may include a point cloud generation device, which can be used to execute the foregoing point cloud generation method; the electronic device can be used to construct map data based on the point cloud generated by the point cloud generation device; and perform corresponding operations according to the map data.
[0149] Exemplarily, the electronic device can be a robot, and the point cloud generation device can be deployed in the robot. Thus, the robot can use the point cloud to construct a map of the current environment, plan an optimal path according to the map, or perform operations such as obstacle avoidance according to the map; the electronic device can also be a car machine in a vehicle, and the point cloud generation device can be deployed in the car machine. Thus, the car machine can use the point cloud generated by the point cloud generation device to perform map construction to achieve high-precision positioning and navigation of the vehicle.
[0150] In some embodiments of the present application, the electronic device may further include a binocular camera and a lidar. The binocular camera can be used to obtain image data, and the lidar can be used to obtain radar data.
[0151] Exemplarily, the electronic device can be a sweeping robot including a binocular camera, a lidar, and a point cloud generation device.
[0152] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can be in the form of a hardware embodiment, a software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can be in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) containing computer-usable program code.
[0153] The present application is described with reference to the schematic flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the schematic flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the processes and / or blocks in the schematic flowchart and / or block diagram can also be implemented. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0154] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more of the flowcharts and / or boxes of the implementation flowchart Figure 1 one or more of the flows and / or boxes Figure 1 of the functions specified in one or more of the boxes.
[0155] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, such that a series of operational steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more of the flowcharts and / or boxes Figure 1 one or more of the flows and / or boxes Figure 1 of the functions specified in one or more of the boxes.
[0156] The above embodiments are merely preferred embodiments given to fully illustrate the present invention, and the scope of protection of the present invention is not limited thereto. Equivalent substitutions or transformations made by those skilled in the art on the basis of the present invention are within the scope of protection of the present invention.
Claims
1. A point cloud generation method, characterized in that, The method includes: Using a point cloud generation model to extract features from the data to be processed, obtaining feature data to be processed; wherein, the data to be processed at least includes image data; the point cloud generation model is obtained by training an initial model using a training data set generated by a verified simulator; the verified simulator is verified based on first data, first internal and external parameters, second data, and second internal and external parameters; the first data is obtained by using a binocular camera and a lidar to collect data from a calibration scene; the first internal and external parameters represent the internal and external parameters of the binocular camera and the lidar obtained using the first data; the second data is obtained by simulating the data collection of the binocular camera and the lidar in the calibration scene using the simulator to be verified; the second internal and external parameters represent the internal and external parameters of the binocular camera and the lidar obtained using the second data; Using the point cloud generation model to extract multi-dimensional disparity features from the feature data to be processed, obtaining disparity feature data; Using the point cloud generation model to perform optimization processing on the multi-scale disparity features of the disparity feature data, obtaining optimized disparity data; Generating depth map data according to the optimized disparity data, and generating a point cloud according to the depth map data.
2. The point cloud generation method according to claim 1, characterized in that The point cloud generation model includes a global feature extractor and a hybrid convolutional feature extractor; the using the point cloud generation model to extract multi-dimensional disparity features from the feature data to be processed, obtaining disparity feature data, includes: Using the global feature extractor to perform convolutional downsampling on the feature data to be processed, obtaining downsampled disparity features, and performing global feature extraction processing according to the downsampled disparity features and the position encoding information corresponding to the data to be processed, obtaining a first disparity feature; Using the hybrid convolutional feature extractor to perform multi-scale local disparity feature extraction on the feature data to be processed, obtaining multi-scale local disparity feature information, and performing hybrid convolutional feature extraction processing according to the multi-scale local disparity feature information, obtaining a second disparity feature; Determining the disparity feature data according to the first disparity feature and the second disparity feature.
3. The point cloud generation method according to claim 2, wherein The point cloud generation model includes an optimization iterator; the using the point cloud generation model to perform optimization processing on the multi-scale disparity features of the disparity feature data, obtaining optimized disparity data, includes: Using the optimization iterator to perform fusion processing on the image features of the image data in the data to be processed and the disparity feature data, obtaining fused features; Determining a disparity correction value according to the first feature information at different scales in the fused features; Performing optimization processing on the initial disparity value according to the disparity correction value, obtaining the optimized disparity data.
4. The point cloud generation method according to claim 3, wherein The performing hybrid convolutional feature extraction processing according to the multi-scale local disparity feature information, obtaining a second disparity feature, includes: Perform three-dimensional convolution processing, spatial convolution processing, and disparity convolution processing on the local disparity feature information at the target scale in the multi-scale local disparity feature information to obtain third feature information; Perform fusion processing on the multi-scale image features of the image data using attention weights to obtain fourth feature information; Determine the second disparity feature based on the third feature information and the fourth feature information.
5. The point cloud generation method according to claim 4, wherein, The determining the disparity correction value according to the first feature information at different scales in the fused feature includes: Perform feature downsampling with multiple resolutions on the first feature information to obtain multiple second feature information; Fuse the multiple second feature information to obtain fused feature information; Determine the disparity correction value according to the fused feature information.
6. The point cloud generation method according to any one of claims 1 to 5, characterized in that The point cloud generation model includes a monocular depth estimation network, a multi-scale feature extraction network, and a point cloud feature extraction network; The using the point cloud generation model to extract features from the data to be processed to obtain the feature data to be processed includes: When the data to be processed includes image data and radar data, use the monocular depth estimation network to extract the monocular depth feature information of the image data, and use the multi-scale feature extraction network to extract the multi-scale depth feature information of the image data; Use the point cloud feature extraction network to extract the point cloud features of the radar data; Determine the monocular depth feature information, the multi-scale depth feature information, and the point cloud features as the feature data to be processed; When the data to be processed includes image data, use the monocular depth estimation network to extract the monocular depth feature information of the image data, and use the multi-scale feature extraction network to extract the multi-scale depth feature information of the image data, and determine the monocular depth feature information and the multi-scale depth feature information as the feature data to be processed.
7. The point cloud generation method according to any one of claims 1 to 5, characterized in that, The method further includes: Obtain first scene data determined by three-dimensional modeling; Obtain second scene data; wherein, the second scene data is scene data obtained by reconstructing a simulation scene according to first image information; Perform sensor data synthesis on the first scene data and the second scene data based on a verified simulator and pre-configured parameters to obtain synthesized image data and synthesized point cloud; wherein, the pre-configured parameters include simulation configuration parameters of a camera and a radar, and environmental parameters; Determine the training dataset based on the synthesized image data and the synthesized point cloud; Use the training dataset to train an initial model to obtain the point cloud generation model.
8. A point cloud generation device, characterized in that, The point cloud generation device includes a first generation unit and a second generation unit; The first generation unit is configured to extract features from the data to be processed using a point cloud generation model to obtain the feature data to be processed; and extract multi-dimensional disparity features from the feature data to be processed using the point cloud generation model to obtain disparity feature data; and perform disparity feature optimization processing on the disparity feature data using the point cloud generation model to obtain optimized disparity data; wherein the data to be processed includes at least image data; the point cloud generation model is obtained by training an initial model using a training data set generated by a verified simulator; the verified simulator is verified based on first data, first internal and external parameters, second data, and second internal and external parameters; the first data is obtained by collecting data on a calibration scene using a binocular camera and a lidar; the first internal and external parameters represent the internal and external parameters of the binocular camera and the lidar obtained using the first data; the second data is obtained by simulating the data collection of the binocular camera and the lidar in the calibration scene using the simulator to be verified; the second internal and external parameters represent the internal and external parameters of the binocular camera and the lidar obtained using the second data; The second generation unit is configured to generate depth map data based on the optimized disparity data, and generate a point cloud based on the depth map data.
9. A point cloud generation device, characterized in that, The point cloud generation device includes a processor and a memory storing instructions executable by the processor; when the instructions are executed by the processor, the method according to any one of claims 1 to 7 is implemented.
10. An electronic device, characterized in that, The electronic device includes the point cloud generation device; The electronic device is configured to construct map data based on the point cloud generated by the point cloud generation device; and perform corresponding operations based on the map data.
Citation Information
Patent Citations
Visual positioning method and device and computer readable medium
CN110568447A
Three-dimensional point cloud reconstruction method, device and electronic equipment
CN112907730A
Network training and scene reconstruction method, device, machine, system and equipment
CN115512042A
Depth estimation model training method and system based on random activation and evaluation method
CN115760949A
Millimeter wave radar dense point cloud generation method and system based on transfer learning
CN116664970A