Three-dimensional reconstruction method and device, vehicle, storage medium and program product
By combining the image frames of binocular cameras and monocular cameras, three-dimensional reconstruction is carried out using depth maps and neural network models, the problem of insufficient correspondence between environmental reconstruction and real environment is solved, and efficient and accurate environmental reconstruction is achieved.
Patent Information
- Application Number
- CN202510804733.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-17
AI Technical Summary
The existing environmental reconstruction technology cannot accurately correspond to the real driving environment, resulting in limited environmental reconstruction function.
Using a combination of binocular camera and monocular camera, we use the depth map and neural network model to perform depth estimation, generate environmental point cloud data, and realize three-dimensional reconstruction.
It improves the accuracy and real-time nature of environmental reconstruction, enhances the efficiency and detailed information of three-dimensional reconstruction, and ensures the correspondence between the reconstruction environment and the real environment.
Smart Images

Figure CN120339524A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of environmental reconstruction, and in particular, to a three-dimensional reconstruction method, device, vehicle, storage medium, and program product. Background Art
[0002] With the development of the automotive industry, current vehicles are generally equipped with various sensors such as cameras, millimeter-wave radars, and lidar, and there are generally environmental reconstruction applications. Currently, the existing environmental reconstruction applications are generally implemented by drawing various traffic elements and functional elements in a computer virtual scene based on the results of sensor perception fusion. However, the driving environment reconstructed in this way cannot accurately correspond to the real driving environment, resulting in limited functionality of the environmental reconstruction. Summary of the Invention
[0003] This application provides a three-dimensional reconstruction method, device, vehicle, storage medium, and program product, which can make the reconstructed driving environment accurately correspond to the real driving environment and improve the accuracy of environmental reconstruction.
[0004] The technical solution of this application is realized as follows: An embodiment of this application provides a three-dimensional configuration method, including: obtaining a first current image frame collected by a binocular camera of a target vehicle and a second current image frame collected by a monocular camera; there is an overlapping area between the first current image frame and the second current image frame; determining a second depth map corresponding to the second current image frame based on a first depth map corresponding to the first current image frame; determining first environmental point cloud data of the target vehicle based on the first depth map and the second depth map, so as to obtain a three-dimensional reconstruction result of the environment around the target vehicle based on the first environmental point cloud data.
[0005] According to the above technical means, since both the first current image frame and the second current image frame are real-time collected complete environmental images around the target vehicle, determining the first environmental point cloud data based on the first depth map corresponding to the first current image frame and the second depth map corresponding to the second current image can improve the real-time performance and richness of the first environmental point cloud data. Moreover, the first depth map can be obtained by depth estimation through the binocular disparity of the first current image frame collected by the binocular camera. Therefore, determining the second depth map corresponding to the second current image frame collected by the monocular camera through the first depth map can improve the accuracy of depth information estimation of the second current image frame, so that the reconstructed driving environment can accurately correspond to the real driving environment and improve the accuracy of environmental reconstruction.
[0006] In some embodiments, two first current images are included in the first current image frame; determining the second depth map corresponding to the second current image frame based on the first depth map corresponding to the first current image frame includes: determining the first depth map based on the disparity information of corresponding pixel points in the two first current images; based on the first depth map, correcting the parameters of a neural network model for estimating image depth information to obtain a target neural network model; and using the target neural network model to perform depth estimation on the second current image frame to obtain the second depth map.
[0007] According to the above technical means, by correcting the parameters of the neural network model for estimating image depth information based on the first depth map determined from the disparity information of corresponding pixel points in the two first current images in the first current image frame, the accuracy of depth estimation of the neural network model can be improved. Therefore, performing depth estimation on the second current image frame based on the corrected target neural network model can improve the accuracy of the determined second depth map.
[0008] In some embodiments, based on the first depth map, correcting the parameters of the neural network model for estimating image depth information to obtain a target neural network model includes: using the neural network model to estimate the depth information of the first current image frame to obtain a reference depth map; based on the reference depth map and the first depth map, performing at least one correction on the parameters of the neural network model until a convergence condition is met to obtain the target neural network model.
[0009] According to the above technical means, by performing at least one correction on the neural network model based on the reference depth map and the first depth map obtained after estimating the first current depth information using the neural network model, the depth information estimation accuracy of the neural network model can be continuously improved. Therefore, using the target neural network model with corrected parameters to estimate the depth information of the second current image frame can improve the accuracy of the determined second depth map.
[0010] In some embodiments, based on the first depth map, correcting the parameters of the neural network model for estimating image depth information to obtain a target neural network model includes: if there is a first region with the same image content in the first current image frame and the second current image frame, using the neural network model to estimate the depth information of the sub-image corresponding to the first region in the second current image frame to obtain the depth map corresponding to the sub-image; based on the depth map of the sub-image and the reference sub-depth map, performing at least one correction on the parameters of the neural network model until a convergence condition is met to obtain the target neural network model; the reference sub-depth is the depth map intercepted from the first depth map corresponding to the first region in the first current image frame.
[0011] According to the above technical means, when it is determined that there is a first region with the same image content in the first current image frame and the second current image frame, based on the depth map obtained after estimating the depth information of the sub-image corresponding to the first region in the second current image frame by a neural network model, and the depth map intercepted from the first depth map corresponding to the first region in the first current image frame, the neural network model is corrected at least once, and the depth information estimation accuracy of the neural network model can be continuously improved. Therefore, using the target neural network model with corrected parameters to estimate the depth information of the second current image frame can improve the accuracy of the second depth map.
[0012] In some embodiments, determining the first depth map based on the disparity information of corresponding pixel points in the two first current images includes: preprocessing the two first current images respectively to obtain two candidate images; determining a second region with the same image content in the two candidate images based on the parameters of the binocular camera; and determining the first depth map based on the disparity information of the corresponding pixel points in the second region of the two candidate images.
[0013] According to the above technical means, by preprocessing the two first current images, the image quality of the first current images can be improved, and the computational amount of three-dimensional reconstruction can be reduced. At the same time, based on the parameters of the binocular camera, a second region with the same image content in the two preprocessed candidate images is determined, and based on the disparity information of the corresponding pixel points in the second region, the first depth map can be accurately and quickly determined.
[0014] In some embodiments, the number of the monocular cameras is multiple, and each monocular camera is arranged at different positions of the target vehicle; using the target neural network model to estimate the depth of the second current image frame to obtain a second depth map includes: performing stitching processing on the second current image frames respectively collected by the multiple monocular cameras to obtain a stitched second current image frame; and using the target neural network model to estimate the depth of the stitched second current image frame to obtain a second depth map.
[0015] According to the above technical means, by performing stitching processing on the second current image frames respectively collected by the multiple monocular cameras and using the target neural network model to estimate the depth of the stitched second current image frame, the depth information of the image frames collected by the monocular cameras can be estimated at one time, thereby improving the efficiency of three-dimensional reconstruction.
[0016] In some embodiments, determining the first environmental point cloud data of the target vehicle based on the first depth map and the second depth map includes: obtaining the current vehicle speed of the target vehicle; based on the current vehicle speed, obtaining the first historical image frame collected by the binocular camera and the second historical image frame collected by the monocular camera; and determining the first environmental point cloud data of the target vehicle based on the environmental point cloud data corresponding to the first historical image frame, the environmental point cloud data corresponding to the second historical image frame, the first depth map, and the second depth map.
[0017] According to the above technical means, based on the current vehicle speed of the target vehicle, the first historical image frame and the second historical image frame adapted to the vehicle speed can be obtained. Further, based on the environmental point cloud data corresponding to the first historical image frame, the environmental point cloud data corresponding to the second historical image frame, the first depth map, and the second depth map, it can be ensured that the first environmental point cloud data of the target vehicle contains more detailed information, thereby improving the effect of 3D reconstruction.
[0018] In some embodiments, determining the first environmental point cloud data of the target vehicle based on the environmental point cloud data corresponding to the first historical image frame, the environmental point cloud data corresponding to the second historical image frame, the first depth map, and the second depth map includes: determining the environmental point cloud data corresponding to the first current image frame based on the first depth map and the first current image frame; and determining the environmental point cloud data corresponding to the second current image frame based on the second depth map and the second current image frame; fusing the environmental point cloud data corresponding to the first historical image frame and the environmental point cloud data corresponding to the first current image frame to obtain a first candidate environmental point cloud data; and fusing the environmental point cloud data corresponding to the second historical image frame and the environmental point cloud data corresponding to the second current image frame to obtain a second candidate environmental point cloud data; and determining the first environmental point cloud data based on the first candidate environmental point cloud data and the second candidate environmental point cloud data.
[0019] According to the above technical means, by fusing the environmental point cloud data corresponding to the first historical image frame and the environmental point cloud data corresponding to the first current image frame to obtain a first candidate environmental point cloud data, and fusing the environmental point cloud data corresponding to the second historical image frame and the environmental point cloud data corresponding to the second current image frame to obtain a second candidate environmental point cloud data, it can be ensured that the first environmental point cloud data finally determined based on the first candidate environmental point cloud data and the second candidate environmental point cloud data contains more details of the environment around the target vehicle, thereby improving the accuracy of 3D reconstruction.
[0020] In some embodiments, second environmental point cloud data determined based on the radar sensor of the target vehicle is obtained; the first environmental point cloud data and the second environmental point cloud data are fused to obtain a three-dimensional reconstruction result of the environment around the target vehicle.
[0021] According to the above technical means, by fusing the first environmental point cloud data and the second environmental point cloud data determined by the radar sensor, the three-dimensional reconstruction result of the environment around the target vehicle can contain more and more accurate point cloud information, thereby improving the accuracy of three-dimensional reconstruction.
[0022] An embodiment of the present application provides a three-dimensional reconstruction device, including: A first acquisition module, configured to acquire a first current image frame collected by a binocular camera of the target vehicle and a second current image frame collected by a monocular camera; there is an overlapping area between the first current image frame and the second current image frame; A first determination module, configured to determine a second depth map corresponding to the second current image frame based on the first depth map corresponding to the first current image frame; A second determination module, configured to determine the first environmental point cloud data of the target vehicle based on the first depth map and the second depth map, so as to obtain a three-dimensional reconstruction result of the environment around the target vehicle based on the first environmental point cloud data.
[0023] An embodiment of the present application provides a vehicle, including a memory and a processor, where the memory stores a computer program that can run on the processor, and is characterized in that when the processor executes the program, the steps in the above method are implemented.
[0024] An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps in the above method are implemented.
[0025] An embodiment of the present application provides a computer program product, including a computer program or instruction, and when the computer program or instruction is executed by a processor, the steps in the above method are implemented.
[0026] The beneficial effects of the present application: The real-time performance and richness of the first environmental point cloud data can be improved, the accuracy of estimating the depth information corresponding to the second current image frame can be improved, so that the reconstructed driving environment can accurately correspond to the real driving environment, and the accuracy of environment reconstruction can be improved.
[0027] The accuracy of depth estimation of the neural network model can be continuously improved. Therefore, based on the corrected target neural network model, depth estimation is performed on the second current image frame, and the accuracy of the determined second depth map can be improved.
[0028] The first depth map can be accurately and quickly determined.
[0029] The efficiency of 3D reconstruction can be improved.
[0030] It can make the first environmental point cloud data of the target vehicle contain more detailed information, thereby improving the effect of 3D reconstruction.
[0031] It can make the first environmental point cloud data finally determined based on the first candidate environmental point cloud data and the second candidate environmental point cloud data contain more details in the environment around the target vehicle, thereby improving the accuracy of 3D reconstruction. Description of the Drawings
[0032] Figure 1 It is a schematic flowchart of a 3D reconstruction method provided by an embodiment of the present application; Figure 2 It is a schematic structural diagram of a vehicle real-scene reconstruction system provided by an embodiment of the present application; Figure 3 It is a schematic flowchart of a 3D reconstruction method based on single-frame forward-looking data provided by an embodiment of the present application; Figure 4 It is a schematic diagram of the principle of determining the depth information of pixel points provided by an embodiment of the present application; Figure 5 It is a schematic flowchart of a 3D reconstruction method based on multi-frame forward-looking data provided by an embodiment of the present application; Figure 6 It is a schematic flowchart of a 3D reconstruction method based on single-frame panoramic data and surround-view data provided by an embodiment of the present application; Figure 7 It is a schematic flowchart of a 3D reconstruction method based on multi-frame panoramic data and surround-view data provided by an embodiment of the present application; Figure 8 It is a schematic diagram of the fusion principle of point cloud data provided by an embodiment of the present application; Figure 9 It is a schematic flowchart of a rendering method of 3D reconstruction point cloud data provided by an embodiment of the present application; Figure 10 It is a schematic structural diagram of a 3D reconstruction device provided by an embodiment of the present application. Detailed Embodiments
[0033] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application.
[0034] To make the objectives, technical solutions, and advantages of this application clearer, the following will further describe this application in conjunction with the accompanying drawings. The described embodiments should not be construed as limitations on this application. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of this application.
[0035] In the following description, reference is made to "some embodiments / other embodiments", which describe subsets of all possible embodiments. However, it can be understood that "some embodiments / other embodiments" can be the same subsets or different subsets of all possible embodiments and can be combined with each other without conflict.
[0036] In the following description, the terms "first / second" are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second" can be interchanged with a specific order or sequence when permitted, so that the embodiments of this application described here can be implemented in an order other than that illustrated or described here.
[0037] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0038] Based on the problems existing in the related art, the embodiments of this application provide a three-dimensional reconstruction method, which can make the reconstructed driving environment accurately correspond to the real driving environment and improve the accuracy of environmental reconstruction. As Figure 1 shown, it is a schematic flowchart of a three-dimensional reconstruction method provided by the embodiments of this application. The method includes the following steps S101 to step S103: Step S101, obtain a first current image frame collected by a binocular camera of a target vehicle and a second current image frame collected by a monocular camera.
[0039] Here, the binocular camera can be arranged on the first side of the target vehicle, and the monocular camera can be arranged on the side adjacent to the first side of the target vehicle. The first side can be the front side of the target vehicle, or at least one of the front side, left side, right side, and rear side of the target vehicle. For example, if the binocular camera is arranged on the front side of the target vehicle, the monocular camera can be arranged on the left side and / or right side of the target vehicle. Another example is that if the binocular camera is arranged on the left side of the target vehicle, the monocular camera can be arranged on the front side and / or rear side of the target vehicle.
[0040] Both the monocular camera and the binocular camera are in-vehicle cameras installed on the target vehicle. For example, the monocular camera can be a surround-view camera or a panoramic camera, and the binocular camera can be a telephoto camera and a wide-angle camera, etc.
[0041] In some embodiments, the monocular camera of the target vehicle may include one or more. The binocular camera may be a combination of two vehicle-mounted cameras. The binocular camera may include at least one group. The installation positions and / or installation angles of multiple monocular cameras disposed on the same side of the target vehicle are different; the installation positions and installation angles of the binocular cameras disposed on the same side of the target vehicle are the same.
[0042] In some embodiments, there is an overlapping area between the first current image frame and the second current image frame, that is, there is an area with the same image content in the first current image frame and the second current image frame. For example, the first current image frame includes the image content collected by the binocular camera for the front, left front, and right front areas of the target vehicle, and the second current image frame includes the image content collected by the monocular camera for the left side, left front, and left rear areas of the target vehicle. Then, the overlapping area between the first current image frame and the second current image frame is the image content corresponding to the left front area of the target vehicle.
[0043] In some embodiments, the first current image frame may include two or more, and the second current image frame may include one or more. The first current image frame may be an environmental image of the first side of the target vehicle collected by the binocular camera at the current moment, and the second current image frame may be an environmental image of the side adjacent to the first side of the target vehicle collected by the monocular camera at the current moment.
[0044] Step S102, determine the second depth map corresponding to the second current image frame based on the first depth map corresponding to the first current image frame.
[0045] In some embodiments, the first depth map corresponding to the first current image frame may be determined based on the internal and external parameters of the binocular camera, and then the second depth map corresponding to the second current image frame may be determined based on the first depth map. The first depth map corresponding to the first current image frame may be determined based on the parallax information of the corresponding pixel points in multiple first current image frames. The depth map corresponding to the second current image frame may be calculated by a depth information estimation algorithm. For example, it may be estimated by a Multilayer Perceptron (MLP) neural network model. The MLP neural network model may be trained based on the depth map corresponding to the first current image frame, and the trained MLP neural network model may be used to estimate the depth information of the second current image frame, so as to obtain the second depth map.
[0046] Step S103, determine the first environmental point cloud data of the target vehicle based on the first depth map and the second depth map, so as to obtain a three-dimensional reconstruction result of the environment around the target vehicle based on the first environmental point cloud data.
[0047] In some embodiments, the first depth map and the second depth map may constitute the depth information of the complete area around the target vehicle. Based on the internal parameters of the binocular camera corresponding to the first current image frame, the pixel coordinates of each pixel point in the first current image frame may be converted into coordinates in the binocular camera coordinate system, and then combined with the depth information of each pixel point in the first depth map to determine the point cloud data corresponding to the first current image frame. Based on the internal parameters of the monocular camera corresponding to the second current image frame, the pixel coordinates of each pixel point in the second current image frame may be converted into coordinates in the monocular camera coordinate system, and then combined with the depth information of each pixel point in the second depth map to determine the point cloud data corresponding to the second current image frame. The point cloud data corresponding to the first current image frame and the point cloud data corresponding to the second current image frame are fused or combined to obtain the first environmental point cloud data of the target vehicle.
[0048] In some embodiments, the first environmental point data may include the position information and color information of each point cloud. The position information may be determined by the depth map and the internal parameters of the monocular camera and the binocular camera, and the color information may be determined by the color information of each pixel point in the first current image frame and the color information of each pixel point in the second current image frame.
[0049] In some embodiments, the first environmental point data may be directly used as the three-dimensional reconstruction result of the environment around the target vehicle; alternatively, based on the point cloud data determined by the radar sensor of the target vehicle and the first environmental point cloud data, the three-dimensional reconstruction result of the environment around the target vehicle may be determined.
[0050] In the embodiments of the present application, the first current image frame collected by the binocular camera of the target vehicle and the second current image frame collected by the monocular camera are obtained; the binocular camera is disposed on the first side of the target vehicle, and the monocular camera is disposed on the side of the target vehicle adjacent to the first side; based on the first depth map corresponding to the first current image frame, the second depth map corresponding to the second current image frame is determined; based on the first depth map and the second depth map, the first environmental point cloud data of the target vehicle is determined to obtain the three-dimensional reconstruction result of the environment around the target vehicle based on the first environmental point cloud data. In this way, since both the first current image frame and the second current image frame are real-time collected complete environment images around the target vehicle, therefore, based on the first depth map corresponding to the first current image frame and the second depth map corresponding to the second current image, the first environmental point cloud data is determined, which can improve the real-time performance and richness of the first environmental point cloud data. Moreover, the first depth map can be obtained by depth estimation through the binocular disparity of the first current image collected by the binocular camera. Therefore, by the first depth map, the second depth map corresponding to the second current image frame collected by the monocular camera is determined, which can improve the accuracy of estimating the depth information corresponding to the second current image frame, so that the reconstructed driving environment can accurately correspond to the real driving environment and improve the accuracy of environment reconstruction.
[0051] In some embodiments, the first current image frame includes two first current images. Determining the second depth map corresponding to the second current image frame based on the first depth map corresponding to the first current image frame in step S102 includes the following steps S1021 to S1023: Step S1021, determine the first depth map based on the disparity information of corresponding pixel points in the two first current images.
[0052] It can be understood that a binocular camera can capture two images of the same area from different perspectives at the same time, obtaining a first current image frame containing two first current images, and the two first current images correspond to different perspectives respectively.
[0053] In some embodiments, the two first current images in the first current image frame can be matched to determine the same area in the two first current images. Then, based on the disparity information of corresponding pixel points in the same area, and the distance and focal length between the two first current images corresponding to the binocular camera, the first depth map is determined, that is, the depth information of corresponding pixel points in the same area is calculated through the binocular depth estimation algorithm, and the depth information of multiple pixel points in the same area can be used to obtain the first depth map.
[0054] Step S1022, based on the first depth map, correct the parameters of the neural network model for estimating image depth information to obtain a target neural network model.
[0055] Here, the neural network model for estimating image depth information can be a pre-trained neural network model or an untrained neural network model. This neural network model can be an MLP neural network model, etc. The neural network model can be self-supervised trained based on the first depth map until the convergence condition is met, and then the target neural network model can be obtained.
[0056] In some embodiments, the neural network model for estimating image depth information can be a model pre-trained using simulated laboratory data or a small number of training samples (image data) in a vehicle environment. However, the accuracy of depth information estimation by this neural network model is relatively low. For example, its estimation accuracy may be less than 50%.
[0057] In some embodiments, the neural network model for estimating image depth information can be used to estimate the depth information of the first current image frame. Based on the difference degree between the estimated depth map and the first depth map, the parameters of the neural network model for estimating image depth information are trained by backpropagation until the difference degree between the estimated depth map and the first depth map is less than a preset threshold, and then the training can be stopped to obtain the target neural network model.
[0058] Step S1023: Use the target neural network model to perform depth estimation on the second current image frame to obtain a second depth map.
[0059] In some embodiments, after obtaining the target neural network model, the second current image frame can be input into the target neural network model, and the target neural network model can perform depth estimation on the second current image to obtain a second depth map.
[0060] In the above embodiments, by correcting the parameters of the neural network model for estimating image depth information based on the first depth map determined by the disparity information of the corresponding pixel points in the two first current image frames, the accuracy of depth estimation of the neural network model can be improved. Therefore, based on the corrected target neural network model to perform depth estimation on the second current image frame can improve the accuracy of the second depth map.
[0061] In some embodiments, the step of correcting the parameters of the neural network model for estimating image depth information based on the first depth map in the above step S1022 to obtain the target neural network model may include the following steps S10221A to step S10222A: Step S10221A: Use the neural network model to estimate the depth information of the first current image frame to obtain a reference depth map.
[0062] In some embodiments, the first current image frame can be input into the neural network model, and the neural network model can estimate the depth information of each pixel point in the first current image frame, and generate a reference depth map based on the depth information of each pixel point.
[0063] Step S10222A: Based on the reference depth map and the first depth map, correct the parameters of the neural network model at least once until the convergence condition is satisfied to obtain the target neural network model.
[0064] In some embodiments, meeting the convergence condition may be that the difference degree value between the reference depth map and the first depth map is less than the depth information difference degree threshold. First, the difference degree value between the reference depth map and the first depth map can be determined, and the parameters of the neural network model are corrected based on the difference degree value between the two depth images. Then, the corrected neural network model is used to estimate the depth of the first current image frame again, and the difference degree value between the depth information of the pixel points in the reference depth map obtained by the re - estimation and the depth information of the corresponding pixel points in the first depth map is determined. Based on this difference degree value, the parameters of the neural network model are corrected again until the difference degree value is less than the depth information difference degree threshold, and then the correction of the parameters of the neural network model can be stopped to obtain the target neural network model. Among them, the difference degree value between the two depth images can be determined based on the average value or weighted average value of the difference degree values of the depth information of each pixel point in the reference depth map and the depth information of the corresponding pixel point in the first depth map.
[0065] In some embodiments, methods such as cosine similarity and Euclidean distance can be used to determine the difference degree value between the depth information of each pixel point in the reference depth map and the depth information of the corresponding pixel point in the first depth map. For example, the average value or weighted average value of the cosine similarities (or Euclidean distances) of the depth information corresponding to each pixel point is determined as the difference degree value between the reference depth map and the first depth map.
[0066] In the above - mentioned embodiments, after estimating the first current depth information using the neural network model, the reference depth map and the first depth map are used to correct the neural network model at least once, which can continuously improve the depth information estimation accuracy of the neural network model. Therefore, using the target neural network model with corrected parameters to estimate the depth information of the second current image frame can improve the accuracy of the second depth map.
[0067] In some embodiments, the step of correcting the parameters of the neural network model for estimating the image depth information based on the first depth map in step S1022 to obtain the target neural network model may include the following steps S10221B to step S10222B: Step S10221B, if there is a first region with the same image content in the first current image frame and the second current image frame, use the neural network model to estimate the depth information of the sub - image corresponding to the first region in the second current image frame to obtain the depth map corresponding to the sub - image.
[0068] In some embodiments, the images collected by the binocular camera and the monocular camera may have overlapping regions, such as the left - front region, right - front region, left - rear region, and right - rear region of the target vehicle. Therefore, there may be regions with the same image content in the first current image frame and the second current image frame.
[0069] In some embodiments, when it is determined that there is a first region with the same image content in the first current image frame and the second current image frame, the image content of the first region in the second current image frame can be intercepted to obtain a corresponding sub-image, and the sub-image is input into a neural network model for depth information estimation to obtain a depth map corresponding to the sub-image.
[0070] Step S10222B: Based on the depth map of the sub-image and the reference sub-depth map, the parameters of the neural network model are corrected at least once until the convergence condition is met, and a target neural network model is obtained.
[0071] Here, the reference sub-depth is the depth map intercepted from the first depth map corresponding to the first region in the first current image frame.
[0072] In some embodiments, the parameters of the neural network model can be corrected based on the difference degree value between the depth map of the sub-image estimated by the neural network model and the reference sub-depth map. For example, the cosine similarity between the depth information of each pixel point in the depth map of the sub-image and the depth information of the corresponding pixel point in the reference sub-depth map can be calculated, and the average value of each cosine similarity is used as the difference degree value between the depth map of the sub-image and the reference sub-depth map.
[0073] In some embodiments, if it is determined that the difference degree value between the depth map of the sub-image and the reference sub-depth map is greater than or equal to the depth information difference threshold, it is considered that the convergence condition is not met. At this time, the parameters of the neural network model can be corrected, and then the corrected neural network model is used to estimate the depth information of the sub-image corresponding to the first region in the second current image frame again. This process is repeated until the difference degree value between the depth map of the sub-image estimated by the neural network model and the reference sub-depth map is less than the depth information difference threshold, that is, it is considered that the convergence condition is met, and the correction of the parameters of the neural network model can be stopped, and thus the target neural network model can be obtained.
[0074] In the above embodiments, when it is determined that there is a first region with the same image content in the first current image frame and the second current image frame, based on the depth map obtained after estimating the depth information of the sub-image corresponding to the first region in the second current image frame by the neural network model, and the depth map intercepted from the first depth map corresponding to the first region in the first current image frame, the neural network model is corrected at least once, which can continuously improve the depth information estimation accuracy of the neural network model. Therefore, using the target neural network model with corrected parameters to estimate the depth information of the second current image frame can improve the accuracy of the second depth map.
[0075] In some embodiments, determining the first depth map based on the disparity information of corresponding pixel points in the two first current images in step S1021 may include steps S10211 to S10213: Step S10211, preprocess the two first current images respectively to obtain two candidate images.
[0076] In some embodiments, preprocessing the two first current images may include performing distortion correction, removing the background, etc. on the two first current images.
[0077] In some embodiments, the binocular camera may be a fish-eye camera. Therefore, the first current image may be a fish-eye image. So, correct the first current image to eliminate the distortion in the first current image. After performing distortion correction on the first current image, the background in the distortion-corrected first current image may be removed. For example, remove target objects that are far from the target vehicle in the first current image, such as the sky, clouds, etc.
[0078] In some embodiments, when removing the background from the first current image, the target objects in the first current image may be recognized. When a specific target object is recognized, the specific target object in the first current image may be removed. Among them, the specific target object may be a target object in a far scene such as the sky, mountains, clouds, etc.
[0079] In other embodiments, when removing the background from the first current image, the distances between various target objects (such as other vehicles, pedestrians, road signs, etc.) around the target vehicle may also be determined first, and the target objects with distances greater than the distance threshold may be removed from the first current image.
[0080] Step S10212, based on the parameters of the binocular camera, determine the second region with the same image content in the two candidate images.
[0081] In some embodiments, based on the internal parameters of the binocular camera, the pixel position coordinates of each pixel point in the respective candidate images may be converted into the position coordinates in the corresponding camera coordinate system, and then based on the external parameters of the binocular camera, the position coordinates in the camera coordinate system may be converted into the position coordinates in the world coordinate system (or the body coordinate system), that is, the pixel points in the two candidate images are unified into the same coordinate system, and pixel point matching is performed based on the position coordinates in the same coordinate system, so as to determine the second region with the same image content in the two candidate images.
[0082] Step S10213, based on the disparity information of the corresponding pixel points in the second region of the two candidate images, determine the first depth map.
[0083] Here, the disparity information of the pixel points in the second region of the candidate image can be the disparity of the same pixel point in the two candidate images. This disparity can be the absolute value of the difference in the horizontal coordinate values of the same pixel point in the pixel coordinate systems (with the origin at the upper left corner of the candidate image, the horizontal direction being the u-axis, and the vertical direction being the v-axis) of the two candidate images.
[0084] In some embodiments, after determining the disparity information of a pixel point, the depth information corresponding to the pixel point can be calculated based on the distance between the centers of the binocular cameras, the focal length (the focal lengths of the binocular cameras are the same), and the disparity information. The depth information of other pixel points in the second region of the candidate image can also be calculated through this process. Finally, the depth information of each pixel point in the second region of the candidate image is obtained, and based on the depth information of each pixel point and the position information in the candidate image, a first depth map is constructed.
[0085] In the above embodiments, by preprocessing the two first current images, the image quality of the first current images can be improved, and the computational amount of 3D reconstruction can be reduced. At the same time, based on the binocular camera parameters, the second region with the same image content in the preprocessed candidate image is determined, and based on the disparity information of the corresponding pixel points in the second region, the first depth map can be accurately and quickly determined.
[0086] In some embodiments, the number of monocular cameras is multiple, and each monocular camera is set at different positions of the target vehicle; the depth estimation of the second current image frame using the target neural network model in step S1023 to obtain the second depth map can include the following steps S10231 to step S10232: Step S10231, perform stitching processing on the second current image frames respectively collected by the multiple monocular cameras to obtain the stitched second current image frame.
[0087] In some embodiments, before performing stitching processing on the second current image frames respectively collected by the multiple monocular cameras, preprocessing operations such as distortion correction and background removal can also be performed on the second current image frames. Then, operations such as feature extraction, feature matching, image registration, and image fusion can be performed on the multiple second current image frames in sequence, so that the stitched second current image frame can be obtained. Color correction, removal of stitching traces, etc. can also be performed on the stitched second current image frame to further improve the quality of the stitched second current image frame.
[0088] In some embodiments, there are regions with the same image content in the second current image frames respectively captured by multiple monocular cameras. The regions with the same image content can be fused, and the regions with different image content can be directly stitched. If monocular cameras are arranged on the left and right sides of the target vehicle, and two monocular cameras are arranged on each side, then the second current image frames captured by the two monocular cameras on the same side can be stitched first, and then the second current image frames obtained after stitching on each side can be stitched to obtain the stitched second current image frame.
[0089] Step S10232, use the target neural network model to perform depth estimation on the stitched second current image frame to obtain the second depth map.
[0090] In some embodiments, input the stitched second current image frame into the target neural network model. The target neural network model can estimate the depth information of each pixel point in the stitched second current image frame, so as to obtain the second depth map.
[0091] In the above embodiments, by performing stitching processing on the second current image frames respectively captured by multiple monocular cameras, and using the target neural network model to perform depth estimation on the stitched second current image frame, it is possible to realize a one-time estimation of the depth information of the image frames captured by the monocular cameras, thereby improving the efficiency of 3D reconstruction.
[0092] In some embodiments, the step of determining the first environmental point cloud data of the target vehicle based on the first depth map and the second depth map in step S103 includes steps S1031 to S1033: Step S1031, obtain the current vehicle speed of the target vehicle.
[0093] In some embodiments, the current vehicle speed information of the target vehicle can be the vehicle speed corresponding to the target vehicle when the binocular camera captures the first current image frame and the monocular camera captures the second current image frame.
[0094] Step S1032, based on the current vehicle speed, obtain the first historical image frame captured by the binocular camera and the second historical image frame captured by the monocular camera.
[0095] Here, the first historical image frame can be one or more image frames captured by the binocular camera before the first current image frame. For example, it can include the image frame before the first current image frame, the first two image frames, the first three image frames, etc. Similarly, the second historical image frame can also be one or more image frames captured by the monocular camera before the second current image frame.
[0096] In some embodiments, the number of first historical image frames and the number of second historical image frames can be determined based on the magnitude of the current vehicle speed of the target vehicle. The number of first historical image frames and the number of second historical image frames can be the same or different.
[0097] In some embodiments, the corresponding relationships among the vehicle speed, the number of first historical image frames, and the number of second historical image frames can be established in advance. For example, if the current vehicle speed is greater than the vehicle speed threshold, the number of first historical image frames can be determined as m1, and the number of second historical image frames can be determined as m2; if the current vehicle speed is less than or equal to the vehicle speed threshold, the number of first historical image frames can be determined as n1, and the number of second historical image frames can be determined as n2, where m1, m2, n1, and n2 are all positive integers, and m1 is less than n1, and m2 is less than n2. Alternatively, if the current vehicle speed is within the first preset vehicle speed range, the number of first image frames can be determined as m3, and the number of second historical image frames can be determined as n3; if the current vehicle speed is within the second preset vehicle speed range, the number of first image frames can be determined as m4, and the number of second historical image frames can be determined as n4, where the upper limit of the first preset vehicle speed range is less than the lower limit of the second preset vehicle speed range, and m3, m4, n3, and n4 are all positive integers, and m4 is less than n3, and m4 is less than n3.
[0098] Step S1033: Determine the first environmental point cloud data of the target vehicle based on the environmental point cloud data corresponding to the first historical image frame, the environmental point cloud data corresponding to the second historical image frame, the first depth map, and the second depth map.
[0099] In some embodiments, the environmental point cloud data of the first historical image frame can be calculated through the parallax information of the pixel points in the first historical image frame. The depth information of each pixel point in the second historical image frame can be obtained through a neural network model for estimating depth information, and then the environmental point cloud data corresponding to the second historical image frame can be calculated based on the depth information of each pixel point.
[0100] In some embodiments, the environmental point cloud data corresponding to each first historical image frame can be obtained by fusing the environmental point cloud data of the first historical image frame with the environmental point cloud data of other image frames before the first historical image frame. The environmental point cloud data corresponding to each second historical image frame can be obtained by fusing the environmental point cloud data of the second historical image frame with the environmental point cloud data of other image frames before the second historical image frame.
[0101] In some embodiments, based on the first depth map and the color information in the first current image frame, the environmental point cloud data corresponding to the first current image frame can be determined. Then, the environmental point cloud data corresponding to the first current image frame and the environmental point cloud data corresponding to the first historical image frame are fused to obtain the target environmental point cloud data corresponding to the first current image frame. Similarly, based on the second depth map and the color information in the second current image frame, the environmental point cloud data corresponding to the second current image frame can be determined. The environmental point cloud data corresponding to the second current image frame and the environmental point cloud data corresponding to the second historical image frame are fused to obtain the target environmental point cloud data corresponding to the second current image frame. Thereafter, the target environmental point cloud data corresponding to the first current image frame and the target environmental point cloud data corresponding to the second current image frame are fused to obtain the first environmental point cloud data of the target vehicle.
[0102] In the above embodiments, based on the current vehicle speed of the target vehicle, the first historical image frame and the second historical image frame adapted to the vehicle speed can be obtained. Further, based on the environmental point cloud data corresponding to the first historical image frame, the environmental point cloud data corresponding to the second historical image frame, the first depth map, and the second depth map, it can be ensured that the determined first environmental point cloud data of the target vehicle contains more detailed information, thereby improving the effect of 3D reconstruction.
[0103] In some embodiments, the step of determining the first environmental point cloud data of the target vehicle based on the environmental point cloud data corresponding to the first historical image frame, the environmental point cloud data corresponding to the second historical image frame, the first depth map, and the second depth map in step S1033 may include steps S10331 to S10333: Step S10331, based on the first depth map and the first current image frame, determine the environmental point cloud data corresponding to the first current image frame; and based on the second depth map and the second current image frame, determine the environmental point cloud data corresponding to the second current image frame.
[0104] In some embodiments, the depth information of each pixel point in the first depth map and the color information of the corresponding pixel point in the first current image frame can be combined, so as to obtain the environmental point cloud data corresponding to the first current image frame; the depth information of each pixel point in the second depth map and the color information of the corresponding pixel point in the second current image frame can be combined, so as to obtain the environmental point cloud data corresponding to the second current image frame.
[0105] Step S10332, fuse the environmental point cloud data corresponding to the first historical image frame and the environmental point cloud data corresponding to the first current image frame to obtain the first candidate environmental point cloud data; and fuse the environmental point cloud data corresponding to the second historical image frame and the environmental point cloud data corresponding to the second current image frame to obtain the second candidate environmental point cloud data.
[0106] In some embodiments, a weighted fusion method can be adopted to fuse the environmental point cloud data corresponding to the first historical image frame and the environmental point cloud data corresponding to the first current image frame, and to fuse the environmental point cloud data corresponding to the second historical image frame and the environmental point cloud data corresponding to the second current image frame. The weights of the environmental point cloud data corresponding to the first historical image frame, the environmental point cloud data corresponding to the first current image frame, the environmental point cloud data corresponding to the second historical image frame, and the environmental point cloud data corresponding to the second current image frame are not limited in this application according to the determination method.
[0107] Step S10333: Determine the first environmental point cloud data based on the first candidate environmental point cloud data and the second candidate environmental point cloud data.
[0108] In some embodiments, the first candidate environmental point cloud data and the second candidate environmental point cloud data can be fused. The fusion method can be splicing, weighted fusion, etc. By fusing the first candidate point cloud data and the second candidate environmental point cloud data, the obtained first environmental point cloud data can form the point cloud data of the complete area around the target vehicle.
[0109] In the above embodiments, by fusing the environmental point cloud data corresponding to the first historical image frame and the environmental point cloud data corresponding to the first current image frame to obtain the first candidate environmental point cloud data, and fusing the environmental point cloud data corresponding to the second historical image frame and the environmental point cloud data corresponding to the second current image frame to obtain the second candidate environmental point cloud data, the first environmental point cloud data finally determined based on the first candidate environmental point cloud data and the second candidate environmental point cloud data can contain more details in the environment around the target vehicle, thereby improving the accuracy of 3D reconstruction.
[0110] In some embodiments, the above method may further include steps S104 to S105: Step S104: Obtain the second environmental point cloud data determined based on the radar sensor of the target vehicle.
[0111] In some embodiments, the second environmental point cloud data can be obtained by the radar sensor of the target vehicle through calculation based on the data collected at the same time when the binocular camera captures the first current image frame and the monocular camera captures the second image frame, or calculated by the computing device of the target vehicle based on the data collected by the radar sensor.
[0112] Step S105: Fuse the first environmental point cloud data and the second environmental point cloud data to obtain the 3D reconstruction result of the environment around the target vehicle.
[0113] In some embodiments, the first environmental point cloud data and the second environmental point cloud data can be weighted and fused. The weights of the first environmental point cloud data and the second environmental point cloud data can be determined based on the configuration level or accuracy of the radar sensor. The higher the configuration level or accuracy of the radar sensor, the smaller the weight of the first environmental point cloud data and the larger the weight of the second environmental point cloud data; conversely, the larger the weight of the first environmental point cloud data and the smaller the weight of the second environmental point cloud data.
[0114] In the above embodiment, by fusing the first environmental point cloud data and the second environmental point cloud data determined by the radar sensor, the three-dimensional reconstruction result of the environment around the target vehicle can contain more and more accurate point cloud information, thereby improving the accuracy of three-dimensional reconstruction.
[0115] In the embodiment of the present application, a first current image frame collected by a binocular camera of a target vehicle and a second current image frame collected by a monocular camera are obtained; the binocular camera is disposed on a first side of the target vehicle, and the monocular camera is disposed on a side adjacent to the first side of the target vehicle; based on a first depth map corresponding to the first current image frame, a second depth map corresponding to the second current image frame is determined; based on the first depth map and the second depth map, first environmental point cloud data of the target vehicle is determined to obtain a three-dimensional reconstruction result of the environment around the target vehicle based on the first environmental point cloud data. In this way, since both the first current image frame and the second current image frame are complete environmental images around the target vehicle collected in real time, therefore, based on the first depth map corresponding to the first current image frame and the second depth map corresponding to the second current image, the first environmental point cloud data is determined, which can improve the real-time performance and richness of the first environmental point cloud data, and the binocular disparity of the first current image collected by the binocular camera can be used to estimate the depth to obtain the first depth map. Therefore, through the first depth map, the second depth map corresponding to the second current image frame collected by the monocular camera is determined, which can improve the accuracy of estimating the depth information corresponding to the second current image frame, so that the reconstructed driving environment can accurately correspond to the real driving environment and improve the accuracy of environment reconstruction.
[0116] Next, the implementation process of the application embodiment in an actual application scenario will be introduced.
[0117] As Figure 2 shown, it is a schematic structural diagram of a vehicle real-scene reconstruction system provided by an embodiment of the present application. The vehicle real-scene reconstruction system 200 includes a three-dimensional reconstruction unit 201 and an environment reconstruction unit 202. The three-dimensional reconstruction unit 201 uses the camera data (front view, surround view, panoramic view, etc.), millimeter-wave radar data, lidar data, etc. of the vehicle itself to construct three-dimensional scene point cloud data matching the input data, and gives the generated three-dimensional scene point cloud data to the environment reconstruction unit 202 to implement environment reconstruction and rendering.
[0118] As shown Figure 3 in the figure, it is a schematic flowchart of a three-dimensional reconstruction method based on single-frame front view data provided by this application. The method includes: S301, distortion correction; Here, the front view data can be subjected to distortion correction.
[0119] In some embodiments, the front view data can be acquired by a wide-angle camera and a telephoto camera arranged on the front side of the vehicle itself. The output images of the wide-angle camera and the telephoto camera for the front view can be subjected to distortion correction according to the calibration parameters of each camera, and the useless background can be removed.
[0120] S302, background removal; Here, the background in the front view data after distortion correction can be removed.
[0121] In some embodiments, by performing target recognition on the front view data after distortion correction, each target object near the vehicle itself can be recognized, and the useless target objects (such as the sky) can be removed; it is also possible to determine the distance of each target object around the vehicle itself from the vehicle, and remove the target objects with a distance greater than the distance threshold.
[0122] S303, image matching; Here, the wide-angle image and the telephoto image after distortion correction and background removal can be matched to obtain the common area of the two images.
[0123] The wide-angle image and the telephoto image after distortion correction can be quickly subjected to image matching according to the external parameters and internal parameters of each camera. By taking the common area of the two images, based on the internal parameters of the wide-angle camera and the pixel coordinates of the pixel points in the wide-angle image, the position coordinates in the coordinate system of the wide-angle camera can be converted. Then, based on the external parameter matrix of the wide-angle camera, the coordinates in the wide-angle camera coordinate system can be converted into the coordinates in the world coordinate system. The processing process of the telephoto image is similar. After that, the position coordinates of the wide-angle image in the world coordinate system and the position coordinates of the telephoto image in the world coordinate system are matched, so that the common area of the two images can be obtained.
[0124] S304, generating a depth map based on the common area of the wide-angle image and the telephoto image.
[0125] The depth information of each pixel point can be calculated based on the parallax information of the pixel points, so as to obtain a depth map. As shown Figure 4 in the figure, the projection points of point P on the telephoto camera and the wide-angle camera are p and p' respectively. The pixel coordinates of point p are (p u , p v ), and the pixel coordinates of p' are (p' u , p' v), the parallax information of the projection points p and p’ can be calculated (|p u - p’ u |, |p v - p’ v |), and thus the depth information corresponding to point P can be calculated based on the parallax information and the distance between the centers of the two cameras.
[0126] S305. Generate point cloud data based on the depth map.
[0127] Point cloud data based on the front view data can be generated according to the depth information of each pixel point in the depth map and the color information of each pixel point in the common area of the two types of images determined in step S303.
[0128] As Figure 5 shown, it is a schematic flowchart of a three-dimensional reconstruction method based on multi-frame front view data provided by this application. The method includes: S501. Distortion correction and background removal are performed on the output images of the front view wide-angle camera and the telephoto camera of the current frame according to the calibration parameters of each camera.
[0129] S502. Image matching, depth map generation, and point cloud data generation are respectively performed on the distortion-corrected telephoto images and wide-angle images of the current frame and the previous n frames, and the displacement information of each frame of the vehicle itself is combined to generate point cloud data based on the telephoto images and point cloud data based on the wide-angle images.
[0130] Here, n can be determined according to the current vehicle speed of the vehicle itself. For example, if the current vehicle speed of the vehicle itself is greater than the vehicle speed threshold, n can be set to a; if the current vehicle speed of the vehicle itself is less than or equal to the vehicle speed threshold, n can be set to b, where a is less than b. n can also be freely set by the user or developer.
[0131] In some embodiments, the front view data (telephoto images and wide-angle images) corresponding to each moment can be processed according to the methods of the aforementioned steps S403 to S405, so as to obtain point cloud data of multiple frames of telephoto images and point cloud data of multiple frames of wide-angle images. Among them, for the point cloud data of the telephoto images and the point cloud data of the wide-angle images collected at the same moment, the depth information of the same point cloud is the same, but the color information may be different.
[0132] S503. Fuse the point cloud data based on the telephoto images and the point cloud data based on the wide-angle images to obtain the fused point cloud data corresponding to the current frame.
[0133] In some embodiments, the point cloud data based on the telephoto image and the point cloud data based on the wide-angle image can both be converted to the world coordinate system, and then, according to the position coordinates of the point cloud data in the world coordinate system, the point cloud data in the point cloud data based on the telephoto image and the point cloud data based on the wide-angle image are fused. This fusion process may include the fusion of the position coordinates of the point cloud data and the fusion of the color information of the point cloud data. The specific fusion method is not limited in this application. For example, it may be weighted fusion or the like.
[0134] S504. Based on the displacement information of each frame of the host vehicle, fuse the point cloud data of the previous n frames and the fused point cloud data corresponding to the current frame to obtain the point cloud data based on the forward-looking data.
[0135] The point cloud data of each forward-looking data can be obtained by fusing the point cloud data before the forward-looking data and the point cloud data of the forward-looking data before the forward-looking data. If n is 3, the point cloud data of the second forward-looking data can be obtained based on the fusion of the point cloud data of the first forward-looking data and the point cloud data of the second forward-looking data; the point cloud data of the third forward-looking data can be obtained based on the fusion of the point cloud data of the first forward-looking data and the point cloud data of the second forward-looking data. That is, by adopting the sequential fusion method, the fused point cloud data corresponding to the current frame can contain more detailed information such as color and texture through the sequential fusion method.
[0136] As Figure 6 shown, it is a schematic flowchart of a three-dimensional reconstruction method based on single-frame panoramic data and surround-view data provided by this application. The method includes: S601. Perform distortion correction on the surround-view image and the panoramic image according to the calibration parameters of each camera.
[0137] Here, the surround-view image may be an image collected by the surround-view camera of the vehicle, and the panoramic image may be an image collected by the panoramic camera. The surround-view camera and the panoramic camera may be set on the left side, the right side, and the rear side of the vehicle body. The surround-view camera and the panoramic camera are both set on the left side, the right side, and the rear side of the vehicle body. For the surround-view camera and the panoramic camera set on the same side of the vehicle body, the installation positions and installation angles of the surround-view camera and the panoramic camera are different. The surround-view image can be corrected for distortion based on the calibration parameters of the surround-view camera, and the panoramic image can be corrected for distortion based on the calibration parameters of the panoramic camera.
[0138] S602. Stitch the distortion-corrected surround-view image and panoramic image into a panoramic view using the internal parameters of each camera.
[0139] In some embodiments, the surround-view image and the panoramic image can be stitched based on the internal parameters of the surround-view camera and the internal parameters of the panoramic camera, so that a panoramic view can be obtained. The panoramic view may be an image obtained by stitching the left and right sides of the vehicle and the rear side of the vehicle.
[0140] An image stitching algorithm can be used to stitch the surround-view image and the panoramic image, and convert the two images into grayscale images; use feature detection algorithms (such as SIFT, SURF, etc.) to detect feature points in the two images; calculate the similarity of each feature point based on the position information of the feature points (such as coordinates in the camera coordinate system, and the pixel coordinates of the pixel points in the image can be converted into camera coordinates through the internal parameters of the camera) (such as through Euclidean distance, cosine similarity, etc.); match the feature points based on the similarity, such as feature points with a similarity greater than the similarity threshold are mutually matching feature points; estimate the perspective transformation matrix of the two images based on the position information of the mutually matching feature points; perform perspective transformation on one of the images using the estimated perspective transformation matrix to align it with the other image; create a new blank canvas with a size suitable for accommodating the stitching result of the two images; perform perspective transformation on one of the images and map it onto the new canvas; directly copy the other image to the corresponding position on the new canvas; the overlapping area of the two images will be fused or superimposed to form the final panoramic image.
[0141] S603. Clear the background of the stitched panoramic image.
[0142] In some embodiments, target objects that are far from the vehicle in the panoramic image can be cleared, and the target objects to be cleared can be determined through target recognition or the distance between the vehicle and surrounding target objects.
[0143] S604. Use the depth map generated from the front-view data to perform self-supervision on the MLP neural network model used for depth estimation of the panoramic image, and output the depth map corresponding to the panoramic image.
[0144] Here, the MLP neural network model can be a pre-trained neural network model or an untrained neural network model.
[0145] In some embodiments, the common area of the front-view image and the panoramic image can be determined first, and the depth information corresponding to this common area in the panoramic image can be estimated using the MLP neural network algorithm. Compare the estimated depth information with the depth information corresponding to this common area in the depth map of the front-view image. Perform self-supervised training on the MLP neural network algorithm according to the difference between the two depth information until the difference between the two depth information corresponding to the common area is less than the preset threshold, and then a trained MLP neural network model can be obtained; then use the trained MLP neural network model to estimate the depth information of the panoramic image, thereby obtaining the depth map corresponding to the panoramic image.
[0146] S605. Obtain the point cloud data corresponding to the panoramic image based on the depth map corresponding to the panoramic image and the color information in the panoramic image after background clearing.
[0147] In some embodiments, the depth information corresponding to each pixel point in the panoramic image can be converted into position coordinates in the world coordinate system, and then the position coordinates and color information of each pixel point in the panoramic image are combined to obtain the point cloud data corresponding to the panoramic image.
[0148] As Figure 7 shown, it is a schematic flowchart of a three-dimensional reconstruction method based on multi-frame panoramic data and surround-view data provided by the present application. The method includes: S701, distortion correction; Perform distortion correction on the current frame of surround-view image and panoramic image according to the calibration parameters of each camera.
[0149] In some embodiments, the current frame of surround-view image, the current frame of panoramic image, the surround-view images of the previous n frames, and the panoramic images of the previous n frames can be sequentially corrected according to the calibration parameters of each camera, so as to eliminate the distortion in the multi-frame surround-view images and panoramic images and improve the image quality.
[0150] S702, image stitching; Stitch the corrected current frame of surround-view image and panoramic image into a panoramic view using the internal parameters of each camera.
[0151] In some embodiments, the surround-view image and panoramic image collected at the same moment (with the same timestamp) can be stitched, that is, the current frame (the n + 1th frame) of surround-view image and the current frame of panoramic image are stitched to obtain the current frame panoramic view; the nth frame of surround-view image and the nth frame of panoramic image are stitched to obtain the nth frame panoramic image; the n - 1th frame of surround-view image and the n - 1th frame of panoramic image are stitched to obtain the n - 1th frame panoramic image,..., the 1st frame of surround-view image and the 1st frame of panoramic image are stitched to obtain the 1st frame panoramic image, so that multiple frames of panoramic views can be obtained.
[0152] S703, background removal; Remove the background from the stitched current frame panoramic view and the previous n frames of panoramic views.
[0153] In some embodiments, the useless background in the current frame panoramic view can be removed, and the useless background in the previous n frames of panoramic views can be removed.
[0154] S704, image matching, depth map generation, point cloud data generation; Perform image matching, depth map generation, and point cloud data generation on the current frame and the previous n frames of panoramic images after background removal in sequence, and generate the point cloud data corresponding to each frame of panoramic image in combination with the displacement information of each frame of the vehicle itself.
[0155] In some embodiments, image matching may be performed on each frame of panoramic images, and the images obtained after matching are fused to obtain a fused panoramic image. The fused panoramic image is processed according to the methods described in step S605 and step S606 to obtain the point cloud data corresponding to the fused panoramic image, and the point cloud data corresponding to the fused panoramic image is used as the point cloud data corresponding to the current frame of panoramic image. After that, the point cloud data of the first n frames and the point cloud data corresponding to the fused panoramic image are fused to obtain point cloud data based on surround-view data and forward-view data.
[0156] S705, fuse the point cloud data; Fuse the point cloud data corresponding to the first n frames of panoramic images and the point cloud data corresponding to the current frame of panoramic image to obtain point cloud data based on surround-view data and forward-view data.
[0157] In some embodiments, the position information of the point cloud data corresponding to the first n frames of panoramic images and the position information of the point cloud data corresponding to the current frame of panoramic image may be fused in combination with the displacement information of each frame of the host vehicle, and the color information of the point cloud data corresponding to the first n frames of panoramic images and the color information of the point cloud data corresponding to the current frame of panoramic image may be fused, so as to obtain point cloud data based on surround-view data and forward-view data.
[0158] In some embodiments, after obtaining the point cloud data based on forward-view data and the point cloud data based on surround-view data and forward-view data, the point cloud data based on forward-view data and the point cloud data based on surround-view data and forward-view data may be stitched or combined to obtain point cloud data based on image data. As Figure 8 shown, the point cloud data 801 based on image data and the point cloud data 802 output by the vehicle-mounted radar at the same moment may be fused to obtain target point cloud data 803.
[0159] After obtaining the target point cloud data, the target point cloud data may be rendered. As Figure 9 shown, it is a schematic flowchart of a rendering method for three-dimensional reconstruction point cloud data provided by the present application. The method includes: S901, Morton sorting; Perform Morton encoding on the target point cloud data and sort the obtained Morton encoding.
[0160] By performing Morton encoding on each point cloud in the target point cloud data and sorting the obtained Morton encoding in the Morton encoding sorting manner, adjacent points in the three-dimensional space can be arranged in adjacent positions.
[0161] S902, re-sort using the shuffle algorithm; Group the sorted Morton encoding and re-sort it in groups.
[0162] The Morton encoding sorting result in step S901 can be split into groups of m data (m can be the group size selected according to the target graphics card, for example, m can be 128), and then a shuffling algorithm is performed in units of groups to shuffle the data in units of groups, so as to reorder each group.
[0163] S903, upload it to the GPU and output the point cloud rendering image; Upload the Morton encoding obtained after reordering in units of groups to the GPU to obtain the rendering image corresponding to the target point cloud data.
[0164] After uploading the Morton encoding obtained after reordering in units of groups to the GPU, the GPU can use a compute shader for custom sampling, and use the method of depth clipping to render the target point cloud data into an image and output it to the in-vehicle display screen for the user to refer to for driving operations.
[0165] The 3D reconstruction method provided by this application can solve the problem that the current environment reconstruction function cannot accurately correspond to the real environment. Based on the 3D reconstruction of the real scene point cloud, the environment reconstruction uses the camera screen to construct an environment reconstruction picture closer to the real environment, and can quickly select and locate the corresponding camera screen area from the 3D environment reconstruction scene. This application uses the data of each vehicle sensor to render the environment reconstruction in the way of 3D reconstruction of the real scene point cloud, making the environment reconstruction of functions such as low-speed driving, camp guarding, and recorder playback more accurate, convenient, and corresponding to the real scene and the camera screen.
[0166] An embodiment of this application provides a 3D reconstruction device, as Figure 10 shown, the 3D reconstruction device 1000 includes: A first acquisition module 1001, configured to acquire a first current image frame collected by a binocular camera of the target vehicle and a second current image frame collected by a monocular camera; there is an overlapping area between the first current image frame and the second current image frame; A first determination module 1002, configured to determine a second depth map corresponding to the second current image frame based on the first depth map corresponding to the first current image frame; A second determination module 1003, based on the first depth map and the second depth map, determines the first environmental point cloud data of the target vehicle, so as to obtain a 3D reconstruction result of the environment around the target vehicle based on the first environmental point cloud data.
[0167] In some embodiments, the first current image frame includes two first current images; the first determination module 1002 includes: A first determination sub-module, configured to determine the first depth map based on the parallax information of corresponding pixel points in the two first current images; A first correction sub-module, configured to correct parameters of a neural network model for estimating image depth information based on the first depth map, and obtain a target neural network model; A first depth estimation sub-module, configured to perform depth estimation on the second current image frame by using the target neural network model, and obtain a second depth map.
[0168] In some embodiments, the first correction sub-module includes: A first depth estimation unit, configured to estimate depth information of the first current image frame by using the neural network model, and obtain a reference depth map; A first correction unit, configured to correct the parameters of the neural network model at least once based on the reference depth map and the first depth map until a convergence condition is met, so as to obtain the target neural network model.
[0169] In some embodiments, the first correction sub-module includes: A second depth estimation unit, configured to, if there is a first region with the same image content in the first current image frame and the second current image frame, estimate depth information of a sub-image corresponding to the first region in the second current image frame by using the neural network model, and obtain a depth map corresponding to the sub-image; A second correction unit, configured to correct the parameters of the neural network model at least once based on the depth map of the sub-image and a reference sub-depth map until a convergence condition is met, so as to obtain the target neural network model; the reference sub-depth is a depth map intercepted from the first depth map and corresponding to the first region in the first current image frame.
[0170] In some embodiments, the first determination sub-module includes: A preprocessing unit, configured to preprocess the two first current images respectively, and obtain two candidate images; A first determination unit, configured to determine a second region with the same image content in the two candidate images based on parameters of the binocular camera; A second determination unit, configured to determine the first depth map based on disparity information of pixel points corresponding to the second region in the two candidate images.
[0171] In some embodiments, the number of the monocular cameras includes a plurality, and different monocular images are collected by different monocular cameras; each monocular camera is disposed at a different position of the target vehicle; the first depth estimation sub-module includes: A splicing unit, configured to perform splicing processing on second current image frames respectively collected by the plurality of monocular cameras, and obtain a spliced second current image frame; A second depth estimation unit, configured to perform depth estimation on the spliced second current image frame by using the target neural network model to obtain a second depth map.
[0172] In some embodiments, the second determination module 1003 includes: A first acquisition sub-module, configured to acquire the current vehicle speed of the target vehicle; A second acquisition sub-module, configured to acquire a first historical image frame collected by the binocular camera and a second historical image frame collected by the monocular camera based on the current vehicle speed; A second determination sub-module, configured to determine first environmental point cloud data of the target vehicle based on the environmental point cloud data corresponding to the first historical image frame, the environmental point cloud data corresponding to the second historical image frame, the first depth map, and the second depth map.
[0173] In some embodiments, the second determination sub-module includes: A third determination unit, configured to determine environmental point cloud data corresponding to the first current image frame based on the first depth map and the first current image frame; and determine environmental point cloud data corresponding to the second current image frame based on the second depth map and the second current image frame; A first fusion unit, configured to fuse the environmental point cloud data corresponding to the first historical image frame and the environmental point cloud data corresponding to the first current image frame to obtain first candidate environmental point cloud data; and fuse the environmental point cloud data corresponding to the second historical image frame and the environmental point cloud data corresponding to the second current image frame to obtain second candidate environmental point cloud data; A second fusion unit, configured to determine the first environmental point cloud data based on the first candidate environmental point cloud data and the second candidate environmental point cloud data.
[0174] In some embodiments, the 3D reconstruction device 1000 further includes: A second acquisition module, configured to acquire second environmental point cloud data determined based on a radar sensor of the target vehicle; A first fusion module, configured to fuse the first environmental point cloud data and the second environmental point cloud data to obtain a 3D reconstruction result of the environment around the target vehicle.
[0175] An embodiment of the present application provides a vehicle, including a memory and a processor, where the memory stores a computer program that can run on the processor, and when the processor executes the program, some or all of the steps in the above method are implemented.
[0176] An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, some or all of the steps in the above method are implemented. The computer-readable storage medium can be transient or non-transient.
[0177] An embodiment of the present application provides a computer program product, including a computer program or instruction. When the computer program or instruction is executed by a processor, the steps in the method described in the above embodiment are implemented.
[0178] An embodiment of the present application provides a computer program product. The computer program product includes a non-transient computer-readable storage medium storing a computer program. When the computer program is read and executed by a computer, some or all of the steps in the above method are implemented. The computer program product can be specifically implemented by means of hardware, software, or a combination thereof. In some embodiments, the computer program product is specifically embodied as a computer storage medium. In other embodiments, the computer program product is specifically embodied as a software product, such as a Software Development Kit (SDK), etc.
[0179] It should be noted here that the descriptions of the above embodiments tend to emphasize the differences between the embodiments, and their similarities can be referred to each other. The descriptions of the above embodiments of the apparatus, vehicle, storage medium, and program product are similar to the descriptions of the above method embodiments and have beneficial effects similar to those of the method embodiments. For the technical details not disclosed in the embodiments of the apparatus, device, vehicle, storage medium, and program product of the present application, please refer to the descriptions of the method embodiments of the present application for understanding.
[0180] It should be understood that the "one embodiment" or "an embodiment" mentioned throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of the present application. Therefore, the "in one embodiment" or "in an embodiment" that appears throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures, or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the magnitudes of the serial numbers of the above processes do not mean the order of execution is prior or subsequent. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application. The serial numbers of the embodiments of the present application above are only for description and do not represent the advantages or disadvantages of the embodiments.
[0181] It should be noted that in this text, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including such element.
[0182] As described above, it is only to fully illustrate the embodiments of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present application should be covered within the protection scope of the present application.
Claims
1. A three-dimensional reconstruction method, characterized in that, Including: Obtaining a first current image frame collected by a binocular camera of a target vehicle and a second current image frame collected by a monocular camera; There is an overlapping area between the first current image frame and the second current image frame; Based on a first depth map corresponding to the first current image frame, determining a second depth map corresponding to the second current image frame; Based on the first depth map and the second depth map, determining first environmental point cloud data of the target vehicle, so as to obtain a three-dimensional reconstruction result of the environment around the target vehicle based on the first environmental point cloud data.
2. The three-dimensional reconstruction method according to claim 1, wherein The first current image frame includes two first current images; the determining the second depth map corresponding to the second current image frame based on the first depth map corresponding to the first current image frame includes: Determining the first depth map based on the disparity information of corresponding pixel points in the two first current images; Based on the first depth map, correcting the parameters of a neural network model for estimating image depth information to obtain a target neural network model; Using the target neural network model to perform depth estimation on the second current image frame to obtain a second depth map.
3. The three-dimensional reconstruction method according to claim 2, wherein The correcting the parameters of the neural network model for estimating image depth information based on the first depth map to obtain a target neural network model includes: Using the neural network model to estimate the depth information of the first current image frame to obtain a reference depth map; Based on the reference depth map and the first depth map, performing at least one correction on the parameters of the neural network model until a convergence condition is met to obtain the target neural network model.
4. The three-dimensional reconstruction method according to claim 2, wherein The correcting the parameters of the neural network model for estimating image depth information based on the first depth map to obtain a target neural network model includes: If there is a first area with the same image content in the first current image frame and the second current image frame, using the neural network model to estimate the depth information of a sub-image corresponding to the first area in the second current image frame to obtain a depth map corresponding to the sub-image; Based on the depth map of the sub-image and a reference sub-depth map, performing at least one correction on the parameters of the neural network model until a convergence condition is met to obtain the target neural network model; the reference sub-depth is a depth map intercepted from the first depth map and corresponding to the first area in the first current image frame.
5. The three-dimensional reconstruction method according to claim 2, wherein The determining the first depth map based on the disparity information of corresponding pixel points in the two first current images includes: Preprocessing the two first current images respectively to obtain two candidate images; Based on the parameters of the binocular camera, determining a second area with the same image content in the two candidate images; Based on the disparity information of corresponding pixel points in the second area of the two candidate images, determining the first depth map.
6. The three-dimensional reconstruction method according to claim 2, characterized in that, The number of the monocular cameras is multiple, and each monocular camera is arranged at a different position of the target vehicle; The using the target neural network model to perform depth estimation on the second current image frame to obtain a second depth map includes: Perform stitching processing on the second current image frames respectively collected by the multiple monocular cameras to obtain the stitched second current image frame; Use the target neural network model to perform depth estimation on the stitched second current image frame to obtain a second depth map.
7. The three-dimensional reconstruction method according to claim 1, wherein, Determining the first environmental point cloud data of the target vehicle based on the first depth map and the second depth map includes: Obtain the current vehicle speed of the target vehicle; Based on the current vehicle speed, obtain the first historical image frame collected by the binocular camera and the second historical image frame collected by the monocular camera; Based on the environmental point cloud data corresponding to the first historical image frame, the environmental point cloud data corresponding to the second historical image frame, the first depth map, and the second depth map, determine the first environmental point cloud data of the target vehicle.
8. The three-dimensional reconstruction method according to claim 7, wherein Determining the first environmental point cloud data of the target vehicle based on the environmental point cloud data corresponding to the first historical image frame, the environmental point cloud data corresponding to the second historical image frame, the first depth map, and the second depth map includes: Based on the first depth map and the first current image frame, determine the environmental point cloud data corresponding to the first current image frame; and based on the second depth map and the second current image frame, determine the environmental point cloud data corresponding to the second current image frame; Fuse the environmental point cloud data corresponding to the first historical image frame and the environmental point cloud data corresponding to the first current image frame to obtain first candidate environmental point cloud data; and fuse the environmental point cloud data corresponding to the second historical image frame and the environmental point cloud data corresponding to the second current image frame to obtain second candidate environmental point cloud data; Based on the first candidate environmental point cloud data and the second candidate environmental point cloud data, determine the first environmental point cloud data.
9. The three-dimensional reconstruction method according to any one of claims 1 to 8, characterized in that, The method further includes: Obtain second environmental point cloud data determined based on the radar sensor of the target vehicle; Fuse the first environmental point cloud data and the second environmental point cloud data to obtain a three-dimensional reconstruction result of the environment around the target vehicle.
10. A three-dimensional reconstruction device, characterized in that, Includes: A first acquisition module, configured to acquire a first current image frame collected by a binocular camera of a target vehicle and a second current image frame collected by a monocular camera; There is an overlapping area between the first current image frame and the second current image frame; A first determination module, configured to determine a second depth map corresponding to the second current image frame based on the first depth map corresponding to the first current image frame; A second determination module, configured to determine the first environmental point cloud data of the target vehicle based on the first depth map and the second depth map, so as to obtain a three-dimensional reconstruction result of the environment around the target vehicle based on the first environmental point cloud data.
11. A vehicle, comprising a memory and a processor, the memory storing a computer program that can run on the processor, characterized in that, When the processor executes the program, it implements the steps in the method according to any one of claims 1 to 9.
12. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps in the method according to any one of claims 1 to 9.
13. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or instruction is executed by the processor, it implements the steps in the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Scene reconstruction method and device, electronic equipment, program and medium
CN108230437A
Depth information determination method and device based on panoramic look-around system
CN112380963A
Binocular vision three-dimensional reconstruction method based on high-precision positioning and deep learning
CN114255279A
Depth estimation model training method, depth estimation method and electronic equipment
CN117314993A
Image processing method and device and storage medium
CN118279465A