Three-dimensional reconstruction method, device, vehicle, storage medium and program product
By combining binocular and monocular cameras with a neural network model to estimate image depth information, the problem of inaccurate correspondence between environment reconstruction and the real environment in existing technologies is solved, and efficient three-dimensional reconstruction effects are achieved.
Patent Information
- Application Number
- CN202510804733.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2045-06-17
AI Technical Summary
Existing environmental reconstruction technology cannot accurately correspond to the real driving environment, resulting in limited effectiveness of the environmental reconstruction function.
By acquiring image frames captured by the target vehicle's binocular camera and monocular camera, the image depth information is estimated using depth maps and neural network models, and three-dimensional reconstruction is performed in combination with environmental point cloud data to improve the accuracy of reconstruction.
The reconstructed driving environment is accurately matched to the real driving environment, which improves the accuracy of environmental reconstruction and the efficiency of three-dimensional reconstruction.
Smart Images

Figure CN120339524B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of environmental reconstruction, and in particular to a three-dimensional reconstruction method, device, vehicle, storage medium, and program product. Background Art
[0002] With the development of the automobile industry, cars are now generally equipped with various sensors such as cameras, millimeter-wave radars, and lidars, and generally have environmental reconstruction applications. However, current environmental reconstruction applications are generally based on the results of individual sensor perception fusion to draw various traffic product elements and functional elements in a computer virtual scene. The reconstructed driving environment in this way cannot accurately correspond to the real driving environment, which limits the role of the environmental reconstruction function. Summary of the Invention
[0003] The present application provides a three-dimensional reconstruction method, device, vehicle, storage medium and program product. The method can make the reconstructed driving environment accurately correspond to the real driving environment, thereby improving the accuracy of environment reconstruction.
[0004] The technical solution of this application is achieved as follows:
[0005] An embodiment of the present application provides a three-dimensional configuration method, including: obtaining a first current image frame captured by a binocular camera and a second current image frame captured by a monocular camera of a target vehicle; there is an overlapping area between the first current image frame and the second current image frame; based on a first depth map corresponding to the first current image frame, determining a second depth map corresponding to the second current image frame; based on the first depth map and the second depth map, determining first environmental point cloud data of the target vehicle, so as to obtain a three-dimensional reconstruction result of the environment surrounding the target vehicle based on the first environmental point cloud data.
[0006] According to the above technical means, since the first current image frame and the second current image frame are both complete environmental images around the target vehicle collected in real time, the first environmental point cloud data is determined based on the first depth map corresponding to the first current image frame and the second depth map corresponding to the second current image, which can improve the real-time and richness of the first environmental point cloud data, and the binocular parallax of the first current image frame collected by the binocular camera can be used to perform depth estimation to obtain the first depth map. Therefore, the second depth map corresponding to the second current image frame collected by the monocular camera is determined through the first depth map, which can improve the accuracy of the depth information estimation corresponding to the second current image frame, so that the reconstructed driving environment can accurately correspond to the real driving environment, thereby improving the accuracy of environmental reconstruction.
[0007] In some embodiments, the first current image frame includes two first current images; determining the second depth map corresponding to the second current image frame based on the first depth map corresponding to the first current image frame includes: determining the first depth map based on the disparity information of corresponding pixel points in the two first current images; based on the first depth map, correcting the parameters of the neural network model used to estimate the image depth information to obtain a target neural network model; and using the target neural network model to perform depth estimation on the second current image frame to obtain a second depth map.
[0008] According to the above-mentioned technical means, by correcting the parameters of the neural network model used to estimate image depth information based on the first depth map determined based on the disparity information of the two corresponding pixel points of the first current image in the first current image frame, the accuracy of the depth estimation of the neural network model can be improved. Therefore, by performing depth estimation on the second current image frame based on the corrected target neural network model, the accuracy of the determined second depth map can be improved.
[0009] In some embodiments, based on the first depth map, the parameters of the neural network model for estimating image depth information are corrected to obtain a target neural network model, including: using the neural network model to estimate the depth information of the first current image frame to obtain a reference depth map; based on the reference depth map and the first depth map, the parameters of the neural network model are corrected at least once until the convergence conditions are met to obtain the target neural network model.
[0010] According to the above technical means, the reference depth map and the first depth map obtained after estimating the first current depth information using the neural network model are corrected at least once, which can continuously improve the depth information estimation accuracy of the neural network model. Therefore, the depth information of the second current image frame is estimated using the target neural network model after parameter correction, which can improve the accuracy of the determined second depth map.
[0011] In some embodiments, based on the first depth map, the parameters of the neural network model for estimating image depth information are corrected to obtain a target neural network model, including: if there is a first area with the same image content in the first current image frame and the second current image frame, using the neural network model to estimate the depth information of the sub-image corresponding to the first area in the second current image frame to obtain a depth map corresponding to the sub-image; based on the depth map of the sub-image and the reference sub-depth map, the parameters of the neural network model are corrected at least once until the convergence conditions are met to obtain the target neural network model; the reference sub-depth is a depth map corresponding to the first area in the first current image frame intercepted from the first depth map.
[0012] According to the above-mentioned technical means, when it is determined that there is a first area with the same image content in the first current image frame and the second current image frame, the depth information of the sub-image corresponding to the first area in the second current image frame is estimated based on the neural network model. The depth map is obtained, and the depth map corresponding to the first area in the first current image frame is intercepted from the first depth map. The neural network model is corrected at least once, and the depth information estimation accuracy of the neural network model can be continuously improved. Therefore, the depth information of the second current image frame is estimated using the target neural network model with corrected parameters, so that the accuracy of the second depth map can be improved.
[0013] In some embodiments, determining the first depth map based on the disparity information of corresponding pixel points in the two first current images includes: preprocessing the two first current images respectively to obtain two candidate images; determining the second area with the same image content in the two candidate images based on the parameters of the binocular camera; and determining the first depth map based on the disparity information of corresponding pixel points in the second area in the two candidate images.
[0014] According to the above technical means, by preprocessing the two first current images, the image quality of the first current images can be improved and the computational complexity of three-dimensional reconstruction can be reduced. At the same time, based on the binocular camera parameters, the second area with the same image content in the two preprocessed candidate images is determined, and based on the disparity information of the corresponding pixel points in the second area, the first depth map can be accurately and quickly determined.
[0015] In some embodiments, there are multiple monocular cameras, and each monocular camera is set at a different position of the target vehicle; using the target neural network model to perform depth estimation on the second current image frame to obtain a second depth map includes: stitching the second current image frames respectively captured by the multiple monocular cameras to obtain a stitched second current image frame; using the target neural network model to perform depth estimation on the stitched second current image frame to obtain a second depth map.
[0016] According to the above technical means, by stitching the second current image frames captured by multiple monocular cameras and using the target neural network model to perform depth estimation on the stitched second current image frames, a one-time estimation of the depth information of the image frames captured by the monocular cameras can be achieved, thereby improving the efficiency of three-dimensional reconstruction.
[0017] In some embodiments, determining the first environmental point cloud data of the target vehicle based on the first depth map and the second depth map includes: obtaining the current speed of the target vehicle; based on the current speed, obtaining the first historical image frame captured by the binocular camera and the second historical image frame captured by the monocular camera; and determining the first environmental point cloud data of the target vehicle based on the environmental point cloud data corresponding to the first historical image frame, the environmental point cloud data corresponding to the second historical image frame, the first depth map, and the second depth map.
[0018] According to the above-mentioned technical means, based on the current speed of the target vehicle, the first historical image frame and the second historical image frame adapted to the speed can be obtained. Further, based on the environmental point cloud data corresponding to the first historical image frame, the environmental point cloud data corresponding to the second historical image frame, the first depth map and the second depth map, it can be determined that the first environmental point cloud data of the target vehicle contains more detailed information, thereby improving the effect of three-dimensional reconstruction.
[0019] In some embodiments, the first environmental point cloud data of the target vehicle is determined based on the environmental point cloud data corresponding to the first historical image frame, the environmental point cloud data corresponding to the second historical image frame, the first depth map and the second depth map, including: determining the environmental point cloud data corresponding to the first current image frame based on the first depth map and the first current image frame; and determining the environmental point cloud data corresponding to the second current image frame based on the second depth map and the second current image frame; fusing the environmental point cloud data corresponding to the first historical image frame with the environmental point cloud data corresponding to the first current image frame to obtain first candidate environmental point cloud data; and fusing the environmental point cloud data corresponding to the second historical image frame with the environmental point cloud data corresponding to the second current image frame to obtain second candidate environmental point cloud data; and determining the first environmental point cloud data based on the first candidate environmental point cloud data and the second candidate environmental point cloud data.
[0020] According to the above-mentioned technical means, by fusing the environmental point cloud data corresponding to the first historical image frame and the environmental point cloud data corresponding to the first current image frame, the first candidate environmental point cloud data is obtained, and the environmental point cloud data corresponding to the second historical image frame and the environmental point cloud data corresponding to the second current image frame are fused to obtain the second candidate environmental point cloud data. The first environmental point cloud data finally determined based on the first candidate environmental point cloud data and the second candidate environmental point cloud data can contain more details in the environment surrounding the target vehicle, thereby improving the accuracy of three-dimensional reconstruction.
[0021] In some embodiments, second environmental point cloud data determined based on a radar sensor of the target vehicle is obtained; the first environmental point cloud data and the second environmental point cloud data are fused to obtain a three-dimensional reconstruction result of the environment surrounding the target vehicle.
[0022] According to the above technical means, by fusing the first environmental point cloud data and the second environmental point cloud data determined by the radar sensor, the three-dimensional reconstruction result of the target vehicle's surrounding environment can contain more and more accurate point cloud information, thereby improving the accuracy of the three-dimensional reconstruction.
[0023] The present invention provides a three-dimensional reconstruction device, comprising:
[0024] A first acquisition module is configured to acquire a first current image frame captured by a binocular camera and a second current image frame captured by a monocular camera of a target vehicle; wherein the first current image frame and the second current image frame have an overlapping area;
[0025] a first determining module, configured to determine a second depth map corresponding to the second current image frame based on a first depth map corresponding to the first current image frame;
[0026] The second determination module is used to determine first environmental point cloud data of the target vehicle based on the first depth map and the second depth map, so as to obtain a three-dimensional reconstruction result of the surrounding environment of the target vehicle based on the first environmental point cloud data.
[0027] An embodiment of the present application provides a vehicle, comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and is characterized in that the processor implements the steps in the above method when executing the program.
[0028] An embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, which implements the steps in the above method when executed by a processor.
[0029] An embodiment of the present application provides a computer program product, including a computer program or instructions, which implement the steps in the above method when the computer program or instructions are executed by a processor.
[0030] Beneficial effects of this application:
[0031] It can improve the real-time and richness of the first environment point cloud data, and improve the accuracy of the depth information estimation corresponding to the second current image frame, so that the reconstructed driving environment can accurately correspond to the real driving environment, thereby improving the accuracy of environment reconstruction.
[0032] The accuracy of the depth estimation of the neural network model can be continuously improved. Therefore, performing depth estimation on the second current image frame based on the modified target neural network model can improve the accuracy of the determined second depth map.
[0033] The first depth map can be determined accurately and quickly.
[0034] It can improve the efficiency of three-dimensional reconstruction.
[0035] This allows the first environmental point cloud data for determining the target vehicle to contain more detailed information, thereby improving the effect of three-dimensional reconstruction.
[0036] The first environment point cloud data finally determined based on the first candidate environment point cloud data and the second candidate environment point cloud data can contain more details of the environment surrounding the target vehicle, thereby improving the accuracy of three-dimensional reconstruction. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 A schematic diagram of a flow chart of a three-dimensional reconstruction method provided in an embodiment of the present application;
[0038] Figure 2 A schematic structural diagram of a vehicle real-scene reconstruction system provided in an embodiment of the present application;
[0039] Figure 3 A schematic diagram of a flow chart of a 3D reconstruction method based on single-frame forward-looking data provided in an embodiment of the present application;
[0040] Figure 4 A schematic diagram of the principle of determining pixel depth information provided in an embodiment of the present application;
[0041] Figure 5 A schematic diagram of a flow chart of a 3D reconstruction method based on multi-frame forward-looking data provided in an embodiment of the present application;
[0042] Figure 6 A flowchart of a 3D reconstruction method based on single-frame circumferential view data and surround view data provided in an embodiment of the present application;
[0043] Figure 7 A flowchart of a 3D reconstruction method based on multiple frames of circumferential view data and surround view data provided in an embodiment of the present application;
[0044] Figure 8 A schematic diagram of the point cloud data fusion principle provided in an embodiment of the present application;
[0045] Figure 9 A flowchart of a method for rendering three-dimensional reconstructed point cloud data provided in an embodiment of the present application;
[0046] Figure 10 A schematic structural diagram of a three-dimensional reconstruction device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0047] The technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application.
[0048] In order to make the purpose, technical solutions and advantages of this application clearer, this application will be further described below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0049] In the following description, reference is made to “some embodiments\other embodiments”, which describe a subset of all possible embodiments, but it can be understood that “some embodiments\other embodiments” may be the same subset or different subsets of all possible embodiments, and may be combined with each other without conflict.
[0050] In the following description, the terms "first\second" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It is understandable that "first\second" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0051] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0052] Based on the problems existing in the related art, the embodiment of the present application provides a three-dimensional reconstruction method, which can make the reconstructed driving environment accurately correspond to the real driving environment, thereby improving the accuracy of the environment reconstruction. Figure 1 FIG. 1 is a flow chart of a three-dimensional reconstruction method provided in an embodiment of the present application, the method comprising the following steps S101 to S103:
[0053] Step S101 : obtaining a first current image frame captured by a binocular camera and a second current image frame captured by a monocular camera of a target vehicle.
[0054] Here, the binocular camera can be positioned on a first side of the target vehicle, and the monocular camera can be positioned on a side of the target vehicle adjacent to the first side. The first side can be the front side of the target vehicle, or at least one of the front, left, right, and rear sides of the target vehicle. For example, if the binocular camera is positioned on the front side of the target vehicle, the monocular camera can be positioned on the left and / or right side of the target vehicle. For another example, if the binocular camera is positioned on the left side of the target vehicle, the monocular camera can be positioned on the front and / or rear side of the target vehicle.
[0055] Both monocular cameras and binocular cameras are on-board cameras installed on the target vehicle. For example, a monocular camera can be a surround-view camera or a panoramic camera, and a binocular camera can be a telephoto camera or a wide-angle camera.
[0056] In some embodiments, the monocular camera of the target vehicle may include one or more, the binocular camera may be a combination of two vehicle-mounted cameras, and the binocular camera may include at least one group. The installation positions and / or installation angles of multiple monocular cameras arranged on the same side of the target vehicle are different; the installation positions and installation angles of the binocular cameras arranged on the same side of the target vehicle are the same.
[0057] In some embodiments, there is an overlapping area between the first current image frame and the second current image frame, that is, there are areas with the same image content in the first current image frame and the second current image frame. For example, the first current image frame includes the image content captured by the binocular camera of the front area, left front area and right front area of the target vehicle, and the second current image frame includes the image content captured by the monocular camera of the left area, left front area and left rear area of the target vehicle. Then the overlapping area between the first current image frame and the second current image frame is the image content corresponding to the left front area of the target vehicle.
[0058] In some embodiments, the first current image frame may include two or more, and the second current image frame may include one or more. The first current image frame may be an environmental image of the first side of the target vehicle captured by the binocular camera at the current moment, and the second current image frame may be an environmental image of the side adjacent to the first side of the target vehicle captured by the monocular camera at the current moment.
[0059] Step S102 : determining a second depth map corresponding to a second current image frame based on a first depth map corresponding to the first current image frame.
[0060] In some embodiments, a first depth map corresponding to a first current image frame can be determined based on internal and external parameters of a binocular camera, and a second depth map corresponding to a second current image frame can be determined based on the first depth map. The first depth map corresponding to the first current image frame can be determined based on disparity information of corresponding pixels in a plurality of first current image frames, and the depth map corresponding to the second current image frame can be calculated using a depth information estimation algorithm, such as a multilayer perceptron (MLP) neural network model. The MLP neural network model can be trained based on the depth map corresponding to the first current image frame, and the trained MLP neural network model can be used to estimate the depth information of the second current image frame to obtain the second depth map.
[0061] Step S103 : determining first environmental point cloud data of the target vehicle based on the first depth map and the second depth map, so as to obtain a three-dimensional reconstruction result of the surrounding environment of the target vehicle based on the first environmental point cloud data.
[0062] In some embodiments, the first depth map and the second depth map can constitute the depth information of the complete area around the target vehicle. Based on the internal parameters of the binocular camera corresponding to the first current image frame, the pixel coordinates of each pixel point in the first current image frame can be converted into coordinates in the binocular camera coordinate system, and then combined with the depth information of each pixel point in the first depth map to determine the point cloud data corresponding to the first current image frame; based on the internal parameters of the monocular camera corresponding to the second current image frame, the pixel coordinates of each pixel point in the second current image frame can be converted into coordinates in the monocular camera coordinate system, and then combined with the depth information of each pixel point in the second depth map to determine the point cloud data corresponding to the second current image frame; the point cloud data corresponding to the first current image frame and the point cloud data corresponding to the second current image frame are fused or combined to obtain the first environmental point cloud data of the target vehicle.
[0063] In some embodiments, the first environmental point data may include position information and color information of each point cloud. The position information can be determined by a depth map and internal parameters of a monocular camera and a binocular camera. The color information can be determined by the color information of each pixel in the first current image frame and the color information of each pixel in the second current image frame.
[0064] In some embodiments, the first environmental point data can be directly used as the three-dimensional reconstruction result of the target vehicle's surrounding environment; or the three-dimensional reconstruction result of the target vehicle's surrounding environment can be determined based on the point cloud data determined by the target vehicle's radar sensor and the first environmental point cloud data.
[0065] In an embodiment of the present application, a first current image frame captured by a binocular camera and a second current image frame captured by a monocular camera of a target vehicle are obtained; the binocular camera is set on a first side of the target vehicle, and the monocular camera is set on a side of the target vehicle adjacent to the first side; based on a first depth map corresponding to the first current image frame, a second depth map corresponding to the second current image frame is determined; based on the first depth map and the second depth map, first environmental point cloud data of the target vehicle is determined to obtain a three-dimensional reconstruction result of the environment surrounding the target vehicle based on the first environmental point cloud data. In this way, since the first current image frame and the second current image frame are both complete environmental images around the target vehicle collected in real time, the first environmental point cloud data is determined based on the first depth map corresponding to the first current image frame and the second depth map corresponding to the second current image, which can improve the real-time and richness of the first environmental point cloud data, and the binocular parallax of the first current image collected by the binocular camera can be used to perform depth estimation to obtain the first depth map. Therefore, the second depth map corresponding to the second current image frame collected by the monocular camera is determined through the first depth map, which can improve the accuracy of the depth information estimation corresponding to the second current image frame, so that the reconstructed driving environment can accurately correspond to the real driving environment, thereby improving the accuracy of environmental reconstruction.
[0066] In some embodiments, the first current image frame includes two first current images, and the step S102 of determining the second depth map corresponding to the second current image frame based on the first depth map corresponding to the first current image frame includes the following steps S1021 to S1023:
[0067] Step S1021 : determining a first depth map based on disparity information of corresponding pixels in two first current images.
[0068] It is understandable that the binocular camera can capture two images of the same area at different viewing angles at the same time to obtain a first current image frame including two first current images, where the two first current images correspond to different viewing angles respectively.
[0069] In some embodiments, the two first current images in the first current image frame can be matched to determine the same area in the two first current images, and then based on the disparity information of the corresponding pixel points in the same area, and the distance and focal length between the binocular cameras corresponding to the two first current images, the first depth map can be determined. That is, the depth information of the corresponding pixel points in the same area is calculated by the binocular depth estimation algorithm, and the depth information of multiple pixel points in the same area can obtain the first depth map.
[0070] Step S1022: Based on the first depth map, the parameters of the neural network model used to estimate the image depth information are corrected to obtain a target neural network model.
[0071] Here, the neural network model used to estimate image depth information can be a pre-trained neural network model or an untrained neural network model. The neural network model can be an MLP neural network model, etc. The neural network model can be self-supervised trained based on the first depth map until the convergence conditions are met, and the target neural network model can be obtained.
[0072] In some embodiments, the neural network model used to estimate image depth information may be a model that has been pre-trained using simulated laboratory data or a small amount of training samples (image data) in a vehicle environment. However, the accuracy of the depth information estimation by the neural network model is low, for example, the estimation accuracy may be less than 50%.
[0073] In some embodiments, the depth information of the first current image frame can be estimated using a neural network model for estimating image depth information. Based on the degree of difference between the estimated depth map and the first depth map, the parameters of the neural network model for estimating image depth information are back-propagated and trained until the degree of difference between the estimated depth map and the first depth map is less than a preset threshold. The training can then be stopped to obtain the target neural network model.
[0074] Step S1023: Use the target neural network model to perform depth estimation on the second current image frame to obtain a second depth map.
[0075] In some embodiments, after obtaining the target neural network model, the second current image frame can be input into the target neural network model, and the depth of the second current image can be estimated by the target neural network model to obtain a second depth map.
[0076] In the above embodiment, by correcting the parameters of the neural network model used to estimate image depth information based on the first depth map determined based on the disparity information of the corresponding pixel points of the two first current images in the two first current image frames, the accuracy of the depth estimation of the neural network model can be improved. Therefore, by performing depth estimation on the second current image frame based on the corrected target neural network model, the accuracy of the second depth map can be improved.
[0077] In some embodiments, the step S1022 described above of modifying the parameters of the neural network model for estimating image depth information based on the first depth map to obtain a target neural network model may include the following steps S10221A to S10222A:
[0078] Step S10221A: Use a neural network model to estimate the depth information of the first current image frame to obtain a reference depth map.
[0079] In some embodiments, the first current image frame can be input into a neural network model, and the neural network model can estimate the depth information of each pixel in the first current image frame, and generate a reference depth map based on the depth information of each pixel.
[0080] Step S10222A: Based on the reference depth map and the first depth map, the parameters of the neural network model are corrected at least once until the convergence conditions are met to obtain the target neural network model.
[0081] In some embodiments, the convergence condition may be satisfied when the difference between the reference depth map and the first depth map is less than a depth information difference threshold. The difference between the reference depth map and the first depth map may be determined first, and the parameters of the neural network model may be corrected based on the difference between the two depth images. The depth of the first current image frame may be estimated again using the corrected neural network model to determine the difference between the depth information of the pixel in the re-estimated reference depth map and the depth information of the corresponding pixel in the first depth map. The parameters of the neural network model may be corrected again based on the difference until the difference is less than the depth information difference threshold. The correction of the parameters of the neural network model may then be stopped to obtain the target neural network model. The difference between the two depth images may be determined based on the average or weighted average of the difference between the depth information of each pixel in the reference depth map and the depth information of the corresponding pixel in the first depth map.
[0082] In some embodiments, the degree of difference between the depth information of each pixel in the reference depth map and the depth information of the corresponding pixel in the first depth map can be determined based on methods such as cosine similarity and Euclidean distance. For example, the average value or weighted average value of the cosine similarity (or Euclidean distance) of the depth information corresponding to each pixel is determined as the degree of difference between the reference depth map and the first depth map.
[0083] In the above embodiment, the reference depth map and the first depth map obtained after estimating the first current depth information using the neural network model are corrected at least once, which can continuously improve the depth information estimation accuracy of the neural network model. Therefore, the depth information of the second current image frame is estimated using the target neural network model after parameter correction, which can improve the accuracy of the second depth map.
[0084] In some embodiments, the step S1022 described above of modifying the parameters of the neural network model for estimating image depth information based on the first depth map to obtain a target neural network model may include the following steps S10221B to S10222B:
[0085] Step S10221B: If there is a first area with the same image content in the first current image frame and the second current image frame, the neural network model is used to estimate the depth information of the sub-image corresponding to the first area in the second current image frame to obtain a depth map corresponding to the sub-image.
[0086] In some embodiments, the images captured by the binocular camera and the monocular camera may have overlapping areas, such as the left front area, right front area, left rear area, and right rear area of the target vehicle. Therefore, the first current image frame and the second current image frame may have areas with the same image content.
[0087] In some embodiments, when it is determined that there is a first area with the same image content in the first current image frame and the second current image frame, the image content of the first area in the second current image frame can be intercepted to obtain a corresponding sub-image, and the sub-image is input into a neural network model to estimate the depth information to obtain a depth map corresponding to the sub-image.
[0088] Step S10222B: Based on the depth map of the sub-image and the reference sub-depth map, the parameters of the neural network model are corrected at least once until the convergence conditions are met to obtain the target neural network model.
[0089] Here, the reference sub-depth is a depth map corresponding to the first area in the first current image frame that is extracted from the first depth map.
[0090] In some embodiments, the parameters of the neural network model may be modified based on the difference between the depth map of the sub-image estimated by the neural network model and the reference sub-depth map. For example, the cosine similarity between the depth information of each pixel in the sub-image depth map and the depth information of the corresponding pixel in the reference sub-depth map may be calculated, and the average of the cosine similarities may be used as the difference between the depth map of the sub-image and the reference sub-depth map.
[0091] In some embodiments, if the difference value between the depth map of the sub-image and the reference sub-depth map is greater than or equal to the depth information difference threshold, it is considered that the convergence condition is not met. At this time, the parameters of the neural network model can be corrected, and the corrected neural network model can be used to estimate the depth information of the sub-image corresponding to the first area in the second current image frame again. This process is repeated until the difference value between the depth map of the sub-image estimated by the neural network model and the reference sub-depth map is less than the depth information difference threshold, that is, it is considered that the convergence condition is met, and the parameter correction of the neural network model can be stopped, so that the target neural network model can be obtained.
[0092] In the above embodiment, when it is determined that there is a first area with the same image content in the first current image frame and the second current image frame, the neural network model is corrected at least once based on the depth map obtained after estimating the depth information of the sub-image corresponding to the first area in the second current image frame based on the neural network model, and the depth map corresponding to the first area in the first current image frame intercepted from the first depth map. The depth information estimation accuracy of the neural network model can be continuously improved. Therefore, by using the target neural network model with corrected parameters to estimate the depth information of the second current image frame, the accuracy of the second depth map can be improved.
[0093] In some embodiments, determining the first depth map based on disparity information of corresponding pixels in the two first current images in step S1021 may include steps S10211 to S10213:
[0094] Step S10211 , pre-processing the two first current images respectively to obtain two candidate images.
[0095] In some embodiments, pre-processing the two first current images may include performing distortion correction, background removal, etc. on the two first current images.
[0096] In some embodiments, the binocular camera may be a fisheye camera, and thus the first current image may be a fisheye image. Therefore, the first current image is corrected to eliminate distortion in the first current image. After the distortion correction is performed on the first current image, the background in the distortion-corrected first current image may be cleared. For example, to remove target objects that are far from the target vehicle in the first current image, such as the sky or clouds, the background may be cleared.
[0097] In some embodiments, when clearing the background of the first current image, the target object in the first current image can be identified. When a specific target object is identified, the specific target object in the first current image can be cleared, wherein the specific target object can be a target object in a distant scene such as the sky, mountains, clouds, etc.
[0098] In other embodiments, when clearing the background of the first current image, the distance between each target object (such as other vehicles, pedestrians, road signs, etc.) around the target vehicle and the target vehicle can be determined first, and the target objects with a distance greater than a distance threshold can be cleared from the first current image.
[0099] Step S10212: Determine a second region having the same image content in the two candidate images based on the parameters of the binocular camera.
[0100] In some embodiments, the pixel position coordinates of each pixel point in each candidate image can be converted into position coordinates in the corresponding camera coordinate system based on the intrinsic parameters of the binocular camera, and then the position coordinates in the camera coordinates can be converted into position coordinates in the world coordinate system (or position coordinates in the vehicle body coordinate system) based on the extrinsic parameters of the binocular camera. That is, the pixel points in the two candidate images are unified into the same coordinate system, and pixel matching is performed based on the position coordinates in the same coordinate system to determine the second area with the same image content in the two candidate images.
[0101] Step S10213: Determine a first depth map based on disparity information of corresponding pixels in the second area of the two candidate images.
[0102] Here, the disparity information of the pixel point in the second area of the candidate image can be the disparity of the same pixel point in the two candidate images, and the disparity can be the absolute value of the difference in the horizontal coordinate values of the same pixel point in the pixel coordinate system of the two candidate images (the origin is located in the upper left corner of the candidate image, the horizontal direction is the u axis, and the vertical direction is the v axis).
[0103] In some embodiments, after determining the disparity information of a pixel point, the depth information corresponding to the pixel point can be calculated based on the distance between the centers of the binocular cameras, the focal length (the focal lengths of the binocular cameras are the same) and the disparity information. The depth information of other pixel points in the second area of the candidate image can also be calculated using this process. Finally, the depth information of each pixel point in the second area of the candidate image is obtained, and then a first depth map is constructed based on the depth information of each pixel point and the position information of the pixel point in the candidate image.
[0104] In the above embodiment, by preprocessing the two first current images, the image quality of the first current images can be improved and the computational complexity of three-dimensional reconstruction can be reduced. At the same time, based on the binocular camera parameters, a second area with the same image content in the preprocessed candidate image is determined, and based on the disparity information of the corresponding pixel points in the second area, the first depth map can be accurately and quickly determined.
[0105] In some embodiments, there are multiple monocular cameras, each of which is located at a different position on the target vehicle; performing depth estimation on the second current image frame using the target neural network model to obtain a second depth map in step S1023 may include the following steps S10231 to S10232:
[0106] Step S10231: performing stitching processing on the second current image frames respectively captured by the multiple monocular cameras to obtain a stitched second current image frame.
[0107] In some embodiments, before stitching together the second current image frames captured by multiple monocular cameras, pre-processing operations such as distortion correction and background removal can also be performed on the second current image frames. Subsequently, operations such as feature extraction, feature matching, image registration, and image fusion can be sequentially performed on the multiple second current image frames to obtain a stitched second current image frame. Furthermore, color correction and stitching artifact removal can be performed on the stitched second current image frame to further improve the quality of the stitched second current image frame.
[0108] In some embodiments, if regions of the second current image frames captured by multiple monocular cameras contain identical image content, these regions can be fused, and regions of differing image content can be directly spliced. If monocular cameras are located on the left and right sides of the target vehicle, with two monocular cameras on each side, the second current image frames captured by the two monocular cameras on the same side can be spliced together first. The second current image frames obtained from the splicing on each side can then be spliced together to obtain a spliced second current image frame.
[0109] Step S10232: Use the target neural network model to perform depth estimation on the spliced second current image frame to obtain a second depth map.
[0110] In some embodiments, the spliced second current image frame is input into a target neural network model, and the target neural network model can estimate the depth information of each pixel in the spliced second current image frame, thereby obtaining a second depth map.
[0111] In the above embodiment, by stitching the second current image frames captured by multiple monocular cameras and performing depth estimation on the stitched second current image frames using the target neural network model, a one-time estimation of the depth information of the image frames captured by the monocular cameras can be achieved, thereby improving the efficiency of three-dimensional reconstruction.
[0112] In some embodiments, determining the first environmental point cloud data of the target vehicle based on the first depth map and the second depth map in step S103 includes steps S1031 to S1033:
[0113] Step S1031, obtaining the current speed of the target vehicle.
[0114] In some embodiments, the current speed information of the target vehicle may be the speed of the target vehicle when the binocular camera captures the first current image frame and when the monocular camera captures the second current image frame.
[0115] Step S1032 : Based on the current vehicle speed, a first historical image frame captured by the binocular camera and a second historical image frame captured by the monocular camera are acquired.
[0116] Here, the first historical image frame may be one or more image frames captured by the binocular camera before the first current image frame, for example, may include the previous image frame, the previous two image frames, the previous three image frames, etc. of the first current image frame; similarly, the second historical image frame may also be one or more image frames captured by the monocular camera before the second current image frame.
[0117] In some embodiments, the number of the first historical image frames and the number of the second historical image frames may be determined based on the current speed of the target vehicle. The number of the first historical image frames and the number of the second historical image frames may be the same or different.
[0118] In some embodiments, a correspondence between vehicle speed, the number of first historical image frames, and the number of second historical image frames can be pre-established. For example, if the current vehicle speed is greater than a speed threshold, the number of first historical image frames can be determined to be m1, and the number of second historical image frames can be determined to be m2. If the current vehicle speed is less than or equal to the speed threshold, the number of first historical image frames can be determined to be n1, and the number of second historical image frames can be determined to be n2, where m1, m2, n1, and n2 are all positive integers, and m1 is less than n1, and m2 is less than n2. Alternatively, if the current vehicle speed is within a first preset speed range, the number of first image frames can be determined to be m3, and the number of second historical image frames can be n3. If the current vehicle speed is within a second preset speed range, the number of first image frames can be determined to be m4, and the number of second historical image frames can be n4, where the upper limit of the first preset speed range is less than the lower limit of the second preset speed range, and m3, m4, n3, and n4 are all positive integers, and m4 is less than n3, and m4 is less than n3.
[0119] Step S1033 : determining first environmental point cloud data of the target vehicle based on the environmental point cloud data corresponding to the first historical image frame, the environmental point cloud data corresponding to the second historical image frame, the first depth map, and the second depth map.
[0120] In some embodiments, the environmental point cloud data of the first historical image frame can be calculated by using the disparity information of the pixel points in the first historical image frame, the depth information of each pixel point in the second historical image frame can be obtained by using a neural network model for estimating depth information, and then the environmental point cloud data corresponding to the second historical image frame can be calculated based on the depth information of each pixel point.
[0121] In some embodiments, the environmental point cloud data corresponding to each first historical image frame can be obtained by fusing the environmental point cloud data of the first historical image frame with the environmental point cloud data of other image frames before the first historical image frame, and the environmental point cloud data corresponding to each second historical image frame can be obtained by fusing the environmental point cloud data of the second historical image frame with the environmental point cloud data of other image frames before the second historical image frame.
[0122] In some embodiments, the environmental point cloud data corresponding to the first current image frame can be determined based on the first depth map and the color information in the first current image frame, and then the environmental point cloud data corresponding to the first current image frame and the environmental point cloud data corresponding to the first historical image frame are fused to obtain the target environmental point cloud data corresponding to the first current image frame. Similarly, the environmental point cloud data corresponding to the second current image frame can be determined based on the second depth map and the color information in the second current image frame, and then the environmental point cloud data corresponding to the second current image frame and the environmental point cloud data corresponding to the second historical image frame are fused to obtain the target environmental point cloud data corresponding to the second current image frame. Thereafter, the target environmental point cloud data corresponding to the first current image frame and the target environmental point cloud data corresponding to the second current image frame are fused to obtain the first environmental point cloud data of the target vehicle.
[0123] In the above embodiment, based on the current speed of the target vehicle, a first historical image frame and a second historical image frame adapted to the speed can be obtained. Further, based on the environmental point cloud data corresponding to the first historical image frame and the environmental point cloud data corresponding to the second historical image frame, the first depth map and the second depth map, it can be determined that the first environmental point cloud data of the target vehicle contains more detailed information, thereby improving the effect of three-dimensional reconstruction.
[0124] In some embodiments, determining the first environmental point cloud data of the target vehicle based on the environmental point cloud data corresponding to the first historical image frame, the environmental point cloud data corresponding to the second historical image frame, the first depth map, and the second depth map in step S1033 may include steps S10331 to S10333:
[0125] Step S10331: Determine the environmental point cloud data corresponding to the first current image frame based on the first depth map and the first current image frame; and determine the environmental point cloud data corresponding to the second current image frame based on the second depth map and the second current image frame.
[0126] In some embodiments, the depth information of each pixel in the first depth map and the color information of the corresponding pixel in the first current image frame can be combined to obtain the environmental point cloud data corresponding to the first current image frame; the depth information of each pixel in the second depth map and the color information of the corresponding pixel in the second current image frame can be combined to obtain the environmental point cloud data corresponding to the second current image frame.
[0127] Step S10332: fuse the environmental point cloud data corresponding to the first historical image frame and the environmental point cloud data corresponding to the first current image frame to obtain first candidate environmental point cloud data; and fuse the environmental point cloud data corresponding to the second historical image frame and the environmental point cloud data corresponding to the second current image frame to obtain second candidate environmental point cloud data.
[0128] In some embodiments, a weighted fusion method can be used to fuse the environmental point cloud data corresponding to the first historical image frame and the environmental point cloud data corresponding to the first current image frame, and to fuse the environmental point cloud data corresponding to the second historical image frame and the environmental point cloud data corresponding to the second current image frame. The weight of the environmental point cloud data corresponding to the first historical image frame, the weight of the environmental point cloud data corresponding to the first current image frame, the weight of the environmental point cloud data corresponding to the second historical image frame, and the weight of the environmental point cloud data corresponding to the second current image frame can be determined according to the method not limited in this application.
[0129] Step S10333 : Determine the first environment point cloud data based on the first candidate environment point cloud data and the second candidate environment point cloud data.
[0130] In some embodiments, the first candidate environment point cloud data and the second candidate environment point cloud data can be fused, and the fusion method can be splicing, weighted fusion, etc. By fusing the first candidate point cloud data and the second candidate environment point cloud data, the first environment point cloud data obtained can constitute the point cloud data of the complete area around the target vehicle.
[0131] In the above embodiment, by fusing the environmental point cloud data corresponding to the first historical image frame and the environmental point cloud data corresponding to the first current image frame, first candidate environmental point cloud data are obtained, and by fusing the environmental point cloud data corresponding to the second historical image frame and the environmental point cloud data corresponding to the second current image frame, second candidate environmental point cloud data are obtained. This allows the first environmental point cloud data ultimately determined based on the first candidate environmental point cloud data and the second candidate environmental point cloud data to include more details in the environment surrounding the target vehicle, thereby improving the accuracy of three-dimensional reconstruction.
[0132] In some embodiments, the above method may further include steps S104 to S105:
[0133] Step S104: Acquire second environment point cloud data determined by the radar sensor of the target vehicle.
[0134] In some embodiments, the second environment point cloud data can be obtained by calculating the data collected by the radar sensor of the target vehicle based on the data collected at the same time when the binocular camera collects the first current image frame and the monocular camera collects the second image frame, or by calculating the data collected by the radar sensor by the computing device of the target vehicle.
[0135] Step S105 : fusing the first environment point cloud data and the second environment point cloud data to obtain a three-dimensional reconstruction result of the environment surrounding the target vehicle.
[0136] In some embodiments, the first environment point cloud data and the second environment point cloud data can be weightedly fused, and the weight of the first environment point cloud data and the weight of the second environment point cloud data can be determined based on the configuration level or accuracy of the radar sensor. The higher the configuration level or the higher the accuracy of the radar sensor, the smaller the weight of the first environment point cloud data and the greater the weight of the second environment point cloud data; conversely, the greater the weight of the first environment point cloud data, the smaller the weight of the second environment point cloud data.
[0137] In the above embodiment, by fusing the first environment point cloud data and the second environment point cloud data determined by the radar sensor, the three-dimensional reconstruction result of the target vehicle's surrounding environment can contain more and more accurate point cloud information, thereby improving the accuracy of the three-dimensional reconstruction.
[0138] In an embodiment of the present application, a first current image frame captured by a binocular camera and a second current image frame captured by a monocular camera of a target vehicle are obtained; the binocular camera is set on a first side of the target vehicle, and the monocular camera is set on a side of the target vehicle adjacent to the first side; based on a first depth map corresponding to the first current image frame, a second depth map corresponding to the second current image frame is determined; based on the first depth map and the second depth map, first environmental point cloud data of the target vehicle is determined to obtain a three-dimensional reconstruction result of the environment surrounding the target vehicle based on the first environmental point cloud data. In this way, since the first current image frame and the second current image frame are both complete environmental images around the target vehicle collected in real time, the first environmental point cloud data is determined based on the first depth map corresponding to the first current image frame and the second depth map corresponding to the second current image, which can improve the real-time and richness of the first environmental point cloud data, and the binocular parallax of the first current image collected by the binocular camera can be used to perform depth estimation to obtain the first depth map. Therefore, the second depth map corresponding to the second current image frame collected by the monocular camera is determined through the first depth map, which can improve the accuracy of the depth information estimation corresponding to the second current image frame, so that the reconstructed driving environment can accurately correspond to the real driving environment, thereby improving the accuracy of environmental reconstruction.
[0139] The following describes the implementation process of the application embodiment in actual application scenarios.
[0140] like Figure 2 Figure 2 is a schematic diagram of the structure of a vehicle real scene reconstruction system provided by an embodiment of the present application. The vehicle real scene reconstruction system 200 includes a 3D reconstruction unit 201 and an environment reconstruction unit 202. The 3D reconstruction unit 201 uses the vehicle's camera data (forward view, surround view, perimeter view, etc.), millimeter wave radar data, lidar data, etc. to construct 3D scene point cloud data that matches the input data, and then provides the generated 3D scene point cloud data to the environment reconstruction unit 202 for environment reconstruction and rendering.
[0141] like Figure 3 FIG. 1 is a flow chart of a 3D reconstruction method based on single-frame forward-view data provided by the present application, the method comprising:
[0142] S301, distortion correction;
[0143] Here, the forward-looking data can be subjected to distortion correction.
[0144] In some embodiments, the forward-looking data may be acquired by a wide-angle camera and a telephoto camera disposed at the front of the vehicle. The output images of the forward-looking wide-angle camera and the telephoto camera may be subjected to distortion correction according to calibration parameters of each camera, and useless background may be removed.
[0145] S302, clearing the background;
[0146] Here, the background in the distortion-corrected front view data can be cleared.
[0147] In some embodiments, target recognition can be performed on the distortion-corrected forward-view data to identify various target objects near the vehicle and remove useless target objects (such as the sky). The distance of each target object around the vehicle from the vehicle can also be determined, and target objects with a distance greater than a distance threshold can be removed.
[0148] S303, image matching;
[0149] Here, the wide-angle image and the telephoto image after distortion correction and background removal can be matched to obtain the common area of the two images.
[0150] The distortion-corrected wide-angle and telephoto images can be quickly matched based on the extrinsic and intrinsic parameters of each camera. The common area between the two images is converted to positional coordinates in the wide-angle camera's coordinate system based on the wide-angle camera's intrinsic parameters and the pixel coordinates of the pixels in the wide-angle image. The wide-angle camera coordinates are then converted to world coordinates based on the wide-angle camera's extrinsic matrix. The processing for the telephoto image is similar, and the wide-angle image's positional coordinates in the world coordinate system are then matched with the telephoto image's positional coordinates in the world coordinate system to obtain the common area between the two images.
[0151] S304 : Generate a depth map based on the common area of the wide-angle image and the telephoto image.
[0152] The depth information of each pixel can be calculated based on the disparity information of the pixel to obtain a depth map. Figure 4 As shown, the projection points of point P on the telephoto camera and the wide-angle camera are p and p' respectively, and the pixel coordinates of point P are (p u , p v ), the pixel coordinates of p' are (p' u , p' v ), the disparity information of the projected points p and p' can be calculated (|p u -p' u |,|p v -p' v |), so that the depth information corresponding to point P can be calculated based on the disparity information and the distance between the two camera centers.
[0153] S305: Generate point cloud data based on the depth map.
[0154] Point cloud data based on the front view data may be generated according to the depth information of each pixel in the depth map and the color information of each pixel in the common area of the two images determined in step S303 .
[0155] like Figure 5 FIG. 1 is a flow chart of a 3D reconstruction method based on multi-frame forward-view data provided by the present application, the method comprising:
[0156] S501 , performing distortion correction and background removal on the output images of the front wide-angle camera and the telephoto camera of the current frame according to the calibration parameters of each camera.
[0157] In step S502 , the distortion-corrected telephoto image and wide-angle image of the current frame and the previous n frames are matched, depth maps are generated, and point cloud data is generated. Furthermore, the displacement information of each frame of the ego vehicle is combined to generate point cloud data based on the telephoto image and point cloud data based on the wide-angle image.
[0158] Here, n can be determined by the vehicle's current speed. For example, if the vehicle's current speed is greater than a threshold, n can be set to a. If the vehicle's current speed is less than or equal to the threshold, n can be set to b, where a is less than b. n can also be freely set by the user or developer.
[0159] In some embodiments, the forward-looking data (telephoto image and wide-angle image) corresponding to each moment can be processed according to the method of the aforementioned steps S403 to S405 to obtain point cloud data of multiple frames of telephoto images and point cloud data of multiple frames of wide-angle images. Specifically, for the point cloud data of the telephoto image and the point cloud data of the wide-angle image collected at the same moment, the depth information of the same point cloud is the same, but the color information may be different.
[0160] S503 , fusing the point cloud data based on the telephoto image and the point cloud data based on the wide-angle image to obtain fused point cloud data corresponding to the current frame.
[0161] In some embodiments, the point cloud data based on the telephoto image and the point cloud data based on the wide-angle image can be converted into the world coordinate system, and then the point cloud data based on the telephoto image and the point cloud data based on the wide-angle image can be fused according to the position coordinates of the point cloud data in the world coordinate system. The fusion process can include the fusion of the position coordinates of the point cloud data and the fusion of the color information of the point cloud data. The specific fusion method is not limited in this application, for example, it can be weighted fusion, etc.
[0162] S504 , based on the displacement information of each frame of the ego vehicle, the point cloud data of the previous n frames and the fused point cloud data corresponding to the current frame are fused to obtain point cloud data based on the forward-looking data.
[0163] The point cloud data of each forward-view data can be obtained by fusing the point cloud data before it and the point cloud data of the forward-view data before it. If n is 3, the point cloud data of the second forward-view data can be obtained by fusing the point cloud data of the first and second forward-view data; and the point cloud data of the third forward-view data can be obtained by fusing the point cloud data of the first and second forward-view data. That is, the sequential fusion method is adopted, and the sequential fusion method can make the fused point cloud data corresponding to the current frame contain more detailed information such as color and texture.
[0164] like Figure 6 FIG. 1 is a flow chart of a 3D reconstruction method based on single-frame circumferential view data and surround view data provided by the present application, the method comprising:
[0165] S601, performing distortion correction on the surround view image and the panoramic view image according to the calibration parameters of each camera.
[0166] Here, the surround view image may be an image captured by the vehicle's surround-view camera, and the circumferential view image may be an image captured by the surround-view camera. The surround view camera and the circumferential view camera may be installed on the left, right, and rear sides of the vehicle body. The surround view camera and the circumferential view camera may be installed on the left, right, and rear sides of the vehicle body. The surround view camera and the circumferential view camera installed on the same side of the vehicle body may have different installation positions and installation angles. Distortion correction may be performed on the surround view image based on the calibration parameters of the surround view camera, and distortion correction may be performed on the circumferential view image based on the calibration parameters of the surround view camera.
[0167] S602 , stitching the distortion-corrected surround view image and the panoramic view image into a panoramic image using the intrinsic parameters of each camera.
[0168] In some embodiments, the surround view image and the periscopic view image can be stitched together based on the internal parameters of the surround view camera and the internal parameters of the periscopic camera to obtain a panoramic image. The panoramic image can be an image obtained by stitching together the left and right sides of the vehicle and the rear side of the vehicle.
[0169] An image stitching algorithm can be used to stitch the surround view image and the panoramic view image, and convert the two images into grayscale images; a feature detection algorithm (such as SIFT, SURF, etc.) is used to detect feature points in the two images; the similarity of each feature point is calculated based on the position information of the feature point (such as the coordinates in the camera coordinate system, the pixel coordinates of the pixel points in the image can be converted into camera coordinates through the camera's internal parameters) (such as through Euclidean distance, cosine similarity, etc.); the feature points are matched based on the similarity, such as feature points with similarity greater than the similarity threshold are mutually matched feature points; the perspective transformation matrix of the two images is estimated based on the position information of the mutually matched feature points; the estimated perspective transformation matrix is used to perform a perspective transformation on one of the images and align it with the other image; a new blank canvas is created with a canvas size suitable for accommodating the stitching result of the two images; a perspective transformation is performed on one of the images and mapped to the new canvas; the other image is directly copied to the corresponding position on the new canvas; the overlapping area of the two images will be fused or superimposed to form the final panoramic image.
[0170] S603, clearing the background of the stitched panoramic image.
[0171] In some embodiments, target objects that are far away from the vehicle in the panoramic image may be cleared, and the target objects to be cleared may be determined by target recognition or the distance between the vehicle and surrounding target objects.
[0172] S604 , using the depth map generated by the front view data to perform self-supervision on the MLP neural network model used for depth estimation of the panoramic image, and outputting a depth map corresponding to the panoramic image.
[0173] Here, the MLP neural network model can be a pre-trained neural network model or an untrained neural network model.
[0174] In some embodiments, the common area of the front view image and the panoramic image can be determined first, and the depth information corresponding to the common area in the panoramic image can be estimated using an MLP neural network algorithm. The estimated depth information is compared with the depth information corresponding to the common area in the depth map corresponding to the front view image, and the MLP neural network algorithm is self-supervised trained based on the difference between the two depth information until the difference between the two depth information corresponding to the common area is less than a preset threshold, so that a trained MLP neural network model can be obtained; the trained MLP neural network model is then used to estimate the depth information of the panoramic image, thereby obtaining a depth map corresponding to the panoramic image.
[0175] S605 , obtaining point cloud data corresponding to the panoramic image based on the depth map corresponding to the panoramic image and the color information in the panoramic image after clearing the background.
[0176] In some embodiments, the depth information corresponding to each pixel in the panoramic image can be converted into position coordinates in the world coordinate system, and then the position coordinates and color information of each pixel in the panoramic image are combined to obtain point cloud data corresponding to the panoramic image.
[0177] like Figure 7 FIG. 1 is a flow chart of a 3D reconstruction method based on multiple frames of circumferential view data and surround view data provided by the present application, the method comprising:
[0178] S701, distortion correction;
[0179] The surround view image and the panoramic view image of the current frame are subjected to distortion correction according to the calibration parameters of each camera.
[0180] In some embodiments, distortion correction can be performed on the surround view image of the current frame, the panoramic view image of the current frame, the surround view image of the previous n frames, and the panoramic view image of the previous n frames in sequence according to the calibration parameters of each camera, thereby eliminating distortion in the multi-frame surround view image and the panoramic view image and improving image quality.
[0181] S702, image stitching;
[0182] The surround view image and the surrounding view image of the current frame after distortion correction are stitched into a panoramic image using the internal parameters of each camera.
[0183] In some embodiments, the surround view images and the panoramic view images collected at the same time (with the same timestamp) can be spliced, that is, the surround view image of the current frame (the n+1th frame) and the panoramic view image of the current frame are spliced to obtain the current frame panoramic image; the surround view image of the nth frame and the panoramic view image of the nth frame are spliced to obtain the nth frame panoramic image; the surround view image of the n-1th frame and the panoramic view image of the n-1th frame are spliced to obtain the n-1th frame panoramic image, ..., the surround view image of the 1st frame and the panoramic view image of the 1st frame are spliced to obtain the 1st frame panoramic image, so that multiple frames of panoramic images can be obtained.
[0184] S703, clear background;
[0185] Clear the background of the stitched panorama of the current frame and the previous n frames.
[0186] In some implementations, useless background in the current panoramic image frame may be cleared, and useless background in the previous n panoramic images may also be cleared.
[0187] S704, image matching, depth map generation, and point cloud data generation;
[0188] The panoramic images after clearing the background of the current frame and the previous n frames are sequentially matched, depth maps are generated, and point cloud data are generated. The point cloud data corresponding to each frame of the panoramic image is generated in combination with the displacement information of each frame of the vehicle.
[0189] In some implementations, image matching can be performed on each frame of the panoramic image, and the resulting images can be fused to obtain a fused panoramic image. The fused panoramic image is then processed according to the methods described in steps S605 and S606 to obtain point cloud data corresponding to the fused panoramic image. The point cloud data corresponding to the fused panoramic image is used as the point cloud data corresponding to the current panoramic image. Subsequently, the point cloud data of the previous n frames is fused with the point cloud data corresponding to the fused panoramic image to obtain point cloud data based on the surround view data and the forward view data.
[0190] S705, fusion of point cloud data;
[0191] The point cloud data corresponding to the previous n frames of panoramic images are fused with the point cloud data corresponding to the current frame of panoramic image to obtain point cloud data based on surround view data and forward view data.
[0192] In some embodiments, the displacement information of each frame of the vehicle can be combined to fuse the position information of the point cloud data corresponding to the panoramic image of the previous n frames and the position information of the point cloud data corresponding to the panoramic image of the current frame, and the color information of the point cloud data corresponding to the panoramic image of the previous n frames and the color information of the point cloud data corresponding to the panoramic image of the current frame can be fused to obtain point cloud data based on the surround view data and the forward view data.
[0193] In some embodiments, after obtaining point cloud data based on forward-view data and point cloud data based on surround-view data and forward-view data, the point cloud data based on forward-view data and the point cloud data based on surround-view data and forward-view data can be spliced or combined to obtain point cloud data based on image data, such as Figure 8 As shown, the point cloud data 801 based on the image data and the point cloud data 802 output by the vehicle-mounted radar at the same time can be fused to obtain the target point cloud data 803.
[0194] After obtaining the target point cloud data, the target point cloud data can be rendered, such as Figure 9 FIG. 1 is a flow chart of a method for rendering three-dimensional reconstructed point cloud data provided by the present application, the method comprising:
[0195] S901, Morton sort;
[0196] Perform Morton encoding on the target point cloud data and sort the obtained Morton codes.
[0197] By performing Morton coding on each point cloud in the target point cloud data and sorting the obtained Morton codes according to the Morton coding sorting method, adjacent points in the three-dimensional space can be arranged in similar positions.
[0198] S902, shuffling algorithm reordering;
[0199] The sorted Morton codes are grouped and re-sorted in groups.
[0200] The Morton coding sorting result in step S901 can be split into groups of m data (m can be a group size selected according to the target graphics card, for example, m can be 128), and then a shuffling algorithm is performed on the groups to shuffle the data in groups and re-sort the groups.
[0201] S903, upload to GPU and output point cloud rendering image;
[0202] After reordering, the Morton codes in groups are uploaded to the GPU to obtain the rendered image corresponding to the target point cloud data.
[0203] After uploading the reordered Morton codes in groups to the GPU, the GPU can use compute shader custom sampling and depth clipping to render the target point cloud data into an image, which is then output to the in-vehicle display screen for the user's reference in driving operations.
[0204] The 3D reconstruction method provided by this application can solve the problem that the current environmental reconstruction function cannot accurately correspond to the real environment. The environmental reconstruction based on the 3D reconstruction of the real-world scenic spot cloud uses the camera image to construct an environmental reconstruction image that is closer to the real environment. The corresponding camera image area can be quickly selected and located from the environmental reconstruction 3D scene. This application uses data from various on-board sensors to render the environmental reconstruction in the manner of 3D reconstruction of the real-world scenic spot cloud, making the environmental reconstruction of functions such as low-speed driving, campsite protection, and recorder playback more accurate and convenient, and corresponding to the real scene and camera image.
[0205] The present application provides a three-dimensional reconstruction device. Figure 10 As shown, the 3D reconstruction device 1000 includes:
[0206] A first acquisition module 1001 is configured to acquire a first current image frame captured by a binocular camera and a second current image frame captured by a monocular camera of a target vehicle; wherein the first current image frame and the second current image frame have an overlapping area;
[0207] A first determining module 1002 is configured to determine a second depth map corresponding to the second current image frame based on a first depth map corresponding to the first current image frame;
[0208] The second determining module 1003 determines first environmental point cloud data of the target vehicle based on the first depth map and the second depth map, so as to obtain a three-dimensional reconstruction result of the surrounding environment of the target vehicle based on the first environmental point cloud data.
[0209] In some embodiments, the first current image frame includes two first current images; and the first determining module 1002 includes:
[0210] A first determining submodule, configured to determine the first depth map based on disparity information of corresponding pixels in two first current images;
[0211] A first correction submodule, configured to correct parameters of a neural network model for estimating image depth information based on the first depth map to obtain a target neural network model;
[0212] The first depth estimation submodule is used to use the target neural network model to perform depth estimation on the second current image frame to obtain a second depth map.
[0213] In some embodiments, the first correction submodule includes:
[0214] a first depth estimation unit, configured to estimate depth information of the first current image frame using the neural network model to obtain a reference depth map;
[0215] A first correction unit is used to correct the parameters of the neural network model at least once based on the reference depth map and the first depth map until a convergence condition is met to obtain the target neural network model.
[0216] In some embodiments, the first correction submodule includes:
[0217] a second depth estimation unit, configured to, if there is a first region with the same image content in the first current image frame and the second current image frame, estimate depth information of a sub-image corresponding to the first region in the second current image frame using the neural network model to obtain a depth map corresponding to the sub-image;
[0218] A second correction unit is used to correct the parameters of the neural network model at least once based on the depth map of the sub-image and the reference sub-depth map until the convergence conditions are met to obtain the target neural network model; the reference sub-depth is a depth map corresponding to the first area in the first current image frame intercepted from the first depth map.
[0219] In some embodiments, the first determining submodule includes:
[0220] a preprocessing unit, configured to preprocess the two first current images respectively to obtain two candidate images;
[0221] A first determining unit is configured to determine, based on the parameters of the binocular camera, a second region having the same image content in the two candidate images;
[0222] The second determining unit is configured to determine the first depth map based on disparity information of pixels corresponding to the second area in the two candidate images.
[0223] In some embodiments, the number of the monocular cameras includes multiple, and different monocular images are collected by different monocular cameras; each monocular camera is set at a different position of the target vehicle; the first depth estimation submodule includes:
[0224] a stitching unit, configured to stitch the second current image frames respectively captured by the plurality of monocular cameras to obtain a stitched second current image frame;
[0225] The second depth estimation unit is used to use the target neural network model to perform depth estimation on the spliced second current image frame to obtain a second depth map.
[0226] In some embodiments, the second determining module 1003 includes:
[0227] A first acquisition submodule is used to acquire the current speed of the target vehicle;
[0228] a second acquisition submodule, configured to acquire, based on the current vehicle speed, a first historical image frame captured by the binocular camera and a second historical image frame captured by the monocular camera;
[0229] The second determination submodule is used to determine the first environmental point cloud data of the target vehicle based on the environmental point cloud data corresponding to the first historical image frame, the environmental point cloud data corresponding to the second historical image frame, the first depth map and the second depth map.
[0230] In some embodiments, the second determining submodule includes:
[0231] a third determining unit, configured to determine, based on the first depth map and the first current image frame, environmental point cloud data corresponding to the first current image frame; and to determine, based on the second depth map and the second current image frame, environmental point cloud data corresponding to the second current image frame;
[0232] A first fusion unit is configured to fuse the environmental point cloud data corresponding to the first historical image frame with the environmental point cloud data corresponding to the first current image frame to obtain first candidate environmental point cloud data; and to fuse the environmental point cloud data corresponding to the second historical image frame with the environmental point cloud data corresponding to the second current image frame to obtain second candidate environmental point cloud data;
[0233] The second fusion unit is configured to determine the first environment point cloud data based on the first candidate environment point cloud data and the second candidate environment point cloud data.
[0234] In some embodiments, the 3D reconstruction apparatus 1000 further includes:
[0235] A second acquisition module is used to acquire second environment point cloud data determined based on the radar sensor of the target vehicle;
[0236] The first fusion module is used to fuse the first environment point cloud data and the second environment point cloud data to obtain a three-dimensional reconstruction result of the environment around the target vehicle.
[0237] An embodiment of the present application provides a vehicle, comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the program, some or all of the steps in the above method are implemented.
[0238] The present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements some or all of the steps in the above method. The computer-readable storage medium may be transient or non-transient.
[0239] An embodiment of the present application provides a computer program product, including a computer program or instructions, which, when executed by a processor, implements the steps in the method described in the above embodiment.
[0240] An embodiment of the present application provides a computer program product, comprising a non-transitory computer-readable storage medium storing a computer program. When the computer program is read and executed by a computer, the computer program implements some or all of the steps of the above-described method. The computer program product can be implemented in hardware, software, or a combination thereof. In some embodiments, the computer program product is embodied as a computer storage medium. In other embodiments, the computer program product is embodied as a software product, such as a software development kit (SDK).
[0241] It should be noted that the above descriptions of the various embodiments tend to emphasize the differences between the various embodiments, and reference can be made to the similarities or similarities between them. The descriptions of the above device, vehicle, storage medium, and program product embodiments are similar to the descriptions of the above method embodiments and have similar beneficial effects as the method embodiments. For technical details not disclosed in the device, equipment, vehicle, storage medium, and program product embodiments of this application, please refer to the description of the method embodiments of this application for understanding.
[0242] It should be understood that "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. The above-mentioned serial numbers of the embodiments of the present application are for description only and do not represent the advantages and disadvantages of the embodiments.
[0243] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0244] The above description is only to fully illustrate the implementation mode of the present application, but the scope of protection of the present application is not limited thereto. Any changes or replacements that can be easily thought of by any technician familiar with this technical field within the technical scope disclosed in this application should be covered by the scope of protection of the present application.
Claims
1. A three-dimensional reconstruction method, characterized in that: include: Obtaining a first current image frame captured by a binocular camera and a second current image frame captured by a monocular camera of a target vehicle; There is an overlapping area between the first current image frame and the second current image frame; determining, based on a first depth map corresponding to the first current image frame, a second depth map corresponding to the second current image frame; Determining first environmental point cloud data of the target vehicle based on the first depth map and the second depth map, so as to obtain a three-dimensional reconstruction result of the surrounding environment of the target vehicle based on the first environmental point cloud data; The first current image frame includes two first current images; and determining the second depth map corresponding to the second current image frame based on the first depth map corresponding to the first current image frame includes: determining the first depth map based on disparity information of corresponding pixels in two first current images; Based on the first depth map, modifying parameters of a neural network model for estimating image depth information to obtain a target neural network model; Performing depth estimation on the second current image frame using the target neural network model to obtain a second depth map; The modifying, based on the first depth map, parameters of the neural network model for estimating image depth information to obtain a target neural network model includes: If there is a first region with the same image content in the first current image frame and the second current image frame, using the neural network model to estimate depth information of a sub-image corresponding to the first region in the second current image frame to obtain a depth map corresponding to the sub-image; Based on the depth map of the sub-image and the reference sub-depth map, the parameters of the neural network model are corrected at least once until the convergence conditions are met to obtain the target neural network model; the reference sub-depth is a depth map corresponding to the first area in the first current image frame intercepted from the first depth map.
2. The three-dimensional reconstruction method according to claim 1, characterized in that: The modifying, based on the first depth map, parameters of the neural network model for estimating image depth information to obtain a target neural network model includes: estimating depth information of the first current image frame using the neural network model to obtain a reference depth map; Based on the reference depth map and the first depth map, the parameters of the neural network model are corrected at least once until a convergence condition is met, thereby obtaining the target neural network model.
3. The three-dimensional reconstruction method according to claim 1, wherein: The determining the first depth map based on disparity information of corresponding pixels in the two first current images includes: Preprocessing the two first current images respectively to obtain two candidate images; Determining, based on the parameters of the binocular camera, a second region having the same image content in the two candidate images; The first depth map is determined based on disparity information of pixels corresponding to the second area in the two candidate images.
4. The three-dimensional reconstruction method according to claim 1, characterized in that: There are multiple monocular cameras, each of which is arranged at a different position of the target vehicle; The using the target neural network model to perform depth estimation on the second current image frame to obtain a second depth map includes: performing stitching processing on the second current image frames respectively captured by the multiple monocular cameras to obtain a stitched second current image frame; The target neural network model is used to perform depth estimation on the spliced second current image frame to obtain a second depth map.
5. The three-dimensional reconstruction method according to claim 1, characterized in that: The determining, based on the first depth map and the second depth map, first environment point cloud data of the target vehicle includes: Obtaining the current speed of the target vehicle; Based on the current vehicle speed, obtaining a first historical image frame captured by the binocular camera and a second historical image frame captured by the monocular camera; First environmental point cloud data of the target vehicle is determined based on the environmental point cloud data corresponding to the first historical image frame, the environmental point cloud data corresponding to the second historical image frame, the first depth map, and the second depth map.
6. The three-dimensional reconstruction method according to claim 5, characterized in that: The determining, based on the environmental point cloud data corresponding to the first historical image frame, the environmental point cloud data corresponding to the second historical image frame, the first depth map, and the second depth map, first environmental point cloud data of the target vehicle includes: Determining, based on the first depth map and the first current image frame, environmental point cloud data corresponding to the first current image frame; and determining, based on the second depth map and the second current image frame, environmental point cloud data corresponding to the second current image frame; Fusing the environment point cloud data corresponding to the first historical image frame and the environment point cloud data corresponding to the first current image frame to obtain first candidate environment point cloud data; and fusing the environment point cloud data corresponding to the second historical image frame and the environment point cloud data corresponding to the second current image frame to obtain second candidate environment point cloud data; The first environment point cloud data is determined based on the first candidate environment point cloud data and the second candidate environment point cloud data.
7. The three-dimensional reconstruction method according to any one of claims 1 to 6, characterized in that: The method further comprises: Acquire second environment point cloud data determined based on a radar sensor of the target vehicle; The first environmental point cloud data and the second environmental point cloud data are fused to obtain a three-dimensional reconstruction result of the environment surrounding the target vehicle.
8. A three-dimensional reconstruction device, characterized in that: include: A first acquisition module is used to acquire a first current image frame captured by the binocular camera of the target vehicle and a second current image frame captured by the monocular camera; There is an overlapping area between the first current image frame and the second current image frame; a first determining module, configured to determine a second depth map corresponding to the second current image frame based on a first depth map corresponding to the first current image frame; a second determining module, configured to determine first environmental point cloud data of the target vehicle based on the first depth map and the second depth map, so as to obtain a three-dimensional reconstruction result of the surrounding environment of the target vehicle based on the first environmental point cloud data; Wherein, the first current image frame includes two first current images; The first determination module is further configured to: determine the first depth map based on disparity information of corresponding pixels in the two first current images; modify parameters of a neural network model for estimating image depth information based on the first depth map to obtain a target neural network model; and perform depth estimation on the second current image frame using the target neural network model to obtain a second depth map; The first determination module is also used to: if there is a first area with the same image content in the first current image frame and the second current image frame, use the neural network model to estimate the depth information of the sub-image corresponding to the first area in the second current image frame to obtain a depth map corresponding to the sub-image; based on the depth map of the sub-image and the reference sub-depth map, correct the parameters of the neural network model at least once until the convergence conditions are met to obtain the target neural network model; the reference sub-depth is a depth map corresponding to the first area in the first current image frame intercepted from the first depth map.
9. A vehicle comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, characterized in that: When the processor executes the program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
11. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Depth information determination method and device based on panoramic look-around system
CN112380963A
Binocular vision three-dimensional reconstruction method based on high-precision positioning and deep learning
CN114255279A
Method for estimating monocular depth, apparatus and device therefor, and storage medium
WO2019223382A1