Methods and apparatus for constructing three-dimensional models

By registering feature points between images and 3D point cloud maps, the problems of illumination variations and scale information loss in 3D reconstruction algorithms are solved, and high-precision 3D model construction is achieved.

CN117808979BActive Publication Date: 2025-10-28AGIBOT INNOVATION (SHANGHAI) TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311184152.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-13
Publication Date
2025-10-28
Estimated Expiration
2043-09-13

AI Technical Summary

Technical Problem

Existing 3D reconstruction algorithms are susceptible to changes in lighting, lack scale information, suffer from severe positioning drift, and have limitations in methods that simultaneously acquire images and LiDAR data.

Method used

By acquiring images and 3D point cloud maps, and utilizing the depth information of common feature points and camera pose information, a 3D model is constructed. The model is then supplemented and optimized through the registration of image data and point cloud data.

Benefits of technology

It enables accurate construction of 3D models without the need for simultaneous acquisition of images and point cloud data, thus improving the accuracy and completeness of model construction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117808979B_ABST
    Figure CN117808979B_ABST
Patent Text Reader

Abstract

This disclosure relates to a method and apparatus for constructing a three-dimensional model. The method includes: acquiring a first image, at least one second image, and a three-dimensional point cloud map, wherein the first image and the second image have a first common feature point; determining the depth information of the first common feature point based on the first common feature point in the first image and the three-dimensional point cloud map; constructing a first three-dimensional model based on the depth information of the first common feature point, the position information of the first common feature point in the camera coordinate system of the first image, and the camera pose information of the first image; determining the camera pose information of the corresponding second image based on the depth information of the first common feature point and the position information of the first common feature point in the camera coordinate system of the corresponding second image; and obtaining a supplemented second three-dimensional model based on the depth information of the first common feature point, the position information of the first common feature point in the camera coordinate system of the corresponding second image, and the camera pose information of the corresponding second image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to image positioning technology, and more specifically, to a method and apparatus for constructing a three-dimensional model. Background Art

[0002] Image localization is a fundamental task in robotics and augmented reality, and is the cornerstone of realistic scene reconstruction. Accurate localization is a necessary condition to ensure the realism and reliability of the reconstructed 3D results.

[0003] Currently, 3D reconstruction algorithms generally use two-dimensional feature points to calculate the relative pose between images. This type of algorithm is easily affected by environmental changes such as lighting. Furthermore, monocular images lack scale information in localization and reconstruction, and as the number of images involved in the calculation increases, localization drift becomes a significant problem.

[0004] In addition, combining high-precision lidar ranging and two-dimensional images can achieve better mapping results, but this method requires the simultaneous acquisition of lidar data and image data, which limits its effectiveness.

[0005] Therefore, there is a need to provide a method for accurately constructing 3D models. Summary of the Invention

[0006] One objective of this invention is to provide a new technical solution for a method of constructing three-dimensional models.

[0007] According to a first aspect of the present invention, a method for constructing a three-dimensional model is provided, comprising:

[0008] Acquire a first image, at least one second image, and a 3D point cloud map; wherein the first image and the at least one second image have a first common feature point;

[0009] Based on the first common feature point in the first image and the three-dimensional point cloud map, determine the depth information of the first common feature point;

[0010] A first 3D model is constructed based on the depth information of the first common feature point, the position information of the first common feature point in the camera coordinate system of the first image, and the camera pose information of the first image.

[0011] Based on the depth information of the first common feature point and the position information of the first common feature point in the camera coordinate system of the corresponding second image, the camera pose information of the corresponding second image is determined.

[0012] Based on the depth information of the first common feature point, the position information of the first common feature point in the camera coordinate system of the corresponding second image, and the camera pose information of the corresponding second image, the first three-dimensional model is supplemented and constructed to obtain the second three-dimensional model.

[0013] Optionally, determining the depth information of each group of first common feature points in the first image based on each group of first common feature points in the first image and the three-dimensional point cloud map includes:

[0014] Obtain the camera pose information of the first image;

[0015] Based on the camera pose information of the first image, the three-dimensional point cloud map is projected onto the camera imaging plane corresponding to the first image to obtain the projection result;

[0016] If the projection result is that the projection point of a radar point in the three-dimensional point cloud map coincides with a first common feature point, the distance between the radar point and the first common feature point is determined as the depth information of the first common feature point.

[0017] Optionally, the method further includes:

[0018] When the projection result is that the projection points of multiple radar points in the three-dimensional point cloud map coincide with the same first common feature point, the distance between each radar point and the same first feature point is determined, and the minimum distance is used as the depth information of the same common feature point.

[0019] Optionally, the method further includes:

[0020] For every two second images, obtain the second common feature point;

[0021] The depth information of the second common feature point is determined based on the position information of the second common feature point in the camera coordinate system of the corresponding second image and the camera pose information of the corresponding second image.

[0022] Based on the depth information of the second common feature point, the position information of the second common feature point in the camera coordinates of the corresponding second image, and the camera pose information of the corresponding second image, the second three-dimensional model is supplemented and constructed to obtain the third three-dimensional model.

[0023] Optionally, the method further includes:

[0024] When the first three-dimensional model is supplemented and constructed based on each second image, the camera pose information of the image participating in the construction of the three-dimensional model, the coordinate information of the three-dimensional points of the supplemented and constructed three-dimensional model, the position information of the common feature points participating in the construction of the three-dimensional model in the camera coordinate system of the corresponding image, and the three-dimensional point cloud information of the radar points corresponding to the three-dimensional points are obtained.

[0025] The camera pose information of the images involved in model construction, the coordinate information of the 3D points of the supplemented 3D model, the position information of the common feature points involved in the 3D model construction in the camera coordinate system of the corresponding images, and the 3D point information of the radar points corresponding to the 3D points are input into the factor graph to obtain the optimized camera pose information and the optimized coordinate information of the 3D points of each image.

[0026] Based on the optimized camera pose information of each image and the optimized coordinate information of the 3D points, the optimized 3D model is obtained.

[0027] Optionally, before obtaining the 3D point cloud information of the radar points corresponding to the 3D points, the method further includes:

[0028] Obtain all second images corresponding to each 3D point;

[0029] In the case that all second images corresponding to the first three-dimensional point include the last image involved in the construction of the three-dimensional model, a target image is extracted from all second images corresponding to the first three-dimensional point; wherein the number of common feature points between the target image and the last image involved in the construction of the three-dimensional model exceeds a first preset threshold.

[0030] Based on the position and attitude information of the camera corresponding to the target image in the 3D point cloud map, the 3D point cloud map is projected onto the imaging plane of the camera corresponding to the target image to obtain the projection result;

[0031] If the projection result is that the projection point of a radar point in the three-dimensional point cloud map coincides with a common feature point, and the common feature point is a pixel point corresponding to the first three-dimensional point, then the three-dimensional point cloud information of the radar point is used as the three-dimensional point cloud information of the radar point corresponding to the first three-dimensional point.

[0032] Optionally, the method further includes:

[0033] When the projection result is that multiple radar points of the three-dimensional point cloud map coincide with the same common feature point, and the common feature point is a pixel point corresponding to the first three-dimensional point, multiple included angle values ​​are obtained based on the line connecting each radar point to the same common feature point and the normal vector of the local plane corresponding to each radar point.

[0034] Select the three-dimensional point cloud information of the radar point corresponding to the minimum included angle value, and use the three-dimensional point cloud information of the radar point corresponding to the minimum included angle value as the three-dimensional point cloud information of the radar point corresponding to the first three-dimensional point.

[0035] Optionally, the method further includes:

[0036] If the second image corresponding to the second three-dimensional point does not include the last image involved in the construction of the three-dimensional model, the radar point closest to the second three-dimensional point is determined, and the three-dimensional point cloud information of the radar point closest to the second three-dimensional point is used as the three-dimensional point cloud information of the radar point corresponding to the second three-dimensional point.

[0037] Optionally, the step of inputting the camera pose information of the images participating in model construction, the coordinate information of the 3D points of the supplemented 3D model, the position information of the common feature points participating in the 3D model construction in the camera coordinate system of the corresponding images, and the 3D point information of the radar points corresponding to the 3D points into the factor graph to obtain the optimized camera pose information and optimized coordinate information of the 3D points of each image includes:

[0038] The camera pose information of the images participating in the model construction is divided into a first group of camera pose information and a second group of camera pose information; wherein, the number of common feature points between the image corresponding to the first group of camera pose information and the last image participating in the 3D model construction exceeds a second preset threshold, and the number of common feature points between the image corresponding to the second group of camera pose information and the last image participating in the model construction does not exceed the second preset threshold.

[0039] The coordinate information of the three-dimensional points of the supplemented three-dimensional model is divided into the coordinate information of the first group of three-dimensional points and the coordinate information of the second group of three-dimensional points; wherein, the number of images corresponding to the first group of three-dimensional points exceeds the third preset threshold, and the number of images corresponding to the second group of three-dimensional points does not exceed the third preset threshold.

[0040] Determine the distance values ​​between each 3D point and its corresponding radar point;

[0041] The first set of camera pose information, the first set of 3D point coordinate information, and the position information of common feature points participating in the 3D model construction in the camera coordinate system of the corresponding image are used as fixed factors. The second set of camera pose information and the second set of 3D point coordinate information are used as variable factors. The distance values ​​between each 3D point and the corresponding radar point are used as the edges of the factors. These are input into the factor graph to obtain the optimized camera pose information and optimized 3D point coordinate information of each image.

[0042] Optionally, the method further includes:

[0043] When a preset number of second images complete the supplementary construction of the first three-dimensional model, the camera pose information of the images located within a specific spatial range, the coordinate information of the three-dimensional points determined based on the images located within the specific spatial range, and the three-dimensional point cloud information of the radar points corresponding to the three-dimensional points are obtained; wherein, the specific spatial range is obtained with the center of the camera corresponding to the last image participating in the construction of the three-dimensional model as the center of the sphere and a preset radius as the radius;

[0044] The camera pose information of the image within a specific spatial range, the coordinate information of the three-dimensional points determined based on the image within a specific spatial range, and the three-dimensional point cloud information of the radar points corresponding to the three-dimensional points are input into the factor map to obtain the optimized camera pose information of the image within a specific spatial range and the optimized coordinate information of the three-dimensional points within a specific spatial range.

[0045] Based on the camera pose information of the image within the optimized specific spatial range and the coordinate information of the three-dimensional points within the optimized specific spatial range, the optimized three-dimensional model is obtained.

[0046] According to a second aspect of the present invention, an apparatus for constructing a three-dimensional model is provided, comprising:

[0047] An acquisition module is used to acquire a first image, at least one second image, and a three-dimensional point cloud map; wherein the first image and the at least one second image have a first common feature point.

[0048] The depth information determination module is used to determine the depth information of the first common feature point based on the first common feature point in the first image and the three-dimensional point cloud map;

[0049] The first three-dimensional model construction module is used to construct a first three-dimensional model based on the depth information of the first common feature points, the position information of the first common feature points in the camera coordinate system of the first image, and the camera pose information of the first image.

[0050] The camera pose information determination module is used to determine the camera pose information of the corresponding second image based on the depth information of the first common feature point and the position information of the first common feature point in the camera coordinate system of the corresponding second image.

[0051] The second 3D model construction module is used to supplement and construct the first 3D model based on the depth information of the first common feature point, the position information of the first common feature point in the camera coordinate system of the corresponding second image, and the camera pose information of the corresponding second image, so as to obtain the second 3D model.

[0052] According to a third aspect of the present invention, a three-dimensional model construction apparatus is provided, comprising a memory and a processor, the memory storing a computer program for controlling the processor to operate in order to perform the method according to any one of the first aspects of the present invention.

[0053] The method for constructing a 3D model provided by this invention no longer requires the simultaneous acquisition of image data and 3D point cloud data. By registering the acquired image data and 3D point cloud data, the accurate construction of the 3D model can be achieved.

[0054] The features and advantages of the embodiments of this specification will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Attached Figure Description

[0055] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments of this specification and, together with their description, serve to explain the principles of these embodiments.

[0056] Figure 1 This is a flowchart illustrating a method for constructing a three-dimensional model according to an embodiment of the present invention.

[0057] Figure 2 This is a schematic diagram of a three-dimensional point cloud map projected onto the camera imaging plane corresponding to a first image according to an embodiment of the present invention.

[0058] Figure 3 This is a schematic diagram illustrating the determination of a specific spatial range according to an embodiment of the present invention.

[0059] Figure 4 This is a schematic diagram of a three-dimensional model construction device according to an embodiment of the present invention.

[0060] Figure 5 This is a schematic diagram of the hardware structure of a three-dimensional model construction device according to an embodiment of the present invention. Detailed Implementation

[0061] Various exemplary embodiments of this specification will now be described in detail with reference to the accompanying drawings.

[0062] The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit the embodiments of this specification or their application or use.

[0063] It should be noted that like reference numerals and letters refer to like items in the following figures, and therefore, once an item is defined in one figure, it need not be further discussed in subsequent figures.

[0064] <Method Example>

[0065] One embodiment of the present invention provides a method for constructing a three-dimensional model. According to... Figure 1 As shown, the method for constructing the three-dimensional model in this embodiment may include the following steps S110 to S150.

[0066] S110, acquire a first image, at least one second image, and a three-dimensional point cloud map; wherein the first image and at least one second image have a first common feature point.

[0067] A 3D point cloud map is a map generated based on 3D point cloud data obtained from panoramic scanning using a LiDAR sensor.

[0068] The first and second images are obtained by taking pictures of different scenes in a panorama or the same scene from different angles using a camera.

[0069] Common feature points are imaging points in different images that correspond to the same scene.

[0070] In the case where there is only one second image, the first image and the second image share a first set of common feature points.

[0071] When there are multiple second images, the first image and each second image share a set of first common feature points. Each set of first common feature points can contain multiple points. The first common feature points of any two sets can be different.

[0072] S120, determine the depth information of the first common feature point based on the first common feature point in the first image and the three-dimensional point cloud map.

[0073] In one specific embodiment, S120 includes S121 to S123.

[0074] S121, Obtain the camera pose information of the first image.

[0075] The camera pose information for the first image is known and can be obtained directly.

[0076] S122, Based on the camera pose information of the first image, project the 3D point cloud map onto the camera imaging plane corresponding to the first image to obtain the projection result.

[0077] Figure 2 This is a schematic diagram of a three-dimensional point cloud map projected onto the camera imaging plane corresponding to a first image according to an embodiment of the present invention.

[0078] Based on the camera pose information of the first image, the position and orientation of the camera corresponding to the first image in the 3D point cloud map can be determined.

[0079] After determining the position and orientation of the camera corresponding to the first image in the 3D point cloud map, refer to... Figure 2A quadrangular pyramid is formed with the camera's center O1 as its vertex and the camera's imaging plane S corresponding to the first image as its base. The depth of the quadrangular pyramid is a pre-calibrated parameter.

[0080] By projecting the radar points enclosed in the quadrangular pyramid onto the camera imaging plane corresponding to the first image along a direction perpendicular to the camera imaging plane of the first image, the projection results of each radar point are obtained.

[0081] S123, when the projection point of a radar point in the projection result is a three-dimensional point cloud map coincides with a first common feature point, determine the distance between the radar point and the first common feature point as the depth information of the first common feature point.

[0082] When the projected points of multiple radar points in the three-dimensional point cloud map coincide with the same first common feature point, the distance between each radar point and the same first common feature point is determined, and the minimum distance is taken as the depth information of the same common feature point.

[0083] It should be noted that not every common feature point can be used to determine a corresponding depth information.

[0084] S130, construct a first three-dimensional model based on the depth information of the first common feature point, the position information of the first common feature point in the camera coordinate system of the first image, and the camera pose information of the first image.

[0085] When there are multiple second images, a 3D model in the camera coordinate system of the first image can be constructed based on the depth information of a set of first common feature points in the first image, the position information of the set of first common feature points in the first image in the camera coordinate system, and the camera pose information of the first image. Then, the 3D model in the camera coordinate system of the first image is transformed into a 3D model in the world coordinate system to serve as the first 3D model.

[0086] Specifically, based on the depth information of each common feature point, the position information of each common feature point in the camera coordinate system of the first image, and the camera pose information of the first image, the three-dimensional position information of the first common feature point in the camera coordinate system of the first image can be obtained. Then, the three-dimensional position information of the first common feature point in the camera coordinate system of the first image is converted into the three-dimensional position information in the world coordinate system.

[0087] The first 3D model information includes the coordinate information of multiple 3D points and the camera pose information of each image.

[0088] The number of first common feature points participating in the construction of the first 3D model exceeds a preset threshold. Preferably, the group with the largest number of first common feature points is used in the construction of the first 3D model.

[0089] When there are multiple second images, a 3D model in the camera coordinate system of the first image can be constructed based on the depth information of each group of first common feature points in the first image, the position information of each group of first common feature points in the first image in the camera coordinate system, and the camera pose information of the first image. Then, the 3D model in the camera coordinate system of the first image is transformed into a 3D model in the world coordinate system to serve as the first 3D model.

[0090] S140, based on the depth information of the first common feature point and the position information of the first common feature point in the camera coordinate system of the corresponding second image, determine the camera pose information of the corresponding second image.

[0091] In the case of a single second image, the camera pose information of the second image is determined based on the Perspective-n-Point (PnP) algorithm, according to the depth information of the first common feature point and the position information of the first common feature point in the camera coordinate system of the corresponding second image.

[0092] When there are multiple second images, the PnP algorithm is used to determine the camera pose information of the corresponding second image based on the depth information of a set of first common feature points and the position information of the set of first common feature points in the camera coordinate system of the corresponding image for each second image.

[0093] S150, based on the depth information of the first common feature point, the position information of the first common feature point in the camera coordinate system of the corresponding second image, and the camera pose information of the corresponding second image, the first three-dimensional model is supplemented and constructed to obtain the second three-dimensional model.

[0094] Based on the depth information of the first common feature point, the position information of the first common feature point in the camera coordinate system of the corresponding second image, and the camera pose information of the corresponding second image, a 3D model in the camera coordinate system of each second image can be constructed. Then, the 3D model in the camera coordinate system of each second image is transformed into a 3D model in the world coordinate system to supplement the first 3D model and obtain the second 3D model.

[0095] Specifically, based on the depth information of each first common feature point, the position information of each first common feature point in the camera coordinate system of the second image, and the camera pose information of the second image, the three-dimensional position information of the first common feature point in the camera coordinate system of the second image can be obtained. Then, the three-dimensional position information of the first common feature point in the camera coordinate system of the second image is converted into the three-dimensional position information in the world coordinate system.

[0096] The second 3D model information includes the coordinate information of multiple 3D points and the camera pose information of each image. Compared with the first 3D model, the second 3D model includes more 3D points and more camera pose information for each image.

[0097] The method for constructing a 3D model provided in this embodiment of the invention no longer requires the simultaneous acquisition of image data and 3D point cloud data. The accurate construction of the 3D model is achieved through the registration of the acquired image data and 3D point cloud data.

[0098] When there are multiple second images, since different second images may also include imaging points of the same scene, these imaging points can participate in the construction of the 3D model to improve the completeness of the 3D model construction. In one embodiment of the present invention, the 3D model construction method further includes: for every two second images, obtaining second common feature points; determining the depth information of the second common feature points based on the position information of the second common feature points in the camera coordinate system of the corresponding second image and the camera pose information of the corresponding second image; and supplementing the second 3D model with the depth information of the second common feature points, the position information of the second common feature points in the camera coordinate system of the corresponding second image, and the camera pose information of the corresponding second image to obtain a third 3D model.

[0099] The number of second common feature points in each group can be multiple.

[0100] The depth information of the second common feature point can be determined using the Direct Linear Transform (DLT) triangulation method.

[0101] Based on each pair of second images, according to the depth information of the second common feature points, the position information of the second common feature points in the camera coordinate system of the corresponding second image, and the camera pose information of the corresponding second image, a 3D model in the camera coordinate system of each of the two second images can be constructed. Then, the 3D models in the camera coordinate system of each of the two second images are transformed into 3D models in the world coordinate system to supplement the second 3D model and obtain a third 3D model.

[0102] Specifically, based on the depth information of each second common feature point, the position information of each second common feature point in the camera coordinate system of the corresponding second image, and the camera pose information of the corresponding second image, the three-dimensional position information of the second common feature point in the camera coordinate system of the second image can be obtained. Then, the three-dimensional position information of the second common feature point in the camera coordinate system of the second image is converted into the three-dimensional position information in the world coordinate system.

[0103] In one embodiment of the present invention, after supplementing the first three-dimensional model based on each second image, the supplemented three-dimensional model is optimized to improve the construction accuracy of the three-dimensional model. This optimization operation specifically includes steps S211 to S213.

[0104] S211, acquire the camera pose information of the image participating in the construction of the 3D model, the coordinate information of the 3D points of the constructed 3D model, the position information of the common feature points participating in the construction of the 3D model in the camera coordinate system of the corresponding image, and the 3D point cloud information of the radar points corresponding to the 3D points.

[0105] The 3D point cloud information of the radar point corresponding to the 3D point can be determined according to the following steps.

[0106] To distinguish whether the second image corresponding to a 3D point includes the last image used in the 3D model construction, this embodiment uses "first" and "second" to differentiate them. Both the first and second 3D points represent a class of 3D points. The second image corresponding to the first 3D point includes the last image used in the 3D model construction. The second image corresponding to the second 3D point does not include the last image used in the 3D model construction. Both the first and second 3D points represent a class of 3D points.

[0107] S211a, obtain all the second images corresponding to each 3D point.

[0108] The second image corresponding to a 3D point refers to the 3D point being determined based on the second image according to the above-mentioned method for supplementing and constructing the first 3D model.

[0109] The second image corresponding to this 3D point can be one or more.

[0110] S211b, when all the second images corresponding to the first three-dimensional point include the last image involved in the construction of the three-dimensional model, extract the target image from all the second images corresponding to the first three-dimensional point; wherein the number of common feature points between the target image and the last image involved in the construction of the three-dimensional model exceeds a first preset threshold.

[0111] If there are multiple target images, one can be selected at will.

[0112] The last image involved in the construction of the 3D model in this step refers to the current image involved in the construction of the 3D model.

[0113] S211c: Based on the camera pose information of the target image, the 3D point cloud map is projected onto the camera imaging plane corresponding to the target image to obtain the projection result.

[0114] The projection method can be referred to above. Figure 2The projection method shown will not be elaborated further here.

[0115] S211d, where the projection point of a radar point in the projection result of a three-dimensional point cloud map coincides with a common feature point, and the common feature point is a pixel corresponding to the first three-dimensional point, the three-dimensional point cloud information of the radar point is used as the three-dimensional point cloud information of the radar point corresponding to the first three-dimensional point.

[0116] The pixel corresponding to the first 3D point refers to the pixel that determines the first 3D point. This 3D point is determined based on this pixel using the aforementioned method for supplementing and constructing the first 3D model.

[0117] S211e, when multiple radar points in the projection result of the three-dimensional point cloud map coincide with the same common feature point, and the common feature point is the pixel point corresponding to the first three-dimensional point, multiple included angle values ​​are obtained based on the line connecting each radar point to the same common feature point and the normal vector of the local surface corresponding to each radar point; the three-dimensional point cloud information of the radar point corresponding to the smallest included angle value is selected, and the three-dimensional point cloud information of the radar point corresponding to the smallest included angle value is used as the three-dimensional point cloud information of the radar point corresponding to the first three-dimensional point.

[0118] Based on a single radar point, multiple neighboring radar points are acquired. The distance between each neighboring radar point and the given radar point must not exceed a preset distance threshold. The number of neighboring radar points can be two or three, without specific limitation. Then, a local surface is fitted using the given radar point and its multiple neighboring radar points. When the curvature of the local surface is 0, the local surface is a plane.

[0119] An angle value is the angle between a line and the normal vector of the local surface corresponding to a radar point. A line is a line connecting a radar point to a point sharing the same common feature.

[0120] It should be noted that not every common feature point can be used to identify a corresponding radar point.

[0121] S211f, if the second image corresponding to the second three-dimensional point does not include the last image involved in the construction of the three-dimensional model, determine the radar point closest to the second three-dimensional point, and use the three-dimensional point cloud information of the radar point closest to the second three-dimensional point as the three-dimensional point cloud information of the radar point corresponding to the second three-dimensional point.

[0122] S212, the camera pose information of the images involved in model construction, the coordinate information of the 3D points of the supplemented 3D model, the position information of the common feature points involved in the 3D model construction in the camera coordinate system of the corresponding images, and the 3D point information of the radar points corresponding to the 3D points are input into the factor graph to obtain the optimized camera pose information of each image and the optimized coordinate information of the 3D points.

[0123] In one specific embodiment, S212 includes S212a to S212d.

[0124] S212a, the camera pose information of the images participating in model construction is divided into a first group of camera pose information and a second group of camera pose information; wherein, the number of common feature points between the image corresponding to the first group of camera pose information and the last image participating in 3D model construction exceeds a second preset threshold, and the number of common feature points between the image corresponding to the second group of camera pose information and the last image participating in model construction does not exceed the second preset threshold.

[0125] The last image involved in the construction of the 3D model refers to the current image involved in the construction of the 3D model.

[0126] S212b, the coordinate information of the three-dimensional points of the supplemented three-dimensional model is divided into the coordinate information of the first group of three-dimensional points and the coordinate information of the second group of three-dimensional points; wherein, the number of images corresponding to the first group of three-dimensional points exceeds the third preset threshold, and the number of images corresponding to the second group of three-dimensional points does not exceed the third preset threshold.

[0127] The image corresponding to a 3D point refers to the image upon which the 3D point is determined.

[0128] S212c determines the distance between each 3D point and its corresponding radar point.

[0129] Specifically, a distance value is determined based on the position coordinates of the three-dimensional point and the corresponding three-dimensional coordinates of the radar point.

[0130] S212d uses the first set of camera pose information, the first set of 3D point coordinate information, and the position information of common feature points participating in the 3D model construction in the corresponding image's camera coordinate system as fixed factors, the second set of camera pose information and the second set of 3D point coordinate information as variable factors, and the distance values ​​between each 3D point and the corresponding radar point as the edges of the factors. These are input into the factor graph to obtain the optimized camera pose information and optimized 3D point coordinate information for each image.

[0131] S213. Based on the optimized camera pose information of each image and the optimized coordinate information of the 3D points, the optimized 3D model is obtained.

[0132] In one embodiment of the present invention, after supplementing the first three-dimensional model based on a preset number of second images, the supplemented three-dimensional model is optimized to improve the construction accuracy of the three-dimensional model. This optimization operation specifically includes steps S214 to S216.

[0133] S214, when the first three-dimensional model is supplemented by a preset number of second images, the camera pose information of the images located within a specific spatial range, the coordinate information of the three-dimensional points determined based on the images located within the specific spatial range, and the three-dimensional point cloud information of the radar points corresponding to the three-dimensional points are obtained; wherein, the specific spatial range is obtained with the center of the camera corresponding to the last image participating in the construction of the three-dimensional model as the center of the sphere and a preset radius as the radius.

[0134] See Figure 3 Using the camera center O2 corresponding to the last image involved in the 3D model construction as the center of the sphere and a preset radius as the radius, a sphere is obtained.

[0135] S215, the camera pose information of the image within a specific spatial range, the coordinate information of the three-dimensional points determined based on the image within the specific spatial range, and the three-dimensional point cloud information of the radar points corresponding to the three-dimensional points are input into the factor map to obtain the optimized camera pose information of the image within the specific spatial range and the optimized coordinate information of the three-dimensional points within the specific spatial range.

[0136] Based on each image located within a specific spatial range, the corresponding three-dimensional points can be determined by following the above-described method for constructing the first three-dimensional model.

[0137] The 3D point cloud information of the radar point corresponding to the 3D point can be determined in the following way: determine the radar point that is closest to the 3D point, and use the 3D point cloud information of the radar point that is closest to the 3D point as the 3D point cloud information of the radar point corresponding to the 3D point.

[0138] S216. Based on the camera pose information of the image within the optimized specific spatial range and the coordinate information of the three-dimensional points within the optimized specific spatial range, the optimized three-dimensional model is obtained.

[0139] <Device Example>

[0140] One embodiment of the present invention provides a three-dimensional model construction apparatus, such as... Figure 4 As shown. The 3D model construction device 400 includes an acquisition module 410, a depth information determination module 420, a first 3D model construction module 430, a camera pose information determination module 440, and a second 3D model construction module 450.

[0141] The acquisition module 410 is used to acquire a first image, at least one second image, and a three-dimensional point cloud map; wherein the first image and at least one second image have a first common feature point.

[0142] The depth information determination module 420 is used to determine the depth information of the first common feature point based on the first common feature point in the first image and the three-dimensional point cloud map.

[0143] The first 3D model construction module 430 is used to construct a first 3D model based on the depth information of the first common feature points, the position information of the first common feature points in the camera coordinate system of the first image, and the camera pose information of the first image.

[0144] The camera pose information determination module 440 is used to determine the camera pose information of the corresponding second image based on the depth information of the first common feature point and the position information of the first common feature point in the camera coordinate system of the corresponding second image.

[0145] The second 3D model construction module 450 is used to supplement and construct the first 3D model based on the depth information of the first common feature point, the position information of the first common feature point in the camera coordinate system of the corresponding second image, and the camera pose information of the corresponding second image, so as to obtain the second 3D model.

[0146] In one embodiment of the present invention, the depth information determination module 420 is further configured to acquire camera pose information of the first image; project the three-dimensional point cloud map onto the camera imaging plane corresponding to the first image based on the camera pose information of the first image to obtain the projection result; and determine the distance between the radar point and the first common feature point when the projection result is that the projection point of a radar point in the three-dimensional point cloud map coincides with a first common feature point, so as to use the depth information of the first common feature point.

[0147] In one embodiment of the present invention, the depth information determination module 420 is further configured to determine the distance between each radar point and the same first feature point when the projection points of multiple radar points in the projection result of the three-dimensional point cloud map coincide with the same first common feature point, and to use the minimum distance as the depth information of the same common feature point.

[0148] In one embodiment of the present invention, the three-dimensional model construction device further includes a third three-dimensional model construction module.

[0149] The third 3D model construction module is used to obtain a set of second common feature points for every two second images; determine the depth information of each set of second common feature points based on the position information of each set of second common feature points in the camera coordinate system of the corresponding second image and the camera pose information of the corresponding second image; and supplement the second 3D model based on the depth information of each set of second common feature points, the position information of each set of second common feature points in the camera coordinate system of the corresponding second image and the camera pose information of the corresponding second image to obtain the third 3D model.

[0150] In one embodiment of the present invention, the three-dimensional model construction apparatus further includes a first three-dimensional model optimization module.

[0151] The first 3D model optimization module is used to acquire, when supplementing the first 3D model based on each second image, the camera pose information of the images participating in the 3D model construction, the coordinate information of the 3D points of the supplemented 3D model, the position information of the common feature points participating in the 3D model construction in the camera coordinate system of the corresponding images, and the 3D point cloud information of the radar points corresponding to the 3D points; inputting the camera pose information of the images participating in the model construction, the coordinate information of the 3D points of the supplemented 3D model, the position information of the common feature points participating in the 3D model construction in the camera coordinate system of the corresponding images, and the 3D point information of the radar points corresponding to the 3D points into the factor map to obtain the optimized camera pose information and the optimized coordinate information of the 3D points of each image; and obtaining the optimized 3D model based on the optimized camera pose information and the optimized coordinate information of the 3D points of each image.

[0152] In one embodiment of the present invention, the first three-dimensional model optimization module is further configured to acquire all second images corresponding to each three-dimensional point; when all second images corresponding to the first three-dimensional point include the last image involved in the construction of the three-dimensional model, extract a target image from all second images corresponding to the first three-dimensional point; wherein the number of common feature points between the target image and the last image involved in the construction of the three-dimensional model exceeds a first preset threshold; based on the position and attitude information of the camera corresponding to the target image in the three-dimensional point cloud map, project the three-dimensional point cloud map onto the camera imaging plane corresponding to the target image to obtain a projection result; when the projection point of a radar point in the three-dimensional point cloud map coincides with a common feature point, and the common feature point is a pixel point corresponding to the first three-dimensional point, use the three-dimensional point cloud information of the radar point as the three-dimensional point cloud information of the radar point corresponding to the first three-dimensional point.

[0153] In one embodiment of the present invention, the first three-dimensional model optimization module is further configured to, when multiple radar points in the projection result of a three-dimensional point cloud map coincide with the same common feature point, and the common feature point is a pixel point corresponding to the first three-dimensional point, obtain multiple included angle values ​​based on the line connecting each radar point to the same common feature point and the normal vector of the local plane corresponding to each radar point; select the three-dimensional point cloud information of the radar point corresponding to the smallest included angle value, and use the three-dimensional point cloud information of the radar point corresponding to the smallest included angle value as the three-dimensional point cloud information of the radar point corresponding to the first three-dimensional point.

[0154] In one embodiment of the present invention, the first three-dimensional model optimization module is further configured to determine the radar point closest to the second three-dimensional point when the second image corresponding to the second three-dimensional point does not include the last image involved in the construction of the three-dimensional model, and use the three-dimensional point cloud information of the radar point closest to the second three-dimensional point as the three-dimensional point cloud information of the radar point corresponding to the second three-dimensional point.

[0155] In one embodiment of the present invention, the first 3D model optimization module is further configured to divide the camera pose information of the images participating in model construction into a first group of camera pose information and a second group of camera pose information; wherein, the number of common feature points between the image corresponding to the first group of camera pose information and the last image participating in 3D model construction exceeds a second preset threshold, and the number of common feature points between the image corresponding to the second group of camera pose information and the last image participating in model construction does not exceed the second preset threshold; divide the coordinate information of the 3D points of the supplemented 3D model into a first group of 3D point coordinate information and a second group of 3D point coordinate information; wherein, the number of images corresponding to the first group of 3D points exceeds a third preset threshold, and the number of images corresponding to the second group of 3D points does not exceed the third preset threshold; determine the distance value between each 3D point and the corresponding radar point; and use the position information of the first group of camera pose information, the coordinate information of the first group of 3D points, and the common feature points participating in 3D model construction in the camera coordinate system of the corresponding image as fixed factors, the second group of camera pose information and the coordinate information of the second group of 3D points as variable factors, and the distance value between each 3D point and the corresponding radar point as the edge of the factor, inputting them into the factor graph to obtain the optimized camera pose information of each image and the optimized coordinate information of the 3D points.

[0156] In one embodiment of the present invention, the three-dimensional model construction apparatus further includes a second three-dimensional model optimization module.

[0157] The second 3D model optimization module is used to acquire camera pose information of images within a specific spatial range, coordinate information of 3D points determined based on images within the specific spatial range, and 3D point cloud information of radar points corresponding to 3D points within the specific spatial range, after supplementing the first 3D model with a preset number of second images. The specific spatial range is defined as a sphere centered on the camera center of the last image participating in the 3D model construction, with a preset radius. The module inputs the camera pose information of images within the specific spatial range, the coordinate information of 3D points determined based on images within the specific spatial range, and the 3D point cloud information of radar points corresponding to 3D points into a factor graph to obtain optimized camera pose information of images within the specific spatial range and optimized coordinate information of 3D points within the specific spatial range. Based on the optimized camera pose information of images within the specific spatial range and optimized coordinate information of 3D points within the specific spatial range, an optimized 3D model is obtained.

[0158] One embodiment of the present invention provides a three-dimensional model construction apparatus, such as... Figure 5As shown. The 3D model construction apparatus 500 includes a memory 520 and a processor 510. The memory 520 stores a computer program that controls the processor 510 to operate and execute the 3D model construction method in any of the above embodiments.

[0159] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. For the electric vehicle embodiments, relevant parts can be found in the descriptions of the method embodiments.

[0160] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0161] Embodiments of this specification may be systems, methods, and / or computer program products. A computer program product may include a computer-readable storage medium having computer instructions stored thereon for causing a processor to implement various aspects of the embodiments of this specification.

[0162] Computer-readable storage media can be tangible devices capable of holding and storing computer instructions for use by computer instruction execution devices. Computer-readable storage media can be, for example—but not limited to—electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing computer instructions thereon, and any suitable combination thereof. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.

[0163] The computer instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper cables, fiber optic cables, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer instructions from the network and forwards them to computer-readable storage media within the respective computing / processing device.

[0164] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this specification. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of computer instructions, which contains one or more executable computer instructions for implementing a specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions. It will be known to those skilled in the art that implementation in hardware, implementation in software, and implementation using a combination of software and hardware are equivalent.

[0165] Various embodiments of this specification have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or improvement of the technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A method for constructing a three-dimensional model, characterized in that, include: Acquire a first image, at least one second image, and a 3D point cloud map; wherein the first image and the at least one second image both have a first common feature point; Determining the depth information of the first common feature point based on the first common feature point in the first image and the three-dimensional point cloud map includes: acquiring camera pose information of the first image; projecting the three-dimensional point cloud map onto the camera imaging plane corresponding to the first image based on the camera pose information of the first image to obtain a projection result; and determining the distance between the radar point and the first common feature point when the projection result is that the projection point of a radar point in the three-dimensional point cloud map coincides with a first common feature point, as the depth information of the first common feature point. A first 3D model is constructed based on the depth information of the first common feature point, the position information of the first common feature point in the camera coordinate system of the first image, and the camera pose information of the first image. Based on the depth information of the first common feature point and the position information of the first common feature point in the camera coordinate system of the corresponding second image, the camera pose information of the corresponding second image is determined. Based on the depth information of the first common feature point, the position information of the first common feature point in the camera coordinate system of the corresponding second image, and the camera pose information of the corresponding second image, the first three-dimensional model is supplemented and constructed to obtain the second three-dimensional model.

2. The method according to claim 1, characterized in that, The method further includes: When the projection result is that the projection points of multiple radar points in the three-dimensional point cloud map coincide with the same first common feature point, the distance between each radar point and the same first common feature point is determined, and the minimum distance is used as the depth information of the same first common feature point.

3. The method according to claim 1, characterized in that, The method further includes: For every two second images, obtain the second common feature point; The depth information of the second common feature point is determined based on the position information of the second common feature point in the camera coordinate system of the corresponding second image and the camera pose information of the corresponding second image. Based on the depth information of the second common feature point, the position information of the second common feature point in the camera coordinates of the corresponding second image, and the camera pose information of the corresponding second image, the second three-dimensional model is supplemented and constructed to obtain the third three-dimensional model.

4. The method according to claim 1, characterized in that, The method further includes: When the first three-dimensional model is supplemented and constructed based on each second image, the camera pose information of the image participating in the construction of the three-dimensional model, the coordinate information of the three-dimensional points of the supplemented and constructed three-dimensional model, the position information of the common feature points participating in the construction of the three-dimensional model in the camera coordinate system of the corresponding image, and the three-dimensional point cloud information of the radar points corresponding to the three-dimensional points are obtained. The camera pose information of the images involved in model construction, the coordinate information of the 3D points of the supplemented 3D model, the position information of the common feature points involved in the 3D model construction in the camera coordinate system of the corresponding images, and the 3D point information of the radar points corresponding to the 3D points are input into the factor graph to obtain the optimized camera pose information and the optimized coordinate information of the 3D points of each image. Based on the optimized camera pose information of each image and the optimized coordinate information of the 3D points, the optimized 3D model is obtained.

5. The method according to claim 4, characterized in that, Before acquiring the 3D point cloud information of the radar points corresponding to the 3D points, the method further includes: Obtain all second images corresponding to each 3D point; In the case that all second images corresponding to the first three-dimensional point include the last image involved in the construction of the three-dimensional model, a target image is extracted from all second images corresponding to the first three-dimensional point; wherein the number of common feature points between the target image and the last image involved in the construction of the three-dimensional model exceeds a first preset threshold. Based on the position and attitude information of the camera corresponding to the target image in the 3D point cloud map, the 3D point cloud map is projected onto the imaging plane of the camera corresponding to the target image to obtain the projection result; If the projection result is that the projection point of a radar point in the three-dimensional point cloud map coincides with a common feature point, and the common feature point is a pixel point corresponding to the first three-dimensional point, then the three-dimensional point cloud information of the radar point is used as the three-dimensional point cloud information of the radar point corresponding to the first three-dimensional point.

6. The method according to claim 5, characterized in that, The method further includes: When the projection result is that multiple radar points of the three-dimensional point cloud map coincide with the same common feature point, and the common feature point is a pixel point corresponding to the first three-dimensional point, multiple included angle values ​​are obtained based on the line connecting each radar point to the same common feature point and the normal vector of the local plane corresponding to each radar point. Select the three-dimensional point cloud information of the radar point corresponding to the minimum included angle value, and use the three-dimensional point cloud information of the radar point corresponding to the minimum included angle value as the three-dimensional point cloud information of the radar point corresponding to the first three-dimensional point.

7. The method according to claim 5, characterized in that, The method further includes: If the second image corresponding to the second three-dimensional point does not include the last image involved in the construction of the three-dimensional model, the radar point closest to the second three-dimensional point is determined, and the three-dimensional point cloud information of the radar point closest to the second three-dimensional point is used as the three-dimensional point cloud information of the radar point corresponding to the second three-dimensional point.

8. The method according to claim 4, characterized in that, The process of inputting the camera pose information of the images participating in model construction, the coordinate information of the 3D points of the supplemented 3D model, the position information of the common feature points participating in the 3D model construction in the camera coordinate system of the corresponding images, and the 3D point information of the radar points corresponding to the 3D points into the factor graph to obtain the optimized camera pose information and optimized coordinate information of the 3D points of each image includes: The camera pose information of the images participating in the model construction is divided into a first group of camera pose information and a second group of camera pose information; wherein, the number of common feature points between the image corresponding to the first group of camera pose information and the last image participating in the 3D model construction exceeds a second preset threshold, and the number of common feature points between the image corresponding to the second group of camera pose information and the last image participating in the model construction does not exceed the second preset threshold. The coordinate information of the three-dimensional points of the supplemented three-dimensional model is divided into the coordinate information of the first group of three-dimensional points and the coordinate information of the second group of three-dimensional points; wherein, the number of images corresponding to the first group of three-dimensional points exceeds the third preset threshold, and the number of images corresponding to the second group of three-dimensional points does not exceed the third preset threshold. Determine the distance values ​​between each 3D point and its corresponding radar point; The first set of camera pose information, the first set of 3D point coordinate information, and the position information of common feature points participating in the 3D model construction in the camera coordinate system of the corresponding image are used as fixed factors. The second set of camera pose information and the second set of 3D point coordinate information are used as variable factors. The distance values ​​between each 3D point and the corresponding radar point are used as the edges of the factors. These are input into the factor graph to obtain the optimized camera pose information and optimized 3D point coordinate information of each image.

9. The method according to claim 1, characterized in that, The method further includes: When a preset number of second images complete the supplementary construction of the first three-dimensional model, the camera pose information of the images located within a specific spatial range, the coordinate information of the three-dimensional points determined based on the images located within the specific spatial range, and the three-dimensional point cloud information of the radar points corresponding to the three-dimensional points located within the specific spatial range are obtained; wherein, the specific spatial range is obtained with the center of the camera corresponding to the last image participating in the construction of the three-dimensional model as the center of the sphere and a preset radius as the radius; The camera pose information of the image within a specific spatial range, the coordinate information of the three-dimensional points determined based on the image within a specific spatial range, and the three-dimensional point cloud information of the radar points corresponding to the three-dimensional points are input into the factor map to obtain the optimized camera pose information of the image within a specific spatial range and the optimized coordinate information of the three-dimensional points within a specific spatial range. Based on the camera pose information of the image within the optimized specific spatial range and the coordinate information of the three-dimensional points within the optimized specific spatial range, the optimized three-dimensional model is obtained.

10. A device for constructing a three-dimensional model, characterized in that, The device includes: An acquisition module is used to acquire a first image, at least one second image, and a three-dimensional point cloud map; wherein the first image and the at least one second image have a first common feature point. The depth information determination module is used to determine the depth information of the first common feature point based on the first common feature point in the first image and the three-dimensional point cloud map; The first three-dimensional model construction module is used to construct a first three-dimensional model based on the depth information of the first common feature points, the position information of the first common feature points in the camera coordinate system of the first image, and the camera pose information of the first image. The camera pose information determination module is used to determine the camera pose information of the corresponding second image based on the depth information of the first common feature point and the position information of the first common feature point in the camera coordinate system of the corresponding second image. The second 3D model construction module is used to supplement and construct the first 3D model based on the depth information of the first common feature points, the position information of the first common feature points in the camera coordinate system of the corresponding second image, and the camera pose information of the corresponding second image, to obtain a second 3D model; wherein, The depth information determination module is further configured to acquire camera pose information of the first image; based on the camera pose information of the first image, project the three-dimensional point cloud map onto the camera imaging plane corresponding to the first image to obtain a projection result; and when the projection result is that the projection point of a radar point in the three-dimensional point cloud map coincides with a first common feature point, determine the distance between the radar point and the first common feature point to use as the depth information of the first common feature point.

11. A device for constructing a three-dimensional model, characterized in that, It includes a memory and a processor, the memory storing a computer program for controlling the processor to operate in order to perform the method according to any one of claims 1-9.

Citation Information

Patent Citations

  • Method and device for constructing three-dimensional point cloud map by multi-machine cooperation and storage medium

    CN111951397A

  • Map construction method, system and device based on visual laser fusion and medium

    CN115342796A