Method, apparatus, and device for constructing a point cloud map
By using the scale reference image of the second camera in the point cloud map construction to calculate the mapping scale of the first camera, the problem of inefficiency in the prior art is solved, and more efficient point cloud map construction and more accurate intelligent robot positioning are achieved.
Patent Information
- Application Number
- CN202011154289.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-10-26
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2040-10-26
AI Technical Summary
The existing technology for building point cloud maps is inefficient, especially when the number of cameras increases, the map is time-consuming and cannot effectively constrain the map size.
By acquiring the map-building keyframe image from the image sequence of the first camera and acquiring the synchronously captured scale reference image from the image sequence of the second camera, the spatial location of the common feature points is determined, the mapping scale of the first camera is calculated, and then a point cloud map is constructed.
While maintaining map construction performance, it improves the efficiency of building point cloud maps and improves the accuracy of intelligent robot positioning.
Smart Images

Figure CN114494612B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer vision technology, and in particular, to a method, device, equipment, and storage medium for constructing a point cloud map. Background Art
[0002] With the development of computer vision technology, a technology for constructing a point cloud map based on computer vision has emerged. Through this technology, the three-dimensional spatial structure of a real scene can be restored from an image sequence, and this technology is also one of the key technologies used in application scenarios such as intelligent robot navigation.
[0003] Structure From Motion (SFM) can gradually use multi-view geometric information to restore the image position and key point coordinates based on the association information of two-dimensional key points between images. For robustness considerations, an incremental mapping framework is widely used.
[0004] Currently, the technology used to construct a point cloud map usually treats the images captured by multiple cameras each time as a series of discrete images for incremental mapping, and then adds loop constraints according to the fixed positions between the cameras for modeling adjustment. However, if, for example, a binocular camera takes 100 shots, then 200 images need to be used for individual composition in the mapping process. And since incremental composition will perform beam adjustment multiple times, and the complexity of beam adjustment with the number of images is O(n 3 ), then this technology will obviously greatly increase the mapping time and cannot effectively constrain the map scale when the number of cameras increases, and the efficiency of constructing the point cloud map is low. Summary of the Invention
[0005] Based on this, it is necessary to provide a method, device, equipment, and storage medium for constructing a point cloud map to solve the technical problem of low efficiency in constructing a point cloud map in the traditional technology.
[0006] A method for constructing a point cloud map, the method includes:
[0007] Obtain at least two key mapping frame images from a first image sequence; the first image sequence is an image sequence captured by a first camera for constructing a point cloud map;
[0008] Obtain an image captured synchronously with a target key mapping frame image in a second image sequence as a scale reference image; the second image sequence is an image sequence captured by a second camera for constructing the point cloud map; the second camera is at a preset distance from the first camera when capturing the image sequence synchronously; the target key mapping frame image is one of the at least two key mapping frame images;
[0009] Based on the at least two mapping key-frame images, obtain the first spatial position corresponding to the common feature points in space; the common feature points are the feature points common to the at least two mapping key-frame images and the scale reference image;
[0010] Based on the target mapping key-frame image and the scale reference image, obtain the second spatial position corresponding to the common feature points in the space;
[0011] Determine the mapping scale of the first camera according to the first spatial position and the second spatial position;
[0012] Based on the mapping scale and the first image sequence, construct a point cloud map based on the point cloud corresponding to the feature points on the target mapping key-frame image in space.
[0013] An apparatus for constructing a point cloud map, comprising:
[0014] A first image acquisition module, configured to acquire at least two mapping key-frame images from a first image sequence; the first image sequence is an image sequence captured by a first camera for constructing a point cloud map;
[0015] A second image acquisition module, configured to acquire an image synchronously captured with the target mapping key-frame image in a second image sequence as a scale reference image; the second image sequence is an image sequence captured by a second camera for constructing the point cloud map; the second camera is at a preset distance from the first camera when synchronously capturing the image sequence; the target mapping key-frame image is one of the at least two mapping key-frame images;
[0016] A first position obtaining module, configured to obtain the first spatial position corresponding to the common feature points in space based on the at least two mapping key-frame images; the common feature points are the feature points common to the at least two mapping key-frame images and the scale reference image;
[0017] A second position obtaining module, configured to obtain the second spatial position corresponding to the common feature points in the space based on the target mapping key-frame image and the scale reference image;
[0018] A scale determination module, configured to determine the mapping scale of the first camera according to the first spatial position and the second spatial position;
[0019] A map construction module, configured to construct a point cloud map based on the mapping scale and the first image sequence, based on the point cloud corresponding to the feature points on the target mapping key-frame image in space.
[0020] A device for constructing a point cloud map, including a first camera, a second camera, and a processor; wherein, the processor is configured to obtain image sequences synchronously captured by the first camera and the second camera, and construct a point cloud map according to the method described above; wherein, when the first camera and the second camera synchronously capture image sequences, they are separated by a preset distance.
[0021] An electronic device, including a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the steps of the method described above are implemented.
[0022] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method described above are implemented.
[0023] The above method, device, equipment and storage medium for constructing a point cloud map include obtaining at least two key mapping frame images from the first image sequence captured by the first camera, and then obtaining an image synchronously captured with the target key mapping frame image from the second image sequence captured by the second camera as a scale reference image, and the target key mapping frame image can be one of the above at least two key mapping frame images. Then, based on the at least two key mapping frame images, the first spatial positions corresponding to the common feature points in space are obtained, and based on the target key mapping frame image and the scale reference image, the second spatial positions corresponding to the common feature points in space are obtained, so as to determine the mapping scale of the first camera according to the first spatial position and the second spatial position. Finally, based on the mapping scale and the first image sequence captured by the first camera, a point cloud map is constructed based on the point cloud corresponding to the feature points on the target key mapping frame in space. This solution can directly add the second camera observation to constrain the map scale of the first camera with almost no increase in complexity, so that the point cloud map can be constructed based on the images captured by the first camera with the real mapping scale obtained, improving the efficiency of constructing the point cloud map while ensuring the mapping performance, and also being beneficial to improving the positioning accuracy of intelligent robots. Description of the Drawings
[0024] Figure 1 It is an application environment diagram of the method for constructing a point cloud map in an embodiment;
[0025] Figure 2 It is a flow schematic diagram of the method for constructing a point cloud map in an embodiment;
[0026] Figure 3 It is a schematic diagram of the principle for determining the mapping scale of the first camera in an embodiment;
[0027] Figure 4 It is a flow schematic diagram of the steps for obtaining key mapping frame images in an embodiment;
[0028] Figure 5 It is a schematic flow chart of the steps for optimizing the mapping scale in an embodiment;
[0029] Figure 6 It is a schematic flow chart of the steps for constructing a point cloud map in an embodiment;
[0030] Figure 7 It is a schematic flow chart of the steps for determining the next key mapping frame image in an embodiment;
[0031] Figure 8 It is a schematic flow chart of the steps for superimposing the newly generated point cloud onto the initial point cloud map in an embodiment;
[0032] Figure 9 It is a schematic flow chart of the front - end mapping process in an application example;
[0033] Figure 10 It is a schematic flow chart of the back - end mapping process in an application example;
[0034] Figure 11 It is a comparative schematic diagram of the mapping effect in an application instance;
[0035] Figure 12 It is a structural block diagram of a device for constructing a point cloud map in an embodiment;
[0036] Figure 13 It is a structural block diagram of a device for constructing a point cloud map in an embodiment;
[0037] Figure 14 It is an internal structure diagram of an electronic device in an embodiment. Detailed implementation manners
[0038] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0039] The method for constructing a point cloud map provided by the present application can be applied to, for example Figure 1In the application environment shown. Among them, the application environment may include a mapping front-end device 110 and a mapping back-end device 120. Among them, the mapping front-end device 110 can be communicatively connected to the mapping back-end device 120 through a network. The mapping front-end device 110 can also be connected to a first camera and a second camera. The mapping front-end device can be used to control the first camera and the second camera to capture real-scene images. The first camera and the second camera can be configured to synchronously capture real-scene images. When capturing real-scene images, the first camera and the second camera are separated by a preset distance. The mapping front-end device 110 can also obtain the image sequences captured by the first camera and the second camera in real time or non-real time. The image sequence captured by the first camera can be recorded as the first image sequence, and the image sequence captured by the second camera can be recorded as the second image sequence.
[0040] The method for constructing a point cloud map provided by this application can be executed alone by the mapping front-end device 110 or by the mapping back-end device 120, or can be executed in cooperation by the mapping front-end device 110 and the mapping back-end device 120.
[0041] In an exemplary embodiment, the method for constructing a point cloud map provided in this application is described by taking the case where the map construction front-end device 110 executes it alone as an example: The map construction front-end device 110 obtains at least two map construction key frame images from the first image sequence, and obtains the images taken synchronously with the target map construction key frame image in the second image sequence as scale reference images; wherein, the target map construction key frame image can be one of the foregoing at least two map construction key frame images; Then, on the one hand, the map construction front-end device 110 obtains the first spatial positions corresponding to the common feature points in the space based on the at least two map construction key frame images, and on the other hand, can obtain the second spatial positions corresponding to the common feature points in the space based on the target map construction key frame image and the scale reference image; wherein, the common feature points are the feature points common to the foregoing at least two map construction key frame images and the scale reference image; wherein, the map construction scale is the scale adopted in the process of constructing the point cloud map. Since the two map construction key frame images in the foregoing first image sequence are both taken by the first camera, and there is a scale scaling effect when constructing a point cloud map through a monocular camera, resulting in the depth of the point cloud determined by it having uncertainty in the actual physical space. Therefore, the above-mentioned first spatial positions calculated only based on the at least two map construction key frame images in the first image sequence usually cannot reflect the true positions of these common feature points in the actual physical space due to the lack of actual physical scale information. For example, although it can be calculated that the distance between a certain common feature point and the first camera in the space is 1, it is impossible to determine whether these common feature points are 1 meter or 1 decimeter away from the first camera in the space, that is, the map construction scale is missing. However, there exists a real number (i.e., the map construction scale) that can align the first spatial positions of these common feature points calculated through the at least two map construction key frame images with the actual spatial positions of the common feature points.
[0042] Next, the map construction front-end device 110 determines the map construction scale of the first camera according to the first spatial position and the second spatial position. Based on this, the map construction front-end device 110 can construct a point cloud map based on the point cloud corresponding to the feature points on the target map construction key frame image in the space based on the map construction scale and the first image sequence.
[0043] In this exemplary embodiment, the method for constructing the point cloud map is executed alone by the map construction front-end device 110, and the real-time construction of the point cloud map of the actual physical scene can be realized when the map construction front-end device 110 is equipped with the first camera and the second camera.
[0044] In an exemplary embodiment, taking the method for constructing a point cloud map provided by this application being executed independently by the mapping backend device 120 as an example for illustration: The mapping backend device 120 can store the first image sequence and the second image sequence sent by the mapping frontend device 110. When it is necessary to construct a point cloud map, the mapping backend device 120 extracts the first image sequence and the second image sequence, and then the mapping backend device 120 can obtain a point cloud map by taking similar step processes as the above method. In this exemplary embodiment, by the mapping backend device 120 independently executing the method for constructing a point cloud map, it is possible to perform offline construction of a point cloud map including but not limited to the point cloud map corresponding to a public image dataset in the backend.
[0045] In an exemplary embodiment, taking the method for constructing a point cloud map provided by this application being executed in cooperation between the mapping frontend device 110 and the mapping backend device 120 as an example for illustration: The mapping frontend device 110 can obtain at least two mapping key frame images from the first image sequence and scale reference images synchronously captured with the target mapping key frame image from the second image sequence, and based on the at least two mapping key frame images, obtain the first spatial positions corresponding to the common feature points in space, and based on the target mapping key frame image and the scale reference images, obtain the second spatial positions corresponding to the common feature points in space. Thus, the mapping frontend device 110 can determine the mapping scale of the first camera. The mapping frontend device 110 can further transmit the mapping scale and the first image sequence to the mapping backend device 120. The mapping backend device 120 constructs a point cloud map based on the point cloud corresponding to the feature points on the target mapping key frame image in space based on the mapping scale and the first image sequence. In this exemplary embodiment, the step of determining the mapping scale of the first camera can be executed by the mapping frontend device 110, and the subsequent step of constructing a point cloud map based on the mapping scale of the first camera and the first image sequence captured by the first camera is executed by the mapping backend device 120 which may have more computing resources, improving the efficiency and accuracy of constructing a point cloud map.
[0046] In the above application scenario, the mapping frontend device 110 can be but not limited to various personal computers, laptop computers, smart phones, and tablet computers, and can also be other electronic devices having at least two cameras, such as a mobile robot. The mapping backend device 120 can be implemented by an independent server or a server cluster composed of multiple servers.
[0047] In one embodiment, as Figure 2 shown, a method for constructing a point cloud map is provided. Taking this method being applied to Figure 1 the mapping backend device 120 therein as an example for illustration, this method may include the following steps:
[0048] Step S201: Obtain at least two mapping key-frame images from the first image sequence.
[0049] In this step, the mapping back-end device 120 can obtain the image sequence captured by the first camera for constructing a point cloud map, which is called the first image sequence. This first image sequence can include multiple frames of images arranged in chronological order obtained by the first camera shooting the spatial real scene. The mapping back-end device 120 obtains at least two mapping key-frame images from the first image sequence. The mapping key-frame image refers to the key-frame image for constructing the point cloud map. One of the at least two mapping key-frame images can be used as the first frame image for constructing the point cloud map in the subsequent steps of the mapping back-end device 120, and one of the frames of images can be denoted as the target mapping key-frame image.
[0050] Step S202: Obtain the image synchronously captured with the target mapping key-frame image in the second image sequence as the scale reference image.
[0051] In this step, after the mapping back-end device 120 obtains at least two mapping key-frame images from the first image sequence, it can further obtain the image synchronously captured with the target mapping key-frame image from the second image sequence as the scale reference image. Specifically, the second image sequence refers to the image sequence captured by the second camera for constructing the aforementioned point cloud map, and it is also multiple frames of images arranged in chronological order obtained by the second camera shooting the spatial real scene. Among them, the second camera and the first camera are at a preset distance when synchronously shooting the corresponding image sequences. That is to say, during the process of shooting the spatial real scene, the distance between the second camera and the first camera is known. Exemplarily, the distance between the first camera and the second camera can be marked by the optical center distance or the baseline length between the two cameras. For example, when the first camera and the second camera synchronously shoot the image sequence, the baseline length can be maintained at about 15 centimeters, and the two cameras can be pre-calibrated to clarify the internal parameters of the two cameras. The exposure time and shooting timestamp of the two cameras can also be further synchronized so that the mapping back-end device 120 can accurately extract the image synchronously captured with the target mapping key-frame image from the second image sequence.
[0052] Step S203: Based on at least two mapping key-frame images, obtain the corresponding first spatial positions of the common feature points in space.
[0053] In this step, the mapping backend device 120 first determines at least two mapping key-frame images and a scale reference image. The feature points common to these at least three images are called common feature points, and these feature points can be determined by means of feature point matching in these at least three images. Specifically, for the calculation process of the above-mentioned first spatial position corresponding to the common feature points in space, the feature points common to the aforementioned at least two mapping key-frame images and the scale reference image can be first determined. Exemplarily, each pixel point in these at least three images can be matched according to features such as the gray value, gray gradient value, etc. that each pixel point has in the corresponding image to determine the common feature points of these at least three images; then, based on the two-dimensional position coordinates of the common feature points on these at least two mapping key-frame images and the rotation and translation relationship of the first camera when these at least two mapping key-frame images are captured, the aforementioned first spatial position corresponding to the common feature points can be calculated.
[0054] After determining the common feature points, the mapping backend device 120 further obtains the spatial positions corresponding to these common feature points in space based on at least two mapping key-frame images, denoted as the first spatial positions. The first spatial positions can be represented by three-dimensional spatial coordinates and are used to characterize the positions of the common feature points in the real scene space calculated based on these mapping key-frame images. These positions can be represented by three-dimensional spatial coordinates. However, since these at least two mapping key-frame images are all captured by the first camera, the first spatial positions calculated therefrom lack actual physical scale information and thus cannot reflect the real positions of these common feature points in the real scene space. For example, if the calculated distance between a certain common feature point in space and the first camera is 1, it is impossible to clarify whether these common feature points are 1 meter or 1 decimeter away from the first camera in space. However, there is a real number, or mapping scale, that can align the first spatial positions of these common feature points calculated from these at least two mapping key-frame images with the actual spatial positions of the common feature points. This alignment process can be performed by steps S204 to S205.
[0055] Step S204: Based on the target mapping key-frame image and the scale reference image, obtain the second spatial position corresponding to the common feature points in space.
[0056] In this step, the mapping backend device 120 can also obtain the corresponding second spatial positions of the common feature points in space based on the target mapping key-frame image and the scale reference image. Specifically, the first camera and the second camera can be regarded as a binocular camera, and the target mapping key-frame image and the scale reference image can be regarded as the images synchronously captured by the binocular camera, and the distance relationship between the two cameras is known. Therefore, the mapping backend device 120 can calculate the spatial positions of the common feature points in the real scene space through binocular matching and triangulation of feature points, which are denoted as the second spatial positions. The second spatial positions can also be represented by three-dimensional spatial coordinates. Since the physical scale between the binocular cameras is known, the obtained second spatial positions have a 1:1 relationship with the real physical world, that is, the second spatial positions can accurately reflect the actual spatial positions of the common feature points in the real scene space.
[0057] Step S205: Determine the mapping scale of the first camera according to the first spatial position and the second spatial position.
[0058] As Figure 3 shown, in this step, the mapping backend device 120 can perform single-point consistency sampling on the common feature points. For example, take the common feature points P1, P2, and P3, obtain the first three-dimensional spatial coordinates used to represent the corresponding first spatial positions of these common feature points, and obtain the second three-dimensional spatial coordinates used to represent the corresponding second spatial positions of the common feature points. Dividing these two sets of spatial coordinates can obtain the mapping scale of the first camera, that is, the process of aligning the first spatial position with the actual spatial position of the common feature points is completed.
[0059] Step S206: Based on the mapping scale and the first image sequence, construct a point cloud map based on the point cloud corresponding to the feature points on the target mapping key-frame image in space.
[0060] In this step, after obtaining the mapping scale of the first camera, the mapping backend device 120 can solve and calculate information such as the camera pose when shooting each frame of image based on the mapping scale of the first camera. The information obtained, including the camera pose, the point cloud corresponding to the feature points in space, etc., can all correspond to the actual physical scale in the real scene. Thus, the mapping backend device 120 can, based on the mapping scale and the first image sequence, use the target mapping key frame image as the first mapping key frame image, and construct a point cloud map based on the point cloud corresponding to the feature points on the target mapping key frame image in space. Specifically, it can use the point cloud corresponding to the feature points on the target mapping key frame image in space as the initial point cloud, find the second frame of image from the remaining images in the first image sequence, calculate the corresponding point cloud and add it to the initial point cloud, and so on until all the images in the first image sequence are used for constructing the point cloud map. Then it can be considered that the construction process of the point cloud map of the real scene is completed. That is to say, the acquisition of the mapping scale only needs to be carried out once, and the mapping backend device 120 can complete the subsequent mapping process based on the obtained mapping scale and the image sequence captured by the first camera to achieve the efficient construction of the point cloud map on the premise of accurately restoring the true scale of the map.
[0061] The above method for constructing a point cloud map includes obtaining at least two mapping key frame images from the first image sequence captured by the first camera. Then, obtain the image captured synchronously with the target mapping key frame image from the second image sequence captured by the second camera as the scale reference image. The target mapping key frame image can be one of the above at least two mapping key frame images. Next, obtain the first spatial position corresponding to the common feature points in space based on the at least two mapping key frame images, and obtain the second spatial position corresponding to the common feature points in space based on the target mapping key frame image and the scale reference image. Thus, determine the mapping scale of the first camera according to the first spatial position and the second spatial position. Finally, based on this mapping scale and the first image sequence captured by the first camera, construct a point cloud map based on the point cloud corresponding to the feature points on the target mapping key frame in space. This solution can directly add the second camera observation to constrain the map scale of the first camera with almost no increase in complexity, so that the point cloud map can be constructed based on the images captured by the first camera that obtains the true mapping scale, improve the efficiency of constructing the point cloud map while ensuring the mapping performance, and is also beneficial to improving the accuracy of intelligent robot positioning.
[0062] For the method of determining at least two key mapping frames in step S201 above, in some embodiments, at least two key mapping frames can be selected from consecutive multiple frames of the first image sequence, that is, the at least two key mapping frames can be consecutive frames in the first image sequence. Exemplarily, taking the selection of two key mapping frames as an example, each group of adjacent two frames of the first image sequence can be used as candidate images for the key mapping frames, and then the group with the largest number of feature point matches among these adjacent two frames is determined as the two key mapping frames.
[0063] In some other embodiments, the at least two key mapping frames selected from the first image sequence may not be consecutive frames in the first image sequence. Among them, one image can be first selected from the first image sequence as a key mapping frame, and then other key mapping frames are determined according to the number of feature point matches between other images in the first image sequence and this key mapping frame. Specifically, in combination with Figure 4 it is described as follows. Obtaining at least two key mapping frames from the first image sequence in step S201 may include:
[0064] Step S401, obtaining the number of feature point matches between each frame of the first image sequence and its adjacent frame;
[0065] In this step, the mapping backend device 120 can match the feature points of each frame of the first image sequence with its adjacent frame in the order of image capture time, and obtain the number of feature point matches between each frame and its adjacent frame.
[0066] Step S402, taking the frame image with the largest number of feature point matches between each frame of the first image sequence and its adjacent frame as the target key mapping frame among the at least two key mapping frames;
[0067] Specifically, the mapping backend device 120 counts the number of feature point matches between each frame in the first image sequence and its adjacent frame, and takes the frame with the largest number of feature point matches as the target key mapping frame among the at least two key mapping frames. This target key mapping frame can be used as the first frame image for constructing the aforementioned point cloud map.
[0068] Step S403, according to the number of feature point matches between other frames of the first image sequence and the target key mapping frame, selecting at least one frame image that meets the preset feature point match number condition from the other frame images as the other key mapping frames among the at least two key mapping frames, so as to obtain at least two key mapping frames.
[0069] In this step, after obtaining the target mapping key-frame image, the mapping backend device 120 can select at least one image from other frame images or remaining frame images of the first image sequence as other mapping key-frame images among at least two mapping key-frame images. Among them, the mapping backend device 120 can select the at least one image based on the number of feature point matches to more accurately construct a point cloud map. Specifically, the mapping backend device 120 determines the number of feature point matches between other frame images of the first image sequence and the target mapping key-frame image, and then can use at least one image that meets the preset feature point match number condition as the above-mentioned other mapping key-frame images. The preset feature point match number condition can be that the number of feature point matches is greater than or equal to a preset feature point match number threshold, or other conditions such as having the largest number of feature point matches with the target mapping key-frame image. In one embodiment, the mapping backend device 120 can select one image with the largest number of feature point matches with the target mapping key-frame image from other frame images as one of the at least two mapping key-frame images, so as to perform mapping initialization to obtain the mapping scale, etc. in combination with the target mapping key-frame image.
[0070] In the technical solution of the above embodiment, the mapping backend device 120 can select the first mapping key-frame image based on the number of feature point matches between adjacent frames in the first image sequence, and select the second mapping key-frame image with the largest number of feature point matches with the first mapping key-frame, so that in the subsequent steps, the mapping scale of the first camera can be more accurately determined based on the feature point matches of the two mapping key-frame images, further improving the mapping accuracy.
[0071] In one embodiment, as Figure 5 shown, determining the mapping scale of the first camera according to the first spatial position and the second spatial position in step S205 may include:
[0072] Step S501, obtaining the initial mapping scale of the first camera according to the first spatial position and the second spatial position;
[0073] Step S502, based on the minimization of the first reprojection error of the first camera and the second reprojection error of the second camera for the point cloud corresponding to the common feature points in space, optimizing the initial mapping scale to obtain the mapping scale.
[0074] In this embodiment, mainly after the mapping backend device 120 obtains the mapping scale of the first camera according to the first and second spatial positions corresponding to the common feature points in space, it first uses it as the initial mapping scale to be optimized, and then optimizes the initial mapping scale according to the reprojection error of the binocular camera to improve the accuracy of the mapping scale of the first camera. Specifically, the mapping backend device 120 can optimize the initial mapping scale s in the form of a unit vector, and the formula is as follows:
[0075]
[0076] wherein, R and T represent the camera poses, s represents the initial mapping scale to be optimized, d(*) represents the normalization operator, and represent the observed coordinates of the 3D point cloud numbered i under the first camera and the second camera, represents the 3D coordinates of the 3D point cloud numbered i in the space under the world coordinate system, h l (*) and h r (*) are the camera observation model equations. The above formula can characterize the difference between the projected values d(h(sRX + T)) calculated from the states of the camera and the point cloud under the observation model and the actual observed values uv.
[0077] In one embodiment, as Figure 6 shown, the construction of the point cloud map based on the mapping scale and the first image sequence in step S206, which is based on the point cloud corresponding to the feature points on the target mapping key frame image in the space, may include:
[0078] Step S601, according to the mapping scale, the target mapping key frame image, and at least one other mapping key frame image among at least two mapping key frame images, determine the point cloud corresponding to the feature points on the target mapping key frame image in the space as the initial point cloud, and obtain the initial point cloud map;
[0079] In this step, mainly after obtaining the mapping scale, the mapping backend device 120 determines the point cloud corresponding to the feature points on the target mapping key frame image in the space as the initial point cloud according to the mapping scale, the target mapping key frame image, and at least one other mapping key frame image among at least two mapping key frame images, and obtains the initial point cloud map. Specifically, the mapping backend device 120 can use the homography matrix decomposition to obtain the camera pose estimation and the feature point matching relationship when shooting the corresponding images according to the mapping scale, the target mapping key frame image, and the at least one other mapping key frame image, and calculate the 3D spatial coordinates of the point cloud corresponding to the feature points on the target mapping key frame image in the real scene space by triangulation. This point cloud is called the initial point cloud, and the corresponding point cloud map is denoted as the initial point cloud map.
[0080] Exemplarily, assume that at least two key mapping frames in the above steps include two key mapping frames, namely the first key mapping frame and the second key mapping frame, and the first key mapping frame can be used as the target key mapping frame. Among them, the first key mapping frame includes feature points A1, B1, C1, and D1. The matching points A2, B2, C2, and D2 that match the feature points A1 to D1 can be determined in the second key mapping frame by means of feature point matching between images. Specifically, in this step, the three-dimensional spatial coordinates of the feature points A1 to D1 can be calculated by triangulation. Triangulation refers to a method of calculating the three-dimensional spatial coordinates of feature points based on the pixel coordinates of the feature points matched in two frames of images and the relative pose relationship of the camera when shooting these two frames. The relative pose relationship may include rotation R and translation t. Among them, after determining the mapping scale, the rotation R and translation t of the camera when shooting these two frames also have actual physical scales accordingly. Therefore, the three-dimensional spatial coordinates of the feature points A1 to D1 characterized by rotation R and translation t also have actual physical scales. Specifically, taking the feature point A1 as an example, assume that the camera coordinate system when the first camera shoots the first key mapping frame is the first coordinate system, and the camera coordinate system when shooting the second key mapping frame is the second coordinate system. Thus, the spatial coordinates of the three-dimensional spatial point P1 corresponding to the feature point A1 in the first coordinate system can be expressed as: s1K -1 p1 = P1. Where s1 represents the depth of the three-dimensional spatial point P1 corresponding to the first coordinate system, K represents the camera internal parameter matrix, and p1 is the coordinate value of the projection point of the three-dimensional spatial point P1 on the first key mapping frame, that is, the feature point A1. Among them, for the depth s1 corresponding to the first coordinate system, the equation can be solved by least squares. In this equation, represents the skew-symmetric matrix of x2, x2 = K -1 p2, where p2 represents the coordinate value of the projection point of the three-dimensional spatial point P1 on the second key mapping frame, that is, the matching point A2, and x1 = K -1 p1. In this way, the three-dimensional coordinates of the point cloud corresponding to the feature points such as A1, B1, C1, and D1 on the first key mapping frame in the real scene space can be obtained. The point cloud corresponding to the feature points A1, B1, C1, and D1 on the first key mapping frame in the real scene space can be used as the initial point cloud, thereby constructing an initial point cloud map containing these point clouds.
[0081] Step S602: Obtain the observed feature points corresponding to the initial point cloud on each of the other frames of the first image sequence;
[0082] In this step, the mapping backend device 120 can obtain the observed feature points corresponding to the aforementioned initial point cloud on each of the other frames of the first image sequence; wherein, the other frames of images may include the frames of the first image sequence other than the aforementioned target mapping key frame image. That is, after obtaining the initial point cloud, the mapping backend device 120 can read all the known three-dimensional point clouds on the point cloud map, and calculate the observations of the known three-dimensional point clouds by each of the other frames of the first image sequence through the feature point matching relationship between the images of the first image sequence, so as to obtain the observed feature points corresponding to the initial point cloud on each of the other frames of images. Among them, calculating the observation of an image on a known three-dimensional point cloud is to obtain the two-dimensional pixel points projected by the known three-dimensional point cloud on the image, and the projected two-dimensional pixel points are the observed feature points. Specifically in this step, the mapping backend device 120 obtains the two-dimensional pixel points projected by the initial point clouds corresponding to the aforementioned feature points A1 to D1 in the real scene space on each of the other frames of images.
[0083] Step S603: Select, according to the uniformity of the distribution of the observed feature points on their respective frame images, the frame images that meet the preset uniformity condition from the other frame images as the next mapping key frame image;
[0084] In this step, the mapping backend device 120 can calculate the uniformity of the distribution of the observed feature points on their respective frame images, and an image with good uniformity can ensure more accurate calculation of the pose of the camera itself. Exemplarily, for the method of calculating the uniformity, the image can be evenly divided into grids, for example, the image is evenly divided into 16×16 grids, and then the number of feature points in each grid is counted, and the variance is calculated. The smaller the variance, the more uniform it is. Based on this, the mapping backend device 120 can select the frame images that meet the preset uniformity condition from the other frame images as the next mapping key frame image, and the preset uniformity condition can be that the aforementioned variance is less than the preset variance threshold, etc.
[0085] Step S604: Based on the target mapping key frame image and the next mapping key frame image, determine the point cloud corresponding to the feature points on the next mapping key frame image in space, and superimpose it on the initial point cloud map to construct a point cloud map.
[0086] In this step, the mapping backend device 120 can use the point cloud corresponding to the feature points on the next mapping key frame image in space as the newly generated point cloud, and superimpose it on the aforementioned initial point cloud map to gradually construct a point cloud map.
[0087] In the technical solution of the above embodiment, based on the initial point cloud map, the mapping back-end device 120 may select the next key mapping frame image that meets the uniformity condition based on the distribution uniformity of the observation feature points corresponding to the known point clouds in the remaining frames of the first image sequence on their respective images, so that the new point cloud accurately generated based on the next key mapping frame image with better uniformity is superimposed on the existing point cloud to gradually construct the point cloud map.
[0088] In one embodiment, as Figure 7 shown, further, the step of selecting a frame image that meets the preset uniformity condition from the other frame images as the next key mapping frame image in step S603 specifically includes:
[0089] Step S701, if there are at least two frame images that meet the preset uniformity condition in the other frame images, use the at least two frame images that meet the preset uniformity condition as candidate images for the next key mapping frame image to obtain at least two candidate images;
[0090] Step S702, determine the number of feature point matches between the candidate image and the synchronously captured image; wherein, the synchronously captured image refers to the image synchronously captured with the candidate image in the second image sequence;
[0091] Step S703, use the candidate image with the largest number of feature point matches among the at least two candidate images as the next key mapping frame image.
[0092] In this embodiment, mainly when the mapping back-end device 120 detects that there are at least two frame images that meet the preset uniformity condition in the other frame images, that is, if there are at least two images with comparable uniformity in the other frame images, these two images can be used as candidate images for the next key mapping frame image, so that at least two candidate images can be obtained. Then, the mapping back-end device 120 further determines the number of feature point matches between the candidate image and the synchronously captured image, that is, compares the binocular feature point matching quantities corresponding to each candidate image. Then, the mapping back-end device 120 uses the candidate image with the largest binocular feature point matching number among these candidate images as the next key mapping frame image. By adopting the technical solution of this embodiment, binocular feature point matching can be further performed when the distribution uniformity of the feature points is comparable. The more the number of binocular feature point matches, the more stable the image shooting can be characterized, and the accuracy of the feature points on the image is relatively high. Therefore, the candidate image with the largest binocular feature point matching number is used as the next key mapping frame image to improve the accuracy of constructing the point cloud map.
[0093] In one embodiment, based on the target mapping key-frame image and the next mapping key-frame image in step S604, determining the point cloud corresponding in space to the feature points on the next mapping key-frame image specifically includes:
[0094] According to the initial point cloud and the observed feature points on the next mapping key-frame image, determining the camera pose corresponding to when the first camera captures the next mapping key-frame image; based on the camera pose, the target mapping key-frame image, and the next mapping key-frame image, determining the point cloud corresponding in space to the feature points on the next mapping key-frame image.
[0095] This embodiment can provide a method for determining the point cloud corresponding in space to the feature points on the next mapping key-frame image based on the target mapping key-frame image and the next mapping key-frame image. Specifically, after the mapping backend device 120 determines the next mapping key-frame image, it can add the next mapping key-frame image to the current map as the input for the camera pose solution corresponding to this image. When inputting the next mapping key-frame image for camera pose solution, the PnP (Perspective-n-Point) method can be used to perform camera pose solution to obtain the camera pose corresponding to when the first camera captures the next mapping key-frame image. After obtaining the camera pose, new point clouds are generated by triangulation again. Specifically, the new point clouds can be generated through the following formula:
[0096]
[0097] where, R represents the rotation of the camera, T represents the translation of the camera, uv represents the observed coordinates of the point cloud in the camera, P represents the three-dimensional coordinates of the new point cloud to be solved, ||h(RP + T) - uv|| 2 represents the difference between the theoretical observed coordinates and the actual observed coordinates of the point cloud projected onto the camera, M represents the new point cloud map corresponding to the newly generated point cloud, and this new point cloud map can be superimposed onto the aforementioned initial point cloud map to gradually complete the construction of the point cloud map.
[0098] In one embodiment, further, as Figure 8 shown, superimposing onto the initial point cloud map to construct the point cloud map in step S604 may include:
[0099] Step S801, obtaining multiple co-visible images of the point cloud corresponding in space to the feature points on the next mapping key-frame image from the first image sequence;
[0100] During the process of gradually building a map, the map building backend device 120 may use the point cloud corresponding to the feature points on the aforementioned next map building key frame image determined in step S604 in the real scene space as the newly generated point cloud, which is used to be superimposed on the aforementioned initial point cloud map to gradually build the point cloud map. In this step, after determining the newly generated point cloud, the map building backend device 120 may obtain multiple co-visible images for the newly generated point cloud from the first image sequence. Among them, the co-visible image refers to the image in the first image sequence that has a co-visible relationship with the newly generated point cloud, and this co-visible relationship can be determined based on the number of pixel points matched by the feature points corresponding to the newly generated point cloud on the next map building key frame image on each frame image in the first image sequence. Exemplarily, the image in which the number of pixel points matched by the feature points corresponding to the newly generated point cloud on the next map building key frame image on each frame image is greater than a preset number threshold can be used as the co-visible image, so that in this way, multiple co-visible images for the newly generated point cloud can be obtained from the first image sequence.
[0101] Step S802, based on the point cloud corresponding to the feature points on the next map building key frame image in space, minimize the reprojection error of each co-visible image to optimize the point cloud corresponding to the feature points on the next map building key frame image in space;
[0102] In this step, the point cloud corresponding to the feature points on the next map building key frame image in space can be used as the newly generated point cloud, and the map building backend device 120 minimizes the reprojection error of each co-visible image based on the newly generated point cloud to optimize the newly generated point cloud. Among them, during the map building process, the map building backend device 120 can perform local optimization on the camera pose and the three-dimensional coordinates of the newly generated point cloud involved in the local map formed by each co-visible image. This local optimization is to minimize the reprojection error of each co-visible image based on the newly generated point cloud, and the optimized newly generated point cloud can be re-superimposed on the existing point cloud map, so that the construction of the point cloud map is made more accurate through this local optimization method.
[0103] Step S803, superimpose the optimized point cloud corresponding to the feature points on the next map building key frame image in space on the initial point cloud map to build the point cloud map.
[0104] In this embodiment, mainly during the process of the map building backend device 120 building the point cloud map, it can perform local optimization on the camera pose and the three-dimensional coordinates of the point cloud involved in the local map, and the optimized point cloud can be re-superimposed on the existing point cloud map, so that the construction of the point cloud map is made more accurate through this local optimization method.
[0105] Specifically, the foregoing local optimization process is described as follows. The mapping backend device 120 can obtain multiple co-visual images of the points in space corresponding to the feature points on the next mapping key-frame image from the first image sequence. That is, for the point cloud newly generated by the next mapping key-frame image, the mapping backend device 120 can determine the co-visual images of the newly generated point cloud from the first image sequence by means of, for example, feature point matching. The number of co-visual images is generally multiple, and the newly generated point cloud corresponding to these co-visual images constitutes a local map of the entire point cloud map. After obtaining this local map, the mapping backend device 120 can input the camera poses of the first camera corresponding to the co-visual images and the three-dimensional coordinates of the point cloud corresponding to the local map involved in all local maps into the following local optimization function for local optimization:
[0106]
[0107] where L represents local, j represents the number of the point cloud, and i represents the number of the co-visual image. In this embodiment, the newly generated point cloud after local optimization can be superimposed on the existing point cloud map to gradually construct the point cloud map.
[0108] In one embodiment, further, the above method can also perform global optimization processing of the point cloud map by the following steps:
[0109] Determine the outliers in the existing point cloud according to the reprojection error of each frame image in the first image sequence with respect to the existing point cloud in the point cloud map; if the ratio of the observed feature points of the outliers in each frame image in the first image sequence to the observed feature points of the existing point cloud in each frame image in the first image sequence is greater than or equal to a preset ratio, then optimize the existing point cloud based on the minimization of the reprojection error of each frame image in the first image sequence with respect to the existing point cloud.
[0110] In this embodiment, the mapping backend device 120 can perform global optimization during the process of constructing the point cloud to make the constructed point cloud more accurate and smooth. Among them, the mapping backend device 120 can remove outliers through the reprojection error. Specifically, the mapping backend device 120 can determine the outliers in the existing point cloud according to the reprojection error of each frame image in the first image sequence with respect to the existing point cloud in the point cloud map. Exemplarily, the mapping backend device 120 can calculate the reprojection error using the following formula:
[0111] reproj = h(RP + T) - uv
[0112] Among them, reproj represents the difference between the theoretical observation coordinates and the actual observation coordinates of the point cloud projected onto the camera, that is, the reprojection error. An outlier can be a point cloud with a reprojection error reproj greater than or equal to a certain threshold. If the mapping backend device 120 determines that the ratio of the observed feature points corresponding to each frame of the first image sequence of the outlier to the observed feature points corresponding to each frame of the first image sequence of the existing point cloud is greater than or equal to a certain ratio, that is, the aforementioned preset ratio, then the existing point cloud can be optimized, and the following formula can be used for optimization:
[0113]
[0114] Among them, G represents global, j represents the number of the point cloud, and i represents the number of each frame of the image. In this embodiment, global optimization can be performed during the process of constructing the point cloud to construct a more accurate point cloud map.
[0115] To more clearly illustrate the method for constructing a point cloud map provided by this application, this method is applied to the construction of a point cloud map by a binocular camera for illustration. The above method can include a front-end mapping process and a back-end mapping process in this application example. The following combines Figure 9 and Figure 10 to make a detailed description as follows:
[0116] For the front-end mapping process, as Figure 9As shown, in step S901, the mapping front-end device 110 can select each frame of binocular images in the order of image capture time. Binocular images refer to the images captured by the left-eye camera and the right-eye camera during the shooting process of the binocular camera. In step S902, the mapping front-end device 110 can perform feature point matching between each frame of image and its adjacent frame of image. After the mapping front-end device 110 completes the adjacent-frame feature point matching of all images, it can count the number of feature point matches between each frame of image in the entire image sequence and its adjacent frame of image, and use the frame with the largest number of feature point matches with its adjacent frame of image as the first frame of image for constructing the point cloud map. In this regard, the mapping front-end device 110 can also perform step S903 while performing feature point matching, that is, the mapping front-end device 110 performs loop detection every certain number of images. When a closed loop is detected, step S904 is performed to perform closed-loop image matching, so that the newly captured image performs feature point matching with the historical frame image with a long time interval. After each loop detection, step S905 needs to be performed, that is, the mapping front-end device 110 continues to perform binocular feature point matching to ensure that the images captured by the camera all perform binocular feature matching. Among them, binocular feature point matching refers to the feature point matching between two images captured synchronously by the left-eye camera and the right-eye camera. After binocular matching, the mapping front-end device 110 can then execute step S906 to determine whether there are remaining images that have not been subjected to feature point matching. If not, it is determined that the front-end mapping process ends. If so, the mapping front-end device 110 returns to step S901 to continue feature point matching.
[0117] For the back-end mapping process, such as Figure 10As shown in the figure, in step S1001, the mapping backend device 120 can select, according to the foregoing first frame image, the frame image with the largest number of feature point matches with the first frame image as the second frame image, use the first frame image and the second frame image as the mapping key frame images, and adopt steps S201 to S206 in the above embodiment to perform mapping initialization to obtain the mapping scale of the camera, align the three-dimensional spatial position coordinates of the common feature points, obtain the initial point cloud, and generate the initial point cloud map. Further, after obtaining the initial point cloud map, the mapping backend device 120 can perform step S1002, that is, select the next mapping key frame image for mapping. The mapping backend device 120 can read all the known three-dimensional point clouds in the initial point cloud map and enter step S1003. That is, the mapping backend device 120 calculates the observed feature points of the known three-dimensional point cloud under all images through the feature point matching relationship between images, and uses a grid to calculate the uniformity of the distribution of each observed feature point in its respective image. The mapping backend device 120 enters step S1003 to determine whether the distribution of the observed feature points in their respective images is uniform. If so and the uniformity of the two images is comparable, the mapping backend device 120 compares the number of binocular feature point matches of the two images, then selects the image with the largest number of binocular feature point matches as the next mapping image, and adds it to the current map as the input for the camera pose solution of the next image. When inputting the next mapping image for camera pose solution, the PnP method is directly used to perform camera pose solution to obtain the camera pose to be optimized. If the mapping backend device 120 determines in step S1003 that the distribution of the observed points is uneven, it returns to step S1002 to reselect the next frame for mapping processing. After obtaining the camera pose in step S1004, the mapping backend device 120 enters step S1005. That is, the mapping backend device 120 can triangulate again to generate a new point cloud. Specifically, the new point cloud can be generated through the following formula:
[0118]
[0119] After generating the new point cloud, the mapping backend device 120 enters step S1006, that is, performs local smoothing processing on the map. The mapping backend device 120 uses the newly generated point cloud and the camera pose to be optimized as the input again, and finds the co-visible images corresponding to the newly generated point cloud through feature point matching. These co-visible images can form a local map. After obtaining the local map, the camera poses and the three-dimensional spatial position coordinates of the point cloud involved in all local maps are put into the following local optimization function for local optimization:
[0120]
[0121] Further, the mapping backend device 120 can also enter step S1007, that is, perform outlier rejection through the reprojection error. The reprojection error can be calculated using the following formula:
[0122] reproj = h(RP + T) - uv.
[0123] Next, the mapping backend device 120 enters step S1008 to detect whether the outlier ratio exceeds the threshold. If the ratio of the observed feature points corresponding to the outliers detected by the mapping backend device 120 to the observed feature points corresponding to the existing point cloud is greater than or equal to the ratio threshold, the mapping backend device 120 enters step S1009 to perform global smoothing or global optimization processing on the constructed point cloud map. After the global optimization processing, the mapping backend device 120 returns to step S1002 to select the next frame for mapping processing; otherwise, the mapping backend device 120 can directly return to step S1002 to select the next frame for mapping processing. Among them, the formula for global optimization can be:
[0124]
[0125] As Figure 11 shown is a comparison diagram of the effects of the mapping solution provided by this embodiment and the mapping solution provided by the traditional technology. It shows the effect display of mapping for the public dataset. The public dataset can be regarded as a series of images taken by a calibrated binocular camera. After obtaining this series of images, the series of images can also be sampled at intervals of 5 Hz to ensure that there are as few redundant key-frame images as possible. Then, mapping can be completed based on the mapping solution provided by this embodiment. Among them, the first effect diagram 1110 shows the effect diagram of the mapping solution provided by the traditional technology, and the second effect diagram 1120 is the effect diagram of the mapping solution provided by this embodiment. It can be seen that the mapping solution provided by this embodiment has better mapping performance than the mapping solution provided by the traditional technology. Specifically, the difference between the solution provided by this embodiment and the mapping solution provided by the traditional technology is that the solution of this embodiment can regard each pair of binocular images as a single image with sparse depth, thereby greatly reducing its algorithm complexity, accelerating the algorithm process, and at the same time using the binocular camera to perform mapping initialization, binocular mapping order and other links to ensure mapping accuracy, which can effectively balance mapping time and mapping performance, and is beneficial to greatly improving the positioning accuracy of intelligent robots.
[0126] It should be understood that although Figures 1 to 10 the steps in the flowchart are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, Figures 1 to 10At least some of the steps may include multiple steps or multiple stages, which do not necessarily need to be executed and completed at the same moment, but can be executed at different moments, and the execution order of these steps or stages does not necessarily need to be sequential, but can be executed alternately or in turn with at least some of the steps or stages in other steps or other steps.
[0127] In one embodiment, as Figure 12 shown, a device for constructing a point cloud map is provided. The device 1200 may include:
[0128] A first image acquisition module 1201, configured to acquire at least two key mapping frame images from a first image sequence; the first image sequence is an image sequence captured by a first camera for constructing a point cloud map;
[0129] A second image acquisition module 1202, configured to acquire an image captured synchronously with a target key mapping frame image in a second image sequence as a scale reference image; the second image sequence is an image sequence captured by a second camera for constructing the point cloud map; the second camera is preset distance away from the first camera when capturing the image sequence synchronously; the target key mapping frame image is one of the at least two key mapping frame images;
[0130] A first position obtaining module 1203, configured to obtain a first spatial position corresponding to a common feature point in space based on the at least two key mapping frame images; the common feature point is a feature point common to the at least two key mapping frame images and the scale reference image;
[0131] A second position obtaining module 1204, configured to obtain a second spatial position corresponding to the common feature point in the space based on the target key mapping frame image and the scale reference image;
[0132] A scale determination module 1205, configured to determine a mapping scale of the first camera according to the first spatial position and the second spatial position;
[0133] A map construction module 1206, configured to construct a point cloud map based on the mapping scale and the first image sequence, with the point cloud corresponding to the feature points on the target key mapping frame image in the space as the basis.
[0134] In one embodiment, the first image acquisition module 1201 is further configured to obtain the number of feature point matches between each frame image of the first image sequence and adjacent frame images; use the frame image with the largest number of feature point matches between adjacent frame images among each frame image of the first image sequence as the target mapping key frame image among the at least two mapping key frame images; and select at least one frame image that meets the preset feature point match number condition from the other frame images according to the number of feature point matches between the other frame images of the first image sequence and the target mapping key frame image as the other mapping key frame images among the at least two mapping key frame images, so as to obtain the at least two mapping key frame images.
[0135] In one embodiment, the preset feature point match number condition includes having the largest number of feature point matches with the target mapping key frame image.
[0136] In one embodiment, the map construction module 1206 is further configured to determine, according to the mapping scale, the target mapping key frame image, and at least one other mapping key frame image among the at least two mapping key frame images, the point cloud corresponding to the feature points on the target mapping key frame image in the space as the initial point cloud, so as to obtain an initial point cloud map; obtain the observed feature points corresponding to the initial point cloud on each of the other frame images of the first image sequence; the other frame images include the frame images in the first image sequence except the target mapping key frame image; select, according to the uniformity of the distribution of the observed feature points on their respective frame images, the frame images that meet the preset uniformity condition from the other frame images as the next mapping key frame image; and determine the point cloud corresponding to the feature points on the next mapping key frame image in the space based on the target mapping key frame image and the next mapping key frame image, and superimpose it on the initial point cloud map to construct the point cloud map.
[0137] In one embodiment, the map construction module 1206 is further configured to, if there are at least two frame images that meet the preset uniformity condition among the other frame images, use the at least two frame images that meet the preset uniformity condition as candidate images for the next mapping key frame image to obtain at least two candidate images; determine the number of feature point matches between the candidate images and the synchronously captured images; the synchronously captured images are the images synchronously captured with the candidate images in the second image sequence; and use the candidate image with the largest number of feature point matches among the at least two candidate images as the next mapping key frame image.
[0138] In one embodiment, the map construction module 1206 is further configured to determine the camera pose corresponding to the first camera when capturing the next key mapping frame image according to the initial point cloud and the observed feature points on the next key mapping frame image; based on the camera pose, the target key mapping frame image, and the next key mapping frame image, determine the point cloud corresponding to the feature points on the next key mapping frame image in the space.
[0139] In one embodiment, the map construction module 1206 is further configured to obtain multiple co-visual images of the point cloud corresponding to the feature points on the next key mapping frame image in the space from the first image sequence; optimize the point cloud corresponding to the feature points on the next key mapping frame image in the space based on the minimization of the reprojection error of each co-visual image with respect to the point cloud corresponding to the feature points on the next key mapping frame image in the space; and superimpose the optimized point cloud corresponding to the feature points on the next key mapping frame image in the space onto the initial point cloud map to construct the point cloud map.
[0140] In one embodiment, the above device 1200 may further include: a global optimization unit configured to determine outliers in the existing point cloud according to the reprojection error of each frame image in the first image sequence with respect to the existing point cloud in the point cloud map; if the ratio of the observed feature points corresponding to each frame image in the first image sequence for the outliers to the observed feature points corresponding to each frame image in the first image sequence for the existing point cloud is greater than or equal to a preset ratio, optimize the existing point cloud based on the minimization of the reprojection error of each frame image in the first image sequence with respect to the existing point cloud.
[0141] In one embodiment, the scale determination module 1205 is further configured to: obtain the initial mapping scale of the first camera according to the first spatial position and the second spatial position; optimize the initial mapping scale to obtain the mapping scale based on the minimization of the first reprojection error of the first camera and the second reprojection error of the second camera with respect to the point cloud corresponding to the common feature points in the space.
[0142] For the specific limitations on the device for constructing the point cloud map, reference may be made to the limitations on the method for constructing the point cloud map described above, which will not be elaborated here. Each module in the above device for constructing the point cloud map can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the electronic device in hardware form or be independent of it, or be stored in the memory of the electronic device in software form for the processor to call and execute the operations corresponding to the above respective modules.
[0143] In one embodiment, a device for constructing a point cloud map is provided. As Figure 13 shown, the device may include: a first camera, a second camera, and a processor. The processor is configured to obtain image sequences synchronously captured by the first camera and the second camera, and construct a point cloud map according to the method described in any of the foregoing embodiments. When the first camera and the second camera synchronously capture image sequences, they are separated by a preset distance.
[0144] In one embodiment, an electronic device is provided. The electronic device can be a terminal of a map building front-end device or a server of a map building back-end device. Its internal structure diagram can be as Figure 14 shown. The electronic device includes a processor, a memory, and a communication interface connected through a system bus. The processor of the electronic device is configured to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the electronic device is configured to communicate with external devices in a wired or wireless manner. The wireless manner can be implemented through WIFI, a carrier network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, a method for constructing a point cloud map is implemented.
[0145] Those skilled in the art can understand that Figure 14 the structure shown in
[0146] is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the electronic device to which the solution of the present application is applied. A specific electronic device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.
[0147] In one embodiment, an electronic device is further provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the steps in the foregoing method embodiments are implemented.
[0148] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above various methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0149] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the various technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0150] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.
Claims
1. A method for constructing a point cloud map, characterized in that The method includes: Obtaining at least two mapping key-frame images from a first image sequence; the first image sequence is an image sequence captured by a first camera for constructing a point cloud map; Obtaining an image captured synchronously with a target mapping key-frame image in a second image sequence as a scale reference image; the second image sequence is an image sequence captured by a second camera for constructing the point cloud map; the second camera is at a preset distance from the first camera when capturing the image sequence synchronously; the target mapping key-frame image is one of the at least two mapping key-frame images; Based on the at least two mapping key-frame images, obtaining a first spatial position corresponding to a common feature point in space; the common feature point is a feature point common to the at least two mapping key-frame images and the scale reference image; Based on the target mapping key-frame image and the scale reference image, obtaining a second spatial position corresponding to the common feature point in the space; Obtaining the mapping scale of the first camera according to the division processing of the first three-dimensional spatial coordinates of the first spatial position and the second three-dimensional spatial coordinates of the second spatial position; Based on the mapping scale and the first image sequence, constructing a point cloud map based on the point cloud corresponding to the feature points on the target mapping key-frame image in space.
2. The method according to claim 1, characterized in that The obtaining at least two mapping key-frame images from the first image sequence includes: Obtaining the number of feature point matches between each frame image of the first image sequence and its adjacent frame image; Taking the frame image with the largest number of feature point matches between each frame image of the first image sequence and its adjacent frame image as the target mapping key-frame image among the at least two mapping key-frame images; According to the number of feature point matches between the other frame images of the first image sequence and the target mapping key-frame image respectively, selecting at least one frame image that meets the preset feature point match number condition from the other frame images as the other mapping key-frame images among the at least two mapping key-frame images, to obtain the at least two mapping key-frame images.
3. The method according to claim 2, wherein The preset feature point match number condition includes having the largest number of feature point matches with the target mapping key-frame image.
4. The method according to claim 1, characterized in that The constructing a point cloud map based on the mapping scale and the first image sequence, including: According to the mapping scale, the target mapping key-frame image and at least one other mapping key-frame image among the at least two mapping key-frame images, determining the point cloud corresponding to the feature points on the target mapping key-frame image in space as the initial point cloud, to obtain an initial point cloud map; Obtaining the observed feature points corresponding to the initial point cloud on the other each frame image of the first image sequence; the other each frame image includes the frame images in the first image sequence except the target mapping key-frame image; According to the evenness of the distribution of the observed feature points on their respective frame images, selecting the frame images that meet the preset evenness condition from the other each frame image as the next mapping key-frame image; Based on the target mapping key-frame image and the next mapping key-frame image, determine the point cloud corresponding to the feature points on the next mapping key-frame image in the space, and superimpose it on the initial point cloud map to construct the point cloud map.
5. The method according to claim 4, characterized in that The selecting, from the other frame images, the frame image that meets the preset uniformity condition as the next mapping key-frame image includes: If there are at least two frame images among the other frame images that meet the preset uniformity condition, use the at least two frame images that meet the preset uniformity condition as candidate images for the next mapping key-frame image, obtaining at least two candidate images; Determine the number of feature point matches between the candidate image and the synchronously captured image; the synchronously captured image is the image synchronously captured with the candidate image in the second image sequence; Use the candidate image with the largest number of feature point matches among the at least two candidate images as the next mapping key-frame image.
6. The method according to claim 4, characterized in that, The determining, based on the target mapping key-frame image and the next mapping key-frame image, the point cloud corresponding to the feature points on the next mapping key-frame image in the space includes: According to the initial point cloud and the observed feature points on the next mapping key-frame image, determine the camera pose corresponding to the first camera when capturing the next mapping key-frame image; Based on the camera pose, the target mapping key-frame image, and the next mapping key-frame image, determine the point cloud corresponding to the feature points on the next mapping key-frame image in the space.
7. The method according to any one of claims 4 to 6, wherein The superimposing on the initial point cloud map to construct the point cloud map includes: From the first image sequence, obtain multiple co-visible images of the point cloud corresponding to the feature points on the next mapping key-frame image in the space; Based on the minimization of the reprojection error of each co-visible image by the point cloud corresponding to the feature points on the next mapping key-frame image in the space, optimize the point cloud corresponding to the feature points on the next mapping key-frame image in the space; Superimpose the optimized point cloud corresponding to the feature points on the next mapping key-frame image in the space on the initial point cloud map to construct the point cloud map; The method further includes: According to the reprojection error of each frame image in the first image sequence by the existing point cloud in the point cloud map, determine the outliers in the existing point cloud; If the ratio of the observed feature points corresponding to each frame image in the first image sequence of the outliers to the observed feature points corresponding to each frame image in the first image sequence of the existing point cloud is greater than or equal to a preset ratio, optimize the existing point cloud based on the minimization of the reprojection error of each frame image in the first image sequence by the existing point cloud.
8. The method according to claim 1, wherein The obtaining the mapping scale of the first camera by performing a division process on the first three-dimensional space coordinates of the first spatial position and the second three-dimensional space coordinates of the second spatial position includes: The initial mapping scale of the first camera is obtained by dividing the first three-dimensional spatial coordinates of the first spatial position by the second three-dimensional spatial coordinates of the second spatial position. Based on the minimization of the first reprojection error of the first camera and the second reprojection error of the second camera with respect to the point cloud corresponding to the common feature points in the space, the initial mapping scale is optimized to obtain the mapping scale.
9. An apparatus for constructing a point cloud map, characterized in that, It includes: A first image acquisition module for acquiring at least two key mapping frame images from a first image sequence; the first image sequence is an image sequence captured by a first camera for constructing a point cloud map. A second image acquisition module for acquiring an image synchronously captured with a target key mapping frame image in a second image sequence as a scale reference image. The second image sequence is an image sequence captured by a second camera for constructing the point cloud map; the second camera is at a preset distance from the first camera when synchronously capturing the image sequence; the target key mapping frame image is one of the at least two key mapping frame images. A first position obtaining module for obtaining a first spatial position corresponding to common feature points in the space based on the at least two key mapping frame images; the common feature points are feature points common to the at least two key mapping frame images and the scale reference image. A second position obtaining module for obtaining a second spatial position corresponding to the common feature points in the space based on the target key mapping frame image and the scale reference image. A scale determination module for obtaining the mapping scale of the first camera by dividing the first three-dimensional spatial coordinates of the first spatial position by the second three-dimensional spatial coordinates of the second spatial position. A map construction module for constructing a point cloud map based on the mapping scale and the first image sequence, with the point cloud corresponding to the feature points on the target key mapping frame image in the space as the basis.
10. A device for constructing a point cloud map, characterized in that, It includes a first camera, a second camera, and a processor; wherein, the processor is configured to acquire the image sequences synchronously captured by the first camera and the second camera, and construct a point cloud map according to the method of any one of claims 1 to 8; wherein, the first camera is at a preset distance from the second camera when synchronously capturing the image sequences.
Citation Information
Patent Citations
Scale reduction method and system, three-dimensional reconstruction method and system, storage medium and equipment
CN111402429A
Construction method and device of visual point cloud map
CN111795704A