Map incremental reconstruction method and device applied to flight scene
By dividing the image area and extracting feature points in the flight scene, combining the optimization of lens pose and map points, the problem of inaccurate matching of feature points in SLAM and SFM methods is solved, and efficient and accurate incremental map reconstruction is achieved.
Patent Information
- Application Number
- CN202510457666.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-11
AI Technical Summary
The existing SLAM and SFM methods have problems in the flight scenarios with inaccurate feature point matching and in real-time map reconstruction, resulting in high or sparse map error rates, making it impossible to quickly build high-precision maps.
By collecting images in the flying state, dividing and extracting feature points, combining existing feature points for incremental reconstruction, optimizing lens position and map points, improving feature point matching accuracy and map reconstruction efficiency.
It improves the density and accuracy of feature point clouds, enhances the adaptability to scene changes, ensures the accuracy and reliability of map reconstruction, supports multi-lens and local optimization, reduces the amount of calculation, and improves map matching.
Smart Images

Figure CN120298600A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of map construction, and particularly to a method and device for incremental reconstruction of a map applied to a flight scenario. Background Art
[0002] Generally, existing methods for visual three-dimensional reconstruction are based on SLAM (Simultaneous Localization and Mapping) or SFM (Structure from Motion). Among them, SLAM uses visual odometry to track an image sequence and uses loop detection to eliminate cumulative errors to achieve three-dimensional incremental reconstruction. SFM analyzes feature points and their motion relationships in multi-view images to restore the three-dimensional scene structure and the camera's motion trajectory to achieve three-dimensional incremental reconstruction.
[0003] However, since SLAM emphasizes real-time performance, it is necessary to use a feature point method or a sparse optical flow method with a small amount of computation for feature point matching. Among them, although the amount of computation of feature points is small, usually the feature point matching is inaccurate, resulting in a high error rate in the obtained map; and the map point cloud obtained by the sparse optical flow method is very sparse and cannot represent a complex three-dimensional surface; and the SLAM visual odometry often fails to track, resulting in the inability to achieve three-dimensional reconstruction. And SFM belongs to post-shooting reconstruction, that is, reconstruction starts after all images to be processed are collected, which is non-incremental and not applicable to application scenarios that hope to quickly construct a map. Therefore, how to provide a new method for incremental three-dimensional reconstruction of a map to improve the matching accuracy of feature points and be applicable to application scenarios that hope to quickly construct a map is particularly important. Summary of the Invention
[0004] The present invention provides a method and device for incremental reconstruction of a map applied to a flight scenario to improve the matching accuracy of feature points and be applicable to application scenarios that hope to quickly construct a map.
[0005] To solve the above technical problems, in a first aspect, the present invention discloses a method for incremental reconstruction of a map applied to a flight scenario, the method including:
[0006] After an image processing device completes lens calibration and obtains an initial map of the current scene, an image acquisition operation is performed based on at least one exposure point of the image processing device in the current scene; wherein, the image processing device is in a flight state, and the initial map of the current scene is determined by first images collected by the image processing device through the lens of the image processing device at at least two of the exposure points when the image processing device is in a flight state;
[0007] When the lens of the image processing device captures any second image at at least one of the exposure points, perform an area division operation on the second image to obtain a plurality of sub-regions;
[0008] For any sub-region of the second image, extract a preset number of feature points;
[0009] According to all the feature points of the second image and all the feature points of the plurality of first images determined in advance, perform an incremental reconstruction operation on the initial map of the current scene.
[0010] As an optional implementation manner, in the first aspect of the present invention, the performing an incremental reconstruction operation on the initial map of the current scene according to all the feature points of the second image and all the feature points of the plurality of first images determined in advance includes:
[0011] According to the second image, determine the projection of the second image on the ground;
[0012] According to the projections of the plurality of first images determined in advance on the ground, screen out a third image from the plurality of first images whose projection intersects with the projection of the second image;
[0013] Match the feature points of the second image with the feature points of the third image to obtain a feature point matching result of the second image, and the feature point matching result corresponding to the second image includes a plurality of first feature matching points;
[0014] From all the first feature matching points included in the second image, screen out second feature matching points that do not have corresponding map points in the initial map of the current scene;
[0015] Perform a spatial triangulation operation on each of the second feature matching points to obtain a map point corresponding to the second feature matching point, and add all the map points corresponding to the second feature matching points to the initial map of the current scene.
[0016] As an optional implementation manner, in the first aspect of the present invention, the method further includes:
[0017] Obtain the vertical distance between the image processing device and the ground when the second image is captured, the internal parameter matrix of the lens used to capture the second image, and the pose of the lens when the second image is captured;
[0018] Wherein, the determining the projection of the second image on the ground according to the second image includes:
[0019] Obtain the pixel coordinates of the four corner imaging points in the second image;
[0020] Calculate the projection of the second image on the ground based on the vertical distance corresponding to the second image, the internal parameter matrix corresponding to the second image, the pose corresponding to the second image, and the pixel coordinates of the four corner imaging points in the second image.
[0021] As an optional implementation manner, in the first aspect of the present invention, the method further includes:
[0022] When all the map points corresponding to the second images collected by the image processing device are added to the map of the current scene, for the multiple imaging points of all the fourth images for which the map construction has been completed, obtain the map coordinates of each imaging point and the positions of all the exposure points, where all the fourth images include all the second images;
[0023] Calculate a scaling factor according to the map coordinates of all the imaging points and the positions of all the exposure points, and use the scaling factor to scale the position of each exposure point among all the exposure points to obtain the scaled position of each exposure point;
[0024] Calculate the target rotation matrix and the target translation vector between the three-dimensional model and the geospatial based on the scaled positions of all the exposure points and the imaging point center coordinates of all the fourth images;
[0025] Use the calculated target rotation matrix and the target translation vector to perform transformation and translation operations on the map of the current scene that has been reconstructed, so that all the map points corresponding to the current scene are aligned with the geographical coordinates.
[0026] As an optional implementation manner, in the first aspect of the present invention, the method further includes:
[0027] Optimize the pose of the lens when collecting the second image and the map points corresponding to each second feature matching point based on each second feature matching point of the second image, the map point corresponding to the second feature matching point, and the focal length of the lens that collected the second image, to obtain the optimized pose of the lens and the optimized map points;
[0028] Wherein, adding the map points corresponding to all the second feature matching points to the initial map of the current scene includes:
[0029] Add the map points corresponding to all the second feature matching points of the optimized second image to the initial map of the current scene;
[0030] Wherein, the pose corresponding to each second image is used as the basis for incremental reconstruction of the initial map of the current scene.
[0031] As an alternative implementation, in the first aspect of the present invention, the method further includes:
[0032] Determine whether the current conditions of the current scene satisfy the globally determined optimization conditions;
[0033] When it is determined that the globally determined optimization conditions are satisfied, based on the internal parameter matrix of the lens of the second image obtained by calibration, optimize the pose of the lens and the map points corresponding to each second feature matching point determined in advance when collecting the second image, to obtain the optimized pose of the lens and the optimized map points, until all the map points corresponding to the second images collected by the image processing device are added to the initial map of the current scene.
[0034] As an alternative implementation, in the first aspect of the present invention, before the image acquisition operation is performed by the image processing device at at least one exposure point in the current scene, the method further includes:
[0035] Based on the image acquisition operation performed by the image processing device at at least one exposure point in the current scene, obtain the metadata corresponding to each lens of the image processing device, where the metadata corresponding to each lens includes the first image collected by the lens at at least one exposure point, the position of the exposure point in the current scene, and the pose of the lens at the exposure point when collecting the first image, and the number of lenses of the image processing device is greater than or equal to 1;
[0036] For any one of the lenses, extract the feature points of each first image of the lens, and determine one of the images from all the first images as the reference image; for any one of the first images other than the reference image among all the first images, match the feature points of the first image with the feature points of the reference image to obtain the feature point matching result corresponding to the first image, and each feature point matching result corresponding to each first image includes a plurality of third feature matching points; and according to the feature point matching result corresponding to the first image, calculate the fundamental matrix between the first image and the reference image; according to the fundamental matrix of the first image, determine the projective projection matrix of the first image; according to the properties of the projective projection matrices of all the first images and the absolute dual quadric coefficients, determine the absolute dual quadric coefficients; and calibrate the focal length of the lens according to the absolute dual quadric coefficients and the projective projection matrix of the first image.
[0037] As an alternative implementation, in the first aspect of the present invention, the method further includes:
[0038] From all the first images, determine one of the images as the reference coordinate system image, where the reference coordinate system image includes the first acquired first image or one of the non-first acquired first images;
[0039] For at least one of the first images other than the reference coordinate system image among all the first images, determine the internal parameter matrix corresponding to the first image according to the width and height of the first image and the focal length of the lens; determine the essential matrix of the first image according to the internal parameter matrix corresponding to the first image and the fundamental matrix of the first image, and decompose the essential matrix of the first image to obtain the rotation matrix and translation vector corresponding to the first image; obtain the projection matrix corresponding to the first image according to the rotation matrix and translation vector corresponding to the first image and the internal parameter matrix corresponding to the first image;
[0040] For at least one of the first images other than the reference coordinate system image among all the first images, perform spatial triangulation on each third feature matching point in the first image according to the projection matrix of the first image to obtain the map point corresponding to the third feature matching point in the current scene;
[0041] Determine the initial map of the current scene according to all the map points corresponding to all the third feature matching points corresponding to all the first images and each first image;
[0042] Wherein, the initial map of the current scene is composed of a plurality of the map points and the first images acquired by all the lenses.
[0043] As an optional implementation manner, in the first aspect of the present invention, the method further includes:
[0044] Based on a pre-determined projection error method, optimize the focal length and principal point coordinates of each lens, the rotation matrix and translation vector corresponding to each first image acquired by each lens, and each map point in the map of the current scene to obtain the optimized focal length of each lens, the optimized rotation matrix and translation vector corresponding to each first image acquired by each lens, and the map points in the optimized map of the current scene.
[0045] A second aspect of the present invention discloses a map incremental reconstruction device applied to a flight scene, the device is applied to a map construction device or an image processing device, and the device includes:
[0046] An acquisition module, configured to perform an image acquisition operation based on at least one exposure point of the image processing device in the current scene after the image processing device completes lens calibration and obtains an initial map of the current scene; wherein, the image processing device is in a flying state, and the initial map of the current scene is determined by first images collected by the image processing device in a flying state through the lens of the image processing device at at least two of the exposure points;
[0047] A region division module, configured to perform a region division operation on any second image when the lens of the image processing device collects the second image at at least one of the exposure points, to obtain a plurality of sub-regions;
[0048] An extraction module, configured to extract a preset number of feature points for any sub-region of the second image;
[0049] A reconstruction module, configured to perform an incremental reconstruction operation on the initial map of the current scene according to all the feature points of the second image and all the feature points of a plurality of the first images determined in advance.
[0050] As an optional implementation manner, in the second aspect of the present invention, the specific manner in which the reconstruction module performs an incremental reconstruction operation on the initial map of the current scene according to all the feature points of the second image and all the feature points of a plurality of the first images determined in advance includes:
[0051] Determine the projection of the second image on the ground according to the second image;
[0052] From a plurality of the first images, screen out third images whose projections intersect with the projection of the second image according to the projections of the plurality of the first images determined in advance on the ground;
[0053] Match the feature points of the second image with the feature points of the third image to obtain a feature point matching result of the second image, and the feature point matching result corresponding to the second image includes a plurality of first feature matching points;
[0054] Screen out second feature matching points that do not have corresponding map points in the initial map of the current scene from all the first feature matching points included in the second image;
[0055] Perform a spatial triangulation operation on each of the second feature matching points to obtain a map point corresponding to the second feature matching point, and add all the map points corresponding to the second feature matching points to the initial map of the current scene.
[0056] As an optional implementation manner, in the second aspect of the present invention, the device further includes:
[0057] An acquisition module, configured to acquire the vertical distance between the image processing device and the ground when the second image is acquired, the internal parameter matrix of the lens for acquiring the second image, and the pose of the lens when the second image is acquired;
[0058] Wherein, the specific manner in which the reconstruction module determines the projection of the second image on the ground according to the second image includes:
[0059] Acquire the pixel coordinates of the four corner imaging points in the second image;
[0060] Based on the vertical distance corresponding to the second image, the internal parameter matrix corresponding to the second image, the pose corresponding to the second image, and the pixel coordinates of the four corner imaging points in the second image, calculate the projection of the second image on the ground.
[0061] As an optional implementation manner, in the second aspect of the present invention, the device further includes:
[0062] A calculation module, configured to, when all the map points corresponding to the second images acquired by the image processing device are added to the map of the current scene, for multiple imaging points of all the fourth images for which map construction has been completed, acquire the map coordinates of each imaging point and the positions of all the exposure points, and calculate a scaling factor according to the map coordinates of all the imaging points and the positions of all the exposure points, wherein all the fourth images include all the second images;
[0063] An alignment module, configured to use the scaling factor to scale the position of each exposure point among all the exposure points to obtain the scaled position of each exposure point;
[0064] The calculation module is further configured to calculate a target rotation matrix and a target translation vector between the 3D model and the geospatial based on the scaled positions of all the exposure points and the imaging point center coordinates of all the fourth images;
[0065] A first optimization module, configured to use the calculated target rotation matrix and the target translation vector to perform transformation and translation operations on the map of the current scene that has been reconstructed, so that all the map points corresponding to the current scene are aligned with the geographical coordinates.
[0066] As an optional implementation manner, in the second aspect of the present invention, the device further includes:
[0067] A second optimization module, configured to optimize the pose of the lens when collecting the second image and the map points corresponding to each of the second feature matching points based on each of the second feature matching points of the second image, the map points corresponding to the second feature matching points, and the internal parameter matrix of the lens that collected the second image, so as to obtain the optimized pose of the lens and the optimized map points;
[0068] Wherein, the specific manner in which the reconstruction module adds the map points corresponding to all the second feature matching points to the initial map of the current scene includes:
[0069] Adding the map points corresponding to all the second feature matching points of the optimized second image to the initial map of the current scene;
[0070] Wherein, the pose corresponding to each second image is used as the basis for incremental reconstruction of the initial map of the current scene.
[0071] As an optional implementation manner, in the second aspect of the present invention, the device further includes:
[0072] A judgment module, configured to judge whether the current condition of the current scene meets a globally determined optimization condition;
[0073] A third optimization module, configured to, when it is judged that the globally determined optimization condition is met, optimize the pose of the lens when collecting the corresponding second image and the map points corresponding to each of the second feature matching points based on the focal length of the lens of the calibrated second image, so as to obtain the optimized pose of the lens and the optimized map points, until all the map points corresponding to all the second images collected by the image processing device are added to the initial map of the current scene.
[0074] As an optional implementation manner, in the second aspect of the present invention, the acquisition module is further configured to, before performing an image acquisition operation based on at least one exposure point of the image processing device in the current scene, perform an image acquisition operation based on at least one exposure point of the image processing device in the current scene, so as to obtain metadata corresponding to each lens of the image processing device, wherein the metadata corresponding to each lens includes a first image collected by the lens at at least one exposure point, the position of the exposure point in the current scene, and the pose of the lens at the exposure point when collecting the first image, and the number of lenses of the image processing device is greater than or equal to 1;
[0075] The device further includes:
[0076] The calibration module is used to, for any one of the lenses, extract the feature points of each of the first images of the lens, and determine one of the first images from all the first images as the reference image; for any one of the first images other than the reference image among all the first images, match the feature points of the first image with the feature points of the reference image to obtain the feature point matching result corresponding to the first image, and the feature point matching result corresponding to each first image includes a plurality of third feature matching points; and calculate the fundamental matrix between the first image and the reference image according to the feature point matching result corresponding to the first image; determine the projective projection matrix of the first image according to the fundamental matrix of the first image; determine the absolute dual quadric coefficient according to the projective projection matrices of all the first images and the properties of the absolute dual quadric coefficient; and calibrate the focal length of the lens according to the absolute dual quadric coefficient and the projective projection matrix of the first image.
[0077] As an optional implementation manner, in the second aspect of the present invention, the device further includes:
[0078] The map construction module is used to determine one of the first images as the reference coordinate system image from all the first images, and the reference coordinate system image includes the first acquired first image or one of the non-first acquired first images;
[0079] The map construction module is further used to, for at least one of the first images other than the reference coordinate system image among all the first images, determine the internal parameter matrix corresponding to the first image according to the width and height of the first image and the focal length of the lens; determine the essential matrix of the first image according to the internal parameter matrix corresponding to the first image and the fundamental matrix of the first image, and decompose the essential matrix of the first image to obtain the rotation matrix and translation vector corresponding to the first image; and reconstruct the projective projection matrix of the first image according to the rotation matrix and translation vector corresponding to the first image and the internal parameter matrix corresponding to the first image to obtain the projection matrix of the first image.
[0080] The map construction module is further used to, for at least one of the first images other than the reference coordinate system image among all the first images, perform spatial triangulation on each of the third feature matching points in the first image according to the projection matrix of the first image to obtain the map points of the third feature matching points in the current scene; and determine the initial map of the current scene according to the map points corresponding to all the third feature matching points corresponding to all the first images.
[0081] Wherein, the initial map of the current scene is composed of a plurality of the map points and the first images acquired by all the lenses.
[0082] As an alternative embodiment, in the second aspect of the present invention, the device further includes:
[0083] A fourth optimization module, configured to optimize the focal length and principal point coordinates of each lens, the rotation matrix and translation vector corresponding to each first image collected by each lens, and each map point in the map of the current scene based on a pre-determined projection error method, so as to obtain the optimized focal length of each lens, the optimized rotation matrix and translation vector corresponding to each first image collected by each lens, and the map points in the optimized map of the current scene.
[0084] The third aspect of the present invention discloses another map incremental reconstruction device applied to a flight scene. The device is applied to a map construction device or an image processing device, and the device includes:
[0085] A memory storing executable program code;
[0086] A processor coupled to the memory;
[0087] The processor calls the executable program code stored in the memory and executes some or all of the steps in any one of the map incremental reconstruction methods applied to a flight scene disclosed in the first aspect of the present invention.
[0088] The sixth aspect of the present invention discloses a computer storage medium. The computer storage medium stores computer instructions, which are used to execute some or all of the steps in any one of the map incremental reconstruction methods applied to a flight scene disclosed in the first aspect of the present invention when being called.
[0089] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:
[0090] In an embodiment of the present invention, after the image processing device completes lens calibration and obtains the initial map of the current scene, an image acquisition operation is performed based on at least one exposure point of the image processing device in the current scene; wherein, the image processing device is in a flying state, and the initial map of the current scene is determined by first images collected by the lens of the image processing device at at least two exposure points when the image processing device is in a flying state; when the lens of the image processing device collects any second image at at least one exposure point in the current scene, a region division operation is performed on the second image to obtain a plurality of sub-regions; for any sub-region of the second image, a preset number of feature points are extracted; according to all the feature points of the second image and all the feature points of multiple first images determined in advance, an incremental reconstruction operation is performed on the initial map of the current scene. It can be seen that in the present invention, the image processing device that has completed lens calibration and is in a flying state performs image acquisition on the current scene, and performs region division on each collected image, and extracts the corresponding number of feature points from the images after region division, improving the extraction accuracy and efficiency of the feature points. The obtained point cloud of feature points is denser, more accurate and more uniform, and has higher adaptability to scenes where the size and illumination change in the scene. Finally, all the feature points of the collected images are matched and analyzed with the feature points of the images used for initial map construction, improving the matching accuracy of the feature points. Finally, based on the analysis results of accurate matching, an incremental reconstruction of the initial map is performed, improving the accuracy and reliability of the map incremental reconstruction, so as to obtain a map with a higher matching degree with the current scene. BRIEF DESCRIPTION OF THE DRAWINGS
[0091] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0092] Figure 1 It is a schematic flowchart of a method for incremental map reconstruction applied to a flight scene disclosed in an embodiment of the present invention;
[0093] Figure 2 It is a schematic flowchart of a method for lens calibration of an image processing device applied to a flight scene disclosed in an embodiment of the present invention;
[0094] Figure 3 It is a schematic structural diagram of a device for incremental map reconstruction applied to a flight scene disclosed in an embodiment of the present invention;
[0095] Figure 4 It is a schematic structural diagram of another device for incremental map reconstruction applied to a flight scene disclosed in an embodiment of the present invention;
[0096] Figure 5 It is a schematic structural diagram of another map incremental reconstruction device applied to a flight scenario disclosed in an embodiment of the present invention. Specific embodiments
[0097] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0098] The terms "first", "second", etc. in the specification and claims of the present invention and the above accompanying drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product or terminal that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or terminals.
[0099] Referring to "embodiment" herein means that a specific feature, structure or characteristic described in connection with the embodiment can be included in at least one embodiment of the present invention. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0100] The present invention discloses a map incremental reconstruction method and device applied to a flight scenario. An image processing device that has completed lens calibration and is in a flight state collects images of the current scene, and divides each collected image into regions. Corresponding numbers of feature points are extracted from the images after region division, which improves the accuracy and efficiency of feature point extraction. The obtained point cloud of feature points is denser, more accurate and more uniform, and has higher adaptability to scenes with size changes and lighting changes in the scene. Finally, all the feature points of the collected images are matched and analyzed with the feature points of the images used for initial map construction, which improves the matching accuracy of the feature points. Finally, based on the analysis results of accurate matching, the initial map is incrementally reconstructed, which improves the accuracy and reliability of map incremental reconstruction, so as to obtain a map with a higher matching degree with the current scene. The following will be described in detail respectively.
[0101] Embodiment 1
[0102] See also Figure 1 , Figure 1 : is a flowchart of a method for incremental map reconstruction applied to flight scenes disclosed in an embodiment of the present invention. Figure 1 The described method can be applied to any Figure 3 In the flight scene constructed by the computer, the computer and the computer are connected, and the flight scene has a corresponding target device or a central control center for controlling the target device, wherein the central control center includes a server (local server or cloud server) or a control platform. When the target device has processing functions such as flight, data acquisition and image analysis, the device may include the target device; when the target device has flight and data acquisition functions, the device may include a central control center for controlling the target device, and the target device transmits the collected data to the central control center for analysis and controls the target device. Furthermore, the target device may be an image processing device or a map construction device; the map construction device includes an image processing device and a flight device. At this time, the image processing device is integrated into the flight device with flight function and is used to collect data. Figure 1 As shown, the incremental map reconstruction method applied to the flight scene may include the following operations:
[0103] 101. After the image processing device completes lens calibration and obtains an initial map of the current scene, an image acquisition operation is performed based on at least one exposure point of the image processing device in the current scene; wherein, the image processing device is in a flying state, and the initial map of the current scene is determined by a first image acquired by the lens of the image processing device at at least two exposure points of the current scene when the image processing device is in a flying state.
[0104] In the embodiment of the present invention, optionally, the lens calibration of the image processing device may be pre-calibrated, or may be calibrated when performing a flight mission, that is, the lens calibration of the image processing device may also be calibrated based on the first image collected at all exposure points. Further optionally, the number of lenses of the image processing device is greater than or equal to 1. When the number of lenses of the image processing device is greater than 1, each lens is calibrated separately, and the calibration is performed based on the first image collected by the lens at all exposure points.
[0105] It should be noted that the exposure points used when constructing the initial map of the current scene may be the same as or partially the same as the exposure points used when reconstructing the initial map of the current scene.
[0106] 102. When a lens of an image processing device captures any second image at at least one exposure point of a current scene, a region division operation is performed on the second image to obtain a plurality of sub-regions.
[0107] 103. For any sub-region of the second image, extract a preset number of feature points.
[0108] In an embodiment of the present invention, optionally, each second image can be divided into a preset number of regions, such as 8 sub-regions. And a preset number (such as 1000) of feature points are respectively extracted from each sub-region. Among them, the type of feature points is not limited, such as SIFT feature points. In this way, by extracting feature points with better matching performance than ORB feature points, the accuracy of feature point matching between frames is further improved, and the accuracy of map reconstruction is improved; and compared with the sparse optical flow method used in SLAM, the positions of the obtained feature points are more accurate, support a larger inter-frame interval, and are more suitable for scenes with greater illumination changes.
[0109] 104. According to all the feature points of the second image and all the feature points of a plurality of previously determined first images, perform an incremental reconstruction operation on the initial map of the current scene.
[0110] In an embodiment of the present invention, when constructing the map of the current scene, the feature points of all the first images have been extracted. Among them, the specific extraction process of the feature points of all the first images can refer to the description of the extraction process of the feature points of the second image above, and will not be repeated here.
[0111] It should be noted that as long as a new second image is captured, the incremental reconstruction operation of the map corresponding to this second image is performed. Further, when there are multiple cameras, the incremental reconstruction operations of the maps corresponding to the second images collected by each camera can be performed simultaneously.
[0112] It can be seen that implementing Figure 1 The described method collects images of the current scene through an image processing device that has completed camera calibration and is in a flying state, and divides each captured image into regions. Corresponding numbers of feature points are respectively extracted from the images after region division, which improves the accuracy and efficiency of feature point extraction. The obtained point cloud is denser, more accurate, and more uniform, and has higher adaptability to scenes with size changes and illumination changes in the scene. Finally, all the feature points of the captured images are matched and analyzed with the feature points of the images used for initial map construction, which improves the matching accuracy of the feature points. Finally, based on the analysis results of accurate matching, an incremental reconstruction is performed on the initial map, which improves the accuracy and reliability of the map incremental reconstruction, so as to obtain a map with a higher matching degree with the current scene; and supports the incremental map reconstruction of multi-camera image processing devices, further improving the efficiency and accuracy of the map incremental reconstruction; and by recording the poses of the exposure points, matching in a limited space instead of global matching is realized, further reducing the computational amount, which is more conducive to the incremental reconstruction of the map.
[0113] In an embodiment of the present invention, as an alternative implementation, an incremental reconstruction operation is performed on the initial map of the current scene according to all the feature points of the second image and all the feature points of multiple previously determined first images, including:
[0114] Determine the projection of the second image on the ground according to the second image;
[0115] According to the projections of multiple previously determined first images on the ground, screen out a third image from the multiple first images whose projection intersects with the projection of the second image;
[0116] Match the feature points of the second image with the feature points of the third image to obtain a feature point matching result of the second image, and the feature point matching result corresponding to the second image includes multiple first feature matching points;
[0117] Screen out second feature matching points that do not have corresponding map points in the initial map of the current scene from all the first feature matching points included in the second image;
[0118] Perform a spatial triangulation operation on each second feature matching point to obtain the map point corresponding to the second feature matching point, and add the map points corresponding to all the second feature matching points to the initial map of the current scene.
[0119] In an embodiment of the present invention, optionally, the ground projection of each first image in all the first images can be directly compared with the projection of the second image to obtain a third image; alternatively, all the first images containing the same ground observation object can be screened out from the multiple first images first, and then the second image is compared with each first image containing the same ground observation object respectively to obtain a third image. In this way, by comparing with the images containing the same ground observation object at the same time, the matching calculation amount is greatly reduced, and the matching efficiency is improved while ensuring the matching accuracy, thereby further improving the matching accuracy of the feature matching points.
[0120] In addition, it should be noted that the number of third images may be greater than or equal to 1. When it is greater than 1, the same operation is performed on each third image to reconstruct the map of the current scene.
[0121] It can be seen that the embodiment of the present invention can also perform an incremental reconstruction of the initial map by performing a projection intersection (overlap) analysis on the newly entered image and the image that has already built the initial map, matching the feature points of the newly entered image with the image having a projection intersection, and performing a spatial triangulation on the feature matching points of the map points not included in the initial map, improving the efficiency and accuracy of the incremental reconstruction of the map.
[0122] In an alternative embodiment, the method may further include the following steps: Obtain the vertical distance between the image processing device and the ground when collecting the second image, the internal parameter matrix of the lens used to collect the second image, and the pose of the lens when the second image is collected; Among them, determining the projection of the second image on the ground according to the second image includes: Obtain the pixel coordinates of the four corner imaging points in the second image; Based on the vertical distance corresponding to the second image, the internal parameter matrix corresponding to the second image, the pose corresponding to the second image, and the pixel coordinates of the four corner imaging points in the second image, calculate the projection of the second image on the ground. In this optional embodiment, optionally, the pixel coordinates of each imaging point of each second image are converted into world coordinates through the following conversion formula, where the conversion formula is: ; ; In the formula, M is the world coordinate corresponding to each imaging point; N is the pixel coordinate of each imaging point; R and t are the rotation matrix and translation vector of the lens when collecting the corresponding second image, that is, the pose; K is the internal parameter matrix corresponding to the second image, where the internal parameter matrix of the second image is determined based on the focal length of the corresponding lens and the width and height of the second image; k is the coefficient determined based on the above vertical distance, such as the above vertical distance divided by the Z-axis coordinate value obtained by R -1 (K -1 N - t).
[0123] It can be seen that the embodiment of the present invention can comprehensively analyze the pixel coordinates of the imaging points of the newly entered image, the pose of the lens, the vertical distance between the image processing device and the ground, and the internal parameter matrix of the lens, improving the analysis efficiency and accuracy of the projection of the newly entered image on the ground, which is conducive to improving the analysis accuracy of the intersection of image projections.
[0124] It should be noted that for the analysis method of the projection corresponding to the first image of the already built map, please refer to the analysis method of the projection corresponding to the second image, which will not be elaborated here.
[0125] In another optional embodiment, the method may further include the following steps: When all the map points corresponding to the second images collected by the image processing device are added to the map of the current scene, for multiple imaging points of all the fourth images that have completed map construction, obtain the map coordinates of each imaging point and the positions of all exposure points; According to the map coordinates of all imaging points and the positions of all exposure points, calculate the scaling factor, and use the scaling factor to scale the position of each exposure point among all exposure points to obtain the scaled position of each exposure point; Calculate the target rotation matrix and the target translation vector between the 3D model and the geospatial based on the positions of all scaled exposure points and the center coordinates of the imaging points of all the fourth images; Use the calculated target rotation matrix and the target translation vector to perform transformation and translation operations on the map of the currently reconstructed current scene, so that the poses of all map points and all lenses corresponding to the current scene are aligned with the geographical coordinates.
[0126] In this optional embodiment, optionally, the target rotation matrix and the target translation vector are calculated through the following formula, where the formula is as follows: ; In the formula, is the position of the scaled exposure point corresponding to the i-th fourth image, is the center coordinate of the imaging point of the i-th fourth image, and N is the number of all fourth images. By using the least squares method to find the minimum values of the rotation matrix R and the translation vector t in the above formula as the target rotation matrix and the translation vector.
[0127] It can be seen that in this optional embodiment, after the map points of all images of the current scene are reconstructed, the scaling factor is calculated based on the positions of all exposure points and the map coordinates of the imaging points, the positions of the exposure points are scaled, and based on the positions of the scaled exposure points and the center coordinates of the imaging points of all images, the target rotation matrix and the target translation vector between the 3D model and the geospatial are calculated, and the map is rotated and translated as a whole, so that the converted map is consistent with the actual geographical coordinate system, improving the accuracy and reliability of the incremental 3D construction of the actual map of the current scene.
[0128] In yet another optional embodiment, the method may further include the following steps: Optimize the pose of the lens when collecting the second image and the map points corresponding to each second feature matching point that are pre-determined based on each second feature matching point of the second image, the map point corresponding to the second feature matching point, and the focal length of the lens that collected the second image, to obtain the optimized pose of the lens and the optimized map points; Among them, adding the map points corresponding to all second feature matching points to the initial map of the current scene includes: Adding the map points corresponding to all second feature matching points of the optimized second image to the initial map of the current scene; Among them, the pose corresponding to each second image is used as the basis for the incremental reconstruction of the initial map of the current scene.
[0129] In this optional embodiment, for each lens, its pose is optimized using the second image collected each time, and the optimized pose is used as the pose for the next tracking.
[0130] It can be seen that in this alternative embodiment, by optimizing the incremental part of the map points corresponding to the newly acquired images, the accuracy of map point determination is improved. Also, by simultaneously optimizing the pose of the lens of the newly acquired image, i.e., the extrinsic parameters, there is no need to rely on visual odometry to estimate the pose through motion, reducing the occurrence of 3D incremental reconstruction failures due to tracking failures. And there is no need to reposition the newly acquired images, which is beneficial to improving the robustness of incremental reconstruction. Also, since there is no need to provide the pose through motion tracking, the acquisition interval duration of the images can be increased, i.e., the number of processed frames for the same flight distance is reduced, further improving the analysis efficiency of the pose and map points during the map incremental reconstruction process and improving the loop detection efficiency. Figure 3 In another alternative embodiment, the method may further include the following steps:
[0131] Judge whether the current conditions of the current scene meet the pre-determined global optimization conditions; When it is judged that the global optimization conditions are met, based on the focal length of the lens of the second image obtained by calibration, optimize the pose of the lens and the map points corresponding to each second feature matching point determined when collecting the second image, to obtain the optimized pose of the lens and the optimized map points, until all the map points corresponding to all the second images collected by the image processing device are added to the initial map of the current scene. In this alternative embodiment, optionally, the global optimization conditions may be that the flight device turns around or flies out of the measurement area during the execution of the flight mission, i.e., when it is monitored that the device turns around or flies out of the measurement area, it means that global optimization is required; it may also be that when a global optimization request is detected, it means that global optimization is required; it may also be that when the large screen finishes the flight mission, it means that global optimization is required, or it may be other, without limitation.
[0132] In this alternative embodiment, when it is judged that the global optimization conditions are not met, there is no need to perform global optimization and continue with the map incremental reconstruction operation based on the newly acquired second images.
[0133] It can be seen that in this alternative embodiment, during the process of map incremental reconstruction based on newly acquired images, if it is judged that the global optimization conditions are met, then perform global optimization on all the map points and the poses of all the lenses that have completed map construction or reconstruction in the current flight mission, so that they are more consistent with the current actual scene, further improving the accuracy and reliability of the obtained map; and there is no need to perform optimization after reconstruction at each exposure point, improving the efficiency of global optimization, and further improving the determination efficiency of a more accurate and reliable map.
[0134]
[0135] In yet another alternative embodiment, the method may further include the following steps: Perform an image acquisition operation based on at least one exposure point of the image processing device in the current scene to obtain metadata corresponding to each lens of the image processing device. Each piece of metadata corresponding to a lens includes a first image captured by the lens at at least one exposure point in the current scene, the position of the exposure point in the current scene, and the pose of the lens at the exposure point when capturing the first image. The number of lenses of the image processing device is greater than or equal to 1. For any lens, extract the feature points of each first image of the lens, and determine one of the images from all the first images as the reference image. For any first image other than the reference image among all the first images, match the feature points of the first image with the feature points of the reference image to obtain the feature point matching result corresponding to the first image. Each feature point matching result corresponding to a first image includes a plurality of third feature matching points; and calculate the fundamental matrix between the first image and the reference image according to the feature point matching result corresponding to the first image. Determine the projective projection matrix of the first image according to the fundamental matrix of the first image; determine the absolute dual quadric coefficient according to the projective projection matrices of all the first images and the properties of the absolute dual quadric coefficient. Calibrate the focal length of the lens according to the absolute dual quadric coefficient and the projective projection matrix of the first image.
[0136] In the embodiments of the present invention, optionally, the number of first images captured by each lens at each exposure point may be 1 or greater than 1.
[0137] In the embodiments of the present invention, optionally, after the metadata is acquired, the metadata may be recorded in the storage unit of the image processing device, or the metadata may be transmitted to a ground storage unit, such as a ground control platform communicatively connected to the image processing device.
[0138] In the embodiments of the present invention, optionally, for the reference image of each lens, it may be randomly selected from all the first images of the lens, or specifically selected, such as selecting the first captured first image as the reference image. The embodiments of the present invention do not make any limitations.
[0139] Regarding the extraction method of the feature points of each first image in the embodiments of the present invention, for details, refer to the specific extraction method of the feature points of the second image, which will not be elaborated here. Optionally, the number of extracted feature points may be equal to or different from the number of feature points of the second image.
[0140] It should be noted that when performing lens calibration, the origin of the image coordinate system can be selected as the center point of the first image, or the upper left corner point of the image, or other corner points.
[0141] In an embodiment of the present invention, further optionally, for any first image, the RANSAC algorithm is used to exclude incorrect matches of all its third feature matching points, so as to further improve the acquisition accuracy of the third feature matching points, thereby facilitating further improvement of the calibration accuracy of the lens and the construction accuracy of the initial map.
[0142] In an embodiment of the present invention, for any first image, when analyzing its corresponding fundamental matrix, the coordinate origin used by the third feature matching points is the same as the way of the image coordinate system origin selected during the aforementioned lens calibration, such as both being the image center.
[0143] In an embodiment of the present invention, optionally, the number of first images that each lens needs to collect is related to the position of the origin of the image coordinate system selected for the aforementioned lens calibration and the nature of the absolute dual quadric surface coefficients. Specifically, the absolute dual quadric surface is a 4*4 symmetric matrix, with a total of 10 parameters and 8 degrees of freedom. One image can provide 4 pairs of constraints. For each lens, if the image center is used as the coordinate origin, the number of first images collected is at least 4. If the image center is not used as the coordinate origin, then the number of images required is greater than 4. Among them, optionally, all the first images collected by each lens can be collected at one exposure point or at multiple exposure points. For example, if 1 exposure point corresponds to 1 first image, then at least 4 exposure points need to be collected respectively. It should be noted that it is preferred to select the image center as the coordinate origin, which can calibrate the lens with fewer images, improve the calibration efficiency while achieving accurate calibration, and thus facilitate improving the efficiency of the incremental reconstruction of the map of the current scene.
[0144] It can be seen that this optional embodiment automatically calibrates the lens of the image processing device by collecting images of the current scene during the execution of the flight mission, without the need for (manual) calibration before use, improving the calibration efficiency and accuracy. Even for a zoom camera, the calibrated lens is adapted to the current scene, thereby facilitating improving the accuracy of subsequent image collection, further facilitating improving the efficiency and accuracy of map incremental reconstruction, and improving the operation convenience of the operator; and by recording the position of each exposure point and the attitude of the lens, the calculation amount of subsequent lens calibration is reduced, further improving the lens calibration efficiency; and optionally only calibrating the focal length of the lens, greatly simplifying the analysis of the absolute dual quadric surface coefficients, while ensuring the analysis accuracy of the absolute dual quadric surface coefficients, improving the analysis efficiency of the absolute dual quadric surface coefficients.
[0145] In yet another alternative embodiment, the method may further include the following steps: From all the first images, determine one of the images as a reference coordinate system image, which includes the first image collected for the first time or one of the first images that is not collected for the first time; For any one of the first images other than the reference coordinate system image among all the first images, according to the width and height of the first image and the focal length of the lens, determine the internal parameter matrix corresponding to the first image; according to the internal parameter matrix corresponding to the first image and the fundamental matrix of the first image, determine the essential matrix of the first image, and decompose the essential matrix of the first image to obtain the rotation matrix and translation vector corresponding to the first image; according to the rotation matrix and translation vector corresponding to the first image and the internal parameter matrix corresponding to the first image, reconstruct the projective projection matrix of the first image to obtain the projection matrix of the first image; For any one of the first images other than the reference coordinate system image among all the first images, perform spatial triangulation on each third feature matching point in the first image according to the projection matrix of the first image to obtain the map point of the third feature matching point in the current scene. Among them, the map point corresponding to each third feature matching point is the actual position of the third feature matching point in the current scene; Determine the initial map of the current scene according to all the map points corresponding to all the third feature matching points corresponding to all the first images and each first image; Among them, the initial map of the current scene is composed of multiple map points and the first images collected by all the lenses. Further, each map point in the initial map of the current scene corresponds to at least one third feature matching point.
[0146] In this alternative embodiment, optionally, for any one of the first images, when calculating the corresponding essential matrix, the coordinate origin of the third feature matching point can be the image center, or the upper left corner of the image, or the upper right corner, without limitation. Among them, for the coordinate origin at the upper left corner of the image, the internal parameter matrix corresponding to each first image can be: ; Where K is the internal parameter matrix corresponding to the first image, f is the focal length of the lens for collecting the first image, and w and h are the width and height of the first image.
[0147] In this alternative embodiment, optionally, for any one of the third feature matching points, obtain its coordinates in the first image and perform spatial triangulation to obtain the corresponding map point.
[0148] It can be seen that in this alternative embodiment, for any lens, by separately performing internal parameter matrix, essential matrix, and reconstruction projection matrix on the remaining images and the selected reference image, and then analyzing the map points of the actual positions of the matching feature points in the current scene, and then synthesizing the map points corresponding to all lenses to determine the initial map of the current scene, the accuracy and efficiency of the construction (initialization) of the initial map of the current scene are improved, which is beneficial to improving the accuracy and efficiency of the incremental reconstruction of the initial map.
[0149] In yet another alternative embodiment, the method may further include the following steps: Based on a pre-determined projection error method, optimize the focal length and principal point coordinates of each lens, the rotation matrix and translation vector corresponding to each first image collected by each lens, and each map point in the initial map of the current scene, to obtain the optimized focal length of each lens, the optimized rotation matrix and translation vector corresponding to each first image collected by each lens, and the map points in the optimized initial map of the current scene.
[0150] In this alternative embodiment, for any map point, obtain the feature points projected by the map point on each first image. Optionally, by constructing the following optimization formula for the coordinates of each map point, the focal length and principal point coordinates of each lens, the rotation matrix and translation vector corresponding to each first image collected by each lens, and each map point in the initial map of the current scene, optimize the focal length of each lens, the rotation matrix and translation vector corresponding to each first image collected by each lens, and the map points in the initial map of the current scene, where the optimization formula is as follows: ; In the formula, m is the number of lenses of the image processing device, such as 5; n is the number of first images collected by each lens; t is the number of map points corresponding to each first image; X j is the j-th map point on each first image; x k,i,j is the pixel coordinate of the feature point projected by the map point X j on the i-th first image collected by the k-th lens, as the optimized observation value; z k,i,j is the depth obtained by translation and rotation of the map point X j , that is, the z-axis coordinate of the map point; f k,x , f k,y are the focal length coordinates of the k-th lens, c k,x and c k,y are the principal point coordinates of the k-th lens, R k,i and t k,i are the rotation matrix and translation vector corresponding to the i-th first image of the k-th lens respectively. Among them, X j , f k,x, f k,y , c k,x and c k,y , R k,i and t k,i are respectively the map points of the initial map to be optimized, the focal length of the lens, the pixel coordinates, the rotation matrix and the translation vector of the exposure points.
[0151] Optionally, each lens collects 1 first image at each exposure point, and optimizes the pose of the first image, that is, optimizes the pose of the exposure point corresponding to the first image. Among them, each pose is composed of the corresponding rotation matrix and translation vector.
[0187] In this optional embodiment, optionally, the projection error method can be any optimization method obtained based on the least squares method, such as the Bundle Adjustment (BA) algorithm.
[0188] It can be seen that after obtaining the initial map of the current scene and completing the initial calibration of the lens, this optional embodiment further constructs corresponding optimization formulas for the posture of each lens at each exposure point, the focal length of each lens and the image principal point coordinates and the map points of the initial map, and calculates the minimum value of the optimization formula based on the projection error method obtained based on the least squares principle, thereby optimizing the initial map, exposure points and lens focal lengths, improving the accuracy of initial map construction and the accuracy of exposure point and lens focal length calibration, so as to further adapt to the current real scene, which is beneficial to further improve the accuracy of subsequent image acquisition, and further help to further improve the accuracy and reliability of subsequent incremental map construction.
[0189] Embodiment 2
[0190] See also Figure 2 , Figure 2 1 is a flowchart of a lens calibration method for an image processing device applied to a flight scene disclosed in an embodiment of the present invention. Figure 2 The described method can be applied to any Figure 3 In the flight scene constructed by the computer, the computer and the computer are connected, and the flight scene has a corresponding target device or a central control center for controlling the target device, wherein the central control center includes a server (local server or cloud server) or a control platform. When the target device has processing functions such as flight, data acquisition and image analysis, the device may include the target device; when the target device has flight and data acquisition functions, the device may include a central control center for controlling the target device, and the target device transmits the collected data to the central control center for analysis and controls the target device. Furthermore, the target device may be an image processing device or a map construction device; the map construction device includes an image processing device and a flight device. At this time, the image processing device is integrated into the flight device with flight function and is used to collect data. Figure 2 As shown, the calibration method may include the following operations:
[0191] 201. Perform an image acquisition operation based on at least one exposure point of the image processing device in the current scene to obtain metadata corresponding to each lens of the image processing device, wherein the metadata corresponding to each lens includes a first image acquired by the lens at at least one exposure point of the current scene, a position of the exposure point in the current scene, and a posture of the lens at the exposure point when acquiring the first image, and the number of lenses of the image processing device is greater than or equal to 1.
[0192] In an embodiment of the present invention, optionally, the lens calibration of the image processing device can be pre-calibrated or can be calibrated during the execution of a flight mission, that is, the lens calibration of the image processing device can also be calibrated based on the first images collected at all exposure points. Further optionally, the number of lenses of the image processing device is greater than or equal to 1. When the number of lenses of the image processing device is greater than 1, each lens is calibrated separately and is calibrated based on the first images collected by the lens at all exposure points.
[0193] 202. For any lens, extract the feature points of each first image of the lens, and determine one of the images from all the first images as the reference image.
[0194] 203. For any first image other than the reference image among all the first images, match the feature points of the first image with the feature points of the reference image to obtain the feature point matching result corresponding to the first image. Each feature point matching result corresponding to a first image includes a plurality of third feature matching points; and based on the feature point matching result corresponding to the first image, calculate the fundamental matrix between the first image and the reference image.
[0195] 204. Determine the projective projection matrix of the first image according to the fundamental matrix of the first image; determine the absolute dual quadric coefficient according to the projective projection matrices of all the first images and the properties of the absolute dual quadric coefficient.
[0196] 205. Calibrate the focal length of the lens according to the absolute dual quadric coefficient and the projective projection matrix of the first image.
[0197] It can be seen that the method described in the embodiment Figure 2 collects images of the current scene during the execution of a flight mission to automatically calibrate the lens of the image processing device, without the need for (manual) calibration before use, improving the efficiency and accuracy of calibration. Even for a zoom camera, the calibrated lens is adapted to the current scene, which is beneficial to improving the accuracy of subsequent image collection, and further beneficial to improving the efficiency and accuracy of map incremental reconstruction, and improving the operation convenience of the operator; and by recording the position of each exposure point and the attitude of the lens, the computational amount of subsequent lens calibration is reduced, further improving the efficiency of lens calibration; and optionally, only the focal length of the lens is calibrated, greatly simplifying the analysis of the absolute dual quadric coefficient, while ensuring the accuracy of the analysis of the absolute dual quadric coefficient and improving the analysis efficiency of the absolute dual quadric coefficient.
[0198] It should be noted that for other technical contents regarding the lens calibration of the image processing device, please refer to the specific description of the relevant content in Embodiment 1, which will not be elaborated in the embodiments of the present invention.
[0199] Embodiment 3
[0200] See also Figure 3 , Figure 3 The present invention discloses a structural diagram of a device for incremental reconstruction of a map applied to a flight scene. The device can be applied to any Figure 3 The device is used in a flight scene constructed by a certain dimension, and the device is applied to a target device or a central control center for controlling the target device, wherein the central control center includes a server (local server or cloud server) or a control platform. When the target device has functions such as flight, data acquisition and image analysis, the device may include the target device; when the target device has flight and data acquisition functions, the device may include a central control center for controlling the target device, and the target device transmits the collected data to the central control center for analysis and controls the target device. Furthermore, the target device may be an image processing device or a map construction device; the map construction device includes an image processing device and a flight device, and at this time, the image processing device is integrated in the flight device with a flight function and is used to collect data. Figure 3 As shown, the device comprises:
[0201] The acquisition module 301 is used to perform an image acquisition operation based on at least one exposure point of the current scene by the image processing device after the image processing device completes lens calibration and obtains the initial map of the current scene; wherein the image processing device is in a flying state, and the initial map of the current scene is determined by a first image acquired by the lens of the image processing device at at least two exposure points of the current scene when the image processing device is in a flying state;
[0202] A region division module 302 is configured to perform a region division operation on any second image to obtain a plurality of sub-regions when the lens of the image processing device captures any second image at at least one exposure point of the current scene;
[0203] An extraction module 303 is used to extract a preset number of feature points from any sub-region of the second image;
[0204] The reconstruction module 304 is used to perform an incremental reconstruction operation on the initial map of the current scene according to all feature points of the second image and all feature points of the predetermined plurality of first images.
[0205] It can be seen that implementation Figure 3The described map incremental reconstruction device applied to the flight scenario has completed the lens calibration and is in the flight state. The image processing device performs image acquisition on the current scene, divides each acquired image into regions, extracts a corresponding number of feature points from the images after region division, improves the accuracy and efficiency of feature point extraction, and the obtained point cloud is denser, more accurate and uniform. Moreover, it has higher adaptability to scenes with changing sizes and lighting in the scene. Finally, all feature points of the acquired images are matched and analyzed with the feature points of the images used for initial map construction, improving the matching accuracy of feature points. Finally, based on the analysis results of accurate matching, the initial map is incrementally reconstructed, improving the accuracy and reliability of map incremental reconstruction, thereby obtaining a map with a higher matching degree to the current scene.
[0206] In an embodiment of the present invention, optionally, the specific manner in which the reconstruction module 304 performs an incremental reconstruction operation on the initial map of the current scene according to all feature points of the second image and all feature points of a plurality of predetermined first images includes:
[0207] Determine the projection of the second image on the ground according to the second image;
[0208] According to the projections of a plurality of predetermined first images on the ground, screen out a third image from the plurality of first images whose projection intersects with the projection of the second image;
[0209] Match the feature points of the second image with the feature points of the third image to obtain a feature point matching result of the second image. The feature point matching result corresponding to the second image includes a plurality of first feature matching points;
[0210] Screen out second feature matching points that do not have corresponding map points in the initial map of the current scene from all the first feature matching points included in the second image;
[0211] Perform a spatial triangulation operation on each second feature matching point to obtain the map point corresponding to the second feature matching point, and add the map points corresponding to all second feature matching points to the initial map of the current scene.
[0212] It can be seen that implementing Figure 3 The described device can also perform incremental reconstruction on the initial map by performing projection intersection (overlap) analysis on the newly entered image and the image that has already built the initial map, matching the feature points of the newly entered image with the image whose projection intersects with it, and performing spatial triangulation on the feature matching points of the map points not included in the initial map, improving the efficiency and accuracy of map incremental reconstruction.
[0213] In an optional embodiment, Figure 4It is a schematic structural diagram of another map incremental reconstruction device applied to a flight scenario disclosed in an embodiment of the present invention. As Figure 4 shown, the device may further include:
[0214] An acquisition module 305, configured to acquire the vertical distance between the image processing device and the ground when the second image is acquired, the internal parameter matrix of the lens for acquiring the second image, and the pose of the lens when the second image is acquired;
[0215] Among them, the reconstruction module 304 determines the specific manner of the projection of the second image on the ground according to the second image, including:
[0216] Acquire the pixel coordinates of the four corner imaging points in the second image;
[0217] Based on the vertical distance corresponding to the second image, the internal parameter matrix corresponding to the second image, the pose corresponding to the second image, and the pixel coordinates of the four corner imaging points in the second image, calculate the projection of the second image on the ground.
[0218] It can be seen that implementing Figure 4 The described device can comprehensively analyze the pixel coordinates of the imaging points of the newly entered image, the pose of the lens, the vertical distance between the image processing device and the ground, and the internal parameter matrix of the lens, improving the analysis efficiency and accuracy of the projection of the newly entered image on the ground, and thus being beneficial to improving the analysis accuracy of the intersection of image projections.
[0219] In another optional embodiment, as Figure 4 shown, the device may further include:
[0220] A calculation module 306, configured to, when all the map points corresponding to all the second images acquired by the image processing device are added to the map of the current scene, for multiple imaging points of all the fourth images for which map construction has been completed, acquire the map coordinates of each imaging point and the positions of all exposure points, and calculate a scaling factor according to the map coordinates of all imaging points and the positions of all exposure points, where all the fourth images include all the second images;
[0221] An alignment module 307, configured to calculate a scaling factor according to the map coordinates of all imaging points and the positions of all exposure points, and use the scaling factor to scale the position of each exposure point among all exposure points to obtain the scaled position of each exposure point;
[0222] The calculation module 306 is further configured to calculate a target rotation matrix and a target translation vector between the three-dimensional model and the geospatial based on the scaled positions of all exposure points and the imaging point center coordinates of all the fourth images;
[0223] The first optimization module 308 is configured to use the calculated target rotation matrix and target translation vector to perform transformation and translation operations on the map of the currently reconstructed current scene, so that all map points corresponding to the current scene are aligned with the geographical coordinates.
[0224] It can be seen that implementing Figure 4 the described device can also, after completing the map point reconstruction of all images of the current scene, calculate a scaling factor based on the positions of all exposure points and the map coordinates of the imaging points, scale the positions of the exposure points, and calculate the target rotation matrix and target translation vector between the three-dimensional model and the geographical space based on the scaled positions of the exposure points and the central coordinates of the imaging points of all images, and perform overall rotation and translation on the map, so that the transformed map is consistent with the actual geographical coordinate system, improving the accuracy and reliability of the incremental three-dimensional construction of the actual map of the current scene.
[0225] In yet another alternative embodiment, as Figure 4 shown, the device may further include:
[0226] The second optimization module 309 is configured to optimize the pre-determined pose of the lens when collecting the second image and the map points corresponding to each second feature matching point based on each second feature matching point of the second image, the map point corresponding to the second feature matching point, and the focal length of the lens that collected the second image, to obtain the optimized pose of the lens and the optimized map points;
[0227] Among them, the specific manner in which the reconstruction module 304 adds the map points corresponding to all second feature matching points to the initial map of the current scene includes:
[0228] Adding the map points corresponding to all second feature matching points of the optimized second image to the initial map of the current scene;
[0229] Among them, the pose corresponding to each second image is used as the basis for incremental reconstruction of the initial map of the current scene.
[0230] It can be seen that implementing Figure 4 the described device can also improve the determination accuracy of map points by optimizing the incremental part of the map points corresponding to the newly entered images, and by simultaneously optimizing the pose of the lens of the newly entered image, that is, the external parameters, without relying on visual odometry to estimate the pose through motion, reducing the inability to achieve the ground caused by tracking failure Figure 3The occurrence of dimensional incremental reconstruction and the elimination of the need to reposition newly acquired images contribute to enhancing the robustness of incremental reconstruction. Additionally, since there is no need to provide poses through motion tracking, the acquisition interval of images can be increased, which means reducing the number of processed frames for the same flight distance, further improving the analysis efficiency of poses and map points during the map incremental reconstruction process and enhancing the loop detection efficiency.
[0231] In yet another alternative embodiment, as Figure 4 shown, the device may further include;
[0232] A judgment module 310, configured to judge whether the current conditions of the current scene meet the globally determined optimization conditions;
[0233] A third optimization module 311, configured to, when it is judged that the globally determined optimization conditions are met, optimize the pose of the lens and the map points corresponding to each second feature matching point when acquiring the second image based on the focal length of the lens of the calibrated second image, so as to obtain the optimized pose of the lens and the optimized map points until all the map points corresponding to all the second images acquired by the image processing device are added to the initial map of the current scene.
[0234] It can be seen that when implementing the Figure 4 described device, during the process of map incremental reconstruction based on newly acquired images, if it is judged that the globally determined optimization conditions are met, global optimization is performed on all the map points and the poses of all the lenses that have completed map construction or reconstruction in the current flight mission, so as to make them more consistent with the current actual scene, further improving the accuracy and reliability of the obtained map; and there is no need to perform optimization after reconstruction is completed at each exposure point, improving the efficiency of global optimization, and further enhancing the determination efficiency of a more accurate and reliable map.
[0235] In yet another alternative embodiment, as Figure 4 shown, the acquisition module 301 is further configured to, before performing the image acquisition operation based on at least one exposure point of the image processing device in the current scene, perform the image acquisition operation based on at least one exposure point of the image processing device in the current scene to obtain the metadata corresponding to each lens of the image processing device, where the metadata corresponding to each lens includes the first image acquired by the lens at at least one exposure point in the current scene, the position of the exposure point in the current scene, and the pose of the lens at the exposure point when acquiring the first image, and the number of lenses of the image processing device is greater than or equal to 1;
[0236] And as Figure 4 shown, the device may further include:
[0237] The calibration module 312 is configured to, for any lens, extract the feature points of each first image of the lens and determine one image from all the first images as the reference image;
[0238] For any first image other than the reference image among all the first images, match the feature points of the first image with the feature points of the reference image to obtain the feature point matching result corresponding to the first image. The feature point matching result corresponding to each first image includes multiple third feature matching points; and calculate the fundamental matrix between the first image and the reference image according to the feature point matching result corresponding to the first image;
[0239] Determine the projective projection matrix of the first image according to the fundamental matrix of the first image; determine the absolute dual quadric coefficient according to the projective projection matrices of all the first images and the properties of the absolute dual quadric coefficient;
[0240] Calibrate the focal length of the lens according to the absolute dual quadric coefficient and the projective projection matrix corresponding to the first image.
[0241] It can be seen that the described device can also collect images of the current scene during the execution of the flight mission to automatically calibrate the lens of the image processing device, without the need for (manual) calibration before use, improving the efficiency and accuracy of calibration. Even for a zoom camera, the calibrated lens is adapted to the current scene, which is beneficial to improving the accuracy of subsequent image collection, and further beneficial to improving the efficiency and accuracy of map incremental reconstruction, and improving the operation convenience of the operator; and by recording the position of each exposure point and the attitude of the lens, the computational amount of subsequent lens calibration is reduced, further improving the efficiency of lens calibration; and optionally, only calibrating the focal length of the lens greatly simplifies the analysis of the absolute dual quadric coefficient, improving the analysis efficiency of the absolute dual quadric coefficient while ensuring the analysis accuracy of the absolute dual quadric coefficient. Figure 4
[0242] In another optional embodiment, as Figure 4 shown, the device may further include:
[0243] The map construction module 313 is configured to determine one image from all the first images as the reference coordinate system image. The reference coordinate system image includes the first image collected for the first time or one of the first images collected non-first;
[0244] The map construction module 313 is further configured to, for any first image among all the first images except the reference coordinate system image, determine the internal parameter matrix corresponding to the first image according to the width and height of the first image and the focal length of the lens; determine the essential matrix of the first image according to the internal parameter matrix corresponding to the first image and the fundamental matrix of the first image, and decompose the essential matrix of the first image to obtain the rotation matrix and translation vector corresponding to the first image; and reconstruct the projective projection matrix of the first image according to the rotation matrix and translation vector corresponding to the first image and the internal parameter matrix corresponding to the first image to obtain the projection matrix of the first image.
[0245] The map construction module 313 is further configured to, for any first image among all the first images except the reference coordinate system image, perform spatial triangulation on each third feature matching point in the first image according to the projection matrix of the first image to obtain the map point of the third feature matching point in the current scene. Wherein, the map point corresponding to each third feature matching point is the actual position of the third feature matching point in the current scene.
[0246] Determine the initial map of the current scene according to all the map points corresponding to all the third feature matching points corresponding to all the first images and each first image.
[0247] Wherein, the initial map of the current scene consists of multiple map points and the first images collected by all the lenses. Further, each map point in the initial map of the current scene corresponds to at least one third feature matching point.
[0248] It can be seen that the implemented Figure 4 The described device can also first analyze, for any lens, the internal parameter matrix, essential matrix, and reconstructed projection matrix by respectively comparing the remaining images with the selected reference image, and then analyze the map points of the matching feature points at the actual positions in the current scene, and then synthesize the map points corresponding to all the lenses to determine the initial map of the current scene, improving the accuracy and efficiency of the construction (initialization) of the initial map of the current scene, thereby facilitating the improvement of the accuracy and efficiency of the incremental reconstruction of the initial map.
[0249] In another optional embodiment, as Figure 4 shown, the device may further include:
[0250] A fourth optimization module 314, configured to optimize the focal length and principal point coordinates of each lens, the rotation matrix and translation vector corresponding to each first image collected by each lens, and each map point in the initial map of the current scene based on a pre-determined projection error method, to obtain the optimized focal length of each lens, the optimized rotation matrix and translation vector corresponding to each first image collected by each lens, and the map points in the optimized initial map of the current scene.
[0251] It can be seen that after implementing Figure 4 the described device can further construct corresponding optimization formulas for the pose of each lens at each exposure point, the focal length of each lens, the principal point coordinates of the image, and the map points of the initial map after obtaining the initial map of the current scene and completing the initial calibration of the lens. Then, based on the projection error method obtained by the least squares principle, it calculates the minimum value of the optimization formula, realizes the optimization of the initial map, exposure points, and lens focal length, improves the construction accuracy of the initial map, the calibration accuracy of exposure points and lens focal length, further adapts to the current real scene, is conducive to further improving the accuracy of subsequent image acquisition, and further conducive to improving the accuracy and reliability of subsequent map incremental construction.
[0252] Embodiment 4
[0253] Please refer to Figure 5 , Figure 5 which is a schematic structural diagram of another map incremental reconstruction device applied to a flight scene disclosed in an embodiment of the present invention. This device can be applied to any flight scene that requires 3D construction, and this device is applied to a target device or a centralized control center for controlling the target device. Among them, the centralized control center includes a server (local server or cloud server) or a control platform. Among them, when the target device has functions such as flight, data collection, and image analysis, the device may include the target device; when the target device has functions of flight and data collection, the device may include a centralized control center for controlling the target device. At this time, the target device transmits the collected data to the centralized control center for analysis and controls the target device. Further, the target device may be an image processing device or a map construction device; the map construction device includes an image processing device and a flight device. At this time, the image processing device is integrated on a flight device with flight functions and is used for data collection. As Figure 3 shown, the device may include: Figure 5 A memory 401 storing executable program code;
[0254] A processor 402 coupled to the memory 401;
[0255] Further, it may also include an input interface 403 and an output interface 404 coupled to the processor 402;
[0256] Among them, the processor 402 calls the executable program code stored in the memory 401 and executes some or all of the steps in the map incremental reconstruction method applied to the flight scene disclosed in Embodiment 1 of the present invention.
[0257] Embodiment 5
[0258]
[0259] An embodiment of the present invention discloses a computer storage medium storing computer instructions, which when called, are used to execute some or all of the steps of the map incremental reconstruction method applied to a flight scenario disclosed in Embodiment 1 of the present invention.
[0260] The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative work.
[0261] Through the above specific descriptions of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course also by hardware. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, including read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc memories, magnetic disk memories, tape memories, or any other computer-readable medium capable of carrying or storing data.
[0262] Finally, it should be noted that: The map incremental reconstruction method and device for flight scenarios disclosed in the embodiments of the present invention only disclose the preferred embodiments of the present invention. It is only used to illustrate the technical solutions of the present invention, rather than to limit them; Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for incremental map reconstruction applied to flight scenarios, characterized in that, The method includes: After the image processing device completes lens calibration and obtains the initial map of the current scene, performing an image acquisition operation based on at least one exposure point of the image processing device in the current scene; wherein, the image processing device is in a flying state, and the initial map of the current scene is determined by first images collected at at least two exposure points through the lens of the image processing device when the image processing device is in a flying state; When the lens of the image processing device collects any second image at at least one of the exposure points, performing a region division operation on the second image to obtain a plurality of sub-regions; For any sub-region of the second image, extracting a preset number of feature points; Performing an incremental reconstruction operation on the initial map of the current scene according to all the feature points of the second image and all the feature points of a plurality of previously determined first images.
2. The method for incrementally reconstructing a map applied to a flight scenario according to claim 1, wherein The performing an incremental reconstruction operation on the initial map of the current scene according to all the feature points of the second image and all the feature points of a plurality of previously determined first images includes: Determining the projection of the second image on the ground according to the second image; According to the projections of a plurality of previously determined first images on the ground, screening out a third image whose projection intersects with the projection of the second image from the plurality of first images; Matching the feature points of the second image with the feature points of the third image to obtain a feature point matching result of the second image, and the feature point matching result corresponding to the second image includes a plurality of first feature matching points; Screening out second feature matching points that do not have corresponding map points in the initial map of the current scene from all the first feature matching points included in the second image; Performing a spatial triangulation operation on each of the second feature matching points to obtain a map point corresponding to the second feature matching point, and adding all the map points corresponding to the second feature matching points to the initial map of the current scene.
3. The method for incrementally reconstructing a map applied to a flight scenario according to claim 2, wherein The method further includes: Obtaining the vertical distance between the image processing device and the ground when the second image is collected, the internal parameter matrix of the lens used to collect the second image, and the pose of the lens when the second image is collected; Wherein, the determining the projection of the second image on the ground according to the second image includes: Obtaining the pixel coordinates of four corner imaging points in the second image; Calculating the projection of the second image on the ground based on the vertical distance corresponding to the second image, the internal parameter matrix corresponding to the second image, the pose corresponding to the second image, and the pixel coordinates of four corner imaging points in the second image.
4. The method for incrementally reconstructing a map applied to a flight scenario according to claim 2 or 3, wherein The method further includes: When all the map points corresponding to all the second images collected by the image processing device are added to the map of the current scene, for a plurality of imaging points of all fourth images that have completed map construction, obtaining the map coordinates of each imaging point and the positions of all the exposure points, where all the fourth images include all the second images; Calculate a scaling factor based on the map coordinates of all the imaging points and the positions of all the exposure points, and use the scaling factor to scale the position of each of the exposure points among all the exposure points to obtain the scaled position of each of the exposure points; Based on the positions of all the scaled exposure points and the imaging point center coordinates of all the fourth images, calculate a target rotation matrix and a target translation vector between the three-dimensional model and the geospatial space; Use the calculated target rotation matrix and the target translation vector to perform transformation and translation operations on the map of the currently reconstructed current scene, so that all the map points corresponding to the current scene are aligned with the geographical coordinates.
5. The method for incremental map reconstruction applied to flight scenarios according to claim 2 or 3, characterized in that, The method further includes: Based on each of the second feature matching points of the second image, the map point corresponding to the second feature matching point, and the focal length of the lens that captured the second image, optimize the pose of the lens when capturing the second image and the map point corresponding to each of the second feature matching points that are pre-determined to obtain the optimized pose of the lens and the optimized map points; Wherein, the adding the map points corresponding to all the second feature matching points to the initial map of the current scene includes: Adding the map points corresponding to all the second feature matching points of the optimized second image to the initial map of the current scene; Wherein, the pose corresponding to each second image is used as a basis for incremental reconstruction of the initial map of the current scene; The method further includes: Judge whether the current condition of the current scene meets a pre-determined global optimization condition; When it is judged that the global optimization condition is met, based on the internal parameter matrix of the lens of the second image obtained by calibration, optimize the pose of the lens when capturing the corresponding second image and the map point corresponding to each of the second feature matching points that are pre-determined to obtain the optimized pose of the lens and the optimized map points, until all the map points corresponding to all the second images captured by the image processing device are added to the initial map of the current scene.
6. The method for incrementally reconstructing a map applied to a flight scenario according to any one of claims 1-3, characterized in that Before the image acquisition operation is performed by the image processing device at at least one exposure point in the current scene, the method further includes: Based on the image acquisition operation performed by the image processing device at at least one of the exposure points in the current scene, obtain the metadata corresponding to each lens of the image processing device, wherein the metadata corresponding to each lens includes the first image captured by the lens at at least one of the exposure points, the position of the exposure point in the current scene, and the pose of the lens at the exposure point when capturing the first image, and the number of lenses of the image processing device is greater than or equal to 1; For any one of the lenses, extract the feature points of each of the first images of the lens, and determine one of the images from all the first images as a reference image; for any one of the first images other than the reference image among all the first images, match the feature points of the first image with the feature points of the reference image to obtain the feature point matching result corresponding to the first image. Each feature point matching result corresponding to each of the first images includes a plurality of third feature matching points; and based on the feature point matching result corresponding to the first image, calculate the fundamental matrix between the first image and the reference image; based on the fundamental matrix of the first image, determine the projective projection matrix of the first image; according to the projective projection matrices of all the first images and the properties of the absolute dual quadric coefficients, determine the absolute dual quadric coefficients, and calibrate the focal length of the lens according to the absolute dual quadric coefficients and the projective projection matrix of the first image.
7. The method for incrementally reconstructing a map applied to a flight scenario according to claim 6, wherein The method further includes: Determine one of the images as a reference coordinate system image from all the first images, where the reference coordinate system image includes the first image collected for the first time or one of the first images collected non-first. For at least one of the first images other than the reference coordinate system image among all the first images, determine the internal parameter matrix corresponding to the first image according to the width and height of the first image and the focal length of the lens; based on the internal parameter matrix corresponding to the first image and the fundamental matrix of the first image, determine the essential matrix of the first image, and decompose the essential matrix of the first image to obtain the rotation matrix and translation vector corresponding to the first image; based on the rotation matrix and translation vector corresponding to the first image and the internal parameter matrix corresponding to the first image, reconstruct the projective projection matrix of the first image to obtain the projection matrix of the first image. For at least one of the first images other than the reference coordinate system image among all the first images, perform spatial triangulation on each of the third feature matching points in the first image according to the projection matrix of the first image to obtain the map point corresponding to the third feature matching point in the current scene. Determine the initial map of the current scene according to all the map points corresponding to all the third feature matching points corresponding to all the first images and each of the first images. Wherein, the initial map of the current scene is composed of a plurality of the map points and the first images collected by all the lenses.
8. The method for incrementally reconstructing a map applied to a flight scenario according to claim 6, wherein The method further includes: Based on a pre-determined projection error method, optimize the focal length and principal point coordinates of each lens, the rotation matrix and translation vector corresponding to each of the first images collected by each lens, and each map point in the map of the current scene to obtain the optimized focal length of each lens, the optimized rotation matrix and translation vector corresponding to each of the first images collected by each lens, and the map points in the optimized map of the current scene.
9. A map incremental reconstruction device applied to a flight scenario, characterized in that, The device is applied to a map construction device or an image processing device, and the device includes: An acquisition module, configured to perform an image acquisition operation at at least one exposure point of the current scene based on an image processing device after the image processing device completes lens calibration and obtains an initial map of the current scene; wherein, the image processing device is in a flight state, and the initial map of the current scene is determined by first images acquired by the lens of the image processing device at at least two exposure points when the image processing device is in a flight state; A region division module, configured to perform a region division operation on any second image when the lens of the image processing device acquires the second image at at least one of the exposure points, to obtain a plurality of sub-regions; An extraction module, configured to extract a preset number of feature points for any sub-region of the second image; A reconstruction module, configured to perform an incremental reconstruction operation on the initial map of the current scene according to all the feature points of the second image and all the feature points of a plurality of first images determined in advance; 10. An incremental map reconstruction device applied to a flight scenario, characterized in that, The device is applied to a map construction device or an image processing device, and the device includes: A memory storing executable program code; A processor coupled to the memory; The processor calls the executable program code stored in the memory and executes the map incremental reconstruction method for a flight scene according to any one of claims 1-8.