Incremental motion averaging method and related device

Through the incremental motion averaging method combined with cross-view feature matching and scale separation strategies, the camera position solving problem is solved, the camera position solving is achieved, and the camera position pose is improved, and the accuracy and stability of three-dimensional reconstruction is improved.

CN120339378APending Publication Date: 2025-07-18INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510501363.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the existing motion recovery structure method, the camera position solution has high uncertainty, making it difficult to achieve interactive fusion and robust solution of camera rotation and position information, especially in large-scale scene three-dimensional reconstructions, there are scale ambiguity and information coupling problems.

Method used

The incremental motion averaging method is adopted, combined with cross-view feature matching information, and the scale separation strategy is used to eliminate camera positioning ambiguity, and the camera's multi-type absolute pose parameters are synchronized through the incremental process, including the camera's global relative scale, rotation and local absolute scale.

Benefits of technology

It improves the accuracy and robustness of large-scale motion average solution, realizes the accurate and stable solution of camera position, and improves the accuracy and stability of three-dimensional reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339378A_ABST
    Figure CN120339378A_ABST
Patent Text Reader

Abstract

The invention belongs to a three-dimensional reconstruction method, and provides an incremental motion averaging method and a related device for solving the technical problems that a global method in an existing motion recovery structure method is high in camera position solving uncertainty and is not beneficial to camera rotation, interactive fusion of position information and robust solving. According to an image set of a scene to be reconstructed, cross-view feature matching information is combined, ambiguity existing in camera positioning is eliminated by using a scale separation strategy, and a camera global relative scale is obtained. And then synchronously resolving multi-type absolute pose parameters of the camera by adopting an incremental flow according to the global relative scale of the camera, the rotation and translation of the camera and the local absolute scale of the camera. The solution accuracy and robustness of large-scale motion averaging can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to a 3D reconstruction method, and specifically relates to an incremental motion averaging method and related devices. Background Art

[0002] Large-scale scene 3D reconstruction based on images is an important research direction in the field of computer vision and is widely used in industries such as remote sensing mapping, autonomous driving, and industrial twins. When performing scene 3D reconstruction, it is first necessary to use the Structure from Motion (SfM) method to solve the absolute poses of a large number of cameras in the global unified coordinate system, and at the same time obtain the sparse representation of the reconstructed scene based on image local features, which is the key scientific problem and core technical means in the entire reconstruction process. The existing SfM methods are mainly divided into two types: incremental and global. Among them, the core problem of the global method is motion averaging, which aims to solve the absolute poses of cameras based on the relative motion of cameras. On the one hand, due to the inherent scale ambiguity in two-view reconstruction in multi-view geometry theory, existing methods usually need to synchronously estimate the absolute position of the camera and the relative baseline length, which will increase the uncertainty of camera position solving. On the other hand, due to the tight coupling of information between camera rotation and position, existing methods usually adopt a step-by-step decoupled solution strategy for rotation solving and rotation-known position solving, which is not conducive to the interactive fusion and robust solution of the above two closely related pieces of information. Summary of the Invention

[0003] Aiming at the technical problems existing in the global method of the existing Structure from Motion method, such as high uncertainty in camera position solving, and being not conducive to the interactive fusion and robust solution of camera rotation and position information, this application provides an incremental motion averaging method and related devices.

[0004] To achieve the above object, this application adopts the following technical solutions: In the first aspect, this application proposes an incremental motion averaging method, including: For the image set of the scene to be reconstructed, combined with cross-view feature matching information, use the scale separation strategy to eliminate the ambiguity in camera positioning and obtain the global relative scale of the camera; The scale separation strategy includes: obtaining a seed set and solving based on a distance metric function in the positive real number domain; gradually adding incremental vertices to the seed set and solving; optimizing the local absolute scale parameters of the camera; According to the global relative scale of the camera, camera rotation, camera translation, and the local absolute scale of the camera, use an incremental process to synchronously solve multiple types of absolute pose parameters of the camera.

[0005] Further, the method for obtaining the cross-view feature matching information includes: Based on the relative motion information between cameras, perform two-view triangulation on the first image pair and the second image pair in the image set of the scene to be reconstructed respectively; one image in the first image pair and one image in the second image pair are the same, denoted as the shared image, and the other image in the first image pair and the other image in the second image pair both share multiple feature matching pairs with the shared image; Based on the two-view triangulation results, obtain the depths of the three-dimensional space points corresponding to the first image pair and the second image pair in the coordinate system of the shared image, and obtain a set of depth ratios of multiple feature matching pairs; Perform median processing on the set of depth ratios of multiple feature matching pairs; According to the median processing result, construct a constraint relationship between the depth ratio and the local absolute scale; For all image pairs that share multiple feature matching pairs with the shared image, jointly solve the constraint relationship between the corresponding depth ratio and the local absolute scale, and obtain the local absolute scale of the cameras corresponding to all images that share multiple feature matching relationships with the shared image relative to the shared image as cross-view feature matching information.

[0006] Furthermore, the method for eliminating the ambiguity in camera positioning by using the scale separation strategy includes: Obtain the set of vertex triples in the local absolute scale solution graph of the shared image; Initialize the local absolute scales of the cameras in the triples in the set of vertex triples according to the constraint relationship between the depth ratio and the local absolute scale; Obtain a seed set according to the initialization result; Use the seed set as the set of added vertices, combine the edge selection reward mechanism, expand the set of added vertices, and simultaneously perform incremental solution of the set of local absolute scales of each camera through local optimization and global optimization; Calculate the global relative scale of the cameras according to the incremental solution result of the set of local absolute scales of each camera.

[0007] Furthermore, the method for obtaining the seed set according to the initialization result includes:

[0008] Among them, is the triple with the smallest distance, represents a distance metric function in the positive real number domain, is the set of vertex triples in the local absolute scale solution graph of the shared image, is relative to image i , image k and image l The depth ratio between them, is relative to imagei ,image j With image l The depth ratio between Relative to the image i ,image j With image k The depth ratio between .

[0009] Furthermore, the method of combining the edge selection reward mechanism, expanding the set of added vertices, and simultaneously performing incremental calculation of the local absolute scale set of each camera through local optimization and global optimization includes: The edge set between any vertex that is not currently added and each vertex in the set of added vertices is recorded as ; According to the constraint relationship between the depth ratio and the local absolute scale, the vertices in the added vertex set are used to The local absolute scale estimate of , edge set Middle Edge Upper depth ratio measurement , opposite vertices The local absolute scale Perform pre-calculation; According to the edge set The edge selection reward of each edge in the , and the initialization solution results of the added incremental points and the local absolute scale are obtained; An alternating local and global optimization strategy is used to incrementally solve the local absolute scale set of each camera; The global relative scale of the camera is calculated based on the incremental solution of the local absolute scale set of each camera.

[0010] Furthermore, the method of synchronously solving multiple types of absolute pose parameters of a camera using an incremental process includes: Based on the epipolar geometry corresponding to the image set of the scene to be reconstructed, the camera absolute pose parameters of the camera ternary set in the epipolar geometry and the relative baseline length used to calculate the absolute position are initialized according to the constraint relationship between the relative pose and the absolute pose of the camera and the relative baseline length formula; According to the initialization results, based on the comprehensive ranking and cycle consistency optimization principle, the triples with the best multi-type parameter consistency are obtained as the incremental seed set; Taking the incremental seed set as the incremental added vertex set, using the absolute scale, rotation and position estimation values of any vertex in the incremental added vertex set, the local absolute scale estimation value of any vertex in the incremental unadded vertex set, and the relative scale, rotation and translation measurement values on any edge in the incremental edge set, the camera absolute scale, rotation and position of the vertices in the incremental unadded vertex set are pre-calculated; Based on the pre-computation results, select incremental vertices according to the principle of maximizing the intersection scale of the multi-type parameter support set, obtain the incremental vertex set, and simultaneously complete the estimation of multi-type camera pose parameters; Adopt an optimization mode that alternates between local and global scopes. Use the estimated values of the current multi-type camera pose parameters as the initial optimization values, obtain the constraints provided by the camera relative motion measurement internal value set through reverse calculation, optimize the absolute pose parameters of the camera, and complete the synchronous solution of the camera pose.

[0011] Further, when adopting the optimization mode that alternates between local and global scopes, it also includes: First, optimize the absolute scale and rotation of the camera, then fix the absolute scale and rotation of the camera, and optimize the absolute position of the camera.

[0012] In a second aspect, the present application proposes an incremental motion averaging system, including: An ambiguity elimination module, which is used to combine the cross-view feature matching information for the image set of the scene to be reconstructed, and use the scale separation strategy to eliminate the ambiguity existing in camera positioning, and obtain the global relative scale of the camera; The scale separation strategy includes: obtaining a seed set and solving based on a distance metric function in the positive real number domain; gradually adding incremental vertices to the seed set and solving; tuning the local absolute scale parameter of the camera; A solution module, which is used to synchronously solve the multi-type absolute pose parameters of the camera by an incremental process according to the global relative scale of the camera, the camera rotation and translation, and the local absolute scale of the camera.

[0013] In a third aspect, the present application proposes an electronic device, including: a memory, one or more processors; the memory is coupled to the processor; wherein, computer program code is stored in the memory, and the computer program code includes computer instructions. When the computer instructions are executed by the processor, the electronic device executes the steps of the above-mentioned incremental motion averaging method.

[0014] In a fourth aspect, the present application proposes a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of the above-mentioned incremental motion averaging method are implemented.

[0015] Compared with the prior art, the present application has the following beneficial effects: The present application proposes an incremental motion averaging method. For the image set of the scene to be reconstructed, by combining cross-view feature matching information and using a scale separation strategy to eliminate the ambiguity in camera positioning, the global relative scale of the camera is obtained. Then, based on the global relative scale of the camera, camera rotation, camera translation, and the local absolute scale of the camera, an incremental process is adopted to synchronously solve multiple types of absolute pose parameters of the camera. The method for eliminating camera positioning ambiguity based on scale separation in the present application eliminates the scale ambiguity in camera position calculation, realizes accurate and robust calculation of the local camera scale, and improves the accuracy of solving large-scale motion averaging. In addition, the method for synchronously solving camera pose based on incremental estimation uses pose coupling information to integrate the solution of multiple types of camera poses into a single-process incremental parameter estimation, improving the robustness of solving large-scale motion averaging.

[0016] The present application also proposes an incremental motion averaging system, an electronic device, and a computer storage medium, which possess all the advantages of the above incremental motion averaging method. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments. It should be understood that the following drawings only show some embodiments of the present application, and thus should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other relevant drawings can also be obtained based on these drawings.

[0018] Figure 1 It is a schematic flowchart of an incremental motion averaging method of the present application; Figure 2 It is a schematic diagram of an incremental motion averaging system of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Usually, the components of the embodiments of the present application described and shown in the drawings here can be arranged and designed in various different configurations.

[0020] Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the present application to be protected, but merely represents the selected embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present application.

[0021] It should be noted that like reference numerals and letters refer to like items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0022] In the description of the embodiments of the present application, it should be noted that if terms such as "upper", "lower", "horizontal", "inner", etc. are used to indicate the orientation or positional relationship, it is based on the orientation or positional relationship shown in the figures, or the orientation or positional relationship in which the inventive product is usually placed during use. This is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present application. In addition, terms such as "first", "second", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0023] In addition, if the term "horizontal" appears, it does not mean that the component is required to be absolutely horizontal, but it can be slightly inclined. For example, "horizontal" only means that its direction is more horizontal relative to "vertical", and does not mean that the structure must be completely horizontal, but it can be slightly inclined.

[0024] In the description of the embodiments of the present application, it should also be noted that unless otherwise clearly specified and limited, if terms such as "set", "installed", "connected", "coupled" are used, they should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific circumstances.

[0025] Large-scale scene three-dimensional reconstruction based on images occupies an extremely important research position in the field of computer vision and has extensive and in-depth applications in many industries such as remote sensing mapping, autonomous driving, and industrial twins. When carrying out the work of scene three-dimensional reconstruction, the key first step is to use the Structure from Motion (SfM) method to calculate the large number of absolute poses of cameras in the global unified coordinate system. At the same time, a sparse representation of the reconstructed scene based on image local features is obtained. This link is the core scientific problem and key technical means in the entire reconstruction process.

[0026] Existing SfM methods can be mainly divided into two categories: incremental and global. The incremental method uses an iterative optimization mode to gradually solve the camera pose in a progressive manner; while the global method uses the means of motion averaging to complete the solution of the camera pose at one time. These two types of methods have their own advantages in terms of solution accuracy, robustness, solution efficiency, consistency, etc., and can complement each other's advantages. Compared with the incremental SfM method, the global method has the remarkable characteristics of fewer optimization rounds and fewer parameters to be solved. In the current situation where the scale of the reconstructed scene is constantly expanding, the global method shows more prominent theoretical advantages and broad application potential, and has therefore attracted much attention in recent years.

[0027] As the core problem of the global SfM method, motion averaging aims to solve the absolute camera pose based on the relative camera motion. Although a large amount of research has been carried out on this problem, due to the extremely high difficulty of solving the problem itself and the drawbacks of the existing solution modes, there is an urgent need to make a breakthrough in both the solution theory and the actual application effect of motion averaging. On the one hand, based on the theory of multi-view geometry, there is an inherent scale ambiguity problem in two-view reconstruction, which makes existing methods often have to synchronously estimate the absolute camera position and the relative baseline length, which will undoubtedly further increase the uncertainty of camera position solution. On the other hand, due to the close coupling relationship between the information of camera rotation and position, existing methods generally adopt a step-by-step decoupled solution strategy of first solving the rotation and then solving the position when the rotation is known. This method is not conducive to the mutual interaction and fusion of the two closely related pieces of information of camera rotation and position, nor is it conducive to achieving robust solution.

[0028] Based on the above situation, this application proposes an incremental motion averaging method and related devices, and the following will make a detailed description of this application in combination with embodiments and drawings.

[0029] As Figure 1 shown, it is a schematic flowchart of an incremental motion averaging method of this application, which may include: S101, for the image set of the scene to be reconstructed, combined with the cross-view feature matching information, use the scale separation strategy to eliminate the ambiguity in camera positioning, and obtain the global relative scale of the camera.

[0030] The scale separation strategy includes: obtaining a seed set and solving based on a distance metric function in the positive real number domain; gradually adding incremental vertices to the seed set and solving; optimizing the local absolute scale parameters of the camera.

[0031] The distance metric function in the domain of positive real numbers provides a quantitative way to measure the relationship between different views. In scene reconstruction, the feature matching relationships between different images can be screened and organized through this distance metric. Obtaining the seed set means selecting a representative set of images from all images as the starting point. The images in this seed set are usually those with relatively reliable feature matches and crucial distributions in the scene. Solving the seed set preliminarily determines the relative pose relationships between these images, laying a foundation for subsequent expansion and refinement. After determining the seed set, incremental vertices, that is, more images, equivalent to the corresponding cameras, are gradually added. This process is a step-by-step expansion process. Each time a new vertex is added, the pose relationships between the images in the entire set are recalculated to update. This incremental addition and calculation method helps to use the information of existing images to more accurately locate the newly added images and continuously optimize the pose estimation of the entire image set. What is usually obtained through incremental addition is only the relative scale relationship and a rough pose estimation. Tuning the local absolute scale parameters of the camera is for further precision. The local absolute scale parameters of the camera determine the actual size and distance measurement of the camera in the scene. By tuning these parameters, the pose estimation of the camera can be made more consistent with the physical scale of the actual scene, improving the accuracy of the reconstruction result.

[0032] S102, according to the global relative scale of the camera, the camera rotation, the camera translation, and the local absolute scale of the camera, synchronously solve multiple types of absolute pose parameters of the camera using an incremental process.

[0033] Based on obtaining the global relative scale of the camera, the camera rotation, the camera translation, and the local absolute scale of the camera, synchronously solve multiple types of absolute pose parameters of the camera using an incremental process. The multiple types of absolute pose parameters can include the position (translation parameter) and direction (rotation parameter) of the camera, etc. The incremental process enables continuous use of new information to update and optimize the pose estimation during the solution process, avoiding error accumulation and inaccuracies caused by one-time calculations.

[0034] To process the image set of the scene to be reconstructed, so as to achieve the synchronous solution of multi-type absolute pose parameters of the camera, and then complete the scene reconstruction and 3D modeling. This application combines cross-view feature matching information and adopts a scale separation strategy to solve the key problems in camera positioning, and finally obtains the global relative scale of the camera, thereby providing a basis for subsequent pose parameter solution. This application effectively eliminates the ambiguity of camera positioning through the scale separation strategy, and the incremental processing method can gradually optimize the results, improving the accuracy and stability of the reconstruction. At the same time, the synchronous solution of multi-type absolute pose parameters makes the entire reconstruction process more complete and accurate. In practical applications, this application can be applied to scenarios such as 3D scene reconstruction, SLAM (Simultaneous Localization and Mapping) systems, virtual reality, and augmented reality, providing accurate camera pose information for these applications, so as to achieve high-quality scene modeling and interactive experience.

[0035] The following is another embodiment of this application, which further elaborates on the solution of this application in detail: S201, Elimination of camera positioning ambiguity based on scale separation.

[0036] Given the image set of the scene to be reconstructed, local features are extracted from each image in the image set of the scene to be reconstructed, and feature matching is performed between image pairs. Then, through the estimation and decomposition of the essential matrix, an epipolar geometry graph of this image set is constructed, denoted as , where represents the camera set corresponding to each image, represents the camera pair set corresponding to the images with matching relationships.

[0037] In practical applications, local feature extraction can be performed using SIFT (Scale-Invariant Feature Transform), SURF (Speeded-Up Robust Features), ORB (Oriented FAST and Rotated BRIEF), etc. After extracting local features from each image, feature matching is then performed for all image pairs in the image set. That is, the features of each image are compared with the features of other images to find pairs of matching feature points. The essential matrix is an important matrix that describes the relative pose relationship between two cameras and contains the rotation and translation information between the two cameras. For image pairs with a matching relationship, using their pairs of matching feature points, the essential matrix between the cameras corresponding to these two images can be estimated. If there is a feature matching relationship between image A and image B, then there is an edge between their corresponding cameras (nodes in the figure), and this edge indicates that there is a certain geometric relationship between these two cameras, which can be described by the essential matrix, etc. In this way, the matching relationships between all images in the entire image set and the geometric relationships between cameras are represented in the form of a graph, and the epipolar geometry graph G is constructed.

[0038] The motion averaging problem is to solve for the absolute pose of cameras in the global coordinate system, including rotation and position, based on the given relative motion of cameras, such as rotation and translation, , . Among them, represents the rotation in the given relative motion of cameras, represents the translation in the given relative motion of cameras, represents the rotation in the absolute pose of cameras in the global coordinate system, represents the position in the absolute pose of cameras in the global coordinate system. Therefore, the motion averaging problem is to solve for the absolute pose of each camera, including rotation and position, in a unified global coordinate system based on the known relative motion relationships, rotation and translation, between cameras through certain algorithms and calculations, so as to provide accurate camera pose information for subsequent applications such as 3D reconstruction and multi-camera collaborative work.

[0039] To eliminate the ambiguity in camera localization and enhance the solvability of the motion averaging problem using the scale separation strategy, cross-view feature matching information is required. Suppose the image pair and the image share multiple feature matching pairs. Using the relative motion information between cameras for the image pair and the image pair Perform two-view triangulation separately, and record the depth of the three-dimensional space points corresponding to the above cross-view matching pairs in the camera coordinate system of the image as: , where is the number of shared matching pairs. Based on the multi-view geometry theory, although due to the scale ambiguity of relative translation, that is , making , but in the ideal case, the depth ratio of different matching pairs remains unchanged, that is . This provides a theoretical premise for the separation and estimation of the camera scale parameter. Due to the existence of false matches and calculation errors, it is necessary to first take the median value of the above depth ratio set: , where represents the depth ratio between image and image relative to image , represents the median operation on the set . On this basis, the following constraint relationship between the depth ratio and the local absolute scale can be constructed:

[0040] where, and represent the local absolute scales of camera , camera and camera relative to camera . By jointly solving the constraint relationships provided by all image pairs that share multiple feature matching pairs with image , the local absolute scales of the cameras corresponding to all images with common feature matching relationships with image relative to camera can be obtained. That is, given the depth ratio set (relative to camera ), solve the elements in the local absolute scale set , where represents the set of image pairs that can calculate the depth ratio relative to image , and represents the set of images that have pairwise common feature matching relationships with image .

[0041] Here, a method for eliminating camera positioning ambiguity based on scale separation is adopted to achieve accurate and robust calculation of the local absolute scale of the camera, mainly including key technical steps such as seed set acquisition and calculation, incremental vertex selection and solution, and camera local absolute scale parameter optimization. The specific methods for each step can be as follows: The acquisition and solution of the seed set are realized based on the pre - calculation of the local absolute scale of the camera triple and its loop - consistency verification. For the image of the local absolute scale solution graph the set of vertex triples is denoted as: . The local absolute scale of the camera in is initialized as follows: . On this basis, the seed set is obtained through the following formula:

[0042] where represents the distance metric function in the positive real - number domain. Here, the normalized mean - squared - difference function is selected, that is: . The initial value of the local absolute scale of the camera in the seed set can be determined as: .

[0043] The selection and solution of the incremental vertex are realized based on the pre - calculation of the local absolute scale of the un - added vertex and the construction of its weighted support set. For a certain vertex in the current un - added vertex set and the set of added vertices the edge set between each vertex is denoted as . Similarly, according to Equation , using the estimated value of the local absolute scale of vertex in , and the depth - ratio measurement value on the edge in , the local absolute scale of vertex is pre - calculated: . Since different selections of and will result in different pre - calculation results of , to realize the effective selection of incremental vertices, the edge - selection reward of each edge in is obtained through the following formula:

[0044] where is the weighted coordination coefficient. Then, using the edge - selection reward obtained above, the selection of the dominant edge , the incremental point , and the initialization and solution of the local absolute scale are realized through the following formula :

[0045] The optimization of the local absolute scale parameters of the camera adopts an alternating local and global optimization strategy. The main difference between local and global optimization lies in that local optimization only faces the currently added single vertex, while global optimization faces all currently estimated vertices. To improve the solution efficiency, the entire parameter solution process is mainly based on local optimization, and the global optimization module is only called when the vertex increment reaches a certain proportion. When performing parameter optimization, the current known local absolute scale is used to back-calculate the depth ratio according to Equation and calculate its distance from the corresponding measured value to obtain the inlier set. Furthermore, the current estimated value of the parameter is used as the initial value, and the local absolute scale is optimized (locally / globally) using the constraints provided by the inlier set.

[0046] Incremental vertex selection and solution are iteratively performed for local / global parameter optimization until all vertices are added or there are no vertices to add, and then the incremental solution of the local absolute scale set with respect to each camera can be achieved. On this basis, the global relative scale can be calculated as follows: and it is used for the camera global absolute scale in the subsequent camera pose synchronization solution based on incremental estimation, the relative baseline length and the calculation of the camera absolute position .

[0047] S202, Camera pose synchronization solution based on incremental estimation Given the camera global relative scale, rotation, translation , and the local absolute scale , an incremental process is adopted to achieve camera pose solution. According to the camera pose constraints and the locality characteristics of incremental parameter solution, as well as the aforementioned incremental local scale averaging method, a new theory and method for pose synchronization solution are developed. There is no need to solve the camera pose step by step (scale, rotation, position), and multiple types of camera absolute pose parameters can be synchronously solved in a single incremental parameter estimation process , and it simultaneously has the characteristics of high accuracy and strong robustness of the incremental method and high efficiency and good consistency of the global method.

[0048] An incremental reference is constructed using camera triples, but multiple types of camera pose information (scale, rotation, translation) need to be used simultaneously. Specifically, the set of camera triples in the epipolar geometry graph is denoted as . According to the constraint relationship between the relative pose and absolute pose of the camera: , and the calculation formula for the relative baseline length: , the following formula is used for Initialize the four key parameters: the absolute pose parameters of the camera and the relative baseline length used to calculate the absolute position.

[0049] Since the above four key parameters are in different parameter spaces, it is impossible to use a processing method similar to the incremental local scale averaging method, that is, Equation , directly use the distance between the inverse calculation value of the multi-type relative motion parameters based on the above initialization results and the given measurement value to obtain the seed set for camera pose synchronization calculation. Based on the principle of comprehensive ranking and loop consistency optimization, it is achieved through the following formula:

[0050] Among them, the comprehensive ranking function is:

[0051] Among them, represents in the set the ranking from small to large. ; , represents the spatial distance metric function, which is the rotation angle distance here, that is: ; , ; , represents the spatial distance metric function, which is the normalized vector angle distance here, that is: . Based on the above two formulas (the formula for initializing the four key parameters of the relative baseline length for calculating the absolute position and the formula implemented based on the principle of comprehensive ranking and loop consistency optimization), the initial value of the camera absolute pose parameters of the seed set can be determined.

[0052] When performing incremental vertex selection and solution, the basic principle is still similar to the previous method. Denote the edge set between a certain vertex in the current unadded vertex set and each vertex in the added vertex set as . Similar to obtaining the seed set, according to the previous formula (the formula for initializing the four key parameters of the relative baseline length for calculating the absolute position), using the (local and global) absolute scale, rotation, and position estimation values of the vertex in : , the local absolute scale estimation value of the vertex in , and relative scale, rotation, and translation measurements on the middle edge : , and pre-compute the absolute scale, rotation, and position of the camera for the vertices .

[0053] Due to the diverse characteristics of the spaces where the above-mentioned multi-type parameters are located, it is also impossible to adopt a weighted summation method similar to that in the incremental local scale averaging method (see the formula for edge selection rewards) here. Instead, based on the principle of maximizing the intersection scale of the multi-type parameter support sets, the selection of incremental vertices is realized. Specifically, based on the above pre-computation results, reverse calculation is performed for each edge in to obtain the relative motion parameter information, compare it with its corresponding measurement value, and obtain the intersection of the multi-type parameter support sets that simultaneously satisfy the following four conditions

[0054] where represents the distance threshold for determining whether each relative motion measurement value is an inlier, , . Then, the support set of the edge obtained from the formula of the intersection of the multi-type parameter support sets can be used to select the dominant edge , the incremental point , and initialize the solution of the absolute pose through the following formula:

[0055] The optimization of the absolute pose parameters of the camera also adopts an optimization mode that alternates between local and global scopes. The estimated value of the current multi-type camera pose parameters is used as the initial value of the optimization, and the constraints provided by the set of inlier camera relative motion measurements obtained through reverse calculation are used to optimize the absolute pose parameters of the camera, thereby realizing the synchronous solution of the camera pose based on incremental estimation. It should be noted that when optimizing the absolute pose parameters of the camera, due to the information coupling of the pose parameters and the diverse characteristics of the parameter space, a parameter step-by-step optimization strategy is adopted here, that is, first optimize the absolute scale and rotation of the camera, and then fix the above parameters and optimize the absolute position of the camera.

[0056] To verify the technical effect of the incremental motion averaging method of the present application, the following verification is carried out in the present application: The method of the present application is experimentally analyzed on the 1D SfM dataset. Since the method of the present application can simultaneously solve the absolute rotation and absolute position of the camera, the method of the present application is compared with the current mainstream rotation averaging methods (including LThe solution accuracies are compared using 1GM, MPLS, IRA, HARA) and translational averaging methods (SATA, BATA, ITA, CReTA - BATA) respectively.

[0057] For the accuracy comparison metrics, the median and mean errors of the camera's absolute rotation / position obtained after rotation / translational averaging are used, denoted as respectively, and the median error of the camera's absolute rotation / position optimized after global bundle adjustment is denoted as respectively. Since the mean error may be too large due to the presence of some outliers after global bundle adjustment, it is not considered here. Additionally, when comparing the estimation accuracies of the camera's rotation / position, method BATA / MPLS is used to provide the solution results of the camera's position / rotation. The comparison experimental results are shown in Tables 1 and 2.

[0058] Table 1 Comparison Results of Camera Absolute Rotation Estimation Accuracy

[0059] Table 2 Comparison Results of Camera Absolute Pose Estimation Accuracy

[0060] As can be seen from Table 1, the method of this application (denoted as IMA) achieved the best overall performance in the three metrics of camera absolute rotation estimation accuracy ( ). Since methods L 1GM and HARA are not very robust to high noise and large - scale scenes, the above two methods performed mediocrely in the metric. As can be seen from Table 2, the method of this application also achieved the best overall performance in the three metrics of camera absolute position estimation accuracy ( ). For method ITA, although it performed excellently in the metric , due to the degradation phenomenon of this method for the collinear motion configuration of the camera, its performance in the metric was relatively poor. For method CReTA - BATA, although the advantage of the method of this application over this method is not very significant, the method of this application performed better on the test data with higher noise (GDM) and larger scale (TFG), and there was no abnormal output (see test data ROF). Therefore, compared with method CReTA - BATA, the method of this application has better robustness and scalability.

[0061] As Figure 2 shown, it is a schematic diagram of an incremental motion averaging system of this application, which may include: An ambiguity elimination module is used to eliminate the ambiguity in camera positioning for the image set of the scene to be reconstructed, combine cross-view feature matching information, and utilize a scale separation strategy to obtain the global relative scale of the camera. The scale separation strategy includes: obtaining a seed set and resolving it based on a distance metric function in the positive real number domain; gradually adding incremental vertices to the seed set and resolving; optimizing the local absolute scale parameters of the camera. A resolution module is used to synchronously resolve multiple types of absolute pose parameters of the camera in an incremental process according to the global relative scale of the camera, the camera rotation, the camera translation, and the local absolute scale of the camera.

[0062] It should be noted that in several embodiments provided in the present application, it should be understood that the disclosed device and method can be implemented in other ways. For example, the system embodiments described above are only illustrative. For example, the division of each module is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules can be combined or integrated into another device, or some features can be ignored or not executed. The modules described as separate components may or may not be physically separated. The components shown as modules can be a physical unit or multiple physical units, that is, they can be located in one place or distributed to multiple different places. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0063] In addition, in each embodiment of the present invention, each module can be integrated in a processing unit, or each module can exist physically alone, or two or more modules can be integrated in a unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0064] The embodiment of the present application also provides an electronic device, which may include one or more processors, a memory, and a communication interface.

[0065] Among them, the memory and the communication interface are coupled to the processor. For example, the memory and the communication interface can be coupled together through a bus.

[0066] Among them, the communication interface is used to transmit data with other devices. The memory stores computer program code. The computer program code includes computer instructions. When the computer instructions are executed by the processor, the electronic device is caused to execute the steps of the above incremental motion averaging method.

[0067] Among them, the processor can be a processor or a controller. For example, it can be a Central Processing Unit (CPU), a general-purpose processor, a Digital Signal Processor (DSP), an Application-Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logic blocks, modules, and circuits described in conjunction with the present disclosure. The processor can also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and so on. The processor can be used to support the electronic device to execute the method steps provided in the above embodiments.

[0068] Among them, the bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The above bus can be divided into an address bus, a data bus, a control bus, etc.

[0069] A computer-readable storage medium provided by an embodiment of the present application stores a computer program, and when the computer program is executed by a processor, the steps of the above incremental motion averaging method are implemented.

[0070] The computer-readable storage medium involved in the present application includes a Random Access Memory (RAM), a memory, a Read-Only Memory (ROM), an Electrically Programmable ROM, an Electrically Erasable Programmable ROM, a register, a hard disk, a removable disk, a CD ROM, or any other form of storage medium well-known in the technical field.

[0071] The above are only the preferred embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. An incremental moving average method, characterized in that, Including: For the image set of the scene to be reconstructed, combining cross-view feature matching information, using a scale separation strategy to eliminate the ambiguity in camera positioning, and obtaining the global relative scale of the camera; The scale separation strategy includes: obtaining a seed set and solving it based on a distance metric function in the positive real number domain; gradually adding incremental vertices to the seed set and solving; optimizing the local absolute scale parameters of the camera; According to the global relative scale of the camera, the camera rotation and translation, and the local absolute scale of the camera, an incremental process is used to synchronously solve multiple types of absolute pose parameters of the camera.

2. The incremental motion averaging method according to claim 1, wherein The method for obtaining the cross-view feature matching information includes: According to the relative motion information between cameras, two-view triangulation is respectively performed on the first image pair and the second image pair in the image set of the scene to be reconstructed; one image in the first image pair and one image in the second image pair are the same, denoted as the shared image, and the other image in the first image pair and the other image in the second image pair both share multiple feature matching pairs with the shared image; Based on the two-view triangulation results, the depths of the three-dimensional space points corresponding to the first image pair and the second image pair in the coordinate system of the shared image are obtained, and a depth ratio set of multiple feature matching pairs is obtained; Median processing is performed on the depth ratio set of multiple feature matching pairs; According to the median processing result, a constraint relationship between the depth ratio and the local absolute scale is constructed; For all image pairs that share multiple feature matching pairs with the shared image, the constraint relationships between the corresponding depth ratios and the local absolute scale are jointly solved, and the local absolute scale of the cameras corresponding to all images that share multiple feature matching relationships with the shared image relative to the shared image is obtained as the cross-view feature matching information.

3. The incremental motion averaging method according to claim 2, wherein The method for using the scale separation strategy to eliminate the ambiguity in camera positioning includes: Obtaining the vertex triple set in the local absolute scale solution graph of the shared image; Initializing the local absolute scale of the cameras in the triples in the vertex triple set according to the constraint relationship between the depth ratio and the local absolute scale; Obtaining a seed set according to the initialization result; Using the seed set as the set of added vertices, combining an edge selection reward mechanism, expanding the set of added vertices, and synchronously performing incremental solution of the local absolute scale sets of each camera through local optimization and global optimization; According to the incremental solution results of the local absolute scale sets of each camera, the global relative scale of the camera is calculated.

4. The incremental motion averaging method according to claim 3, wherein The method for obtaining a seed set according to the initialization result includes: Among them, is the triple with the minimum distance, represents a distance metric function within the positive real number domain, is the set of vertex triples in the graph for solving the local absolute scale of the shared image, is relative to the image i , the image k and the image l is the depth ratio between them, is relative to the image i , the image j and the image l is the depth ratio between them, is relative to the image i , the image j and the image k is the depth ratio between them.

5. The incremental motion averaging method according to claim 3, wherein The method for combining an edge selection reward mechanism, expanding the set of added vertices, and synchronously performing incremental solution of the local absolute scale sets of each camera through local optimization and global optimization includes: Denote the edge set between any vertex that has not been added currently and each vertex in the set of added vertices as ; Based on the constraint relationship between the depth ratio and the local absolute scale, using the estimated values of the local absolute scale of the vertices in the set of added vertices and the depth ratio measurement values on the edges in the edge set to pre-compute the local absolute scale of the vertex ; ​​ According to the edge selection rewards of each edge in the edge set the initialization solution results of the added incremental points and the local absolute scale are obtained; An alternating local and global optimization strategy is used to perform incremental solution of the local absolute scale sets of each camera; According to the incremental solution results of the local absolute scale sets of each camera, the global relative scale of the camera is calculated.

6. The incremental motion averaging method according to claim 3, characterized in that The method for using an incremental process to synchronously solve multiple types of absolute pose parameters of the camera includes: Based on the epipolar geometry graph corresponding to the image set of the scene to be reconstructed, according to the constraint relationship between the relative pose and the absolute pose of the camera, as well as the relative baseline length formula, initialize the absolute pose parameters of the cameras in the camera triple set in the epipolar geometry graph and the relative baseline length used to calculate the absolute position; According to the initialization result, based on the principle of the best overall ranking and cycle consistency, obtain the triple with the best consistency of multiple types of parameters as the incremental seed set; Using the incremental seed set as the set of incrementally added vertices, utilize the absolute scale, rotation, and position estimation values of any vertex in the set of incrementally added vertices, the local absolute scale estimation value of any vertex in the set of incrementally unadded vertices, and the relative scale, rotation, and translation measurement values on any edge in the incremental edge set to pre-calculate the absolute scale, rotation, and position of the vertices in the set of incrementally unadded vertices; According to the pre-calculation result, based on the principle of maximizing the intersection scale of the support sets of multiple types of parameters, complete the selection of incremental vertices to obtain the incremental vertex set, and simultaneously complete the estimation of multiple types of camera pose parameters; Adopt an optimization mode that alternates between local and global scopes. Use the estimated values of the current multiple types of camera pose parameters as the initial values for optimization, and obtain the constraints provided by the set of internal values of the camera relative motion measurements through reverse calculation to optimize the absolute pose parameters of the camera and complete the synchronous solution of the camera pose.

7. The incremental moving average method according to claim 6, wherein When adopting the optimization mode that alternates between local and global scopes, it also includes: First optimize the absolute scale and rotation of the camera, then fix the absolute scale and rotation of the camera, and optimize the absolute position of the camera.

8. An incremental moving average system, characterized in that, It includes: An ambiguity elimination module, which is used to combine the cross-view feature matching information with the image set of the scene to be reconstructed, and use the scale separation strategy to eliminate the ambiguity in camera positioning to obtain the global relative scale of the camera; The scale separation strategy includes: obtaining a seed set and solving based on a distance metric function in the positive real number domain; gradually adding incremental vertices to the seed set and solving; tuning the local absolute scale parameters of the camera; A solution module, which is used to synchronously solve multiple types of absolute pose parameters of the camera by an incremental process according to the global relative scale of the camera, the rotation and translation of the camera, and the local absolute scale of the camera.

9. An electronic device, characterized in that, It includes: A memory and one or more processors; the memory is coupled to the processor; wherein, computer program code is stored in the memory, and the computer program code includes computer instructions. When the computer instructions are executed by the processor, the electronic device executes the steps of the incremental motion averaging method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, and when the computer program is executed by the processor, it realizes the steps of the incremental motion averaging method according to any one of claims 1-7.