Motion averaging method and system, electronic equipment and readable storage medium
Through the partitioning and consolidation solution strategy of adaptive grouping and reference alignment, the accuracy and robustness of the global motion recovery structure method in large-scale scene data is solved, and efficient and accurate camera pose solution is achieved.
Patent Information
- Application Number
- CN202510501377.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-07-18
AI Technical Summary
When facing large-scale scenario data, the existing global motion recovery structure method is restricted, has insufficient robustness and is inefficient in solving the upper limit of solution accuracy.
The adaptive grouping method is used to group the external polar geometric maps. Based on the cross-group support set maximization and group scale equalization, combined with the comprehensive ranking optimal principle and vertex coverage information, an alignment reference set is constructed to realize the synchronous position solution and global alignment of the camera group.
It improves the accuracy and robustness of camera group estimation, reduces error accumulation, improves calculation efficiency, is highly adaptable, and can handle complex scenarios.
Smart Images

Figure CN120339379A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to a solution method, and specifically relates to a motion averaging method, system, electronic device, and readable storage medium. Background Art
[0002] In the 3D reconstruction work of large-scale scenes based on images, the Structure from Motion (SfM) technology is extremely crucial. The SfM technology uses the feature matching relationship between images to solve the absolute poses of a large number of cameras in a globally unified coordinate system, becoming the core step of 3D reconstruction work.
[0003] The existing SfM methods mainly fall into two categories: (1) The incremental method uses iterative optimization to progressively solve the camera pose. (2) The global method uses motion averaging to solve the camera pose at once. The two methods complement each other in terms of solution accuracy and robustness, as well as solution efficiency and consistency. Compared with the incremental method, the global method has the characteristics of fewer optimization rounds and solution parameters.
[0004] As the core problem of the global method, motion averaging aims to solve the absolute pose of the camera based on the relative motion of the cameras. Although there has been relatively sufficient research, when facing the increasingly expanding scene and data scale, the existing methods for solving the motion averaging problem on the entire epipolar geometry graph at once still need to be improved in terms of efficiency, scalability, etc. To address these problems, relevant scholars have proposed a motion averaging method based on a divide-and-conquer solution strategy, that is, using the strategies of problem decomposition, grouped solution, and solution merging to improve the solution efficiency and scalability of the method. However, when performing problem decomposition, using a pre-grouping method that has nothing to do with the problem to be solved cannot fully exploit the information related to the current problem to be solved on the epipolar geometry graph, restricting the upper limit of the solution accuracy of the motion averaging problem. In addition, when performing solution merging, existing methods mostly use the relative motion between groups or the common visible space points with higher solution uncertainty to estimate the transformation of the local coordinate systems between groups, so there is still room for further improvement in the robustness of problem solving. Summary of the Invention
[0005] This application aims at the technical problems of the global method in the existing Structure from Motion technology, such as the restricted upper limit of solution accuracy and insufficient solution robustness, and provides a motion averaging method, system, electronic device, and readable storage medium.
[0006] To achieve the above object, this application adopts the following technical solutions: In the first aspect, this application proposes a motion averaging method, including: Performing adaptive grouping on the epipolar geometry graph to obtain an initial camera grouping; Based on maximizing the cross-group support set and equalizing the group size, combined with the initial camera grouping, camera group estimation and the synchronous pose solution of each camera are completed; Based on the principle of the optimal comprehensive ranking, an alignment reference seed set is constructed; Combined with the vertex coverage information, vertices are gradually added to the alignment reference seed set to obtain an alignment reference set for aligning each camera group; and taking the local coordinate system where the alignment reference set is located as the global coordinate system, the absolute poses of each camera in the alignment reference set in the global coordinate system are calculated synchronously; Each camera group is respectively aligned with the alignment reference set as the alignment standard in the global coordinate system to complete motion averaging.
[0007] Further, the method for adaptively grouping the epipolar geometry graph includes: using the community discovery algorithm to adaptively group the large-scale epipolar geometry graph to obtain multiple subgraphs; Each of the multiple subgraphs is processed respectively to obtain the incremental seed set of each subgraph and the initial solution result of the camera pose as the initial camera grouping result.
[0008] Further, the method for obtaining the incremental seed set of each subgraph and the initial solution result of the camera pose includes: All triple sets are searched respectively in each subgraph; The initial absolute pose values of each camera in the local coordinate system of the triple are calculated respectively through the relative poses on the edges of each triple in the triple set; The relative poses on the edges of the triple are inversely calculated respectively through the initial absolute pose values of each camera in the local coordinate system of the triple to obtain the inverse calculated relative poses; The deviations between the relative poses on the edges of the triple and the inverse calculated relative poses are calculated respectively, and the triple with the smallest deviation is used as the incremental seed set, and the initialization result of the camera pose of the incremental seed set is used as the initial solution result of the camera pose.
[0009] Further, after obtaining the initial camera grouping, it further includes: According to the absolute pose estimation values of each vertex in the vertex set of the camera group for which the absolute pose of the camera has been estimated, and the relative pose measurement values on each edge in the edge set, the absolute pose of any vertex in the vertex set of the camera group for which the absolute pose of the camera has not been estimated is pre-calculated to obtain the intersection of the multi-type parameter support sets of the pose pre-calculation results of any vertex in the vertex set of the camera group for which the absolute pose of the camera has not been estimated; the edge set is the edge set between any vertex in the vertex set of the camera group for which the absolute pose of the camera has not been estimated and any vertex in the vertex set of the camera group for which the absolute pose of the camera has been estimated.
[0010] Further, the method for completing camera group estimation and synchronous pose solution of each camera by combining the initial camera grouping based on maximizing the cross-group support set and balancing the group size includes: All vertices in the vertex set of the camera group for which the absolute pose of the camera has not been estimated are successively calculated using the following objective function:
[0011] where, is the vertex number at the other end of the incremental point in the leading edge, is the added group, is the serial number of the incremental point, is the intersection of multi-type parameter support sets, is,[[]] is,[[]] is any vertex set in the vertex set of the camera for which the absolute pose of the camera has been estimated; Through , , the leading edge, incremental point and added group are determined, the camera group of all vertices in the vertex set of the camera for which the absolute pose of the camera has not been estimated is estimated, and the synchronous pose of each camera is solved.
[0012] Further, the method for constructing an alignment reference seed set based on the principle of optimal comprehensive ranking includes: Each triple is calculated respectively through the following formula:
[0013] where, represents the rank of sorted from small to large in the set ; , represents the spatial distance metric function, which is the rotation angle distance here, that is: ; , , ; , represents the spatial distance metric function, , represents the sum of the degrees of the three vertices in the triple , represents the comprehensive ranking function; The triple with the smallest comprehensive ranking function value is selected as the alignment reference seed set.
[0014] Further, the method of gradually adding vertices to the alignment reference seed set by combining vertex coverage information includes: Determine the set of potentially addable vertices; the elements in the set of potentially addable vertices include vertices in the entire vertex set that do not belong to the set of estimated poses but have a connection relationship with the vertices in the set of estimated poses. For the points in the set of potentially addable vertices, successively select the triple with the largest weighted sum according to the intersection scale of the multi-type parameter support sets and the edge set scale between the vertices in the intersection of the multi-type parameter support sets and other vertices in the covered vertex set, and add it to the alignment reference seed set until the union of the added vertices and the covered vertex set is the entire vertex set.
[0015] Further, the method of aligning each camera group with the alignment reference set as the alignment standard in the global coordinate system includes: If there are common vertices between the vertex set of the camera group and the alignment reference set, calculate the similarity transformation for the global alignment of the vertex set of the camera group through the following formula:
[0016] where, is the result of the scale similarity transformation, is the result of the rotation similarity transformation, is the result of the position similarity transformation, are respectively the camera absolute scale, rotation, and position of the vertex in the vertex set of the camera group, are respectively the camera absolute scale, rotation, and position of the vertex in the alignment reference set; If there are common vertices between the vertex set of the camera group and the alignment reference set, calculate the similarity transformation for the global alignment of the vertex set of the camera group through the following formula: .
[0017] In a second aspect, the present application proposes a motion averaging system, including: An initial grouping module for adaptively grouping the epipolar geometry graph to obtain an initial camera grouping; A grouping estimation module for completing camera group estimation and synchronous pose calculation of each camera by combining the initial camera grouping based on maximizing the cross-group support set and balancing the grouping scale; An alignment seed module for constructing an alignment reference seed set based on the principle of the optimal comprehensive ranking; An alignment standard module, which is used to combine vertex coverage information, gradually add vertices to the alignment reference seed set, and obtain an alignment reference set for aligning each camera group; and use the local coordinate system where the alignment reference set is located as the global coordinate system to synchronously calculate the absolute poses of each camera in the alignment reference set in the global coordinate system; An alignment module, which is used to align each camera group in the global coordinate system with the alignment reference set as the alignment standard to complete motion averaging.
[0018] In a third aspect, the present application proposes an electronic device, including: a memory and one or more processors; the memory is coupled to the processor; wherein, computer program code is stored in the memory, and the computer program code includes computer instructions. When the computer instructions are executed by the processor, the electronic device executes the steps of the above-mentioned motion averaging method.
[0019] In a fourth aspect, the present application proposes a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of the above-mentioned motion averaging method are implemented.
[0020] Compared with the prior art, the present application has the following beneficial effects: The present application proposes a motion averaging method. After obtaining the initial camera grouping, based on the maximization of the cross-group support set and the balance of the grouping scale, camera group estimation and synchronous pose solution of each camera are completed. Then, based on the principle of the optimal comprehensive ranking, an alignment reference set for aligning each camera group is constructed by combining vertex coverage information, and then each camera group is aligned with the alignment reference set as the alignment standard in the global coordinate system to complete motion averaging. The present application adopts a task-driven online camera grouping and alignment reference construction method to synchronously realize group estimation, reference construction, and the corresponding camera pose parameter estimation. In addition, to improve the calculation efficiency and accuracy, when solving and merging sub-problems, only camera pose parameters are used without involving the feature point spatial coordinates with larger parameter scale and uncertainty. Based on the divide-and-conquer solution strategy of benchmark alignment, the large-scale original problem is decomposed into several sub-problems and solved separately. At the same time, an alignment reference is additionally constructed, and the global alignment of the solution results of each sub-problem is realized by means of this reference, so as to realize the efficient and consistent solution of the large-scale original problem.
[0021] The present application also proposes a motion averaging system, an electronic device, and a computer storage medium, which have all the advantages of the above-mentioned motion averaging method. Description of the Drawings
[0022] To more clearly illustrate the technical solutions of the embodiments of the present application, the accompanying drawings required for the embodiments will be briefly introduced below. It should be understood that the following accompanying drawings only show some embodiments of the present application, and thus should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related accompanying drawings can also be obtained based on these drawings.
[0023] Figure 1 It is a schematic flowchart of a moving average method of the present application; Figure 2 It is a schematic diagram of a moving average system of the present application. Detailed implementation manners
[0024] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Usually, the components of the embodiments of the present application described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations.
[0025] Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application claimed, but merely represents the selected embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0026] It should be noted that similar reference numerals and letters denote similar items in the following accompanying drawings. Therefore, once an item is defined in one accompanying drawing, it does not need to be further defined and explained in subsequent accompanying drawings.
[0027] In the description of the embodiments of the present application, it should be noted that if terms such as "upper", "lower", "horizontal", "inner", etc. are used to indicate the orientation or positional relationship, it is based on the orientation or positional relationship shown in the accompanying drawings, or the orientation or positional relationship in which the invention product is usually placed when in use. It is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation of the present application. In addition, terms such as "first", "second", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0028] In addition, if the term "horizontal" appears, it does not mean that the component is required to be absolutely horizontal, but it can be slightly inclined. For example, "horizontal" only means that its direction is more horizontal relative to "vertical", and does not mean that the structure must be completely horizontal, but it can be slightly inclined.
[0029] In the description of the embodiments of the present application, it should also be noted that unless otherwise clearly specified and limited, if the terms "set", "installed", "connected", and "connected" appear, they should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected, or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific circumstances.
[0030] SfM technology is mainly divided into two categories: (1) Incremental SfM By gradually adding images and iteratively optimizing the camera poses and 3D points. For example, select two images, set the pose of one as the identity matrix, estimate the fundamental matrix or essential matrix through feature matching, decompose to obtain the pose of the other image, and triangulate to generate the initial 3D points. Each time, select the new image with the most matching points, estimate the pose through PnP (Perspective-n-Points), triangulate the new points, and perform local Bundle Adjustment (BA) optimization. Finally, perform global BA to eliminate the cumulative error.
[0031] (2) Global SfM Process all images at once, and jointly estimate all camera poses through a motion averaging strategy. Match all images pairwise to construct a view graph. Solve the camera directions by optimizing the global rotation consistency. Estimate the camera positions based on the rotation results. Finally, optimize all camera poses and 3D points.
[0032] Motion averaging plays a key role in smoothing noise and establishing consistency in the global method, but its efficiency and adaptability face challenges in complex scenarios. Regarding the bottleneck of the divide-and-conquer strategy for the motion averaging problem in large-scale epipolar geometry graphs, the core contradiction lies in the insufficient coupling between the decomposition-solution-merging process and the geometric structure, resulting in double losses in accuracy and robustness.
[0033] Based on the above situation, the present application proposes a motion averaging method, system, electronic device, and readable storage medium. The present application will be described in detail below with reference to the embodiments and the accompanying drawings.
[0034] As Figure 1 described, it is a schematic flowchart of a motion averaging method of the present application, which may include: S101, adaptively group the epipolar geometry graph to obtain an initial camera grouping.
[0035] Epipolar geometry is an important concept in computer vision that describes the geometric relationship between two camera views and reflects the projection relationship of spatial points in different views. Adaptive grouping is to preliminarily classify cameras according to certain characteristics of the epipolar geometry graph, such as the relative position between cameras, attitude differences, image feature matching situations, etc. It provides a basic framework for subsequent more precise processing, enabling cameras within the same group to have relatively similar geometric characteristics and facilitating subsequent synchronization and alignment operations.
[0036] S102, based on maximizing the cross-group support set and balancing the group size, combined with the initial camera grouping, complete the camera group estimation and the synchronous pose calculation of each camera.
[0037] Maximizing the cross-group support set is to find a set that can provide the maximum support and constraints between different camera groups. For example, the corresponding point relationships obtained through feature matching, etc. Maximizing the cross-group support set helps to establish a more accurate relative pose relationship between cameras and enhance the stability and reliability of the entire system. In addition, in order to avoid the situation where the number of cameras in some groups is too large or too small, affecting the efficiency and accuracy of subsequent processing, the group size can be balanced. This can make the computational amount within each group relatively uniform and also better utilize the information interaction between different groups.
[0038] Combined with the initial camera grouping for camera group estimation and pose calculation, based on the initial grouping, further complete the optimal grouping of all cameras through the above two principles, and calculate the relative pose of each camera within its respective group.
[0039] S103, based on the principle of the optimal comprehensive ranking, construct an alignment reference seed set.
[0040] The principle of the optimal comprehensive ranking can involve the consideration of multiple factors. For example, the observation quality of the camera, such as image clarity, the number of feature points, etc.; the importance of the position in the scene, such as whether it is at a key perspective; the degree of association with other cameras, etc.
[0041] Comprehensively evaluate and rank the cameras according to these factors, and select a part of the cameras with the optimal ranking as the alignment reference seed set. The cameras in these seed sets will serve as the basis for subsequent alignment operations, and their poses and features will be used to guide the alignment of other cameras.
[0042] S104, combined with the vertex coverage information, gradually add vertices to the alignment reference seed set to obtain an alignment reference set for aligning each camera group; and use the local coordinate system where the alignment reference set is located as the global coordinate system, and synchronously calculate the absolute poses of each camera in the alignment reference set in the global coordinate system.
[0043] The vertex coverage information reflects the coverage of different positions in the scene by the camera. By considering the vertex coverage, it can be ensured that the cameras added to the alignment reference set can cover the entire scene as comprehensively as possible, thereby improving the accuracy and integrity of the alignment. The process of gradually adding vertices is a step-by-step optimization process. By continuously selecting appropriate cameras to add to the alignment reference set, a set that can effectively align each camera group is finally obtained.
[0044] Taking the local coordinate system where the alignment reference set is located as the global coordinate system is to unify the coordinate system of the entire system, facilitating subsequent calculations and processing. Synchronously calculate the absolute poses of each camera in the global coordinate system, so that the position and pose of each camera can be described within a unified framework.
[0045] S105, respectively, align each camera group in the global coordinate system with the alignment reference set as the alignment standard to complete the motion averaging.
[0046] In the global coordinate system, each camera group takes the alignment reference set as a reference and adjusts its own pose so that the cameras within the group reach the best alignment state with the cameras in the alignment reference set. This process can be achieved through optimization algorithms (such as the least squares method, etc.). By minimizing the pose difference between the cameras within the group and the alignment reference, the alignment of the camera groups is realized.
[0047] After completing the alignment of all camera groups, the motion averaging of the entire system is realized. The result of the motion averaging can be used for various subsequent applications, such as 3D reconstruction, robot navigation, etc., providing accurate camera pose information for these applications.
[0048] In view of the potential risks such as error accumulation and efficiency decline existing in the existing moving average method when facing ultra-large-scale scenario data, the present application adopts a divide-and-conquer solution strategy based on benchmark alignment, decomposes the large-scale original problem, and additionally constructs an alignment benchmark to achieve global alignment. By finding the maximum support set between different groups, the present application makes full use of the geometric consistency relationship between cameras in different groups, can more comprehensively consider the constraint conditions between cameras, thereby improving the accuracy of camera group estimation and pose calculation. More constraint conditions can make the calculated camera poses more accurate and reduce error accumulation. The balance of the group scale ensures that the number of cameras in each group is relatively balanced, avoids the problem of uneven distribution of computing resources caused by some groups being too large or too small, helps to improve the overall computing efficiency, makes the processing time of each group relatively balanced, and there will be no situation where a certain group has too long a computing time due to too many cameras or insufficient information due to too few cameras. Based on the principle of the optimal comprehensive ranking, an alignment benchmark seed set is constructed, considering multiple factors to evaluate the importance and applicability of cameras, and can select the most representative and stable cameras as the seeds of the alignment benchmark. These seed cameras perform excellently in terms of observation quality, position importance, and degree of association with other cameras, providing a reliable basis for subsequent alignment operations, and helping to improve the accuracy and stability of alignment. When adding vertices to the alignment benchmark seed set, combining the vertex coverage information can ensure that the alignment benchmark set can comprehensively cover the scene. The obtained alignment benchmark set can better represent the geometric structure of the entire scene, making the camera group alignment based on it more accurate, better adapting to complex scene structures, and improving the adaptability of the moving average method to different scenes.
[0049] The present application fully considers various relationships between cameras and the geometric characteristics of the scene, and can effectively handle various changes and uncertainties in camera movement. For example, when facing different shooting scenes, camera layouts, and movement states, through operations such as adaptive grouping, cross-group constraints, and optimized selection of alignment benchmarks, it can automatically adapt and find the optimal moving average result, with strong robustness and adaptability.
[0050] The following is a more detailed embodiment of the moving average method of the present application to further elaborate on the present application: S201, online estimation of camera groups and synchronous calculation of poses.
[0051] When estimating camera groups and synchronously estimating poses for a large-scale epipolar geometry graph, the following method can be specifically adopted: (1) Adaptively construct a multi-group seed set based on the community discovery algorithm.
[0052] Specifically, the community discovery algorithm is used to adaptively group the large-scale epipolar geometry graph to obtain several subgraphs, that is, to group the cameras. Then, each subgraph is processed in turn to achieve the acquisition of the incremental seed set for each subgraph and the initial solution of the camera pose. The following methods can be specifically adopted: For each subgraph, first find all the triple sets in it. A triple refers to a three-camera subset in which three cameras are pairwise connected. Calculate the initial absolute pose of each camera in the local coordinate system of the triple using the relative poses on the edges in each triple, and inversely calculate the relative pose on the edge according to the absolute pose in the local coordinate system of the triple, and calculate the deviation between the inversely calculated relative pose and the relative pose measurement value. The triple with the smallest deviation is the incremental seed set, and the initialization result of its camera pose is the initial solution result of the camera pose.
[0053] It should be noted that the subgraph is used as the initial seed set, and the adaptive grouping result is only used for obtaining the incremental seed set and not for subsequent camera addition and pose solution. The incremental seed set is the basic initial camera grouping, and there are some cameras in each group.
[0054] Based on the principles of maximizing the cross-group support set and balancing the group size, online estimation of the camera group and synchronous solution of the pose are realized.
[0055] In the initial camera grouping result, the cameras within each camera group are the estimated cameras, and the rest are the unestimated cameras.
[0056] The vertex sets of the absolute poses of the currently unestimated and estimated cameras can be respectively denoted as and , where is the number of camera groups, a certain vertex in and a certain vertex set in . On this basis, using the estimated values of the absolute poses of the vertices in and the relative pose measurement values on the edges in , the absolute pose of (in the local coordinate system of each camera group) is pre-calculated, and then the intersection of the multi-type parameter support sets of the pose pre-calculation results is obtained. The specific calculation method can be: Given the absolute pose of each vertex and knowing the relative pose on the edge between the vertex and the vertex for which the absolute pose is to be estimated currently, the absolute pose can be pre-calculated, how many vertices are there in There are as many pre-calculation results as there are absolute poses.
[0057] Ideally, the various pre-calculation results are consistent, but in reality there are differences. Therefore, different pre-calculation results can all be obtained by back-calculating the relative poses on the edges and comparing them with the measured values of the actual relative poses. When the relative pose deviations of different parameter types are all relatively small, this edge is the support edge for the current pre-calculation result of the absolute pose, and all the support edges form the intersection of the support sets. It should be noted that different parameter types refer to scale, position, and rotation.
[0058] It should be noted that not only is it necessary to determine the serial numbers added to the vertices in , but it is also necessary to simultaneously determine the serial numbers of the added camera groups for this vertex . In addition, to prevent the difference in the number of cameras in each group in the online grouping results here from affecting the subsequent combined solution effect, when selecting the leading edge , incremental points , and added groups , in addition to following the principle of maximizing the cross-group support set, it is also necessary to introduce a constraint on the balance of the grouping scale according to the number of vertices already added to each group currently . It should be noted that the subscripts of the leading edge and incremental points represent serial numbers. The leading edge is the edge connecting two vertices, the incremental point represents the point to be added, and the added group represents the group to which it is added.
[0059] Therefore, the objective function for online incremental vertex selection and group estimation here is defined as follows:
[0060] By calculating cyclically according to the above formula (1), the added group and subscript , subscript are obtained, and then the leading edge, added group, and incremental points can be correspondingly obtained, and thus the online estimation of the camera group and the synchronous solution of the pose can be realized.
[0061] According to the above process, the groups of each vertex in are sequentially estimated and the absolute pose is solved until all the vertices are added to each camera group. It should be noted that the method for solving the absolute pose is similar to the previous method.
[0062] S202, Online construction of the reference set and synchronous solution of the pose.
[0063] This step solves the problem of how to construct a reference in the global coordinate system and obtain the absolute pose in the global coordinate system for subsequent alignment of each group.
[0064] When constructing the alignment reference, to balance the reference construction and the global alignment accuracy, a task-driven online solution mode is adopted to synchronously select the reference vertices and solve their poses, so that the alignment reference construction task closely serves the camera pose solving task based on the divide-and-conquer strategy. When selecting the alignment reference seed set, that is, the aforementioned reference vertices, the principle of the best comprehensive ranking is adopted. In practical applications, an optimal triple including three cameras can be selected first as the alignment reference seed set. However, to make the selected alignment reference seed set have a large enough vertex coverage, when calculating the comprehensive ranking, in addition to ranking the loop consistency of four key parameters including the global relative scale , relative rotation , relative baseline length , relative translation , the triple vertex coverage ranking is introduced, that is, the following comprehensive ranking function is adopted:
[0065] where represents the ranking in ascending order in the set , ; , represents the spatial distance metric function, which is the rotation angle distance here, that is: ; , , ; , represents the spatial distance metric function, which is the normalized vector angle distance here, that is: ; represents the sum of the degrees of the three vertices in
[0066] Through the above calculation, the smallest ranking is selected as the alignment reference seed set. Subsequently, based on the alignment reference seed set, more cameras are added to form the alignment reference set.
[0067] When performing incremental selection of the alignment reference set vertices and synchronous pose solution, similar to the selection of the reference seed set, the vertex coverage information of the vertices to be added needs to be considered additionally. Specifically, when selecting the next vertex to be added to the reference set, first, the potential vertex set to be added is no longer the set of all vertices whose poses have not been estimated currently , but the coverage vertex set of the set of vertices whose poses have been estimated currently , that is, the set of vertices that do not belong to in the vertex universe , but the vertex set that has connection relationships with the vertices; in addition, considering the vertex coverage, the incremental point selection function defined here also needs to take into account the intersection scale of the multi-type parameter support sets, as well as the vertices and the set other vertices in the edge set scale. The incremental point selection method for the final alignment benchmark is to maximize the following weighted sum: , where is the weighted harmonic coefficient, and the maximum weighted sum is selected and added to the alignment benchmark set. It should be noted that according to the characteristics of the task, the vertex addition stop condition for incremental alignment benchmark construction here also changes: the addition stops when the union of the added vertices and their covered vertex sets is the vertex universal set .
[0068] S203, Global Alignment of Camera Poses Guided by Common Vertices Based on the group estimation and benchmark construction results, the global alignment of the camera pose grouping solution results is carried out. To achieve simple and accurate global alignment, only the camera pose parameters are relied on here, and the common vertex guidance method is preferred. Specifically, the camera pose estimation results of each group and the benchmark can be respectively denoted as and , represents the vertex set of the th group, represents the alignment benchmark set. respectively refer to the camera absolute scale, rotation, and position of the vertex in the set, respectively refer to the camera absolute scale, rotation, and position of the vertex in the set.
[0069] If and are the same vertex, then the similarity transformation between the local coordinate systems can be calculated by the following formula:
[0070] If and there is an edge between them, then the similarity transformation between the and local coordinate systems can be calculated by the following formula:
[0071] Among them, Indicates in The local coordinate system where And The relative baseline length between. Comparing Equation (3) and Equation (4), it can be seen that the alignment transformation calculation method based on the common vertex only involves absolute pose parameters with higher accuracy and reliability. Therefore, when performing global alignment of the camera pose, the method guided by the common vertex is preferred and implemented based on the maximum voting principle. Specifically, if And There is a common vertex between, for each common vertex, the similarity transformation for Global alignment can be calculated by Equation (3) , in addition, for And For each edge between, the above similarity transformation can also be calculated by Equation (4). By separately calculating the scale, rotation, and translation parameter estimation distances, according to the degree of agreement of the similarity transformation estimated values, the elements of the estimated value set obtained by Equation (3) vote on each estimated value obtained by Equation (3), and The initial value is determined as the one with the most votes, the similarity transformation result calculated by Equation (3). If And There is no common vertex between, then for And For each edge between, the similarity transformation for Global alignment can be calculated by Equation (4) . For each element in the similarity transformation estimated value set, the scale, rotation, and translation parameter estimation distances are calculated separately to measure the degree of conformity of other elements in the set with the current element, realizing the internal voting of each element in the set, and The initial value is determined as the one with the most votes, the similarity transformation result calculated by Equation (4). Based on the above calculation results, global alignment of the camera pose based on the benchmark can be realized.
[0072] To verify the technical effects of this application, the method of this application was evaluated using the 1D SfM dataset. The evaluation task was absolute rotation estimation, that is, solving the rotation averaging problem, and the evaluation metric was the median error of absolute rotation. First, ablation experiments were conducted, and the experimental results are shown in Table 1. The rightmost column in Table 1 corresponds to the method of this application. As can be seen from Table 1, compared with other ablation experiment cases, the method of this application has the best comprehensive performance, and it can also be seen the effectiveness of the online camera group estimation, online construction of the reference set, and global alignment of camera poses proposed by the method of this application. Then, comparative experiments were conducted. The comparative methods included three methods based on Robust Loss (IRLS, MPLS, DESC), three methods based on outlier filtering (OMSTs, HRRA, HARA), and three methods based on deep learning (NeuRoRA, MSP, RAGO). The experimental results are shown in Table 2. As can be seen from Table 2, the method of this application achieved the overall best performance among all comparative methods, which further verified the effectiveness of the method of this invention.
[0073] Table 1 Results of Ablation Experiments
[0074] Table 2 Results of Comparative Experiments
[0075] As Figure 2 shown, it is a schematic diagram of the motion averaging system of this application, which may include: An initial grouping module for adaptively grouping the epipolar geometry graph to obtain an initial camera grouping; A grouping estimation module for completing camera group estimation and synchronous pose calculation of each camera by combining the initial camera grouping based on maximizing the cross-group support set and balancing the group scale; An alignment seed module for constructing an alignment reference seed set based on the principle of the best comprehensive ranking; An alignment standard module for gradually adding vertices to the alignment reference seed set in combination with vertex coverage information to obtain an alignment reference set for aligning each camera group; and using the local coordinate system where the alignment reference set is located as the global coordinate system to synchronously calculate the absolute poses of each camera in the global coordinate system in the alignment reference set; An alignment module for respectively aligning each camera group in the global coordinate system with the alignment reference set as the alignment standard to complete motion averaging.
[0076] It should be noted that in the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of each module is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules can be combined or integrated into another device, or some features can be ignored or not executed. The modules described as separate components may or may not be physically separated. The components shown as modules can be a physical unit or multiple physical units, that is, they can be located in one place or distributed to multiple different places. Part or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0077] In addition, in each embodiment of the present invention, each module can be integrated in a processing unit, or each module can exist physically separately, or two or more modules can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0078] The embodiments of the present application further provide an electronic device, which may include one or more processors, a memory, and a communication interface.
[0079] Among them, the memory and the communication interface are coupled to the processor. For example, the memory and the communication interface can be coupled together through a bus.
[0080] Among them, the communication interface is used for data transmission with other devices. The memory stores computer program code. The computer program code includes computer instructions. When the computer instructions are executed by the processor, the electronic device is caused to execute the steps of the above-mentioned moving average method.
[0081] Among them, the processor can be a processor or a controller. For example, it can be a Central Processing Unit (CPU), a general-purpose processor, a Digital Signal Processor (DSP), an Application-Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logic blocks, modules, and circuits described in connection with the present disclosure. The processor can also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and so on. The processor can be used to support the electronic device in executing the method steps provided in the above embodiments.
[0082] Among them, the bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The above buses can be divided into an address bus, a data bus, a control bus, etc.
[0083] A computer-readable storage medium provided by an embodiment of the present application stores a computer program, and when the computer program is executed by a processor, the steps of the above moving average method are implemented.
[0084] The computer-readable storage medium involved in the present application includes a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD ROM, or any other form of storage medium known in the technical field.
[0085] The above are only the preferred embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A moving average method, characterized in that, Including: Performing adaptive grouping on the epipolar geometry graph to obtain an initial camera grouping; Based on maximizing the cross-group support set and balancing the grouping scale, combined with the initial camera grouping, completing camera group inference and synchronous pose calculation for each camera; Based on the principle of the best comprehensive ranking, constructing an alignment reference seed set; Combined with vertex coverage information, gradually adding vertices to the alignment reference seed set to obtain an alignment reference set for aligning each camera group; And using the local coordinate system where the alignment reference set is located as the global coordinate system, synchronously calculating the absolute poses of each camera in the global coordinate system in the alignment reference set; Respectively making each camera group align with the alignment reference set as the alignment standard in the global coordinate system to complete motion averaging.
2. The moving average method according to claim 1, wherein The method for performing adaptive grouping on the epipolar geometry graph includes: using a community discovery algorithm to perform adaptive grouping on a large-scale epipolar geometry graph to obtain multiple subgraphs; Processing the multiple subgraphs respectively to obtain the incremental seed set and the initial camera pose solution result of each subgraph as the initial camera grouping result.
3. The moving average method according to claim 2, wherein The method for obtaining the incremental seed set and the initial camera pose solution result of each subgraph includes: Searching for all triple sets in each subgraph respectively; Calculating the initial absolute pose of each camera in the local coordinate system of the triple through the relative poses on each edge in the triple set; Respectively back-calculating the relative pose on the triple edge through the initial absolute pose of each camera in the local coordinate system of the triple to obtain the back-calculated relative pose; Calculating the deviation between the relative pose on the triple edge and the back-calculated relative pose respectively, taking the triple with the smallest deviation as the incremental seed set, and taking the camera pose initialization result of the incremental seed set as the initial camera pose solution result.
4. The moving average method according to claim 1, wherein After obtaining the initial camera grouping, it further includes: According to the absolute pose estimation values of each vertex in the vertex set of the camera group for which the absolute pose of the camera has been estimated, and the relative pose measurement values on each edge in the edge set, pre-calculating the absolute pose of any vertex in the vertex set of the camera group for which the absolute pose of the camera has not been estimated, to obtain the intersection of the multi-type parameter support sets of the pose pre-calculation results of any vertex in the vertex set of the camera group for which the absolute pose of the camera has not been estimated; the edge set is the edge set between any vertex in the vertex set of the camera group for which the absolute pose of the camera has not been estimated and any vertex in the vertex set of the camera group for which the absolute pose of the camera has been estimated.
5. The moving average method according to claim 1, wherein The method for completing camera group inference and synchronous pose calculation for each camera based on maximizing the cross-group support set and balancing the grouping scale, combined with the initial camera grouping, includes: Successively calculating all vertices in the vertex set of the camera group for which the absolute pose of the camera has not been estimated using the following objective function: wherein, is the vertex number at the other end of the incremental point in the leading edge, is the addition group, is the serial number of the incremental point, is the intersection of the multi-type parameter support sets, is, is, is any vertex set in the vertex set of the already estimated absolute camera pose; By , , Determine the leading edge, incremental points, and addition groups, complete the camera group inference for all vertices in the vertex set of the cameras whose absolute poses have not been estimated in the camera group, and perform synchronous pose solution for each camera.
6. The moving average method according to claim 1, wherein The method for constructing an alignment reference seed set based on the principle of the best comprehensive ranking includes: Calculating each triple respectively through the following formula: Among them, represents the ranking in ascending order in the set , ; , represents the spatial distance metric function, which is the rotation angle distance here, that is: ; , , ; , represents the spatial distance metric function, , represents the sum of the degrees of the three vertices in the triple , represents the comprehensive ranking function; Selecting the triple with the smallest comprehensive ranking function value as the alignment reference seed set.
7. The moving average method according to claim 1, wherein The method for gradually adding vertices to the alignment reference seed set in combination with vertex coverage information includes: Determine the set of potential added vertices; the elements in the set of potential added vertices include the vertices in the entire vertex set that do not belong to the estimated pose set but have a connection relationship with the vertices in the estimated pose set. For the points in the set of potential added vertices, successively select the triple with the largest weighted sum according to the intersection scale of the multi-type parameter support sets and the edge set scale between the vertices in the intersection of the multi-type parameter support sets and other vertices in the covered vertex set, and add it to the alignment reference seed set until the union of the added vertices and the covered vertex set is the entire vertex set.
8. The moving average method according to claim 1, characterized in that, The method of respectively aligning each camera group with the alignment reference set as the alignment standard in the global coordinate system includes: If there are common vertices between the camera group vertex set and the alignment reference set, calculate the similarity transformation for the global alignment of the camera group vertex set through the following formula: Among them, is the result of scale similarity transformation, is the result of rotation similarity transformation, is the result of position similarity transformation, are respectively the camera absolute scale, rotation, and position of vertex in the vertex set of the camera group, are respectively the camera absolute scale, rotation, and position of vertex in the alignment reference set; If there are common vertices between the camera group vertex set and the alignment reference set, calculate the similarity transformation for the global alignment of the camera group vertex set through the following formula: 。 9. A moving average system, characterized in that, Including: An initial grouping module for adaptively grouping the epipolar geometry graph to obtain an initial camera grouping. A grouping estimation module for completing camera group estimation and synchronous pose calculation of each camera by combining the initial camera grouping based on maximizing the cross-group support set and balancing the grouping scale. An alignment seed module for constructing an alignment reference seed set based on the principle of the best comprehensive ranking. An alignment standard module for gradually adding vertices to the alignment reference seed set in combination with vertex coverage information to obtain an alignment reference set for aligning each camera group. And use the local coordinate system where the alignment reference set is located as the global coordinate system, and synchronously calculate the absolute poses of each camera in the alignment reference set in the global coordinate system. An alignment module for respectively aligning each camera group with the alignment reference set as the alignment standard in the global coordinate system to complete motion averaging.
10. An electronic device, characterized in that, Including: A memory and one or more processors; the memory is coupled to the processor; wherein, computer program code is stored in the memory, and the computer program code includes computer instructions. When the computer instructions are executed by the processor, the electronic device executes the steps of the motion averaging method according to any one of claims 1-8.
11. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, the steps of the motion averaging method according to any one of claims 1-8 are implemented.