Multi-view oriented three-dimensional scene image reconstruction registration and optimization method and system
By employing adaptive feature extraction and cross-view semantic association graph construction, combined with multi-objective joint optimization and dynamic weight adjustment, the problems of poor feature matching accuracy and low registration precision in traditional methods are solved, achieving high-precision 3D scene reconstruction.
Patent Information
- Application Number
- CN202511805963.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-03
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2045-12-03
AI Technical Summary
Traditional 3D scene reconstruction methods struggle to handle registration between images from multiple perspectives in complex environments, especially when there are occlusions, reflections, or areas with simple textures. Feature matching accuracy is poor, and existing optimization methods lack effective feedback mechanisms, resulting in low registration accuracy and reconstruction results that are prone to distortion or breakage.
By employing adaptive feature extraction and cross-perspective semantic association graph construction, semantic consistency constraints are established to achieve adaptive mapping of feature points. Furthermore, through a multi-objective joint optimization function and a dynamic adaptive weight adjustment mechanism, the dynamic optimization and refinement of feature correspondences are driven. Combining spatial geometric constraints and semantic association information forms a closed-loop coupling, thereby improving the accuracy and robustness of feature matching.
It significantly improves the accuracy and robustness of feature matching in complex scenes, accurately captures the relative pose and local deformation between viewpoints, avoids error accumulation, and improves the geometric consistency and spatial accuracy of 3D reconstruction. It is suitable for high-precision 3D scene reconstruction in complex environments.
Smart Images

Figure CN121259059B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to image processing technology, and in particular to a multi-view oriented three-dimensional scene image reconstruction registration and optimization method and system. BACKGROUND
[0002] Three-dimensional scene image reconstruction technology has wide applications in computer vision, virtual reality, augmented reality, intelligent driving and other fields. Multi-view three-dimensional scene reconstruction is to recover the geometric structure and spatial features of the scene through image data obtained from different angles, and then construct a complete three-dimensional model. This technology involves image feature extraction, feature matching, spatial registration and optimization, and is one of the core research directions in computer vision.
[0003] Traditional three-dimensional scene reconstruction methods usually include feature point-based reconstruction methods, voxel-based reconstruction methods and point cloud-based reconstruction methods. These methods have achieved certain results in specific application scenarios, but still have significant challenges in complex environments. With the development of deep learning technology, neural network-based three-dimensional scene reconstruction methods have gradually emerged, which improves the accuracy and robustness of reconstruction by learning feature representation and spatial geometric relationship.
[0004] Traditional feature extraction methods are difficult to cope with complex textures and lighting changes in the scene, resulting in a lack of consistency in features extracted under different angles, affecting the accuracy of subsequent feature matching. Especially in the presence of occlusion, reflection or single texture area, feature point recognition and matching often have serious errors, thereby affecting the overall reconstruction quality.
[0005] Existing technologies often use global transformation models when dealing with registration problems between multi-view images, which are difficult to accurately describe the local non-rigid deformation existing in the scene. This leads to a significant decrease in registration accuracy in scenes with complex structures or dynamic elements, especially in cases where the angle difference is large, the reconstruction result often appears distorted or broken.
[0006] Existing reconstruction optimization methods usually treat feature matching and spatial registration as independent steps, lack effective feedback mechanisms, and cannot realize mutual optimization of registration results and feature correspondence. This one-way processing flow limits the system's ability to correct errors, especially in cases where the initial registration has large errors, the optimization process is easily trapped in local optimal solution, and it is difficult to obtain a globally optimal reconstruction result. SUMMARY
[0007] The embodiments of the present application provide a multi-view oriented three-dimensional scene image reconstruction registration and optimization method and system, which can solve the problems in the prior art.
[0008] The first aspect of the embodiment of the present application provides a multi-view-oriented three-dimensional scene image reconstruction registration and optimization method, comprising:
[0009] Multi-frame three-dimensional scene image data from different spatial positions is acquired, adaptive feature extraction is performed on the multi-frame three-dimensional scene image data, and a multi-level feature set containing spatial positions and geometric structures is obtained;
[0010] A cross-view semantic association graph is constructed based on the multi-level feature set, adaptive mapping of corresponding feature points of the same spatial entity under different views is realized by establishing semantic consistency constraints, and a feature correspondence relationship with confidence evaluation is obtained;
[0011] A multi-view spatial transformation relationship is constructed according to spatial position information and geometric structure information contained in the feature correspondence relationship, and transformation parameters describing relative poses and local deformations between views are obtained through iterative calculation;
[0012] The transformation parameters are used to align coordinate systems of the multi-frame three-dimensional scene image data, an initial three-dimensional reconstruction result under a unified coordinate system is generated, a local area with registration deviation is identified based on a re-projection error distribution of spatial points in the initial three-dimensional reconstruction result, and a multi-objective joint optimization function is constructed for the local area;
[0013] The multi-objective joint optimization function is integrated into a dynamic adaptive weight regulation mechanism to guide evolution of the initial three-dimensional reconstruction result under spatial geometric constraints, and spatial configuration information obtained in the evolution process is coupled with the cross-view semantic association graph to drive dynamic optimization and refinement of the feature correspondence relationship.
[0014] The cross-view semantic association graph is constructed based on the multi-level feature set, adaptive mapping of corresponding feature points of the same spatial entity under different views is realized by establishing semantic consistency constraints, and a feature correspondence relationship with confidence evaluation is obtained, comprising:
[0015] For each feature point in the multi-level feature set, spatial information of the feature point under different scales, including local structural features and global context features, is extracted, the local structural features and the global context features are fused, and a composite feature descriptor of the feature point is obtained;
[0016] The cross-view semantic association graph is constructed based on the composite feature descriptor, semantic propagation constraints are established in the cross-view semantic association graph, unconfirmed candidate correspondence relationships are guided by existing reliable correspondence relationships, and an initial semantic association strength distribution is obtained;
[0017] Based on the initial semantic association strength distribution, the corresponding feature points of the same spatial entity under different perspectives are analyzed, the semantic consistency deviation between each candidate corresponding point and the surrounding confirmed correspondence is calculated, and the semantic consistency deviation is used as a constraint to optimize the initial semantic association strength distribution to obtain the corrected semantic association strength.
[0018] Based on the corrected semantic association strength and the stability of the semantic consistency deviation, a confidence assessment model is established. The reliability of semantic propagation and the consistency of association strength are evaluated through the confidence assessment model, and a confidence assessment value is generated for the corresponding feature point.
[0019] The confidence assessment value is used to update the correspondence in the cross-view semantic association graph. Correspondences that exceed the dynamic threshold are determined as reliable feature point mappings. The confidence assessment value is used to mark the reliability of the correspondence, and finally, feature correspondences with quality assessment are obtained.
[0020] Based on the composite feature descriptor, a cross-perspective semantic association graph is constructed. Semantic propagation constraints are established in the cross-perspective semantic association graph. Existing reliable correspondences guide unconfirmed candidate correspondences, resulting in an initial semantic association strength distribution including:
[0021] For each feature point in the composite feature descriptor, the semantic similarity of the feature points across the cross-view range is analyzed, and a cross-view semantic association graph is generated by combining the spatial location information of the feature points; feature point pairs that satisfy the bidirectional optimal matching criterion are determined in the cross-view semantic association graph, the feature point pairs are defined as reliable correspondences, and the relative positional relationship and local neighborhood topology of the reliable correspondences in three-dimensional space are analyzed to construct the spatial structure features of the reliable correspondences.
[0022] Using the spatial structural features, a constraint propagation network is constructed in the cross-perspective semantic association graph with the reliable correspondence as the propagation source point, and the spatial structural features are used to form topological consistency constraints on the propagation path.
[0023] For the unconfirmed candidate correspondences in the cross-perspective semantic association graph, the topological differences between the candidate correspondences and the reliable correspondences on the propagation path are evaluated and mapped to a propagation attenuation coefficient that characterizes the degree of constraint satisfaction.
[0024] The semantic similarity of the candidate correspondences is optimized by using the propagation attenuation coefficient. By progressively transmitting the spatial structural features of reliable correspondences in the constraint propagation network, the topological constraint modulation of the candidate correspondences is achieved, generating a semantic association strength distribution.
[0025] The transformation parameters are used to align the coordinate systems of the multi-frame 3D scene image data, generating an initial 3D reconstruction result in a unified coordinate system. Based on the reprojection error distribution of spatial points in the initial 3D reconstruction result, local regions with registration deviations are identified. A multi-objective joint optimization function is constructed for these local regions, including:
[0026] The local coordinate system of each frame in the multi-frame 3D scene image data is transformed to a unified coordinate system using the transformation parameters. A feature fusion network is established under the unified coordinate system to fuse and reconstruct spatial points from different frames, generating an initial 3D reconstruction result.
[0027] For each spatial point in the initial 3D reconstruction result, the deviation between the reprojection position and the actual observation position of each spatial point under each observation view is analyzed based on the feature fusion network, and the spatial distribution of the deviation is statistically analyzed to obtain the reprojection error distribution characteristics.
[0028] Based on the reprojection error distribution characteristics, identify the set of spatial points exceeding the error threshold in the initial 3D reconstruction result, construct a spatial structure optimization map, analyze the connectivity and clustering of the set of spatial points, and divide the error spatial points with spatial connectivity into local regions with registration deviations.
[0029] For the local region, a regional optimization objective is constructed in the spatial structure optimization diagram, which includes geometric consistency constraints and minimization of reprojection error. At the same time, the boundary continuity constraint relationship between the local region and the adjacent regions is analyzed.
[0030] Based on the region optimization objective and the boundary continuity constraint, a multi-objective joint optimization function is established to ensure that adjacent local regions maintain spatial continuity and geometric consistency at the boundary.
[0031] Based on the reprojection error distribution characteristics, the set of spatial points exceeding the error threshold in the initial 3D reconstruction result is identified. A spatial structure optimization map is constructed, and the connectivity and clustering of the spatial point set are analyzed. Error spatial points with spatial connectivity are divided into local regions with registration deviations, including:
[0032] Based on the reprojection error distribution characteristics, spatial points whose reprojection errors exceed the error threshold are identified in the initial 3D reconstruction results, and the 3D coordinates, normal vectors and observation view information of the spatial points are extracted to obtain a set of spatial points that exceed the error threshold.
[0033] Using each spatial point in the set of spatial points as a node, the geodesic distance and local curvature consistency between pairs of spatial points are calculated, and weighted edges are established between pairs of spatial points that satisfy the distance and curvature constraints. The weight of the weighted edges is jointly measured by the geodesic distance and curvature consistency to construct a spatial structure optimization graph.
[0034] In the spatial structure optimization graph, the weighted edge connections are traversed using a graph connectivity analysis algorithm to identify spatial point subsets that form connected subgraphs through the weighted edge connections. The cluster density is calculated based on the weighted edge weight distribution within the spatial point subsets to obtain connectivity analysis results and clustering analysis results.
[0035] Based on the connectivity analysis results and the clustering analysis results, the connected subgraphs in the spatial structure optimization graph are screened, and the connected subgraphs with clustering density exceeding the clustering threshold are marked as candidate local regions. Based on the distribution characteristics of the internal spatial points of the candidate local regions, the boundaries are adjusted, and the candidate local regions after boundary adjustment are determined as a set of local regions with registration deviation.
[0036] The multi-objective joint optimization function is integrated into a dynamic adaptive weight control mechanism to guide the evolution of the initial 3D reconstruction result under spatial geometric constraints. Simultaneously, the spatial configuration information obtained during the evolution process is coupled in a closed loop with the cross-view semantic association graph, driving the dynamic optimization and refinement of feature correspondences, including:
[0037] For each optimization objective term in the multi-objective joint optimization function, the convergence rate and contribution weight of each objective term are calculated based on the current residual distribution of the initial 3D reconstruction result, and an adaptive weight adjustment mechanism is constructed.
[0038] The multi-objective joint optimization function is reconstructed using the adaptive weight adjustment mechanism. A multi-scale geometric constraint network is constructed in the reconstructed multi-objective joint optimization function. The multi-scale geometric constraint network limits the adjustment range of spatial point positions and the topological stability of the neighborhood, thereby guiding the structural evolution of the initial three-dimensional reconstruction result.
[0039] Based on the multi-scale geometric constraint network, the positional changes of spatial points and the local topological structure changes during the structural evolution are extracted, and a self-attention aggregation mechanism is introduced to calculate the configurational stability measure of spatial points to obtain spatial configuration information.
[0040] The configuration stability measure in the spatial configuration information is converted into the spatial geometric confidence of feature point pairs through the self-attention aggregation mechanism. The spatial geometric confidence is then used to modulate and weight the semantic association strength of feature point pairs in the cross-view semantic association graph to form a closed-loop coupling.
[0041] The confidence level of the correspondence between feature point pairs is calculated based on the modulated cross-view semantic association graph. Feature point pairs below the confidence threshold are rematched, and the matching results and the spatial configuration information are fed back to the weighted reconstructed multi-objective joint optimization function to drive the dynamic optimization and refinement of feature correspondence.
[0042] The multi-objective joint optimization function is reconstructed using the adaptive weight adjustment mechanism. A multi-scale geometric constraint network is then constructed within the reconstructed multi-objective joint optimization function. This multi-scale geometric constraint network limits the spatial point position adjustment range and neighborhood topological stability, including:
[0043] Obtain the adaptive weight coefficients of each optimization objective term output by the adaptive weight adjustment mechanism, and use the adaptive weight coefficients to dynamically weight and combine the geometric consistency objective term and the topology preservation objective term in the multi-objective joint optimization function to obtain the weighted reconstructed multi-objective joint optimization function;
[0044] A multi-scale geometric constraint network is constructed based on the initial 3D reconstruction results. The fine-grained constraint layer constructs local differential geometric constraints by calculating the local surface tangent plane deviation between spatial points and nearest neighbor points, while the coarse-grained constraint layer constructs global structural geometric constraints by calculating the global position distribution entropy of spatial points and the intersection consistency of multi-view observation rays.
[0045] The local differential geometric constraints and the global structural geometric constraints are embedded into the weighted reconstructed multi-objective joint optimization function. The allowable position adjustment vector of the spatial points is calculated based on the local differential geometric constraints, and the global position drift penalty factor is calculated based on the global structural geometric constraints.
[0046] Based on the allowable position adjustment vector and the global position drift penalty factor, a neighborhood topology connection graph is constructed. By extracting the set of topological connection edges between spatial points and neighborhood points and calculating the edge length change rate and edge angle change rate of the topological connection edges, the neighborhood topology stability index is obtained.
[0047] A second aspect of this invention provides a system for reconstructing, registering, and optimizing three-dimensional scene images from multiple perspectives, comprising:
[0048] The first unit is used to acquire multi-frame 3D scene image data from different spatial locations, and perform adaptive feature extraction on the multi-frame 3D scene image data to obtain a multi-level feature set containing spatial location and geometric structure.
[0049] The second unit is used to construct a cross-perspective semantic association graph based on the multi-level feature set. By establishing semantic consistency constraints, it realizes the adaptive mapping of corresponding feature points of the same spatial entity under different perspectives, and obtains feature correspondence with confidence assessment.
[0050] The third unit is used to construct a multi-view spatial transformation relationship based on the spatial location information and geometric structure information contained in the feature correspondence relationship, and to obtain the transformation parameters describing the relative attitude and local deformation between viewpoints through iterative calculation.
[0051] The fourth unit is used to perform coordinate system alignment on the multi-frame 3D scene image data using the transformation parameters, generate an initial 3D reconstruction result under a unified coordinate system, and identify local regions with registration deviations based on the reprojection error distribution of spatial points in the initial 3D reconstruction result, and construct a multi-objective joint optimization function for the local regions.
[0052] The fifth unit is used to integrate the multi-objective joint optimization function into a dynamic adaptive weight control mechanism, guide the evolution of the initial 3D reconstruction result under spatial geometric constraints, and simultaneously couple the spatial configuration information obtained during the evolution process with the cross-view semantic association graph to drive the dynamic optimization and refinement of feature correspondence.
[0053] A third aspect of the present invention provides an electronic device, comprising:
[0054] processor;
[0055] Memory used to store processor-executable instructions;
[0056] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.
[0057] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.
[0058] The beneficial effects of this application are as follows:
[0059] The adaptive feature extraction and cross-view semantic association graph construction method proposed in this invention can effectively handle the apparent differences of scene features under different viewpoints. By establishing semantic consistency constraints, it achieves accurate correspondence of feature points, significantly improves the accuracy and robustness of feature matching in complex scenes, and solves the problem of feature matching difficulties in traditional methods when the viewpoint changes greatly.
[0060] The multi-view spatial transformation relationship construction and coordinate system alignment method designed in this invention can accurately capture the relative posture and local deformation between viewpoints. By accurately identifying the registration deviation area through reprojection error analysis, it effectively avoids the accumulation of errors in global optimization and improves the geometric consistency and spatial accuracy of 3D reconstruction.
[0061] The multi-objective joint optimization and dynamic adaptive weight control mechanism adopted in this invention forms a closed-loop coupling between spatial geometric constraints and semantic association information, realizing the synergistic optimization of feature correspondence and spatial configuration. This results in higher spatial structure accuracy while maintaining the integrity of details, making it suitable for high-precision 3D scene reconstruction tasks in complex environments. Attached Figure Description
[0062] Figure 1 This is a flowchart illustrating the method for multi-view 3D scene image reconstruction, registration, and optimization according to an embodiment of the present invention.
[0063] Figure 2 This is a flowchart of multi-level spatial constraints and topological stability analysis in an embodiment of the present invention. Detailed Implementation
[0064] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0065] The technical solution of the present invention will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.
[0066] Figure 1 This is a flowchart illustrating the multi-view 3D scene image reconstruction, registration, and optimization method according to an embodiment of the present invention. Figure 1 As shown, the method includes:
[0067] Acquire multi-frame 3D scene image data from different spatial locations, and perform adaptive feature extraction on the multi-frame 3D scene image data to obtain a multi-level feature set containing spatial location and geometric structure;
[0068] Based on the multi-level feature set, a cross-perspective semantic association graph is constructed. By establishing semantic consistency constraints, adaptive mapping of corresponding feature points of the same spatial entity under different perspectives is achieved, resulting in feature correspondence with confidence assessment.
[0069] Based on the spatial location information and geometric structure information contained in the feature correspondence, a multi-view spatial transformation relationship is constructed, and transformation parameters describing the relative attitude and local deformation between viewpoints are obtained through iterative calculation.
[0070] The transformation parameters are used to align the coordinate system of the multi-frame 3D scene image data to generate an initial 3D reconstruction result in a unified coordinate system. Based on the reprojection error distribution of spatial points in the initial 3D reconstruction result, local regions with registration deviations are identified, and a multi-objective joint optimization function is constructed for the local regions.
[0071] The multi-objective joint optimization function is integrated into a dynamic adaptive weight control mechanism to guide the evolution of the initial 3D reconstruction result under spatial geometric constraints. At the same time, the spatial configuration information obtained during the evolution process is coupled with the cross-view semantic association graph to form a closed loop, driving the dynamic optimization and refinement of feature correspondence.
[0072] In one optional implementation, a cross-perspective semantic association graph is constructed based on the multi-level feature set. By establishing semantic consistency constraints, adaptive mapping of corresponding feature points of the same spatial entity under different perspectives is achieved, resulting in feature correspondences with confidence assessment, including:
[0073] For each feature point in the multi-level feature set, the spatial information of the feature point at different scales is extracted, including local structural features and global context features. The local structural features and global context features are then fused to obtain a composite feature descriptor of the feature point.
[0074] Based on the composite feature descriptor, a cross-perspective semantic association graph is constructed. Semantic propagation constraints are established in the cross-perspective semantic association graph. Existing reliable correspondences guide unconfirmed candidate correspondences to obtain the initial semantic association strength distribution.
[0075] Based on the initial semantic association strength distribution, the corresponding feature points of the same spatial entity under different perspectives are analyzed, the semantic consistency deviation between each candidate corresponding point and the surrounding confirmed correspondence is calculated, and the semantic consistency deviation is used as a constraint to optimize the initial semantic association strength distribution to obtain the corrected semantic association strength.
[0076] Based on the corrected semantic association strength and the stability of the semantic consistency deviation, a confidence assessment model is established. The reliability of semantic propagation and the consistency of association strength are evaluated through the confidence assessment model, and a confidence assessment value is generated for the corresponding feature point.
[0077] The confidence assessment value is used to update the correspondence in the cross-view semantic association graph. Correspondences that exceed the dynamic threshold are determined as reliable feature point mappings. The confidence assessment value is used to mark the reliability of the correspondence, and finally, feature correspondences with quality assessment are obtained.
[0078] For each feature point in the multi-level feature set, the feature extraction module first extracts spatial information at three different scale levels. Local structural feature extraction uses circular neighborhood windows with radii of 8 pixels, 16 pixels, and 32 pixels. Within each window, a gradient direction histogram is calculated, dividing the 360-degree direction into 36 intervals, each with a width of 10 degrees. The cumulative sum of gradient magnitudes falling into each interval is used as the local structural description vector. Global context feature extraction recursively divides a 128-pixel × 128-pixel square region centered on the current feature point into minimum 8-pixel × 8-pixel sub-blocks using a quadtree partitioning method. The mean, variance, and principal direction angle of each sub-block are calculated to form a hierarchical global context description vector. The feature fusion process concatenates the local structural feature vector and the global context feature vector using a weighted connection method, with the local feature weight set to 0.6 and the global feature weight set to 0.4, generating a composite feature descriptor with a dimension of 256. The descriptor normalization process uses the L2 norm to ensure that the sum of squares of each dimension component is 1, thereby improving the stability of subsequent similarity calculations.
[0079] The cross-perspective semantic association graph construction module establishes a graph structure based on composite feature descriptors. Each node in the graph represents a feature point and its descriptor, and the edge weight is determined by calculating the cosine similarity between two descriptors. The association graph adopts a sparse storage method, retaining only edges with a similarity greater than 0.3 to reduce storage overhead. The semantic propagation constraint mechanism uses a confidence propagation algorithm, selecting feature point pairs with a similarity greater than 0.85 as the initial set of reliable correspondence seeds. During propagation, for each unconfirmed candidate correspondence, all confirmed correspondences within a 50-pixel radius of its neighborhood are collected, and the consistency scores between the candidate and confirmed correspondences are calculated in three dimensions: spatial location, orientation angle, and scale ratio. Spatial location consistency is measured by Euclidean distance: a score of 1.0 is given when the distance is less than 10 pixels, the score linearly decreases to 0.3 when the distance is between 10 and 30 pixels, and the score is 0 when the distance exceeds 30 pixels. Orientation consistency is calculated by the angle difference. A score of 1.0 is given when the difference is less than 15 degrees, the score decreases according to a cosine function when the difference is between 15 and 45 degrees, and the score is 0 when the difference exceeds 45 degrees. Scale consistency is calculated by comparing the scale parameters of feature points. A score of 1.0 is given when the scale difference is less than 0.2, and the score decreases linearly to 0.1 when the difference is between 0.2 and 0.6. The arithmetic mean of the scores in the three dimensions constitutes the initial semantic association strength distribution, with the strength values limited to the range of 0 to 1.
[0080] The semantic consistency constraint optimization module analyzes the distribution patterns of corresponding feature points of the same spatial entity from different perspectives. For each candidate corresponding point, the semantic consistency deviation between it and the surrounding confirmed correspondences is calculated. The deviation metric includes two components: spatial distribution deviation and feature similarity deviation. Spatial distribution deviation is calculated by constructing a Delaunay triangulation within a 60-pixel radius around the candidate point and comparing the similarity of the triangulation topology between different perspectives. Similarity is calculated by combining the ratio of triangle side lengths and angle differences. A similarity below 0.7 is considered to indicate a significant deviation. Feature similarity deviation is calculated by the average feature distance between the candidate point and the confirmed points in its neighborhood, using Euclidean distance. A deviation is considered to exist when the distance exceeds 1.5 times the average distance between confirmed points in the neighborhood. Based on the semantic consistency deviation, the initial semantic association strength is iteratively optimized using the gradient descent method. The learning rate is set to 0.01, the momentum coefficient is set to 0.9, and the maximum number of iterations is 100. The optimization is terminated early when the strength change is less than 0.001 in 10 consecutive iterations. The objective function is optimized by combining two objectives: maximizing semantic association strength and minimizing consistency deviation, with a weight ratio of 7 to 3, to obtain the corrected semantic association strength.
[0081] The confidence assessment model is based on two input variables: the modified semantic association strength and the stability of semantic consistency deviation. Stability is measured by calculating the standard deviation of the semantic association strength over five consecutive iterations. A standard deviation less than 0.01 indicates high stability, between 0.01 and 0.05 indicates moderate stability, and greater than 0.05 indicates low stability. The confidence assessment uses a piecewise linear function: when the semantic association strength is greater than 0.8 and the stability is high, the confidence level is set to 0.95 times the association strength; when the semantic association strength is between 0.5 and 0.8 and the stability is moderate, the confidence level is set to 0.8 times the association strength; and when the semantic association strength is less than 0.5 or the stability is low, the confidence level is set to 0.5 times the association strength. Semantic propagation reliability is assessed by statistically analyzing the propagation path length and the average confidence level of nodes along the path. A path length of less than 5 hops and an average confidence level greater than 0.7 indicate reliable propagation. The consistency assessment of association strength calculates the strength variance among multiple candidate correspondences for the same feature point. A variance less than 0.05 is considered good consistency, a variance between 0.05 and 0.15 is considered moderate consistency, and a variance greater than 0.15 is considered poor consistency. After the confidence score is calculated, interval mapping is performed to ensure that all values fall within the range of 0 to 1.
[0082] The cross-perspective semantic association graph update mechanism dynamically adjusts the correspondence status using confidence evaluation values. The dynamic threshold is adaptively set based on the confidence distribution of confirmed correspondences in the current graph, specifically the 75th percentile of the confirmed relationship confidence multiplied by 0.9, with a lower limit of 0.5 and an upper limit of 0.95. Candidate correspondences with confidence evaluation values exceeding the dynamic threshold are identified as reliable feature point mappings and marked as confirmed in the association graph, while the propagation weights of their adjacent nodes are updated. Correspondences with confidence evaluation values between 0.7 times and the dynamic threshold are marked as unobservable and given priority in subsequent iterations, increasing their propagation frequency. Correspondences with confidence evaluation values below 0.7 times the dynamic threshold are removed or marked as unreliable, and their corresponding edge connections are deleted from the graph. Each confirmed correspondence is assigned a quality label containing three fields: confidence value, stability level, and consistency level. The confidence value is rounded to three decimal places, and the stability and consistency levels are divided into high, medium, and low. The quality assessment adopts a comprehensive scoring mechanism, with confidence weighted at 0.5, stability weighted at 0.3, and consistency weighted at 0.2. A comprehensive score greater than 0.8 is considered high quality, 0.6 to 0.8 is considered medium quality, 0.4 to 0.6 is considered low quality but usable, and less than 0.4 is considered unusable.
[0083] In one optional implementation, a cross-perspective semantic association graph is constructed based on the composite feature descriptor. Semantic propagation constraints are established within this graph, and unconfirmed candidate correspondences are guided by existing reliable correspondences to obtain an initial semantic association strength distribution, including:
[0084] For each feature point in the composite feature descriptor, the semantic similarity of the feature points across the cross-view range is analyzed, and a cross-view semantic association graph is generated by combining the spatial location information of the feature points; feature point pairs that satisfy the bidirectional optimal matching criterion are determined in the cross-view semantic association graph, the feature point pairs are defined as reliable correspondences, and the relative positional relationship and local neighborhood topology of the reliable correspondences in three-dimensional space are analyzed to construct the spatial structure features of the reliable correspondences.
[0085] Using the spatial structural features, a constraint propagation network is constructed in the cross-perspective semantic association graph with the reliable correspondence as the propagation source point, and the spatial structural features are used to form topological consistency constraints on the propagation path.
[0086] For the unconfirmed candidate correspondences in the cross-perspective semantic association graph, the topological differences between the candidate correspondences and the reliable correspondences on the propagation path are evaluated and mapped to a propagation attenuation coefficient that characterizes the degree of constraint satisfaction.
[0087] The semantic similarity of the candidate correspondences is optimized by using the propagation attenuation coefficient. By progressively transmitting the spatial structural features of reliable correspondences in the constraint propagation network, the topological constraint modulation of the candidate correspondences is achieved, generating a semantic association strength distribution.
[0088] For each feature point in the composite feature descriptor, the semantic similarity analysis module employs cosine similarity calculation to perform L2 norm normalization on each 256-dimensional composite feature descriptor. The normalization formula divides each dimension value by the Euclidean modulus of the vector, ensuring that the vector modulus is 1 after processing. Semantic similarity calculation across viewpoints is achieved by constructing an N×M dimensional similarity matrix, where N is the number of feature points in the first viewpoint and M is the number of feature points in the second viewpoint. Matrix element values are obtained through the dot product of the two normalized descriptors. The similarity matrix is stored in a compressed sparse row format, retaining only elements with a similarity greater than 0.3, achieving a storage space compression rate of 78%. Spatial location information extraction includes three components: image coordinates, scale parameters, and principal direction angle of the feature points. Coordinates are stored as floating-point numbers accurate to one decimal place, the scale parameter ranges from 0.5 to 8.0, and the principal direction angle ranges from 0 to 360 degrees. Spatial similarity is calculated using a Gaussian kernel function, taking the Euclidean distance between feature points as input. The kernel function bandwidth is set to 50 pixels; similarity approaches 1.0 when the distance is less than 50 pixels, and decays to below 0.01 when the distance exceeds 200 pixels. The cross-view semantic association graph is constructed using an adjacency list data structure. Each node stores a feature point identifier, image coordinates, a composite feature descriptor, and a list of adjacent edges. Each adjacent edge contains four attribute fields: target node identifier, semantic similarity, spatial similarity, and comprehensive similarity. Comprehensive similarity is calculated through a weighted sum, with a semantic similarity weight of 0.7 and a spatial similarity weight of 0.3. These weights can be dynamically adjusted according to the application scenario, ranging from 0.5 to 0.9 for the semantic weight.
[0089] The bidirectional optimal matching criterion determination module implements a mutual nearest neighbor matching algorithm, which consists of two stages: forward matching and reverse verification. In the forward matching stage, for each feature point in viewpoint 1, all its adjacent edges in the semantic association graph are traversed, and the edge with the highest comprehensive similarity is selected as the initial match. The matching result is stored in a temporary hash table, with the source feature point identifier as the key and the target feature point identifier and similarity as the value. In the reverse verification stage, a bidirectional consistency check is performed on each initial match, calculating the best matching point for the target feature point in viewpoint 2. If this matching point happens to be the source feature point, it is confirmed as a bidirectional optimal matching pair. Match pair selection uses a similarity threshold filter, with the threshold set to 0.8. Simultaneously, the similarity of the matching pair must rank in the top 10% within its image region to ensure matching quality. After reliable correspondence is determined, geometric consistency verification is performed. The RANSAC algorithm is used to estimate the fundamental matrix, with an interior point threshold set to 2 pixels, 1000 iterations, and a geometric verification pass rate exceeding 60%.
[0090] The relative position relationship analysis module calculates the spatial geometric features of feature point pairs in a reliable correspondence, including four components: position difference, distance ratio, angle difference, and scale ratio. The position difference calculates the coordinate difference between two feature points in their respective image coordinate systems; the difference vector is normalized to avoid the influence of image size. The distance ratio is obtained by calculating the ratio of the distances from the feature points to the image center, with a range of 0.1 to 10.0; outliers outside this range are truncated. The angle difference calculates the angle difference between the line connecting the feature points and the horizontal axis, with a range limited to 0 to 180 degrees. The scale ratio is obtained by comparing the scale parameters of the feature points, with a range of 0.2 to 5.0. The local neighborhood topology analysis uses an improved Delaunay triangulation algorithm, selecting up to 15 neighboring feature points within an 80-pixel radius around each feature point to construct a constrained Delaunay triangulation network. The constraints include an upper limit of 120 pixels for the side length and a lower limit of 15 degrees for the angle. The topological feature extraction of triangulation networks includes seven statistical measures: number of triangles, average side length, standard deviation of side length, average angle, standard deviation of angle, total area, and average area. Spatial structure features are encoded using fixed-length vectors: relative positional relationships occupy 32 dimensions, angle and scale information occupy 16 dimensions, triangulation topological statistical features occupy 48 dimensions, and neighborhood connectivity features occupy 32 dimensions, totaling 128 dimensions. All dimensions are standardized with zero mean and unit variance to ensure numerical stability.
[0091] The constraint propagation network construction module establishes a directed graph topology based on reliable correspondences. Nodes in the graph correspond to feature points, and directed edges represent constraint propagation paths. A hierarchical sampling strategy is used to select propagation source points, dividing the image into an 8×8 grid. Within each grid, the feature point with the highest connectivity is selected as a candidate source point. These candidate source points are further sorted and filtered using a centrality index. Connectivity is calculated by counting the number of adjacent feature points within a 100-pixel radius, and the centrality index is calculated using the betweenness centrality algorithm, reflecting the importance of a node in the network. The final number of propagation source points is controlled between 30% and 40% of the total number of reliable correspondences to ensure a balance between propagation coverage and computational efficiency. The edge weights of the constraint propagation network are set based on the spatial distance and structural similarity between the source and target points. The distance weight uses an exponential decay function with a decay parameter set to 0.02, while the similarity weight uses a linear function with a slope parameter set to 1.5.
[0092] Topological consistency constraints are constructed based on multiple components of spatial structural features. These constraints are categorized into four types: distance preservation constraints, angle preservation constraints, neighborhood structure preservation constraints, and scale preservation constraints. Distance preservation constraints require that the distance ratio between adjacent feature points along the propagation path remain stable under different viewpoints, with an allowed rate of change of 15%. If the rate of change exceeds this range, the degree of constraint violation is calculated linearly based on the magnitude of the exceedance. Angle preservation constraints require that the angular change of the line connecting feature points after viewpoint transformation be controlled within 30 degrees. The angular change is calculated using the dot product of the direction vectors, and the degree of violation is calculated as the square root function of the change. Neighborhood structure preservation constraints require that the k-nearest neighbor topological relationship of feature points remain stable, with k set to 5. Topological stability is measured using Jaccard similarity; a similarity below 0.7 is considered a constraint violation. Scale preservation constraints require that the scale ratio of adjacent feature points remain consistent after viewpoint transformation, with an allowed rate of change of 25%. The degree of violation is calculated as a quadratic function of the rate of change.
[0093] The candidate correspondence topology difference assessment module performs a comprehensive difference analysis on unconfirmed candidate correspondences. The analysis process includes three levels: local difference assessment, global difference assessment, and path difference assessment. Local difference assessment calculates the spatial structural feature differences between a candidate correspondence and its k-nearest neighbor reliable correspondences. The k-value is set to 3, and the difference is measured using weighted Euclidean distance with weights allocated as follows: relative position feature 0.35, angle feature 0.25, triangulation feature 0.25, and connectivity feature 0.15. Global difference assessment analyzes the structural deviation of candidate correspondences throughout the propagation network. This is achieved by calculating the degree of deviation between the local and global feature distributions of the candidate correspondences. The degree of deviation is measured using standardized Z-scores; a Z-score absolute value greater than 2 is considered a significant deviation. Path difference assessment analyzes the shortest path features from candidate correspondences to each propagation source point. Path features include path length, average connectivity of nodes along the path, and cumulative path weights. Path differences are calculated by comparing the distribution differences between the path features of candidate correspondences and those of confirmed correspondences.
[0094] The propagation attenuation coefficient mapping module converts topological differences into attenuation coefficients ranging from 0 to 1. The mapping function employs a piecewise design to accommodate different levels of difference. When the local difference is less than 0.2, the attenuation coefficient is set to 0.95. When the difference is between 0.2 and 0.4, the attenuation coefficient decreases linearly from 0.95 to 0.75. When the difference is between 0.4 and 0.7, the attenuation coefficient decreases exponentially from 0.75 to 0.3. When the difference exceeds 0.7, the attenuation coefficient is set to 0.1. The global difference attenuation coefficient mapping uses a Sigmoid function, with the function's center point set at a difference value of 0.5 and a slope parameter set to 5 to ensure strong discriminative power near the center point. The path difference attenuation coefficient is calculated using a joint function of path length and weight, with a path length weight of 0.6 and a path weight weight of 0.4. The final attenuation coefficient is the weighted geometric mean of the attenuation coefficients at the three levels.
[0095] The semantic similarity optimization module performs topological constraint modulation processing, which consists of three steps: preprocessing, attenuation modulation, and constraint modulation. The preprocessing step checks the initial semantic similarity for range and handles outliers, limiting the similarity value to between 0.1 and 1.0, with outliers truncated. The attenuation modulation step multiplies the preprocessed similarity by a propagation attenuation coefficient to obtain the attenuated similarity value. The constraint modulation step applies penalties for violations of four types of topological consistency constraints based on candidate relationships: a penalty factor of 0.8 for distance constraint violations, 0.75 for angle constraint violations, 0.7 for neighborhood constraint violations, and 0.85 for scale constraint violations. When multiple constraints are violated simultaneously, the penalty factors are accumulated and multiplied.
[0096] The progressive transmission of spatial structural features employs an improved breadth-first propagation algorithm. The propagation process begins at each source node and proceeds with weighted propagation according to the edge weights of the constraint propagation network. The initial propagation strength is set to 1.0, decreasing by a decay factor of 0.8 with each hop, and the propagation depth is limited to 5 hops. During propagation, each node receives structural feature information from multiple sources, fusing the multi-source information using a weighted average method. The weights are proportional to the propagation strength and path reliability. Topology constraint modulation is performed in real-time during propagation; each propagation step checks constraint satisfaction and updates the modulation coefficients.
[0097] The semantic association strength distribution generation module normalizes and performs quality assessment on all optimized similarity values. Normalization uses quantile normalization to map similarity values to the range of 0.1 to 1.0. The mapping function is designed based on the cumulative distribution function to ensure good discriminative power in the strength distribution. Quality assessment is achieved by calculating three indicators for each relation: confidence, consistency, and stability. Confidence is calculated based on the optimized similarity, consistency is calculated based on the degree of constraint satisfaction, and stability is calculated based on the magnitude of numerical changes during propagation. The final semantic association strength distribution is stored in sparse matrix form, containing four fields: relation identifier, strength value, quality indicator, and metadata.
[0098] In one optional implementation, the transformation parameters are used to align the coordinate systems of the multi-frame 3D scene image data to generate an initial 3D reconstruction result in a unified coordinate system. Based on the reprojection error distribution of spatial points in the initial 3D reconstruction result, local regions with registration deviations are identified. A multi-objective joint optimization function is constructed for the local regions, including:
[0099] The local coordinate system of each frame in the multi-frame 3D scene image data is transformed to a unified coordinate system using the transformation parameters. A feature fusion network is established under the unified coordinate system to fuse and reconstruct spatial points from different frames, generating an initial 3D reconstruction result.
[0100] For each spatial point in the initial 3D reconstruction result, the deviation between the reprojection position and the actual observation position of each spatial point under each observation view is analyzed based on the feature fusion network, and the spatial distribution of the deviation is statistically analyzed to obtain the reprojection error distribution characteristics.
[0101] Based on the reprojection error distribution characteristics, identify the set of spatial points exceeding the error threshold in the initial 3D reconstruction result, construct a spatial structure optimization map, analyze the connectivity and clustering of the set of spatial points, and divide the error spatial points with spatial connectivity into local regions with registration deviations.
[0102] For the local region, a regional optimization objective is constructed in the spatial structure optimization diagram, which includes geometric consistency constraints and minimization of reprojection error. At the same time, the boundary continuity constraint relationship between the local region and the adjacent regions is analyzed.
[0103] Based on the region optimization objective and the boundary continuity constraint, a multi-objective joint optimization function is established to ensure that adjacent local regions maintain spatial continuity and geometric consistency at the boundary.
[0104] The processing module for aligning the coordinate systems of multi-frame 3D scene image data using transformation parameters employs a homogeneous coordinate transformation method to convert the local coordinate systems of each frame to a predefined unified coordinate system. The transformation parameters include three components: a rotation matrix, a translation vector, and a scale factor. The rotation matrix is represented by a 3×3 orthogonal matrix, the translation vector is a 3D column vector, and the scale factor is a positive real number. The coordinate transformation process performs homogeneous coordinate expansion on the 3D coordinates of each spatial point, expanding them to four-dimensional homogeneous coordinates. The transformation is calculated using a 4×4 transformation matrix, which is composed of the rotation matrix, translation vector, and scale factor. The transformation calculation uses matrix multiplication, maintaining double-precision floating-point format, with the numerical range limited to within 15 significant digits. The coordinate transformation results are validated to ensure the transformed coordinates are within a reasonable spatial range, set to -1000 to +1000 meters. Coordinate points outside this range are marked as abnormal and processed further.
[0105] The feature fusion network under a unified coordinate system adopts a hierarchical network structure, consisting of four main layers: an input layer, a feature extraction layer, a fusion layer, and an output layer. The input layer receives spatial point coordinates, color information, and feature descriptors from different frames. Coordinate data is stored as a 3D floating-point array, color information as an 8-bit integer with three RGB channels, and feature descriptors as 256-dimensional floating-point vectors. The feature extraction layer encodes the input data. Coordinate features are generated into 64-dimensional feature vectors using a position encoder, and color features are generated into 32-dimensional feature vectors using a color encoder. Descriptor features retain their original 256 dimensions. The fusion layer uses an attention mechanism to weightedly fuse features of the same spatial point from different frames. Attention weights are determined by calculating the similarity between feature vectors using a dot product attention method, and weight normalization uses the softmax function. During the fusion process, feature vectors from each observation frame are collected for each spatial point, and a weighted average is used to generate the fused feature representation. Weight allocation is determined based on a combination of observation angle, distance, and image quality. The output layer generates fused 3D coordinates of spatial points and comprehensive feature descriptors, maintaining millimeter-level coordinate accuracy and compressing the feature descriptor dimension to 128 dimensions to improve processing efficiency.
[0106] The initial 3D reconstruction results are generated using an incremental reconstruction strategy, gradually expanding from a seed point set to the complete scene. Seed point selection is based on two criteria: feature fusion quality and geometric constraint satisfaction. Spatial points with a fusion quality score higher than 0.8 and a geometric constraint satisfaction rate exceeding 90% are selected as seed points. The reconstruction expansion process employs a neighborhood growing algorithm, expanding from the seed points to the surrounding space with an expansion step size set to one-thousandth of the scene scale. The expansion direction is determined based on local normal vectors. Density control is implemented during spatial point generation, ensuring that the point cloud density of the reconstruction results is between 10 and 50 points per cubic centimeter through adaptive sampling. The reconstruction results are stored using an octree data structure with a maximum depth of 12 layers and a maximum leaf node count of 64. Storage efficiency is optimized through compression encoding.
[0107] The reprojection error analysis module performs multi-view reprojection calculations for each spatial point in the initial 3D reconstruction result, using a pinhole camera model. The camera intrinsic parameter matrix contains four components: focal length, principal point coordinates, and distortion coefficients. The focal length is represented as a floating-point number in pixels, the principal point coordinates are the pixel coordinates of the image center, and the distortion coefficients include radial and tangential distortion parameters. The reprojection calculation transforms the 3D spatial points to the camera coordinate system using camera extrinsic parameters, and then projects them onto the image plane using the intrinsic parameter matrix. The projection result is the image pixel coordinates. The deviation between the reprojected position and the actual observation position is calculated using Euclidean distance, with the deviation unit in pixels and the calculation precision maintained to two decimal places. The deviation calculation covers all observable viewpoints of the spatial point, generating multiple reprojection error values for each spatial point. These error values are stored in a dynamic array to accommodate different numbers of observation viewpoints.
[0108] The reprojection error distribution feature extraction employs statistical analysis methods to statistically analyze the spatial distribution of reprojection errors for all spatial points. The statistical features include seven basic statistical measures: mean, standard deviation, maximum, minimum, median, and quantiles, with statistical precision maintained to three decimal places. Spatial distribution analysis utilizes a grid-based approach, dividing the three-dimensional space into 10cm-sided cubic grids and calculating the error statistical features of spatial points within each grid. Error distribution visualization uses a heatmap method, mapping error values to color codes: areas with errors less than 1 pixel are displayed in green, areas with errors between 1 and 3 pixels are displayed in yellow, and areas with errors exceeding 3 pixels are displayed in red. Distribution feature storage uses a sparse matrix format, recording only grid cells with errors exceeding a threshold. The stored content includes three fields: grid coordinates, error statistics, and the number of spatial points.
[0109] The identification of spatial point sets exceeding the error threshold employs an adaptive threshold determination method, with threshold calculation based on the statistical characteristics of the error distribution. The base threshold is set as the mean of the error distribution plus 1.5 times the standard deviation. Adaptive adjustment corrects for the skewness and kurtosis of the error distribution; the threshold increases by 20% when the skewness is greater than 1, and decreases by 10% when the kurtosis is greater than 3. The spatial point filtering process traverses all reconstructed spatial points, calculating the average reprojection error for each point. Spatial points with errors exceeding the adaptive threshold are added to the over-threshold set. The set is stored using a hash table structure, where the key is the unique identifier of the spatial point, and the values include three attributes: spatial point coordinates, average error, and maximum error. Outlier filtering is achieved through geometric consistency checks, examining the geometric relationship between over-threshold spatial points and their neighboring points. The neighborhood radius is set to 5 centimeters, and the geometric relationship includes both distance consistency and normal vector consistency. Isolated points that do not meet the geometric consistency requirements are marked as outliers and removed from the set.
[0110] The spatial structure optimization graph construction is based on a set of spatial points exceeding a threshold to establish the graph topology. Each node in the graph corresponds to a spatial point exceeding the threshold, and edges connect node pairs whose spatial distance is less than the threshold. The distance threshold is set to 5% of the scene feature scale; 50 cm for architectural scenes and 5 cm for object scenes. Graph construction uses the k-nearest neighbor method, with each node connecting to a maximum of 8 nearest neighbors. Connection weights are calculated based on spatial distance and error similarity, with distance accounting for 70% and error similarity accounting for 30%. Graph storage uses an adjacency list format, where each node stores coordinates, error value, and a list of neighboring nodes. The neighboring node list contains two fields: target node identifier and connection weight. Graph optimization is performed using the minimum spanning tree algorithm, retaining the connection edge with the smallest weight and removing redundant connections to simplify the graph structure.
[0111] Connectivity analysis employs a depth-first search algorithm to identify connected components in the spatial structure optimization graph. Each connected component corresponds to a potential registration deviation region. The search process starts from unvisited nodes and recursively visits all reachable nodes. Visit markers are stored in node attributes to avoid repeated visits. The size of a connected component is measured by the number of nodes; small connected components with fewer than 5 nodes are considered noise and filtered out. Clustering analysis is achieved by calculating the spatial density of connected components. Density is defined as the average distance from all nodes within a connected component to the centroid, where the centroid coordinates are the arithmetic mean of all node coordinates. A density threshold is set to 10% of the scene feature scale; connected components exceeding this threshold are considered too dispersed and require further subdivision.
[0112] The local region partitioning process performs spatial clustering on each connected component using the density-based DBSCAN method. Clustering parameters include a neighborhood radius and a minimum number of points. The neighborhood radius is set to twice the average distance between points, and the minimum number of points is set to 10% of the total number of points in the connected component. The clustering results generate multiple local regions, each containing a set of spatially connected points exceeding a threshold. Region quality is assessed by calculating the mean error and spatial distribution uniformity of points within the region. Regions with a mean error exceeding 1.5 times the global mean or exhibiting uneven spatial distribution are marked for special handling. Region identifiers use a hierarchical encoding method, with the encoding format being the connected component index plus the region index, ensuring global uniqueness.
[0113] The regional optimization objective comprises two main objectives: geometric consistency constraints and minimization of reprojection errors. Geometric consistency constraints require spatial points within a local region to satisfy local smoothness and continuity conditions. Smoothness is measured by calculating the angle between the normal vectors of adjacent points, which should be less than 30 degrees. Continuity is measured by the rate of change of distance between checkpoints, which should be less than 20%. The reprojection error minimization objective is established using the least squares method. The objective function is the sum of the squares of all reprojection errors. Weights are assigned based on the observation angle and distance, with frontal observation angles having higher weight than side observation angles, and close-range observations having higher weight than distant observations. Constraints include the effective range of spatial point coordinates and the distance constraints between adjacent points. The coordinate range is determined based on the bounding box of the initial reconstruction result, and the distance constraints are determined based on local geometric features.
[0114] Boundary continuity constraint resolution is achieved by analyzing the spatial relationships between adjacent local regions. Adjacency is determined by calculating the shortest distance between boundary points of a region; regions with a distance less than the region's feature scale are considered adjacent. Boundary point identification uses a convex hull algorithm to calculate the 3D convex hull of each local region, and points on the convex hull surface are marked as boundary points. Continuity constraints require that boundary points of adjacent regions maintain spatial continuity after optimization. Continuity is achieved by limiting the variation in boundary point coordinates, which cannot exceed 50% of the average distance between points. Constraint relationships are stored using a constraint graph structure, where nodes correspond to local regions, edges correspond to constraint relationships, and edge attributes include three fields: constraint type, constraint strength, and violation penalty.
[0115] The multi-objective joint optimization function is established using a weighted multi-objective optimization framework. The optimization function comprises three components: a reprojection error term, a geometric consistency term, and a boundary continuity term. The weight of the reprojection error term is set to 0.6, the weight of the geometric consistency term to 0.3, and the weight of the boundary continuity term to 0.1. These weights can be dynamically adjusted according to application requirements. The optimization variables are the three-dimensional coordinates of all spatial points within the local region. The number of variables equals the number of spatial points multiplied by three. Variable constraints include the effective range of coordinates and distance restrictions between adjacent points. The optimization algorithm employs gradient descent, with a learning rate of 0.001, a momentum coefficient of 0.9, and an upper limit of 1000 iterations. Convergence is determined based on the change in the objective function value; convergence is considered achieved when the change is less than 0.0001 over 10 consecutive iterations.
[0116] In one optional implementation, based on the reprojection error distribution characteristics, a set of spatial points exceeding an error threshold in the initial 3D reconstruction result is identified, a spatial structure optimization map is constructed, and the connectivity and clustering of the spatial point set are analyzed. Error spatial points with spatial connectivity are divided into local regions with registration deviations, including:
[0117] Based on the reprojection error distribution characteristics, spatial points whose reprojection errors exceed the error threshold are identified in the initial 3D reconstruction results, and the 3D coordinates, normal vectors and observation view information of the spatial points are extracted to obtain a set of spatial points that exceed the error threshold.
[0118] Using each spatial point in the set of spatial points as a node, the geodesic distance and local curvature consistency between pairs of spatial points are calculated, and weighted edges are established between pairs of spatial points that satisfy the distance and curvature constraints. The weight of the weighted edges is jointly measured by the geodesic distance and curvature consistency to construct a spatial structure optimization graph.
[0119] In the spatial structure optimization graph, the weighted edge connections are traversed using a graph connectivity analysis algorithm to identify spatial point subsets that form connected subgraphs through the weighted edge connections. The cluster density is calculated based on the weighted edge weight distribution within the spatial point subsets to obtain connectivity analysis results and clustering analysis results.
[0120] Based on the connectivity analysis results and the clustering analysis results, the connected subgraphs in the spatial structure optimization graph are screened, and the connected subgraphs with clustering density exceeding the clustering threshold are marked as candidate local regions. Based on the distribution characteristics of the internal spatial points of the candidate local regions, the boundaries are adjusted, and the candidate local regions after boundary adjustment are determined as a set of local regions with registration deviation.
[0121] The spatial point processing module, which identifies spatial points exceeding an error threshold based on reprojection error distribution characteristics, employs an adaptive threshold determination strategy. The error threshold calculation is based on the statistical characteristics of the reprojection error distribution. The threshold determination process calculates the mean and standard deviation of the reprojection errors for all spatial points. The basic error threshold is set as the mean plus twice the standard deviation. Adaptive adjustments are made based on the skewness and kurtosis of the error distribution. When the skewness is greater than 1.5, the threshold is increased by 25% to accommodate a right-skewed distribution; when the kurtosis is less than 0.5, the threshold is decreased by 15% to accommodate a flat distribution. The spatial point screening process traverses all spatial points in the initial 3D reconstruction results, calculating the reprojection error of each spatial point under all observation views. A weighted average method is used to calculate the comprehensive reprojection error, with weights determined based on the observation angle and image quality. The observation angle weight is calculated using a cosine function, with a weight of 1.0 for frontal observation angles and decreasing to 0.3 for side observation angles. Image quality weights are calculated based on image sharpness and contrast; images with high sharpness have a weight of up to 1.2, while blurry images have a weight reduced to 0.6.
[0122] The spatial point information extraction module performs complete attribute extraction on the selected spatial points exceeding the threshold. The extracted content includes three core attributes: 3D coordinates, normal vectors, and observation viewpoint information. The 3D coordinates are represented in the world coordinate system with millimeter-level precision. The data type is 64-bit double-precision floating-point numbers, and the coordinate range is determined based on the scene bounding box. Points outside the range are marked as anomalies and undergo further verification. Normal vector calculation uses principal component analysis (PCA) to fit a plane within the local neighborhood of each spatial point. The neighborhood radius is set to three times the average point spacing, and at least six neighboring points are required for effective normal vector estimation. Normal vector normalization ensures the vector magnitude is 1. Directional consistency is verified through the dot product of the normal vectors of neighboring points; if the dot product is less than 0.5, it is recalculated or marked as unreliable. The observation viewpoint information records all camera views that can observe the spatial point, including three parameters: camera position, orientation angle, and observation distance. The camera position is represented in 3D coordinates, the orientation angle is represented in Euler angles, and the observation distance is calculated using Euclidean distance.
[0123] The set of spatial points exceeding the error threshold is constructed and stored using a dynamic data structure. The set employs a hash table for fast indexing and lookup. The key is a globally unique identifier for the spatial point, and the value is a structure containing coordinates, normal vectors, observation viewpoints, and error information. The structure design includes six fields: point identifier, 3D coordinate array, normal vector array, list of observation viewpoints, reprojection error value, and quality assessment flag. The point identifier uses 64-bit integer encoding; the first 32 bits represent the source image frame number, and the last 32 bits represent the intra-frame point number, ensuring global uniqueness. The quality assessment flag records the reliability level of the spatial point, categorized into four levels: high reliability, medium reliability, low reliability, and unreliable. The assessment is determined based on a comprehensive evaluation of the reprojection error magnitude, normal vector stability, and the number of observation viewpoints.
[0124] The geodesic distance calculation module employs Dijkstra's shortest path algorithm to calculate the geodesic distance between pairs of spatial points on a local surface. The distance calculation is based on a local mesh structure formed by the spatial points. Mesh construction uses the Delaunay triangulation method, creating a local triangular mesh around each spatial point. The side length of the triangles is limited to within 5 times the average point spacing; excessively long edges are removed to avoid unreasonable connections. The geodesic distance calculation performs a path search along the edges of the triangular mesh, with the path weight being the Euclidean length of the edge. The search depth is limited to 20 hops to control computational complexity. The distance constraint threshold is set to 8% of the scene feature scale; 20 centimeters for indoor scenes and 2 meters for outdoor scenes. Spatial point pairs exceeding the threshold are not connected.
[0125] Local curvature consistency is calculated using a joint measure of Gaussian curvature and mean curvature. Curvature calculation is based on the local neighborhood geometric features of spatial points. Gaussian curvature is calculated by fitting a local quadratic surface using the least squares method, requiring at least nine neighboring points. A fitting error greater than 10% of the average point spacing is marked as unreliable. Mean curvature is obtained by calculating the rate of change of the normal vector within the local neighborhood using the central difference method, with calculation precision maintained to four decimal places. Curvature consistency is evaluated by comparing the curvature values of pairs of spatial points. Curvature consistency is considered achieved when the difference in Gaussian curvature is less than 0.02 and the difference in mean curvature is less than 0.05. The curvature constraint threshold is adaptively set, determined based on the standard deviation of the curvature distribution in the local region. 1.5 times the standard deviation is used as the constraint threshold to ensure that the constraint is neither too strict nor too lenient.
[0126] Weighted edge connections are established using a joint metric method, with edge weights determined by a weighted combination of geodesic distance weights and curvature consistency weights. The geodesic distance weight is calculated using a Gaussian kernel function, with the kernel bandwidth parameter set to half the distance constraint threshold; smaller distances result in larger weights, ranging from 0.1 to 1.0. The curvature consistency weight is calculated using an exponential decay function, with the decay parameter determined based on curvature differences. When curvature is perfectly consistent, the weight is 1.0; as curvature differences increase, the weight exponentially decays to 0.2. The joint weight calculation uses a geometric mean, i.e., the square root of the product of the geodesic distance weight and the curvature consistency weight, ensuring a balanced influence of the two factors. Edge weight standardization maps all weight values to the range of 0 to 1 using a max-min normalization method, with the minimum weight mapped to 0.05 to avoid numerical instability.
[0127] The spatial structure optimization graph construction uses an adjacency list data structure to store the graph's topology information. Each node in the graph corresponds to a spatial point exceeding the threshold, and the node stores complete attribute information of the spatial point. The adjacency list consists of three parts: node identifier, node attributes, and an adjacency edge list. The adjacency edge list records the target nodes and edge weights of the connections. The graph construction process adopts a block-parallel strategy to improve computational efficiency, dividing the space into multiple cubic blocks. Spatial points within each block are processed independently, and overlapping is used between block boundaries to avoid boundary effects. Graph storage optimization uses a compressed sparse matrix format, storing only edge connections with non-zero weights, with a storage compression rate typically exceeding 85%. Graph quality verification is achieved by checking connectivity and the rationality of weight distribution. The number of isolated nodes should be less than 5% of the total number of nodes, and the edge weight distribution should exhibit reasonable statistical characteristics.
[0128] The graph connectivity analysis algorithm employs a depth-first search method to traverse weighted edge connections and identify subsets of spatial points forming connected subgraphs. The search process starts with unvisited nodes and recursively visits all reachable nodes along the weighted edges. Visit markers are stored in node attributes to avoid duplicate visits. The connected subgraph identification process records the set of nodes and edges contained in each connected component. The size of the connected component is statistically analyzed using three metrics: the number of nodes, the number of edges, and the total weight. Small-scale connected component filtering is achieved by setting a minimum node count threshold; connected components with fewer than 8 nodes are considered noise and excluded from subsequent analysis. The connectivity analysis results are stored in a tree structure, with the root node representing the entire graph, child nodes representing individual connected components, and leaf nodes representing spatial points within each connected component.
[0129] Cluster density calculation is based on statistical analysis of the weighted edge weight distribution within connected subgraphs. The density measure includes three statistics: weight mean, weight variance, and weight distribution skewness. The weight mean reflects the overall density of the connected subgraph; a larger mean indicates stronger connections between spatial points within the subgraph. The weight variance reflects the uniformity of density; a smaller variance indicates relatively uniform density across different parts of the subgraph, while a larger variance indicates regions with significant differences in density. The weight distribution skewness reflects the symmetry of the weight distribution; positive skewness indicates the presence of a few high-weight connections, while negative skewness indicates a relatively uniform weight distribution. The overall cluster density score is calculated using a weighted combination method: weight mean (0.5), weight variance (0.3), and weight skewness (0.2), ensuring that the score reflects both overall density and distributional characteristics.
[0130] Clustering analysis results are generated by sorting and classifying the cluster density of all connected subgraphs, with the classification criteria based on the quantiles of the density distribution. Connected subgraphs with a density exceeding the 75th quantile are classified as highly clustered, those with a density between the 25th and 75th quantiles are classified as moderately clustered, and those with a density below the 25th quantile are classified as lowly clustered. The clustering analysis results store four fields: connected subgraph identifier, cluster density score, clustering level, and statistical characteristics. The statistical characteristics include detailed information such as the number of nodes, the number of edges, the average weight, and the standard deviation of the weights.
[0131] The connected subgraph selection process comprehensively evaluates the results of connectivity and clustering analysis. The selection criteria include three conditions: minimum number of nodes, minimum cluster density, and connectivity stability. The minimum number of nodes is set to a threshold of 15 to ensure that candidate local regions have sufficient spatial points to support reliable geometric analysis. The clustering threshold is determined using an adaptive method, calculating the mean and standard deviation of the cluster density of all connected subgraphs. The threshold is set to the mean plus 0.5 times the standard deviation, ensuring that connected subgraphs with significantly higher-than-average clustering are selected. Connectivity stability is assessed by analyzing the topological robustness of the connected subgraphs; subgraphs that remain connected after removing the 10% of edges with the lowest weights are considered to have good stability.
[0132] Candidate local regions are labeled using a hierarchical encoding method, which includes three parts: connected subgraph index, clustering level, and quality assessment. Connected subgraph indices are assigned in the order of discovery, increasing from 1. The clustering level uses letter coding: A for high clustering, B for medium clustering, and C for low clustering. The quality assessment uses numerical coding: 1 for high quality, 2 for medium quality, and 3 for low quality. Candidate local regions are stored using a region tree data structure, supporting fast spatial queries and neighborhood analysis operations.
[0133] Boundary adjustment processing performs local optimization based on the distribution characteristics of spatial points within candidate local regions. The distribution characteristic analysis includes three aspects: point density distribution, normal vector consistency, and geometric shape regularity. Point density distribution is achieved by calculating a local density field, generated using Gaussian kernel interpolation, with the kernel bandwidth parameter set to twice the average point spacing. Density gradient calculation employs the central difference method; regions with large gradients indicate drastic density changes and are considered potential boundary regions. Normal vector consistency is evaluated by calculating the dot product of the normal vectors of adjacent spatial points; regions with a dot product less than 0.8 are considered to have significant geometric feature changes and are located at region boundaries. Geometric shape regularity is evaluated by fitting local geometric primitives, including planes, cylinders, and spheres; regions with large fitting errors are considered to have complex geometries and require special processing.
[0134] The boundary adjustment algorithm employs a region growing method, expanding outwards from the high-density core region of the candidate local area, while checking the geometric consistency of boundary points during the expansion process. The expansion conditions include three aspects: distance constraints, normal vector constraints, and curvature constraints. The distance constraint requires that the shortest distance between the new point and the core region is less than four times the average point spacing. The normal vector constraint requires that the angle between the normal vector of the new point and the average normal vector of the neighborhood is less than 45 degrees. The curvature constraint requires that the curvature characteristics of the new point are consistent with the curvature distribution of the neighborhood. Boundary shrinkage is achieved by removing unstable boundary points. Unstable point identification is based on the rate of change of local geometric features; boundary points with a rate of change exceeding a threshold are removed to obtain a more stable region boundary.
[0135] The set of local regions exhibiting registration bias was determined using a final quality assessment and verification process. The assessment indicators included three aspects: regional compactness, geometric consistency, and the potential for improving reprojection error. Regional compactness was assessed by calculating the coefficient of variation of the distance from spatial points within the region to the centroid; a coefficient of variation less than 0.3 indicated good regional compactness. Geometric consistency was assessed by analyzing the consistency of the normal vectors and curvature distributions of spatial points within the region; a standard deviation of the normal vector angle less than 30 degrees and a curvature distribution skewness less than 1.0 indicated good geometric consistency. The potential for improving reprojection error was assessed by simulating the effects of local optimization; regions with an expected error improvement greater than 20% were considered to have good optimization potential.
[0136] In one optional implementation, the multi-objective joint optimization function is integrated into a dynamic adaptive weight control mechanism to guide the evolution of the initial 3D reconstruction result under spatial geometric constraints. Simultaneously, the spatial configuration information obtained during the evolution process is coupled in a closed loop with the cross-view semantic association graph, driving the dynamic optimization and refinement of feature correspondences, including:
[0137] For each optimization objective term in the multi-objective joint optimization function, the convergence rate and contribution weight of each objective term are calculated based on the current residual distribution of the initial 3D reconstruction result, and an adaptive weight adjustment mechanism is constructed.
[0138] The multi-objective joint optimization function is reconstructed using the adaptive weight adjustment mechanism. A multi-scale geometric constraint network is constructed in the reconstructed multi-objective joint optimization function. The multi-scale geometric constraint network limits the adjustment range of spatial point positions and the topological stability of the neighborhood, thereby guiding the structural evolution of the initial three-dimensional reconstruction result.
[0139] Based on the multi-scale geometric constraint network, the positional changes of spatial points and the local topological structure changes during the structural evolution are extracted, and a self-attention aggregation mechanism is introduced to calculate the configurational stability measure of spatial points to obtain spatial configuration information.
[0140] The configuration stability measure in the spatial configuration information is converted into the spatial geometric confidence of feature point pairs through the self-attention aggregation mechanism. The spatial geometric confidence is then used to modulate and weight the semantic association strength of feature point pairs in the cross-view semantic association graph to form a closed-loop coupling.
[0141] The confidence level of the correspondence between feature point pairs is calculated based on the modulated cross-view semantic association graph. Feature point pairs below the confidence threshold are rematched, and the matching results and the spatial configuration information are fed back to the weighted reconstructed multi-objective joint optimization function to drive the dynamic optimization and refinement of feature correspondence.
[0142] For each optimization objective in the multi-objective joint optimization function, the convergence rate calculation module performs dynamic analysis based on the current residual distribution of the initial 3D reconstruction results. The residual distribution statistics employ a sliding window method with a window length of 50 iterations. The residual values within the window are calculated using an exponentially weighted moving average, with a decay factor set to 0.95 to ensure that recent residual changes have a greater impact on the convergence rate. The convergence rate is determined by calculating the slope of the residual changes between consecutive iterations. The slope calculation uses a least-squares linear fitting method with a fitting window length of 10 iterations. A slope absolute value greater than 0.001 is considered a faster convergence rate, while a slope less than 0.0001 is considered a slower convergence rate. Contribution weight calculation is based on the degree of influence of each objective on the overall optimization effect. The degree of influence is measured by comparing the overall error change before and after optimizing that objective individually. Objectives with an error change greater than 10% are considered to have a higher contribution weight. Contribution weight normalization ensures that the sum of the weights of all objective items is 1. Normalization is implemented using the softmax function, with a temperature parameter set to 0.5 to enhance the discriminative power of weight differences.
[0143] The adaptive weight adjustment mechanism employs a gradient-based weight update strategy, where the weight adjustment is proportional to the product of the convergence rate and the contribution weight. The weight update formula is the current weight plus the learning rate multiplied by the product of the convergence rate and the contribution weight. The initial learning rate is set to 0.01. Using an adaptive adjustment strategy, the learning rate increases by 50% when the weight change is less than 0.001 for five consecutive iterations, and decreases by 30% when the weight oscillation is greater than 0.1. Weight adjustment boundary constraints ensure that the weight of a single target item does not exceed 0.8 and is not lower than 0.05, preventing excessively large weights from causing optimization bias or excessively small weights from causing target item failure. Weight adjustment history is stored in a circular buffer with a size set to 200 iterations. The record includes four fields: iteration number, weight of each target item, convergence rate, and contribution weight. Weight adjustment stability monitoring is achieved by calculating the weight variance. When the variance exceeds 0.05, stabilization processing is triggered, using a moving average method to smooth weight changes.
[0144] The weighted reconstruction process of the multi-objective joint optimization function recombines the objective terms based on the output weights of the adaptive weight adjustment mechanism. The reconstructed optimization function adopts a weighted summation form, where each objective term is multiplied by its corresponding weight and then summed. The weight values are obtained in real time from the weight adjustment mechanism, and the update frequency is set to once every 10 iterations to balance response speed and stability. Objective term standardization ensures the comparability of objective terms with different dimensions. The standardization method uses z-score normalization, i.e., the objective term value is subtracted from the mean and then divided by the standard deviation. The mean and standard deviation are calculated based on statistical data from 100 historical iterations. The gradient calculation of the reconstructed optimization function uses an automatic differentiation method, maintaining the gradient precision in double-precision floating-point format. Gradient norm constraints prevent gradient explosion, with a constraint threshold set to 10.0. Gradients exceeding the threshold are truncated.
[0145] The multi-scale geometric constraint network is constructed based on a hierarchical constraint structure established by the reconstructed multi-objective joint optimization function. The constraint network comprises three levels: point-level constraints, edge-level constraints, and surface-level constraints. Point-level constraints limit the adjustment range of each spatial point's position, determined by the point's local geometric features. The adjustment radius for flat regions is set to twice the average point spacing, while for regions with greater curvature, the radius is set to 0.5 times the average point spacing. Edge-level constraints limit the distance variation between adjacent spatial points, with the rate of change not exceeding 20% of the original distance. If this limit is exceeded, an elastic restoring force is applied, proportional to the variation rate, with a proportionality coefficient of 0.8. Surface-level constraints maintain the continuity and smoothness of local surfaces. Constraints are established by fitting local quadratic surfaces using the least squares method. When the fitting error exceeds 10% of the average point spacing, the constraint is invalidated and refitted. The constraint network is updated every 5 optimization iterations, recalculating local geometric features and constraint parameters during the update process.
[0146] The maintenance of neighborhood topology stability is achieved by monitoring changes in the neighborhood relationships of spatial points. These relationships are represented using a k-nearest neighbor graph, with k set to 8. The rate of change in neighborhood relationships is measured by calculating Jaccard similarity. The similarity calculation compares the k-nearest neighbor sets of spatial points before and after optimization. If the similarity is below 0.7, the topology change is considered too large, triggering topology stabilization. Stabilization is achieved by increasing the weight of the topology constraint terms by 50%. The topology constraint terms penalize drastic changes in the topology structure; the penalty strength is quadratically related to the degree of change, ensuring the smoothness of topology changes. Structural evolution guidance is achieved through flyback control of the constraint network. When the degree of constraint violation exceeds a threshold, the optimization direction is automatically adjusted along the opposite direction of the constraint gradient, and the adjustment magnitude is proportional to the degree of violation.
[0147] The spatial point position change extraction module records the coordinate change trajectory of each spatial point during structural evolution. The change is calculated as the difference between the current coordinates and the initial coordinates, represented by a three-dimensional vector. The vector magnitude represents the magnitude of the change, and the vector direction represents the direction of the change. Position change statistics include the mean, standard deviation, and maximum value of the change magnitude, as well as the principal direction angle of the change direction. The statistical results are used to evaluate the convergence and stability of the evolution process. Local topological structure change feature extraction is based on the evolutionary analysis of spatial point neighborhood relationships. Features include changes in the number of neighborhood points, changes in neighborhood distance, and changes in neighborhood angle. Changes in the number of neighborhood points are calculated by comparing the differences in the k-nearest neighbor sets; spatial points with a change rate exceeding 30% are marked as topologically unstable. Changes in neighborhood distance are calculated by statistically analyzing the rate of change of distance to each neighboring point; a distance change is considered significant when the standard deviation of the distance change is greater than 20% of the average distance. Changes in neighborhood angle are calculated by statistically analyzing the changes in the angles connecting neighboring points; an angle change exceeding 45 degrees is considered a significant change in local geometry.
[0148] The self-attention aggregation mechanism employs the multi-head attention mechanism in the Transformer architecture to calculate the configurational stability metric of spatial points. The input to the attention mechanism is the change in position and topological structure features of the spatial points. Feature encoding uses a joint representation of positional and change encoding, with positional encoding being 64-dimensional, change encoding 32-dimensional, and the joint feature being 96-dimensional. The multi-head attention mechanism comprises eight attention heads, each with 12 dimensions. Attention weights are calculated by the dot product of the query vector, key vector, and value vector, and the dot product result is normalized using the softmax function. The configurational stability metric is calculated by weighted aggregation of the outputs of each attention head. The aggregation weights are determined based on the variance of the attention heads; attention heads with smaller variances have larger weights, indicating that the features they focus on are more stable. The stability metric ranges from 0 to 1; a larger value indicates greater configurational stability, and a smaller value indicates greater configurational change. The metric calculation uses a sliding window averaging method with a window length of 15 iterations to smooth out short-term fluctuations in the metric value.
[0149] The spatial configuration information generation module integrates location changes, topological change characteristics, and configuration stability metrics to form a complete configuration description. Configuration information is stored in a structured data format, containing six fields: spatial point identifier, current coordinates, change vector, topological feature vector, stability metric, and timestamp. The configuration information update frequency is synchronized with the optimization iteration, updating once per iteration. Historical configuration information saves records from the last 50 iterations for trend analysis. Configuration quality is assessed by comprehensively analyzing the distribution characteristics of the stability metric; configurations with a mean stability metric greater than 0.8 are considered high-quality, while those with a mean less than 0.5 are considered low-quality and require further optimization.
[0150] The conversion of configuration stability metrics into spatial geometric confidence scores for feature point pairs is implemented using a mapping function, which is determined based on the configuration stability metric and the spatial relationships between feature point pairs. These spatial relationships include three factors: Euclidean distance, angular relationship, and topological connectivity between feature point pairs. Feature point pairs with closer distances, more stable angular relationships, and stronger topological connectivity have higher geometric confidence scores. Confidence scores are calculated using a weighted geometric average method, with a weight of 0.6 for configuration stability metrics, 0.25 for spatial distance factors, and 0.15 for angular stability. The confidence score range is limited to 0.1 to 1.0 to avoid extreme values affecting subsequent processing. Confidence score updates employ an exponential smoothing method with a smoothing coefficient of 0.3 to maintain the continuity of confidence score changes.
[0151] Cross-perspective semantic association graph modulation and weighting processing utilizes spatial geometric confidence to correct the semantic association strength of feature point pairs. The modulation formula is the original association strength multiplied by the square root of the geometric confidence, then multiplied by the modulation coefficient. The modulation coefficient is adaptively determined based on the distribution characteristics of the geometric confidence: 0.8 is set when the confidence variance is greater than 0.1, and 1.2 is set when the variance is less than 0.05. Boundary constraints are applied to the modulated association strength to ensure that the strength value is within the range of 0.05 to 1.0; values outside this range are truncated. An incremental update strategy is used for the association graph update, updating only feature point pairs with confidence changes exceeding 0.05 to reduce computational overhead. The update process maintains the sparsity of the association graph, removing weak associations with a modulated strength below 0.1 and retaining strong associations with a strength above 0.3.
[0152] Closed-loop coupling is achieved by establishing a bidirectional information flow between spatial configuration information and a cross-perspective semantic association graph. The forward information flow converts configuration stability metrics into geometric confidence and modulates the association strength, while the reverse information flow uses the modulated association strength as guidance for configuration optimization. A synchronization mechanism for the bidirectional information flow ensures the temporal consistency of information transmission, employing a producer-consumer model where configuration information acts as the producer and association graph modulation as the consumer, achieving asynchronous communication through a message queue. Information flow control utilizes a backpressure mechanism, automatically reducing the production frequency when the consumption rate is slower than the production rate to prevent memory overflow. Coupling strength is quantified using mutual information; a mutual information value greater than 0.3 indicates good coupling, while a value less than 0.1 requires adjustment of coupling parameters.
[0153] The confidence score for the correspondence is calculated based on statistical analysis of the modulated cross-perspective semantic association graph. The confidence score comprehensively considers three factors: association strength, association stability, and association consistency. Association strength is directly derived from the modulated semantic association strength. Association stability is measured by calculating the variance of the association strength within a time window; a variance less than 0.02 indicates higher stability. Association consistency is assessed by comparing the consistency of association strength across feature point pairs from different perspectives. Consistency is measured using the correlation coefficient; a correlation coefficient greater than 0.8 indicates better consistency. The confidence score is calculated using a weighted harmonic mean method, with association strength weighted at 0.5, stability weighted at 0.3, and consistency weighted at 0.2, ensuring a balanced impact of the three factors.
[0154] The confidence threshold employs an adaptive setting strategy, determined based on the statistical characteristics of the confidence distribution. The threshold is calculated as the 25th percentile of the confidence distribution, ensuring that the lowest 25% of feature point pairs are filtered out. Feature point pairs below the confidence threshold are re-matched using the Hungarian algorithm to find the optimal allocation, with the objective function being to maximize the overall association strength. The re-matching process limits the matching distance; matches exceeding three times the average feature point distance are rejected to avoid unreasonable long-distance matching. Match result verification is achieved through geometric consistency checks; matches that do not satisfy the epipolar constraint are marked as low-quality matches.
[0155] The matching results and spatial configuration information are fed back to the weighted reconstructed multi-objective joint optimization function via a feedback control mechanism. The fed-back information includes three parts: new feature point correspondences, correspondence quality assessments, and configuration change suggestions. New correspondences update the data terms in the optimization function, the quality assessment adjusts the weights of the correspondences, and the configuration change suggestions affect the parameter settings of the constraint terms. The feeding-back frequency is set to once every 20 optimization iterations to avoid frequent updates that could lead to optimization instability. The validity of the fed-back information is verified through cross-validation; 20% of the correspondences are randomly selected for independent verification, and fed-back information is rejected if the verification pass rate is below 80%.
[0156] The dynamic optimization and refinement process employs a multi-round iterative strategy, with each round comprising five stages: weight adjustment, structural evolution, configuration analysis, correlation graph update, and correspondence optimization. The round interval is set at 50 basic optimization iterations, and the number of rounds is adaptively determined based on convergence, with a maximum limit of 10 rounds. Convergence is determined based on the rate of change of the overall reprojection error; convergence is considered achieved when the error change is less than 1% for three consecutive rounds. The refinement effect is evaluated by comparing the reprojection error, geometric consistency, and feature matching accuracy before and after refinement; refinement is considered effective when all three indicators show improvement.
[0157] In one optional implementation, the adaptive weight adjustment mechanism is used to reconstruct the multi-objective joint optimization function with weights. A multi-scale geometric constraint network is constructed in the reconstructed multi-objective joint optimization function. The multi-scale geometric constraint network limits the spatial point position adjustment range and neighborhood topological stability, including:
[0158] Obtain the adaptive weight coefficients of each optimization objective term output by the adaptive weight adjustment mechanism, and use the adaptive weight coefficients to dynamically weight and combine the geometric consistency objective term and the topology preservation objective term in the multi-objective joint optimization function to obtain the weighted reconstructed multi-objective joint optimization function;
[0159] A multi-scale geometric constraint network is constructed based on the initial 3D reconstruction results. The fine-grained constraint layer constructs local differential geometric constraints by calculating the local surface tangent plane deviation between spatial points and nearest neighbor points, while the coarse-grained constraint layer constructs global structural geometric constraints by calculating the global position distribution entropy of spatial points and the intersection consistency of multi-view observation rays.
[0160] The local differential geometric constraints and the global structural geometric constraints are embedded into the weighted reconstructed multi-objective joint optimization function. The allowable position adjustment vector of the spatial points is calculated based on the local differential geometric constraints, and the global position drift penalty factor is calculated based on the global structural geometric constraints.
[0161] Based on the allowable position adjustment vector and the global position drift penalty factor, a neighborhood topology connection graph is constructed. By extracting the set of topological connection edges between spatial points and neighborhood points and calculating the edge length change rate and edge angle change rate of the topological connection edges, the neighborhood topology stability index is obtained.
[0162] like Figure 2 As shown, the method includes:
[0163] The adaptive weight coefficient acquisition module receives real-time weight outputs for each optimization objective from the adaptive weight adjustment mechanism. Weight acquisition uses a publish-subscribe pattern for asynchronous communication. The weight data structure uses a key-value pair mapping, where the key is the objective identifier and the value is the corresponding weight coefficient. The weight coefficient maintains a double-precision floating-point format, with a numerical range limited to 0.01 to 0.95. The weight acquisition frequency is synchronized with the optimization iteration. The latest weight coefficient is acquired at the start of each iteration, with an acquisition timeout set to 100 milliseconds. If the timeout occurs, cached historical weight values are used. Weight validity verification checks whether the sum of the weight coefficients is close to 1.0. If the deviation exceeds 0.05, weight re-normalization is triggered. Normalization uses a scaling method to ensure that the weight sum is 1.0. Weight change monitoring is achieved by calculating the magnitude of changes in continuously acquired weight coefficients. Changes exceeding 0.1 are recorded as significant change events for subsequent optimization strategy adjustments.
[0164] The dynamic weighted combination process of the geometric consistency objective and the topology preservation objective is adjusted in real time based on the acquired adaptive weight coefficients. The geometric consistency objective includes three sub-items: minimizing reprojection error, consistency of feature point correspondence, and continuity of spatial point positions. The basic weights for each sub-item are set to 0.5, 0.3, and 0.2, respectively. The topology preservation objective includes three sub-items: neighborhood connectivity preservation, local curvature continuity, and global structural stability. The basic weights are set to 0.4, 0.35, and 0.25, respectively. The dynamic weighted calculation multiplies the basic weights by the adaptive weight coefficients to obtain the real-time weighted coefficients. The weighted coefficients adopt an exponential smoothing update strategy, with a smoothing coefficient set to 0.3 to avoid drastic weight changes affecting optimization stability. The objective item combination adopts a weighted linear combination method. The combination function is the sum of each objective item multiplied by its corresponding weight coefficient. The combination result undergoes numerical stability checks to prevent numerical overflow or underflow.
[0165] The weighted reconstructed multi-objective joint optimization function is constructed using a modular design. The function structure comprises three main components: a data fitting term, a geometric constraint term, and a regularization term. The data fitting term accounts for 60% to 70% of the total weight, the geometric constraint term accounts for 20% to 30%, and the regularization term accounts for 5% to 15%, with the specific proportions dynamically adjusted based on adaptive weight coefficients. The function reconstruction process maintains the differentiability of the objective terms, ensuring the continuity and numerical stability of gradient calculations. The computational complexity of the reconstructed function is controlled through sparsity techniques, removing objective terms with weights less than 0.01 to reduce computational overhead. Function verification uses numerical differentiation to check the correctness of gradient calculations, with gradient errors controlled within 1e-6.
[0166] The multi-scale geometric constraint network is constructed based on the spatial point distribution characteristics of the initial 3D reconstruction results, employing a hierarchical design. The network architecture adopts a pyramid structure, comprising three layers: fine-grained constraint layer, medium-grained constraint layer, and coarse-grained constraint layer. The constraint scope of each layer is the local neighborhood, medium-grained region, and global structure, respectively. The constraint network initialization process analyzes the density distribution and geometric characteristics of spatial points; regions with high density are subject to finer-grained constraints, while regions with low density are subject to coarser-grained constraints. The network topology is represented using a graph structure, where nodes correspond to spatial points, edges correspond to constraint relationships, and edge weights represent constraint strength. Constraint strength is calculated based on the local geometric characteristics and global importance of spatial points. Geometric characteristics include curvature, rate of change of the normal vector, and neighborhood density, while global importance is calculated using the PageRank algorithm.
[0167] The construction of local differential geometric constraints in the fine-grained constraint layer is achieved by calculating the deviation of the local surface tangent plane between a spatial point and its nearest neighbor. The k-nearest neighbor method is used for nearest neighbor selection, with the k value adaptively determined based on the local point density. For high-density regions, k is set to 12 to 16, while for low-density regions, k is set to 6 to 10. The local surface fitting employs the moving least squares method, fitting a quadratic surface within the neighborhood of each spatial point, with the fitting radius set to 2.5 times the average point spacing. The tangent plane is calculated based on the first-order partial derivatives of the fitted surface at the spatial point. The tangent plane normal vector is obtained through the cross product of the partial derivatives, and normalization ensures a modulus of 1.0. The tangent plane deviation is calculated as the directed distance from a neighboring point to the fitted tangent plane. The sign of the deviation indicates whether the point is above or below the tangent plane, and the absolute value of the deviation indicates the degree of deviation. The deviation threshold is adaptively set based on the local geometric complexity: 10% of the average point spacing for flat regions and 25% for regions with greater curvature.
[0168] The global structural geometric constraint construction of the coarse-grained constraint layer comprises two parts: global position distribution entropy calculation and multi-view observation ray convergence consistency analysis. Global position distribution entropy calculation divides the 3D space into a cubic mesh, with the mesh side length set to 2% of the scene scale. The number of spatial points within each mesh is counted, and the information entropy of the point number distribution is calculated. A high entropy value indicates a relatively uniform distribution of spatial points, while a low entropy value indicates clustering. The entropy calculation accuracy is maintained to three decimal places. Multi-view observation ray convergence consistency analysis collects all observable rays from each spatial point. The ray equations are expressed parametrically, including two parameters: the origin coordinates and the direction vector. Convergence consistency is evaluated by calculating the nearest distance between all ray pairs. Ray pairs with a distance less than 1 mm are considered convergent, while those with a distance greater than 5 mm are considered incongruent. The consistency score is the ratio of the number of convergent ray pairs to the total number of ray pairs. A score above 0.8 indicates good convergence consistency, while a score below 0.5 indicates significant observational discrepancies.
[0169] The multi-objective joint optimization function, after being embedded and weighted by local differential geometric constraints and global structural geometric constraints, is implemented through constraint terms. Local constraint terms are added to the optimization function as the sum of squares of tangent plane deviations. The constraint strength weights are adaptively adjusted according to the degree of deviation; the weights are smaller when the deviation is less than a threshold and significantly increased when the deviation exceeds the threshold. Global constraint terms include a distribution entropy regularization term and a ray convergence consistency term. The distribution entropy term encourages uniform distribution of spatial points, while the convergence consistency term penalizes observation inconsistencies. The proportion of constraint weights in the total optimization function is dynamically adjusted according to the degree of constraint violation; the constraint weight is 5% when the violation is minor and increases to 30% when the violation is severe. Automatic differentiation techniques are used to calculate the gradients of the constraint terms, ensuring both accuracy and computational efficiency.
[0170] The permissible position adjustment vector calculation is based on the gradient direction and constraint strength of the local differential geometric constraints. The adjustment vector direction is along the negative direction of the constraint gradient, i.e., the direction that reduces constraint violations. The vector magnitude is determined according to the degree of constraint violation and local geometric features. Adjustment magnitude limits ensure that a single adjustment does not exceed 50% of the average point spacing, preventing excessive adjustments from damaging the geometric structure. The adjustment vector calculation considers the joint effect of multiple constraints, using a weighted average method to combine the adjustment vectors generated by each constraint. The weights are allocated according to the importance and degree of violation of the constraints. Vector validity verification is achieved by checking whether the adjusted vector meets the boundary conditions; adjustment vectors exceeding the spatial boundary are truncated. The adjustment vector is stored in a sparse format, recording only non-zero adjustment vectors to save storage space.
[0171] The global position drift penalty factor is calculated based on the degree of violation of global structural geometric constraints and the global importance of spatial points. The penalty factor is proportional to the distribution entropy deviation and the degree of ray intersection inconsistency; the larger the deviation, the larger the penalty factor. The base penalty factor is set to 1.0. For every 0.1 increase in distribution entropy deviation, the penalty factor increases by 0.2; for every 0.1 increase in ray intersection inconsistency rate, the penalty factor increases by 0.3. The upper limit of the penalty factor is set to 5.0 to avoid excessive penalties that could lead to optimization convergence difficulties. A sliding window averaging method is used for penalty factor calculation, with a window length of 20 iterations to smooth out short-term fluctuations in the factor. The penalty factor is applied to position adjustment constraints; the adjustment magnitude is inversely proportional to the penalty factor, with a larger penalty factor allowing for a smaller adjustment magnitude.
[0172] The neighborhood topology connection graph is constructed based on a joint consideration of the allowable position adjustment vector and the global position drift penalty factor to establish topological connections between spatial points. The connection graph adopts an undirected graph structure, where nodes represent spatial points, edges represent topological connections, and edge weights represent connection strength. Connection establishment criteria include two conditions: distance constraints and geometric similarity constraints. The distance constraint requires that the distance between spatial points be less than three times the average point spacing, and the geometric similarity constraint requires that the angle between normal vectors be less than 60 degrees and the curvature difference be less than 50% of the local average curvature. Connection strength is calculated based on the product of the inverse distance and geometric similarity; the closer the point pairs and the more similar their geometric features, the stronger the connection. The connection graph is updated every 10 optimization iterations, recalculating connection relationships and connection strengths during the update process. The graph is stored in an adjacency list format, supporting efficient neighborhood query and traversal operations.
[0173] The extraction of the topological connection edge set is achieved by traversing the edge structure of the neighborhood topological connection graph. The extraction process records the start point, end point, and connection strength information of each edge. The edge set is stored using a dynamic array, supporting dynamic addition and deletion of edges. The edge length is calculated as the Euclidean distance between the two connected spatial points, with distance accuracy maintained at the millimeter level. Calculation results are cached to avoid duplicate calculations. The edge angle is calculated as the angle between the connecting edge and a reference direction, which can be selected as the coordinate axis direction or a local principal direction, with an angle range of 0 to 180 degrees. The angle is calculated using the vector dot product method. Edge set quality is evaluated by statistically analyzing the distribution characteristics of edge lengths and angles. Edge sets with good uniformity of distribution are of higher quality, while edge sets with severe skewness require adjustment of connection criteria.
[0174] The edge length change rate is calculated by comparing the length changes of the connected edges before and after optimization. The change rate is defined as the ratio of the length change to the original length. Change rate statistics include four indicators: mean, standard deviation, maximum, and minimum, with statistical precision maintained to four decimal places. The edge length change rate threshold is set at 20%; edges exceeding this threshold are marked as significantly changed edges requiring special attention. Change rate distribution analysis is performed using histogram statistics, with 20 histogram intervals and an interval width of 10% of the threshold. Change rate monitoring employs a sliding window method with a window length of 15 iterations to monitor continuous change trends.
[0175] The calculation method for the angle change rate is similar to that for the side length change rate, comparing the changes in the angle between the connecting edge and the reference direction before and after optimization. The threshold for the angle change rate is set at 30 degrees; angle changes exceeding this threshold are considered significant. The angle change analysis considers the periodicity of angles, correctly handling jumps between 180 degrees and 0 degrees. The calculated change rate results are used to evaluate the stability of the topology; a small change rate indicates a stable topology, while a large change rate indicates a significant change in the topology.
[0176] The neighborhood topological stability index is calculated by combining the rates of change of side length and the rates of change of corner angles. The index uses a weighted geometric mean method, with a weight of 0.6 for the rate of change of side length and 0.4 for the rate of change of corner angles. The stability index ranges from 0 to 1; values closer to 1 indicate greater topological stability, while values closer to 0 indicate greater topological change. Robust statistical methods are used in the index calculation, employing the median instead of the mean to reduce the impact of outliers. A stability threshold of 0.7 is set; regions below this threshold are marked as topologically unstable and require enhanced topological constraints. The index is updated synchronously with the optimization iterations, updating once per iteration, and historical index values are saved for trend analysis.
[0177] A second aspect of this invention provides a system for reconstructing, registering, and optimizing three-dimensional scene images from multiple perspectives, comprising:
[0178] The first unit is used to acquire multi-frame 3D scene image data from different spatial locations, and perform adaptive feature extraction on the multi-frame 3D scene image data to obtain a multi-level feature set containing spatial location and geometric structure.
[0179] The second unit is used to construct a cross-perspective semantic association graph based on the multi-level feature set. By establishing semantic consistency constraints, it realizes the adaptive mapping of corresponding feature points of the same spatial entity under different perspectives, and obtains feature correspondence with confidence assessment.
[0180] The third unit is used to construct a multi-view spatial transformation relationship based on the spatial location information and geometric structure information contained in the feature correspondence relationship, and to obtain the transformation parameters describing the relative attitude and local deformation between viewpoints through iterative calculation.
[0181] The fourth unit is used to perform coordinate system alignment on the multi-frame 3D scene image data using the transformation parameters, generate an initial 3D reconstruction result under a unified coordinate system, and identify local regions with registration deviations based on the reprojection error distribution of spatial points in the initial 3D reconstruction result, and construct a multi-objective joint optimization function for the local regions.
[0182] The fifth unit is used to integrate the multi-objective joint optimization function into a dynamic adaptive weight control mechanism, guide the evolution of the initial 3D reconstruction result under spatial geometric constraints, and simultaneously couple the spatial configuration information obtained during the evolution process with the cross-view semantic association graph to drive the dynamic optimization and refinement of feature correspondence.
[0183] A third aspect of the present invention provides an electronic device, comprising:
[0184] processor;
[0185] Memory used to store processor-executable instructions;
[0186] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.
[0187] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.
[0188] This invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the invention.
[0189] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for reconstructing, registering, and optimizing 3D scene images from multiple perspectives, characterized in that: include: Acquire multi-frame 3D scene image data from different spatial locations, and perform adaptive feature extraction on the multi-frame 3D scene image data to obtain a multi-level feature set containing spatial location and geometric structure; Based on the multi-level feature set, a cross-perspective semantic association graph is constructed. By establishing semantic consistency constraints, adaptive mapping of corresponding feature points of the same spatial entity under different perspectives is achieved, resulting in feature correspondence with confidence assessment. Based on the spatial location information and geometric structure information contained in the feature correspondence, a multi-view spatial transformation relationship is constructed, and transformation parameters describing the relative attitude and local deformation between viewpoints are obtained through iterative calculation. The transformation parameters are used to align the coordinate system of the multi-frame 3D scene image data to generate an initial 3D reconstruction result in a unified coordinate system. Based on the reprojection error distribution of spatial points in the initial 3D reconstruction result, local regions with registration deviations are identified, and a multi-objective joint optimization function is constructed for the local regions. The multi-objective joint optimization function is integrated into a dynamic adaptive weight control mechanism to guide the evolution of the initial 3D reconstruction result under spatial geometric constraints. Simultaneously, the spatial configuration information obtained during the evolution process is coupled in a closed loop with the cross-view semantic association graph, driving the dynamic optimization and refinement of feature correspondences, including: For each optimization objective term in the multi-objective joint optimization function, the convergence rate and contribution weight of each objective term are calculated based on the current residual distribution of the initial 3D reconstruction result, and an adaptive weight adjustment mechanism is constructed. The multi-objective joint optimization function is reconstructed using the adaptive weight adjustment mechanism. A multi-scale geometric constraint network is constructed in the reconstructed multi-objective joint optimization function. The multi-scale geometric constraint network limits the adjustment range of spatial point positions and the topological stability of the neighborhood, thereby guiding the structural evolution of the initial three-dimensional reconstruction result. Based on the multi-scale geometric constraint network, the positional changes of spatial points and the local topological structure changes during the structural evolution are extracted, and a self-attention aggregation mechanism is introduced to calculate the configurational stability measure of spatial points to obtain spatial configuration information. The configuration stability measure in the spatial configuration information is converted into the spatial geometric confidence of feature point pairs through the self-attention aggregation mechanism. The spatial geometric confidence is then used to modulate and weight the semantic association strength of feature point pairs in the cross-view semantic association graph to form a closed-loop coupling. The confidence level of the correspondence between feature point pairs is calculated based on the modulated cross-view semantic association graph. Feature point pairs below the confidence threshold are rematched, and the matching results and the spatial configuration information are fed back to the weighted reconstructed multi-objective joint optimization function to drive the dynamic optimization and refinement of feature correspondence.
2. The method according to claim 1, characterized in that, Based on the multi-level feature set, a cross-perspective semantic association graph is constructed. By establishing semantic consistency constraints, adaptive mapping of corresponding feature points of the same spatial entity under different perspectives is achieved, resulting in feature correspondences with confidence assessment, including: For each feature point in the multi-level feature set, the spatial information of the feature point at different scales is extracted, including local structural features and global context features. The local structural features and global context features are then fused to obtain a composite feature descriptor of the feature point. Based on the composite feature descriptor, a cross-perspective semantic association graph is constructed. Semantic propagation constraints are established in the cross-perspective semantic association graph. Existing reliable correspondences guide unconfirmed candidate correspondences to obtain the initial semantic association strength distribution. Based on the initial semantic association strength distribution, the corresponding feature points of the same spatial entity under different perspectives are analyzed, the semantic consistency deviation between each candidate corresponding point and the surrounding confirmed correspondence is calculated, and the semantic consistency deviation is used as a constraint to optimize the initial semantic association strength distribution to obtain the corrected semantic association strength. Based on the corrected semantic association strength and the stability of the semantic consistency deviation, a confidence assessment model is established. The reliability of semantic propagation and the consistency of association strength are evaluated through the confidence assessment model, and a confidence assessment value is generated for the corresponding feature point. The confidence assessment value is used to update the correspondence in the cross-view semantic association graph. Correspondences that exceed the dynamic threshold are determined as reliable feature point mappings. The confidence assessment value is used to mark the reliability of the correspondence, and finally, feature correspondences with quality assessment are obtained.
3. The method according to claim 2, characterized in that, Based on the composite feature descriptor, a cross-perspective semantic association graph is constructed. Semantic propagation constraints are established in the cross-perspective semantic association graph. Existing reliable correspondences guide unconfirmed candidate correspondences, resulting in an initial semantic association strength distribution including: For each feature point in the composite feature descriptor, the semantic similarity of the feature points across the cross-view range is analyzed, and a cross-view semantic association graph is generated by combining the spatial location information of the feature points. In the cross-view semantic association graph, feature point pairs that satisfy the bidirectional optimal matching criterion are determined, the feature point pairs are defined as reliable correspondences, and the relative positional relationship and local neighborhood topology of the reliable correspondences in three-dimensional space are analyzed to construct the spatial structure features of the reliable correspondences. Using the spatial structural features, a constraint propagation network is constructed in the cross-perspective semantic association graph with the reliable correspondence as the propagation source point, and the spatial structural features are used to form topological consistency constraints on the propagation path. For the unconfirmed candidate correspondences in the cross-perspective semantic association graph, the topological differences between the candidate correspondences and the reliable correspondences on the propagation path are evaluated and mapped to a propagation attenuation coefficient that characterizes the degree of constraint satisfaction. The semantic similarity of the candidate correspondences is optimized by using the propagation attenuation coefficient. By progressively transmitting the spatial structural features of reliable correspondences in the constraint propagation network, the topological constraint modulation of the candidate correspondences is achieved, generating a semantic association strength distribution.
4. The method according to claim 1, characterized in that, The transformation parameters are used to align the coordinate systems of the multi-frame 3D scene image data, generating an initial 3D reconstruction result in a unified coordinate system. Based on the reprojection error distribution of spatial points in the initial 3D reconstruction result, local regions with registration deviations are identified. A multi-objective joint optimization function is constructed for these local regions, including: The local coordinate system of each frame in the multi-frame 3D scene image data is transformed to a unified coordinate system using the transformation parameters. A feature fusion network is established under the unified coordinate system to fuse and reconstruct spatial points from different frames, generating an initial 3D reconstruction result. For each spatial point in the initial 3D reconstruction result, the deviation between the reprojection position and the actual observation position of each spatial point under each observation view is analyzed based on the feature fusion network, and the spatial distribution of the deviation is statistically analyzed to obtain the reprojection error distribution characteristics. Based on the reprojection error distribution characteristics, identify the set of spatial points exceeding the error threshold in the initial 3D reconstruction result, construct a spatial structure optimization map, analyze the connectivity and clustering of the set of spatial points, and divide the error spatial points with spatial connectivity into local regions with registration deviations. For the local region, a regional optimization objective is constructed in the spatial structure optimization diagram, which includes geometric consistency constraints and minimization of reprojection error. At the same time, the boundary continuity constraint relationship between the local region and the adjacent regions is analyzed. Based on the region optimization objective and the boundary continuity constraint, a multi-objective joint optimization function is established to ensure that adjacent local regions maintain spatial continuity and geometric consistency at the boundary.
5. The method according to claim 4, characterized in that, Based on the reprojection error distribution characteristics, the set of spatial points exceeding the error threshold in the initial 3D reconstruction result is identified. A spatial structure optimization map is constructed, and the connectivity and clustering of the spatial point set are analyzed. Error spatial points with spatial connectivity are divided into local regions with registration deviations, including: Based on the reprojection error distribution characteristics, spatial points whose reprojection errors exceed the error threshold are identified in the initial 3D reconstruction results, and the 3D coordinates, normal vectors and observation view information of the spatial points are extracted to obtain a set of spatial points that exceed the error threshold. Using each spatial point in the set of spatial points as a node, the geodesic distance and local curvature consistency between pairs of spatial points are calculated, and weighted edges are established between pairs of spatial points that satisfy the distance and curvature constraints. The weight of the weighted edges is jointly measured by the geodesic distance and curvature consistency to construct a spatial structure optimization graph. In the spatial structure optimization graph, the weighted edge connections are traversed using a graph connectivity analysis algorithm to identify spatial point subsets that form connected subgraphs through the weighted edge connections. The cluster density is calculated based on the weighted edge weight distribution within the spatial point subsets to obtain connectivity analysis results and clustering analysis results. Based on the connectivity analysis results and the clustering analysis results, the connected subgraphs in the spatial structure optimization graph are screened, and the connected subgraphs with clustering density exceeding the clustering threshold are marked as candidate local regions. Based on the distribution characteristics of the internal spatial points of the candidate local regions, the boundaries are adjusted, and the candidate local regions after boundary adjustment are determined as a set of local regions with registration deviation.
6. The method according to claim 1, characterized in that, The multi-objective joint optimization function is reconstructed using the adaptive weight adjustment mechanism. A multi-scale geometric constraint network is then constructed within the reconstructed multi-objective joint optimization function. This multi-scale geometric constraint network limits the spatial point position adjustment range and neighborhood topological stability, including: Obtain the adaptive weight coefficients of each optimization objective term output by the adaptive weight adjustment mechanism, and use the adaptive weight coefficients to dynamically weight and combine the geometric consistency objective term and the topology preservation objective term in the multi-objective joint optimization function to obtain the weighted reconstructed multi-objective joint optimization function; A multi-scale geometric constraint network is constructed based on the initial 3D reconstruction results. The fine-grained constraint layer constructs local differential geometric constraints by calculating the local surface tangent plane deviation between spatial points and nearest neighbor points, while the coarse-grained constraint layer constructs global structural geometric constraints by calculating the global position distribution entropy of spatial points and the intersection consistency of multi-view observation rays. The local differential geometric constraints and the global structural geometric constraints are embedded into the weighted reconstructed multi-objective joint optimization function. The allowable position adjustment vector of the spatial points is calculated based on the local differential geometric constraints, and the global position drift penalty factor is calculated based on the global structural geometric constraints. Based on the allowable position adjustment vector and the global position drift penalty factor, a neighborhood topology connection graph is constructed. By extracting the set of topological connection edges between spatial points and neighborhood points and calculating the edge length change rate and edge angle change rate of the topological connection edges, the neighborhood topology stability index is obtained.
7. A multi-view 3D scene image reconstruction, registration, and optimization system, used to implement the method of any one of claims 1-6, characterized in that, include: The first unit is used to acquire multi-frame 3D scene image data from different spatial locations, and perform adaptive feature extraction on the multi-frame 3D scene image data to obtain a multi-level feature set containing spatial location and geometric structure. The second unit is used to construct a cross-perspective semantic association graph based on the multi-level feature set. By establishing semantic consistency constraints, it realizes the adaptive mapping of corresponding feature points of the same spatial entity under different perspectives, and obtains feature correspondence with confidence assessment. The third unit is used to construct a multi-view spatial transformation relationship based on the spatial location information and geometric structure information contained in the feature correspondence relationship, and to obtain the transformation parameters describing the relative attitude and local deformation between viewpoints through iterative calculation. The fourth unit is used to perform coordinate system alignment on the multi-frame 3D scene image data using the transformation parameters, generate an initial 3D reconstruction result under a unified coordinate system, and identify local regions with registration deviations based on the reprojection error distribution of spatial points in the initial 3D reconstruction result, and construct a multi-objective joint optimization function for the local regions. The fifth unit is used to integrate the multi-objective joint optimization function into a dynamic adaptive weight control mechanism, guide the evolution of the initial 3D reconstruction result under spatial geometric constraints, and simultaneously couple the spatial configuration information obtained during the evolution process with the cross-view semantic association graph to drive the dynamic optimization and refinement of feature correspondence.
8. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 6.
Citation Information
Patent Citations
Multi-view three-dimensional reconstruction method, system and equipment based on deep learning
CN115170746A
Real scene three-dimensional model reconstruction method and system based on deep learning multi-view dense matching
CN117315169A