Three-dimensional Gaussian visual positioning method for sparse visual angle

Through three-dimensional Gaussian sputtering model and pseudo-view rendering technology, the problem of unstable feature matching under sparse view angle is solved, the positioning accuracy and system stability are improved, and it is suitable for highly robust visual positioning under sparse view angle conditions.

CN120339380AActive Publication Date: 2025-07-18CHINA UNIV OF PETROLEUM (EAST CHINA)

Patent Information

Application Number
CN202510797866.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-07-18
Estimated Expiration
2045-06-16

AI Technical Summary

Technical Problem

The prior art is difficult to build a stable and accurate three-dimensional structure under sparse perspective conditions, resulting in unstable feature matching, low positioning accuracy, and high computational overhead of neural rendering methods and slow response speed, making it difficult to meet the needs of mobile terminals or real-time systems.

Method used

Using a three-dimensional Gaussian sputtering model, color images and depth maps are generated through sparse viewing image training, global and local features are extracted, feature databases are constructed, and matching accuracy and system stability are improved through closest search and pseudo-view rendering, combining dynamic inner point filtering and pose solution.

Benefits of technology

It significantly improves the positioning accuracy and system stability under sparse perspective, achieves high robustness and actual availability under sparse training conditions, completes the lack of viewing angles, and enhances the balance and coverage of feature expression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339380A_ABST
    Figure CN120339380A_ABST
Patent Text Reader

Abstract

The invention discloses a sparse view angle-oriented three-dimensional Gaussian visual positioning method, which belongs to the technical field of positioning, is used for visual positioning, and comprises the following steps of: extracting global image features and local image features of a sparse image, and constructing a feature database; performing nearest neighbor search on the feature database and the RGB image to obtain a sparseness index, calculating a camera pose of a pseudo view angle, rendering a color image and a depth image under the pseudo view angle by using a training main model, extracting global features and local features and storing the global features and the local features into the feature database, constructing a 2D-3D corresponding relation in combination with a target color image, and performing dynamic inner point screening and pose solving. And obtaining a positioning result. According to the method, visual angle deficiency of a key area is intelligently completed, the balance and coverage of feature expression in a three-dimensional scene are effectively improved, and dynamic updating and mismatching elimination of a 2D-3D matching relation are realized, so that the positioning precision and the system stability are remarkably improved, and the method has higher robustness and actual availability under a sparse training condition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention discloses a three-dimensional Gaussian visual localization method for sparse viewpoints, belonging to the field of localization technology. Background Art

[0002] With the growing demand for high-precision three-dimensional perception and pose estimation in scenarios such as augmented reality (AR), robot navigation, and autonomous driving, visual localization technology, as its core support, has received extensive attention. Currently, the mainstream methods include: absolute pose regression (APR) based on deep learning, scene coordinate regression (SCR), and three-dimensional neural rendering methods (NeRF) that have developed rapidly in recent years. The APR method predicts the camera pose of an image end-to-end through a neural network, avoiding explicit geometric modeling; the SCR method estimates the camera pose through algorithms such as PnP by leveraging the regression relationship between image pixels and three-dimensional space points. In addition, neural rendering methods such as NeRF build continuous volume density and color fields to model the implicit three-dimensional structure of a scene, and some studies have attempted to introduce them into visual localization tasks to assist in geometric information recovery.

[0003] However, the above methods still face a series of challenges in actual deployment. They generally rely on a large number of densely collected training images to support the spatial structure modeling and feature generalization capabilities of the model. When the image distribution is sparse and the view coverage is uneven, the model is difficult to capture the true geometric features of the scene, resulting in a decrease in matching accuracy and unstable pose estimation. Especially in the case of sparse viewpoints, the overlapping area between images is significantly reduced, severely restricting the effective matching between 2D key points and 3D points, causing an increase in the proportion of false matches and an increase in pose errors. At the same time, neural rendering methods generally have a large computational cost, a slow convergence speed, and are highly dependent on view density, making it difficult to meet the requirements of mobile terminals or real-time systems for efficiency and response speed.

[0004] Therefore, there is an urgent need for a highly robust visual localization method that can adapt to sparse viewpoint conditions, which can construct a stable and accurate three-dimensional structure under the conditions of limited training samples and weak geometric constraints, and improve the reliability of feature matching and pose estimation to better serve the actual application requirements. Summary of the Invention

[0005] The purpose of the present invention is to provide a three-dimensional Gaussian visual localization method for sparse viewpoints, so as to solve the problem that in the existing technology, it is difficult to construct a three-dimensional structure with strong geometric consistency under the conditions of sparse training images and insufficient geometric constraints, resulting in unstable feature matching, low localization accuracy, or even inability to complete localization.

[0006] A three-dimensional Gaussian visual localization method for sparse perspectives, including training a three-dimensional Gaussian sputtering model using RGB images and camera poses, generating color images and depth maps through two-dimensional projection and differentiable rendering mechanisms, extracting global image features and local image features of sparse images, and constructing a feature database; Perform a nearest neighbor search on the feature database and the RGB image to obtain a sparsity metric, calculate the camera pose of the pseudo-perspective, render the color map and depth map under the pseudo-perspective using the trained main model, extract global features and local features and incorporate them into the feature database, construct a 2D-3D correspondence in combination with the target color map, perform dynamic inlier screening and pose solution to obtain the localization result.

[0007] Training the three-dimensional Gaussian sputtering model includes using the RGB image as a sparse perspective image, using the same sparse perspective image as the input of the three-dimensional Gaussian sputtering model, and constructing the three-dimensional Gaussian sputtering model through a collaborative modeling method of a dual-branch three-dimensional Gaussian point cloud model, including a training main model and an auxiliary supervision model , each three-dimensional Gaussian point cloud model includes multiple Gaussian points, and the Gaussian point is a quadruple , , , , which are the spatial position, the covariance matrix of the spatial scale and direction of the Gaussian point, the color value, and the opacity respectively; Input the same sparse perspective image into the training main model and the auxiliary supervision model , construct a cross-model view Figure 1 consistency loss based on the rendered images under the same pseudo-perspective, and the loss function is: ; ; In the formula, is the balance coefficient, and are respectively and the rendered images under the same pseudo-perspective, represents the L1 norm, represents the structural similarity index; The color reconstruction loss of a single three-dimensional Gaussian point cloud model is: ; In the formula, is the pixel color of the training image, is the corresponding pixel value of the rendered image, is the pixel , is the pixel color; Dual-branch training total loss function is: ; In the formula, is the weight coefficient, which is used to adjust the proportion of the cross-model visual Figure 1 consistency loss in the total loss; Minimize the dual-branch training total loss to complete the training of the three-dimensional Gaussian sputtering model.

[0008] Generating a color image and a depth map includes, from the camera view, projecting the three-dimensional covariance matrix into a two-dimensional covariance matrix according to the Jacobian transformation : ; In the formula, 𝐽 is the Jacobian matrix of the affine approximation of the projection function, is the view transformation matrix of the camera viewing direction; For each pixel in the image plane , according to sort the two-dimensional distances from all visible Gaussian points, and construct a Gaussian set , use the alpha-blending rendering method to accumulate the color values to generate , and use as the color image: ; ; In the formula, is the cumulative transmittance of the Gaussian point , is the projection opacity of the Gaussian point at the pixel , represents the th Gaussian point, represents the th Gaussian point, is the projection opacity of the Gaussian point at the pixel ; Generate a depth map, and the depth of a single pixel is: ; In the formula, represents the absolute depth value of the Gaussian point from the current view, and use as the depth map.

[0009] Constructing a feature database includes constructing a global image feature set and a local image feature set; The global image feature set is: ; wherein, is the th global feature vector extracted by the image encoding network, is the th rotation matrix in the camera pose, is the th translation vector in the camera pose, is 's maximum value; The local image feature set is: ; wherein, is the three-dimensional space coordinates obtained by back-projecting the local key points of the image, is the local feature descriptor bound to .

[0010] Constructing the global image feature set includes using the local image feature extraction network to extract two-dimensional key point coordinates and the corresponding , combining the camera intrinsic matrix corresponding to the sparse view image, the extrinsic matrix , and the depth value of each key point in the depth map, mapping the two-dimensional key points to the three-dimensional space, and calculating the corresponding : ; wherein, is the camera rotation matrix in the extrinsic matrix, is the translation vector in the extrinsic matrix, binding with the corresponding to form a local three-dimensional feature item and storing it in the local image feature set.

[0011] Obtaining the sparsity index includes, for the Gaussian points in the three-dimensional Gaussian sputtering model, counting the visible view set of , is equivalent to all sparse view images that can observe . For the th sparse view image , let The set of observed Gaussian points is , The sparsity index of is: ; The overlapping degree of the viewing angles of two images and is the intersection - union ratio of their visible Gaussian point sets : ; In the formula, is the th sparse - view image, is the set of observed Gaussian points; Use the sparsity index to construct a sparsity - sorted list.

[0012] Calculating the camera pose of the pseudo - view includes selecting the one with the lowest sparsity from the sparsity - sorted list as the target view to be enhanced, and screening out the image set from all sparse - view images whose is greater than the set threshold : : ; is the reference image with the farthest Euclidean distance from in terms of spatial position, its pose is , is the translation vector, is the rotation angle represented by a unit quaternion; Let the camera pose of be ; ; ; ; In the formula, is the pseudo - view translation vector, is the pseudo - view rotation angle, is the interpolation factor, is the intermediate parameter.

[0013] Rendering the color map and depth map under the pseudo - view using the training main model includes using the same method as generating the color image and depth map, and using the training main model Generate corresponding color images and depth maps under the pseudo-view camera pose ; Extract the local key point coordinates from the color image and , combine with the depth map and map to three-dimensional coordinates with the camera pose to form local three-dimensional feature terms ; Extract the global features of the color image , bind with the camera pose to form global feature terms, where is the pseudo-view camera rotation matrix; Add the local three-dimensional feature terms and global feature terms to the feature database: ; ; In the formula, , are respectively the updated , .

[0014] Combining the target color map to construct the 2D-3D correspondence includes calculating the cosine similarity score between the global feature vector of the image to be processed and : ; Sort in descending order according to , select the images with the highest scores, query the local image feature set corresponding to the image to be processed from the feature database, and use the local feature matcher LightGlue to match the two-dimensional key points of the image to be processed with the local feature descriptors bound with three-dimensional coordinates in the reference image to construct a 2D-3D correspondence set : ; Take as the initial inlier set : .

[0015] Performing dynamic inlier filtering and pose solution includes, in the th round of iteration, using the th inlier set as the input, and using the PnP-RANSAC algorithm to estimate the current camera pose matching pairs For each matching pair, calculate the reprojection error : ; In the formula, represents the camera projection function; If is greater than the set threshold , remove the corresponding matching pair from the inlier set to obtain the updated round of inlier set ; After obtaining , based on the matching pair , use the trained main model to render the color image and depth image in the corresponding view, extract local features and rematch them with the local features of the image to be processed, and construct round of enhanced inlier set ; Merge with the current round of inlier set to obtain round of inlier set: ; In the formula, is the inlier set of the th round of iteration; The iteration process terminates when any of the following conditions is met: (1) The intersection over union of the inlier sets for two consecutive rounds exceeds the set threshold : ; (2) The number of iteration rounds exceeds the maximum set value ; Finally, output the current optimal camera pose as the positioning result of the image to be processed.

[0016] Compared with the prior art, the present invention has the following beneficial effects: The present invention intelligently completes the missing perspectives of key regions, effectively improves the balance and coverage of feature expression in the 3D scene, realizes the dynamic update of the 2D-3D matching relationship and the elimination of incorrect matches, thereby significantly improving the positioning accuracy and system stability, and has stronger robustness and practical usability under sparse training conditions. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 is the overall flowchart of the present invention; Figure 2 is the flowchart for constructing the feature database of the present invention; Figure 3 is the flowchart of the vision enhancement method based on sparse perception; Figure 4 It is a flowchart for dynamic inlier screening and pose solution. Specific implementation mode

[0018] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts shall fall within the protection scope of the present invention.

[0019] A three-dimensional Gaussian visual positioning method for sparse viewpoints includes training a three-dimensional Gaussian sputtering model using RGB images and camera poses, generating color images and depth maps through a two-dimensional projection and differentiable rendering mechanism, extracting global image features and local image features of sparse images, and constructing a feature database; Performing a nearest neighbor search on the feature database and the RGB image to obtain a sparsity index, calculating the camera pose of the pseudo-viewpoint, rendering color maps and depth maps at the pseudo-viewpoint using the trained main model, extracting global features and local features and incorporating them into the feature database, constructing a 2D-3D correspondence in combination with the target color map, performing dynamic inlier screening and pose solution, and obtaining a positioning result.

[0020] Training the three-dimensional Gaussian sputtering model includes using the RGB image as a sparse viewpoint image, using the same sparse viewpoint image as the input of the three-dimensional Gaussian sputtering model, and constructing the three-dimensional Gaussian sputtering model through a co-modeling method of a double-branch three-dimensional Gaussian point cloud model, including a training main model and an auxiliary supervision model , each three-dimensional Gaussian point cloud model includes multiple Gaussian points, and the Gaussian point is a quadruple , , , , are the spatial position, the covariance matrix of the spatial scale and direction of the Gaussian point, the color value, and the opacity respectively; Inputting the same sparse viewpoint image into the training main model and the auxiliary supervision model , constructing a cross-model view Figure 1 consistency loss based on the rendered images under the same pseudo-viewpoint, and the loss function is: ; ; In the formula, is a balance coefficient, and are respectively and Rendered images at the same pseudo-viewpoint denotes the L1 norm denotes the structural similarity index The color reconstruction loss of a single 3D Gaussian point cloud model is ; wherein is the pixel color of the training image is the corresponding pixel value of the rendered image is the pixel , is the pixel color The total loss function for dual-branch training is ; wherein is the weight coefficient, used to adjust the proportion of the cross-model view Figure 1 consistency loss in the total loss Minimizing the total loss of dual-branch training completes the training of the 3D Gaussian sputtering model

[0021] Generating a color image and a depth map includes, at the camera viewpoint, projecting the 3D covariance matrix into a 2D covariance matrix according to the Jacobian transformation : ; wherein, 𝐽 is the Jacobian matrix of the affine approximation of the projection function is the view transformation matrix of the camera viewing direction For each pixel in the image plane , according to sorting the 2D distances from all visible Gaussian points, constructing a Gaussian set , using the alpha-blending rendering method to accumulate color values, generating , taking as the color image: ; ; wherein is the cumulative transmittance of the Gaussian point , is the projected opacity of the Gaussian point at the pixel , represents the th Gaussian point represents the th Gaussian point Is the Gaussian point At the pixel Projection opacity; Generate a depth map, the depth of a single pixel Is: ; In the formula, Represents the absolute depth value of the Gaussian point Under the current perspective, take As the depth map.

[0022] Construct a feature database including constructing a global image feature set and a local image feature set; The global image feature set Is: ; In the formula, Is the th global feature vector extracted by the image encoding network, Is the th rotation matrix in the camera pose, Is the th translation vector in the camera pose, Is The maximum value of; The local image feature set Is: ; In the formula, Is the three-dimensional space coordinates obtained by back-projecting the local key points of the image, Is associated with Bound local feature descriptor.

[0023] Constructing the global image feature set includes using the local image feature extraction network to extract the two-dimensional key point coordinates And the corresponding , combined with the camera intrinsic matrix , extrinsic matrix , the depth value of each key point in the depth map , map the two-dimensional key points to three-dimensional space, and calculate the corresponding : ; In the formula, Is the camera rotation matrix in the extrinsic matrix, Is the translation vector in the extrinsic matrix, take And the corresponding Bind to form a local three-dimensional feature item and stored in the local image feature set.

[0024] Obtaining the sparsity index includes, for the Gaussian points in the three-dimensional Gaussian sputtering model , counting the visible view angle set of , equivalent to all the sparse view images that can observe . For the th sparse view image , let the set of Gaussian points observed be , the sparsity index of is: ; The view overlap degree between two images and is the intersection-to-union ratio of their visible Gaussian point sets : ; In the formula, is the th sparse view image, is the set of Gaussian points observed; Using the sparsity index to construct a sparsity sorted list.

[0025] Calculating the camera pose of the pseudo-view includes selecting the one with the lowest sparsity from the sparsity sorted list as the target view to be enhanced, and screening out from all the sparse view images those with the greater than the set threshold image set : ; is the reference image with the farthest Euclidean distance from in terms of spatial position, the pose is , is the translation vector, is the rotation angle represented by the unit quaternion; Let the camera pose of be ; ; ; ; In the formula, is the pseudo-viewpoint translation vector, is the pseudo-viewpoint rotation angle, is the interpolation factor, is the intermediate parameter.

[0026] Rendering the color image and the depth map under the pseudo-viewpoint using the trained main model includes using the same method as generating the color image and the depth map, and using the trained main model at the pseudo-viewpoint camera pose to generate the corresponding color image and the depth map ; Extract the local key point coordinates from the color image and , and combine with the depth map to map to the three-dimensional coordinates , forming the local three-dimensional feature item ; Extract the global feature of the color image , bind it with the camera pose , forming the global feature item, is the pseudo-viewpoint camera rotation matrix; Add the local three-dimensional feature item and the global feature item to the feature database: ; ; In the formula, , are the updated , .

[0027] Combining the target color image to construct the 2D-3D correspondence includes calculating the cosine similarity score of the global feature vector of the image to be processed and : ; Sort in descending order according to , select the images with the highest scores, query the local image feature set corresponding to the image to be processed from the feature database, and use the local feature matcher LightGlue to match the two-dimensional key points of the image to be processed with the local feature descriptors bound with three-dimensional coordinates in the reference image to construct the 2D-3D correspondence set : ; Use as the initial set of inliers : .

[0028] Performing dynamic inlier screening and pose solution includes, in the th iteration, using the th set of inliers as the input, and adopting the PnP-RANSAC algorithm to estimate the current camera pose matching pairs , calculating the reprojection error for each matching pair : ; In the formula, represents the camera projection function; If is greater than the set threshold , the corresponding matching pair is removed from the set of inliers to obtain the updated th set of inliers ; After obtaining , based on the matching pairs , use the trained main model to render the color map and depth map in the corresponding perspective, extract local features and rematch them with the local features of the image to be processed, and construct the th set of enhanced inliers ; Merge with the current th set of inliers to obtain the th set of inliers: ; In the formula, is the set of inliers in the th iteration; The iteration process terminates when any of the following conditions is met: (1) The intersection over union of the sets of inliers in two consecutive rounds exceeds the set threshold : ; (2) The number of iteration rounds exceeds the maximum set value ; Finally, output the current optimal camera pose as the positioning result of the image to be processed.

[0029] The overall process of the present invention is as shown in Figure 1As shown, it includes training a three-dimensional Gaussian sputtering model using RGB images and camera poses, performing sparse view enhancement, extracting global image features and local image features of the sparse images, and constructing a feature database; performing nearest neighbor search on the feature database and the image to be processed to obtain a sparsity index, calculating the camera pose of the pseudo-viewpoint, rendering a color map and a depth map under the pseudo-viewpoint using the trained main model, extracting global features and local features and incorporating them into the feature database, constructing a 2D-3D correspondence in combination with the target color map, performing dynamic inlier screening and pose solution to obtain the positioning result. The process of constructing the feature database of the present invention is as Figure 2 shown. First, obtain RGB images and camera parameters, perform three-dimensional Gaussian point cloud modeling, generate a color map and a depth map after rendering, extract local features from the color map using the NetVLAD network, extract local features from the depth map using the SuperPoint network, and finally obtain the constructed feature database. The process of the visual enhancement method based on sparse perception is as Figure 3 shown. Input the set of candidate viewpoints, calculate the visibility scores of each viewpoint, select the most sparse viewpoint, select a reference viewpoint, interpolate to generate a pseudo-viewpoint, extract the features of the pseudo-viewpoint and add them to the feature database, determine whether the end condition is met. If it is met, end the calculation; if not, return to select the most sparse viewpoint. The interpolated pseudo-viewpoint needs to be returned and added to the set of candidate viewpoints. The process of dynamic inlier screening and pose solution is as Figure 4 shown, including inputting the image to be processed (query image) and the feature database, obtaining the initial pose through image retrieval, extracting the local features of the query image, performing 2D-3D matching, updating the 2D-3D correspondence, and performing camera pose solution. At this time, if it does not converge, return to perform 2D-3D matching; if it converges, output the camera pose.

[0030] In the embodiment of the present invention, the following steps are included: S1: Obtain sparse view images and their camera pose information, construct two three-dimensional Gaussian point cloud models, namely the training main model and the auxiliary supervision model, and construct a cross-model view Figure 1 consistency loss by minimizing the difference in the rendering results of the two models under the same viewpoint to optimize the three-dimensional structure modeling ability of the main model.

[0031] S2: Extract the global feature vector and local feature descriptor of the sparse image, use the depth information, camera intrinsic matrix and pose information corresponding to the image to map the local feature points to the three-dimensional space through the camera back-projection model, construct a feature database including global features, three-dimensional coordinates of key points and local features, and establish an inverted index structure to accelerate image retrieval.

[0032] S3: Based on the sparse perception perspective enhancement strategy proposed in the present invention, statistically analyze the visibility distribution of each point in the three-dimensional Gaussian model from the perspective of the training images, and calculate the sparsity index of each training perspective by combining the distribution density of the camera poses in the three-dimensional space and the overlap degree between perspectives, and generate a sorted list of sparsity based on this for subsequent enhanced modeling of sparse regions.

[0033] S4: Select the training perspectives with higher rankings from the sorted list of sparsity. According to other training images that partially overlap with the perspective coverage area and have significant spatial position differences, calculate and generate the camera poses of the pseudo-perspectives through linear interpolation of the camera positions and quaternion interpolation of the poses. Subsequently, use the training main model to render the color image and depth map from this pseudo-perspective, extract their global and local features and incorporate them into the feature database to enhance the feature expression ability of the sparse regions.

[0034] S5: Extract the global features and local features of the query image, calculate the similarity between the global features of the query image and the global feature set in the database, retrieve several reference images with the highest similarities as candidate perspectives, and then perform descriptor matching between the local features of the query image and the local features with bound three-dimensional coordinates in the candidate images to establish the initial correspondence between the two-dimensional key points and three-dimensional space points in the query image.

[0035] S6: Based on the dynamic inlier screening and iterative optimization strategy proposed in this patent, use the reprojection error and the three-dimensional Gaussian rendering consistency index, and adopt the PnP-RANSAC algorithm to optimize the initial two-dimensional to three-dimensional matching relationship. In each iteration, dynamically eliminate the mismatched points based on the above indexes, update the inlier set and optimize the pose estimation until the inlier set converges, and finally output the accurate camera pose of the query image.

[0036] The present invention first obtains sparse training images and their corresponding camera pose information, constructs a Gaussian point cloud model based on the three-dimensional Gaussian rendering technology, and extracts the global and local visual features of the images to construct a feature database. Subsequently, by analyzing the visibility distribution of the Gaussian points, identify the regions with insufficient perspective coverage in the training set, and generate the camera poses of the pseudo-perspectives by combining reference perspectives that are far apart in spatial distribution but have similar coverage areas, use the Gaussian model to render the pseudo-images and complete the feature database. In the positioning stage, after inputting the query image, extract its global features, perform image retrieval to obtain candidate reference perspectives, and then construct the initial 2D-3D correspondence through local feature matching. Finally, adopt the dynamic inlier screening and PnP-RANSAC optimization strategy guided by the reprojection error and rendering consistency to iteratively estimate and output the final camera pose.

[0037] In the construction of the feature database, first, the training images and their corresponding intrinsic and extrinsic camera information are obtained, and a Gaussian point cloud representation of the scene is constructed based on this image set and 3D Gaussian modeling technology. Each Gaussian point has attributes such as position, covariance matrix, color value, and opacity, which can be used for subsequent differentiable rendering processes. At the image level, the global visual features and local key point information of the training images are extracted respectively. Among them, the global features are extracted by the NetVLAD network for subsequent image retrieval, and the local features are used to extract 2D key points and their descriptors through the SuperPoint network. Combining the camera pose and the depth map obtained by rendering, the key points are back-projected into the 3D space to generate a set of local 3D features that bind 3D coordinates and local descriptors. After construction, the above global and local features are respectively organized into a feature database, where the local features are used for 2D-3D matching and the global features are used for image retrieval.

[0038] The sparse perception enhancement strategy aims to solve the problem of missing geometric information in local regions caused by uneven distribution of the viewing angles of training images. The present invention evaluates the viewing angle sparsity of each training image based on the visibility distribution information of the point cloud in the 3D Gaussian model. Specifically, for each 3D Gaussian point, the set of training images in which it is observed is counted, and then the average visibility score is calculated for each training image to form a sorted list of sparsity. The viewing angle with the highest sparsity is selected as the enhancement target. To expand the spatial information coverage of this sparse viewing angle, the present invention further selects training images that are relatively far away in 3D space position but have a high overlap in the observation area as reference viewing angles, and uses interpolation to generate the camera poses of pseudo-viewing angles. Under the generated pseudo-viewing angles, the color images and depth maps are rendered using the trained 3D Gaussian point cloud model, and their corresponding global and local visual features are extracted. Finally, the feature data items of this pseudo-viewing angle are added to the feature database. The above enhancement process can be iteratively performed until the visibility scores of all viewing angles reach a set threshold or the number of generated pseudo-viewing angles reaches a set upper limit, so as to achieve the structural complementation and expression enhancement of the original sparse viewing angle space.

[0039] The dynamic inlier filtering and pose solution process is mainly used to establish a stable correspondence between 2D key points and 3D spatial points after the query image is input, and improve the accuracy and robustness of camera pose estimation through dynamic optimization. First, the global features of the query image to be located are extracted and similarity-matched with the global features of the training images stored in the feature database, and several images with the highest retrieval scores are selected as candidate reference views. On this basis, the local key points and descriptors of the query image are further extracted and locally matched with the 3D local features bound in the candidate reference images to construct an initial two-dimensional to three-dimensional (2D-3D) matching set. Subsequently, the present invention introduces a dynamic inlier filtering mechanism, and uses an iterative PnP-RANSAC process to estimate the camera pose, and updates the current inlier set by combining the reprojection error and the rendering consistency error in each iteration. After each round of optimization, the view is rendered on the 3D Gaussian model using the estimated camera pose, and the local features in the new view are extracted to guide the generation of new 2D-3D matching points, thereby supplementing potential missing matches. The above process terminates after meeting the inlier set convergence condition or reaching the maximum number of iterations, and finally outputs the optimized camera pose result. This module effectively alleviates the matching failure problem caused by the initial pose error or occlusion interference, and significantly improves the positioning robustness and accuracy under sparse view conditions.

[0040] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or equivalently replace some or all of the technical features, and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A three-dimensional Gaussian visual positioning method for sparse perspectives, characterized in that, Including training a three-dimensional Gaussian sputtering model using RGB images and camera poses, generating color images and depth maps through two-dimensional projection and differentiable rendering mechanisms, extracting global image features and local image features of sparse images, and constructing a feature database; Performing a nearest neighbor search on the feature database and RGB images to obtain a sparsity index, calculating the camera pose of the pseudo-viewpoint, rendering color and depth maps at the pseudo-viewpoint using the trained main model, extracting global and local features and incorporating them into the feature database, constructing a 2D-3D correspondence relationship with the target color map, performing dynamic inlier screening and pose solution, and obtaining the positioning result.

2. The three-dimensional Gaussian visual positioning method for sparse viewpoints according to claim 1, wherein Training a three-dimensional Gaussian sputtering model includes using the RGB image as a sparse view image, using the same sparse view image as the input of the three-dimensional Gaussian sputtering model, and constructing the three-dimensional Gaussian sputtering model through a collaborative modeling method of a two-branch three-dimensional Gaussian point cloud model, including training a main model and an auxiliary supervision model , each three-dimensional Gaussian point cloud model includes multiple Gaussian points, and the Gaussian point is a quadruple , , , , are the spatial position, the covariance matrix of the spatial scale and direction of the Gaussian point, the color value, and the opacity respectively; Input the same sparse perspective images into the training main model and the auxiliary supervision model , construct a cross-model view consistency loss based on the rendered images under the same pseudo-perspective, and the loss function is as follows: ; ; In the formula, is the balance coefficient, and are respectively and rendered images under the same pseudo-viewing angle, represents the L1 norm, represents the structural similarity index; Color reconstruction loss of a single three-dimensional Gaussian point cloud model is as follows: ; In the formula, is the pixel color of the training image, is the corresponding pixel value of the rendered image, is the pixel , is the pixel color; Total loss function for dual-branch training is as follows: ; In the formula, is the weight coefficient, which is used to adjust the proportion of the cross-model view consistency loss in the total loss; Minimizing the total loss of double-branch training to complete the training of the three-dimensional Gaussian sputtering model.

3. A three-dimensional Gaussian vision positioning method for sparse viewpoints according to claim 2, characterized in that, Generating a color image and a depth map includes, from a camera perspective, projecting a three-dimensional covariance matrix into a two-dimensional covariance matrix according to a Jacobian transformation : ; where 𝐽 is the Jacobian matrix of the affine approximation of the projection function, is the view transformation matrix of the camera viewing direction; For each pixel in the image plane , according to the two-dimensional distances to all visible Gaussian points for sorting, a Gaussian set is constructed, and the color values are accumulated using the alpha-blending rendering method to generate , and is used as the color image: ; ; In the formula, is the cumulative transmittance of the Gaussian point , is the projected opacity of the Gaussian point at the pixel ; represents the -th Gaussian point represents the -th Gaussian point is the cumulative transmittance of the Gaussian point at the pixel ; Generate a depth map, the depth of a single pixel is as follows: ; In the formula, represents the Gaussian point The absolute depth value under the current perspective, and is used as the depth map.

4. A three-dimensional Gaussian visual positioning method for sparse perspectives according to claim 3, characterized in that Constructing a feature database includes constructing a global image feature set and a local image feature set; The global image feature set is: ; wherein, is the th global feature vector extracted by the image coding network, is the th rotation matrix in the camera pose, is the th translation vector in the camera pose, is the maximum value; The local image feature set is as follows: ; In the formula, is the three-dimensional space coordinates obtained by back-projecting the local key points of the image, is the local feature descriptor bound to.

5. A three-dimensional Gaussian visual positioning method for sparse perspectives according to claim 4, characterized in that Constructing a global image feature set includes extracting two-dimensional key point coordinates from sparse-view images using an image local feature extraction network and the corresponding , combining the camera intrinsic matrix corresponding to the sparse-view image, the extrinsic matrix , and the depth value of each key point in the depth map , mapping the two-dimensional key points to three-dimensional space, and calculating the corresponding : ; In the formula, is the camera rotation matrix in the external parameter matrix, is the translation vector in the external parameter matrix. Bind with the corresponding to form the local three-dimensional feature item and store it in the local image feature set.

6. The three-dimensional Gaussian vision positioning method for sparse perspectives according to claim 5, characterized in that Obtaining the sparsity index includes, for the Gaussian points in the three-dimensional Gaussian sputtering model , counting the set of visible viewpoints , which is equivalent to all the sparse viewpoint images that can observe . For the th sparse viewpoint image , let the set of Gaussian points observed be , the sparsity index of is: ; Two images and have an overlap degree of the intersection - union ratio of the visible Gaussian point sets of the two : ; In the formula, is the th sparse view image, is the observed Gaussian point set; Use the sparsity metric Construct a sparsity sorted list.

7. A three-dimensional Gaussian vision positioning method for sparse perspectives according to claim 6, characterized in that, Calculating the camera pose of the pseudo-viewpoint includes selecting the one with the lowest sparsity from the sorted list of sparsities as the target viewpoint to be enhanced, and screening out those from all sparse-viewpoint images that with greater than the set threshold to form an image set : ; To be the reference image with the farthest Euclidean distance in spatial position from the reference image, with a pose of where , is the translation vector and is the rotation angle represented by a unit quaternion; Let the camera pose be , and the pseudo-viewpoint pose generated by interpolation is defined as: ; ; ; ; In the formula, is the pseudo-viewpoint translation vector, is the pseudo-viewpoint rotation angle, is the interpolation factor, is the intermediate parameter.

8. A three-dimensional Gaussian vision positioning method for sparse viewpoints according to claim 7, characterized in that, Rendering the color image and depth map in the pseudo-view using the trained main model includes using the same method as for generating the color image and depth map and leveraging the trained main model at the pseudo-view camera pose to generate the corresponding color image and depth map ; Extract local key point coordinates from a color image and combine with a depth map and map the camera pose to three-dimensional coordinates to form a local three-dimensional feature term ; Extract the global features of a color image and bind them to the camera pose to form a global feature item, where is the pseudo-view camera rotation matrix; Adding local three-dimensional feature terms and global feature terms to the feature database: ; ; In the formula, , are respectively the updated , .

9. A three-dimensional Gaussian vision positioning method for sparse viewpoints according to claim 8, characterized in that, Constructing the 2D-3D correspondence in combination with the target color map includes calculating the global feature vector of the image to be processed and cosine similarity score : ; Sort in descending order and select the images with the highest scores. Query the local image feature set corresponding to the image to be processed from the feature database, and use the local feature matcher LightGlue to match the 2D key points of the image to be processed with the local feature descriptors bound with 3D coordinates in the reference image to construct a set of 2D-3D correspondence relationships : ​ ; Take as the initial set of interior points : 。 10. A three-dimensional Gaussian vision positioning method for sparse viewpoints according to claim 9, characterized in that Performing dynamic inlier screening and pose solution includes, in the th iteration, using the th inlier set as input, adopting the PnP-RANSAC algorithm to estimate the current camera pose matching pair , calculating the reprojection error for each matching pair : ; In the formula, represents the camera projection function; If is greater than the set threshold , the corresponding matching pair is removed from the inlier set to obtain the updated inlier set for the current round ; Obtain After that, based on the matching pairs , use the trained main model to render the color image and depth image in the corresponding perspective, extract local features and rematch them with the local features of the image to be processed, and construct rounds of enhanced inlier sets ; Merge with the current set of in - wheel points to obtain the set of in - wheel points: ; In the formula, is the set of interior points for the -th iteration; The iterative process terminates when any of the following conditions is met: (1) The intersection-union ratio of the inner point sets exceeds the set threshold in two consecutive rounds : ; (2) The number of iteration rounds exceeds the maximum set value ; Finally output the current optimal camera pose As the positioning result of the image to be processed.

Citation Information

Patent Citations

  • Visual positioning method of sparse three-dimensional point cloud chart based on VSLAM

    CN110889349A

  • Visual repositioning method and system based on 3D Gaussian scene and storage medium

    CN118941629A

  • Sparse visual angle 3D model construction method and device, equipment and storage medium

    CN119107404A

  • Image rendering method and device based on Gaussian splashing, equipment, storage medium and program product

    CN119295638A

  • Three-dimensional reconstruction method based on three-dimensional Gaussian sputtering technology and related equipment

    CN119478247A

Cited By

  • 3DGS visual repositioning method and system based on intrinsic image decomposition, terminal and storage medium

    CN121544713A

  • Visual precise positioning method based on pixel fusion alignment

    CN121640007A

  • Vector-line-guided three-dimensional Gaussian air-ground cross-view-angle vehicle positioning method and system

    CN121883600A

  • A three-dimensional gauss space-ground cross-view vehicle positioning method and system based on vector line guidance

    CN121883600B

  • Power grid equipment 3D Gaussian splash updating method using unknown view angle image

    CN122368319A