A method and system for enhanced visual positioning based on cross-domain three-dimensional gauss sputtering

By combining image editing models and 3D Gaussian sputtering models, an extended dataset is generated and fine-tuned, which solves the problem of insufficient robustness of visual positioning methods in cross-domain scenarios and achieves higher positioning accuracy and stability.

CN120388074BActive Publication Date: 2025-11-18BEIJING BIG DATA ADVANCED TECH RES INST
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510873576.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-11-18
Estimated Expiration
2045-06-27

AI Technical Summary

Technical Problem

Existing visual localization methods lack robustness in cross-domain scenarios due to long-tailed dataset distribution, device distortion, and changes in ambient lighting, making it difficult to adapt to complex lighting and color variations in sub-domain scenarios.

Method used

An extended dataset is generated by using an image editing model to convert images from the first image domain to images from the second image domain, thereby enriching the training data. The 3D Gaussian sputtering model is then fine-tuned to adapt to different environmental feature distributions, generating a localization image training set and optimizing the visual localization model.

Benefits of technology

It improves the stability and accuracy of the visual positioning model in cross-domain scenarios, reduces positioning errors caused by environmental differences, and enhances the model's adaptability and generalization ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388074B_ABST
    Figure CN120388074B_ABST
Patent Text Reader

Abstract

The application discloses a cross-domain three-dimensional Gaussian sputtering enhanced visual positioning method and system, which comprises the following steps: generating a second image corresponding to a first image in a first image field through an image editing model to obtain an extended data set of the first image data set in the first image field; performing sputtering model fine-tuning on a three-dimensional Gaussian sputtering model suitable for the feature distribution of the second image field according to the corresponding first image and second image in the extended data set; generating a positioning image training set of an image positioning model according to the three-dimensional Gaussian sputtering model after sputtering model fine-tuning, and performing positioning model training on the image positioning model according to the positioning image training set to obtain a trained image positioning model. The application makes the visual positioning system have stronger adaptability and generalization ability, reduces the positioning precision decline phenomenon, and meets the stability and high-precision requirements of cross-domain visual positioning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of image positioning, specifically relating to a method, system, storage medium, device, and computer program product based on cross-domain three-dimensional Gaussian sputtering enhanced visual positioning. Background Technology

[0002] In modern image localization tasks, high-precision visual localization technology is widely used in fields such as autonomous driving, robot navigation, intelligent monitoring, and augmented reality. These applications typically require stable and accurate localization on image data acquired from different devices.

[0003] Currently, visual localization methods based on absolute pose regression have gained widespread attention in cross-domain visual localization tasks due to their simple architecture, ease of deployment, and computational efficiency. Traditional methods mainly rely on a single pinhole camera model for training, achieving image localization within a single image domain.

[0004] However, due to the long-tailed distribution of the training dataset, existing methods struggle to adapt to complex lighting and tone variations in subdomain scenes such as nighttime images. Summary of the Invention

[0005] This application aims to provide a cross-domain 3D Gaussian sputtering-enhanced visual positioning method, system, storage medium, device, and computer program product, which at least solves the problem of insufficient adaptability of image positioning to sub-domain scenes.

[0006] In a first aspect, embodiments of this application disclose a method for enhancing visual localization based on cross-domain 3D Gaussian sputtering, including:

[0007] A second image, corresponding to a first image in a first image domain, is generated using an image editing model to obtain an extended dataset of the first image dataset in the first image domain.

[0008] Based on the first and second images corresponding to the extended dataset, the three-dimensional Gaussian sputtering model adapted to the feature distribution of the second image domain is fine-tuned so that the fine-tuned three-dimensional Gaussian sputtering model adapts to the feature distribution of the second image domain as well as the feature distribution of the first image domain.

[0009] Based on the three-dimensional Gaussian sputtering model that has been fine-tuned by the sputtering model, a training set of positioning images for the image positioning model is generated, and the image positioning model is trained based on the training set of positioning images to obtain the trained image positioning model.

[0010] Secondly, embodiments of this application also disclose a cross-domain three-dimensional Gaussian sputtering enhanced visual positioning system, comprising:

[0011] An extended dataset module is used to generate a second image in a second image domain corresponding to a first image in a first image domain, through an image editing model, so as to obtain an extended dataset of the first image dataset in the first image domain;

[0012] The sputtering fine-tuning module is used to fine-tune the sputtering model of the three-dimensional Gaussian sputtering model adapted to the feature distribution of the second image domain according to the first image and the second image corresponding to the extended dataset, so that the fine-tuned three-dimensional Gaussian sputtering model is adapted to the feature distribution of the second image domain as well as the feature distribution of the first image domain.

[0013] The positioning training module is used to generate a positioning image training set for the image positioning model based on the three-dimensional Gaussian sputtering model that has been fine-tuned by the sputtering model, and to train the image positioning model based on the positioning image training set to obtain the trained image positioning model.

[0014] Thirdly, embodiments of this application also disclose a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps described in the first or second aspect.

[0015] Fourthly, embodiments of this application also disclose an electronic device, including a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, performs the steps as described in the first or second aspect.

[0016] Fifthly, embodiments of this application also disclose a computer program product, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps described in the first or second aspect.

[0017] In summary, in this embodiment, an image editing model is used to convert images from the first image domain to images from the second image domain to expand the dataset, enrich the training data, and improve the adaptability of the visual positioning model to different environmental features. This reduces the robustness problem caused by the long-tail distribution of the dataset in traditional methods. Furthermore, the paired first and second images in the expanded dataset are used to fine-tune the 3D Gaussian sputtering model, making it adaptable to the feature distribution of both the first and second image domains. This optimizes the model's adaptability to changes in features such as illumination and hue in cross-domain images, enabling it to effectively capture visual information in multiple environments and improve the stability and accuracy of cross-domain visual positioning. The fine-tuned sputtering model is used to generate a positioning image training set for training the image positioning model, ensuring that the training data has consistent cross-domain feature distribution and providing good constraints on the model training process. This allows the model to better adapt to cross-domain scenarios in practical applications, improve visual positioning accuracy, and reduce positioning errors caused by device distortion or environmental changes. Finally, the visual positioning model is trained using the positioning image training set, enabling it to achieve stable positioning under various device and scenario conditions, reducing the impact of data scarcity on the model's generalization ability. Therefore, the method based on the embodiments of this application, through data expansion, model fine-tuning, and training optimization, enables the visual positioning system to possess stronger adaptability and generalization ability, while reducing the positioning accuracy decline caused by environmental differences in traditional methods, thereby meeting the practical requirements of stability and high accuracy for cross-domain visual positioning applications. It plays a crucial role in solving the problem of insufficient robustness of existing visual positioning methods due to long-tailed dataset distribution, device distortion, and cross-environmental lighting variations. Attached Figure Description

[0018] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of this application. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:

[0019] Figure 1 This is a flowchart illustrating the steps of a cross-domain 3D Gaussian sputtering-enhanced visual positioning method provided in an embodiment of this application.

[0020] Figure 2 This is a flowchart of another method for enhancing visual localization based on cross-domain 3D Gaussian sputtering, provided in an embodiment of this application.

[0021] Figure 3 This is a schematic diagram of a cross-domain 3D Gaussian sputtering enhanced visual positioning system provided in an embodiment of this application;

[0022] Figure 4This is a block diagram of an electronic device provided in an embodiment of this application. Detailed Implementation

[0023] Exemplary embodiments of the present application will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this application will be thorough and complete, and will fully convey the scope of the present application to those skilled in the art.

[0024] Figure 1 This embodiment provides a cross-domain 3D Gaussian sputtering-enhanced visual localization method, which specifically includes the following steps:

[0025] Step 101: Generate a second image in a second image domain corresponding to the first image in the first image domain using an image editing model, so as to obtain an extended dataset of the first image dataset in the first image domain.

[0026] In some embodiments of this application, a primary domain (second image domain) image is generated based on a secondary domain (first image domain) image to construct an extended dataset for visual localization optimization. Specifically, the secondary domain image can be transformed using an image editing model to generate image data that conforms to the feature distribution of the primary domain. The image editing model utilizes a pre-trained visual feature transformation mechanism, combined with illumination adjustment, tone matching, and dynamic detail restoration methods, to ensure that the transformed primary domain image matches the features of the real scene in terms of visual consistency. In this process, the generated extended dataset not only enhances the adaptability of the visual localization model to different environmental features but also reduces the robustness problem caused by the long-tailed distribution of the dataset in traditional methods, thereby improving the accuracy and stability of the cross-domain visual localization system.

[0027] In a specific example, this is applied to environmental adaptability optimization of autonomous driving systems to improve the vehicle's visual localization capabilities under different lighting and environmental conditions. First, nighttime (secondary domain) road scene images are acquired and converted into daytime (primary domain) road scene images using an image editing model. Simultaneously, illumination intensity, road reflectivity, and dynamic elements are adjusted to match the primary domain feature distribution. Subsequently, the converted primary domain image is combined with the original daytime scene image to construct an expanded dataset for optimizing the visual localization model. During subsequent training, the model can more accurately learn the visual features of the primary domain environment, thereby improving the robustness and localization accuracy of the autonomous driving system in complex environments.

[0028] Step 102: Fine-tune the sputtering model of the three-dimensional Gaussian sputtering model adapted to the feature distribution of the second image domain based on the first and second images in the extended dataset.

[0029] Among them, the finely tuned three-dimensional Gaussian sputtering model is adapted to the feature distribution of the second image domain, and also to the feature distribution of the first image domain.

[0030] In some embodiments of this application, an extended dataset is used to optimize the 3DGaussian Splatting (3DGS) model, making it adaptable in both the primary domain (second image domain) and the secondary domain (first image domain). Specifically, by adjusting the model parameters, it can stably render high-precision 3D scene images under different scene feature distributions. The extended dataset contains paired images of the primary and secondary domains. These images are used to fine-tune the sputtering model, enabling it to adapt not only to the primary domain feature distribution but also to effectively adjust to the secondary domain environment. During the sputtering model fine-tuning process, the model learns the lighting changes, tone shifts, and dynamic scene features between cross-domain images. This allows the fine-tuned model to generate stable, domain-consistent rendering results in cross-domain visual localization tasks, reducing cross-environment visual errors and improving localization accuracy.

[0031] In a specific example, cross-temporal visual positioning optimization was applied to an indoor navigation system to enhance its positioning stability under varying lighting conditions. Nighttime (subdomain) environmental images were acquired, and corresponding daytime (primary domain) images were generated. These paired images were then used to fine-tune a 3D Gaussian sputtering model. During fine-tuning, the model learned the lighting variation characteristics of different time periods, ensuring that the final rendered nighttime scene not only possessed the geometric accuracy of the primary domain but also reasonably reflected the lighting adjustment characteristics of the subdomain. The optimized model maintained high-precision visual positioning in nighttime environments, enabling the navigation system to retain high robustness even in complex lighting conditions.

[0032] Step 103: Based on the three-dimensional Gaussian sputtering model that has been fine-tuned by the sputtering model, generate a training set of positioning images for the image positioning model, and train the image positioning model based on the training set of positioning images to obtain the trained image positioning model.

[0033] In some embodiments of this application, a training set of localization images is generated using fine-tuned 3DGS, and the image localization model is optimized based on this training set to improve its cross-domain visual localization capability. Specifically, a series of images for the localization task can be rendered using a fine-tuned 3D Gaussian sputtering model, ensuring that the generated images cover the visual feature distribution of the primary domain (second image domain) and the secondary domain (first image domain). During this process, the model learns the illumination changes, tone matching, and dynamic disturbance control between cross-domain images, so that the final localization image training set can effectively reflect the visual features of real-world scenes under different times and device conditions. Subsequently, this training set is used to optimize the image localization model, enabling it to maintain stability and high accuracy in complex environments, thereby improving the robustness and generalization ability of cross-domain visual localization.

[0034] In a specific example, visual positioning optimization is applied to an intelligent monitoring system to ensure accurate target location identification even during day-night transitions or equipment switching. First, a training set of positioning images is generated based on a finely tuned 3D Gaussian sputtering model, including indoor environment images under different lighting conditions. Then, this training set is used to train the visual positioning model, enabling it to accurately learn the effects of lighting changes and equipment distortion in the scene. After training, the optimized visual positioning model can stably locate targets in monitoring scenarios across different time periods, while reducing errors caused by lighting or equipment switching, thereby improving the system's stability and reliability in cross-domain scenarios.

[0035] In summary, in this embodiment, an image editing model is used to convert images from the first image domain to images from the second image domain to expand the dataset, enrich the training data, and improve the adaptability of the visual positioning model to different environmental features. This reduces the robustness problem caused by the long-tail distribution of the dataset in traditional methods. Furthermore, the paired first and second images in the expanded dataset are used to fine-tune the 3D Gaussian sputtering model, making it adaptable to the feature distribution of both the first and second image domains. This optimizes the model's adaptability to changes in features such as illumination and hue in cross-domain images, enabling it to effectively capture visual information in multiple environments and improve the stability and accuracy of cross-domain visual positioning. The fine-tuned sputtering model is used to generate a positioning image training set for training the image positioning model, ensuring that the training data has consistent cross-domain feature distribution and providing good constraints on the model training process. This allows the model to better adapt to cross-domain scenarios in practical applications, improve visual positioning accuracy, and reduce positioning errors caused by device distortion or environmental changes. Finally, the visual positioning model is trained using the positioning image training set, enabling it to achieve stable positioning under various device and scenario conditions, reducing the impact of data scarcity on the model's generalization ability. Therefore, the method based on the embodiments of this application, through data expansion, model fine-tuning, and training optimization, enables the visual positioning system to possess stronger adaptability and generalization ability, while reducing the positioning accuracy decline caused by environmental differences in traditional methods, thereby meeting the practical requirements of stability and high accuracy for cross-domain visual positioning applications. It plays a crucial role in solving the problem of insufficient robustness of existing visual positioning methods due to long-tailed dataset distribution, device distortion, and cross-environmental lighting variations.

[0036] Figure 2 This application provides another method for enhancing visual localization based on cross-domain 3D Gaussian sputtering, which specifically includes the following steps:

[0037] Step 201: Train the initialized 3D Gaussian sputtering model using the sputtering image training set containing the second image in the second image domain.

[0038] Among them, the trained 3D Gaussian sputtering model outputs a rendered image with a higher image fidelity than the preset image fidelity threshold in the second image domain.

[0039] In some embodiments of this application, an initialized 3D Gaussian Splatting (3DGS) model will be trained to ensure that its feature adaptation capability in the primary domain (second image domain) meets preset requirements. Specifically, based on a sputtered image training set containing primary domain images, the model parameters can be adjusted to enable accurate rendering of high-precision 3D scene images. The sputtered image training set contains image data from multiple perspectives in the scene and incorporates photometric embedding techniques to maintain lighting consistency and improve rendering quality. During this process, the model gradually learns the lighting changes corresponding to each perspective, ensuring that the generated rendered images meet preset image fidelity threshold requirements after training, thereby optimizing the model's adaptability in cross-domain scenes.

[0040] In a specific example, data preprocessing is applied to autonomous driving scenarios to improve the accuracy of visual localization in the primary domain (e.g., daytime scenes). First, road environment images under different daytime lighting conditions are acquired, and these images are used to construct a sputtering image training set, including multiple viewpoints under different weather conditions. Then, an initialized 3D Gaussian sputtering model is used to train this training set, enabling it to learn the photometric variation patterns under different lighting conditions. After training, the model generates rendered images that meet the image fidelity requirements of the experimental environment and maintain lighting consistency.

[0041] Optionally, a color multilayer perceptron and a photometric compensation multilayer perceptron are established in the three-dimensional Gaussian sputtering model. Step 201 includes the following sub-steps:

[0042] Sub-step 2011 involves inputting the second images from the sputtering image training set into the three-dimensional Gaussian sputtering model to obtain the color vector of each second image from the color multilayer perceptron, the photometric compensation vector and dynamic uncertainty parameter value of each second image from the photometric compensation multilayer perceptron, and the output image of each second image from the sputtering image training set to the three-dimensional Gaussian sputtering model.

[0043] In some embodiments of this application, the color vector, photometric compensation vector, and dynamic uncertainty parameter values ​​of the main domain (second image neighborhood) image in the sputtering image training set are obtained to support subsequent optimization training of the sputtering model. Specifically, each main domain image in the sputtering image training set can be input into 3DGS to extract color vectors from a Color Multi-Layer Perceptron (Color-MLP), photometric compensation vectors and dynamic uncertainty parameter values ​​from a Photometric Compensation Multi-Layer Perceptron (PC-MLP), and the model output image corresponding to each main domain image in the sputtering image training set can be obtained. In this process, the Color Multi-Layer Perceptron is used to encode the color features of the image, so that the generated output image maintains high accuracy in terms of tone consistency, while the Photometric Compensation Multi-Layer Perceptron is used to model the lighting changes and dynamic interference factors of the scene, so that the subsequent sputtering model training can more accurately simulate different lighting environments. Finally, the obtained color vector, photometric compensation vector, and dynamic uncertainty parameter values ​​provide key support for the construction of the sputtering training loss function, ensuring that the 3D Gaussian sputtering model achieves the preset target in cross-domain image adaptability.

[0044] In a specific example, ambient lighting adaptation optimization is applied to indoor positioning scenarios to improve the accuracy of the visual positioning system under different lighting conditions. First, indoor environmental images under different lighting conditions during the day (main domain) are acquired and processed using a 3D Gaussian sputtering model. During processing, color vectors for each main domain image are extracted from a color multilayer perceptron to ensure color consistency in subsequent rendered images. Simultaneously, photometric compensation vectors are obtained from a photometric compensation multilayer perceptron to correct for lighting shifts under different lighting conditions. Furthermore, dynamic uncertainty parameter values ​​are acquired to adjust the model's response to disturbances from moving objects. After processing, the optimized data can be used to construct a sputtering training loss function, ensuring that the visual positioning model maintains stable positioning accuracy under various indoor lighting conditions.

[0045] Sub-step 2012: Determine the pixel color prediction image of each second image in the sputtered image training set based on the color vector and photometric compensation vector of each second image in the sputtered image training set, and determine the dynamic object uncertainty data of each second image in the sputtered image training set based on the dynamic uncertainty parameter value of each second image in the sputtered image training set.

[0046] In some embodiments of this application, pixel color prediction images are generated using the color vectors and photometric compensation vectors of the main domain (second image neighborhood) images in the sputtering image training set. Dynamic object uncertainty data is then calculated based on dynamic uncertainty parameter values ​​to improve the model's adaptability to cross-domain lighting and dynamic scene interference. Specifically, pixel-level color distribution can be predicted by analyzing the color vectors of each main domain image to ensure that the generated images conform to the main domain characteristics in terms of hue consistency. Simultaneously, the image brightness and shadow distribution are adjusted using photometric compensation vectors to optimize the lighting adaptation of the main domain images. Subsequently, dynamic object uncertainty data in the scene is calculated based on the dynamic uncertainty parameter values. This data is used to estimate the degree of influence of moving objects in the scene on the image's visual features. In this process, the pixel color prediction images ensure that the generated training data is consistent in visual feature distribution, while the dynamic object uncertainty data is used to optimize the model's robustness to dynamic scene changes, providing effective support for subsequent sputtering model training.

[0047] In a specific example, this is applied to vision optimization in autonomous driving systems to improve visual recognition accuracy in nighttime driving scenarios (sub-domain). First, road scene images under different lighting conditions during the day (main domain) are acquired, and the color vector and photometric compensation vector for each image are extracted. Then, the color vectors are used to calculate pixel color prediction images, ensuring that the generated images maintain visual consistency under different lighting conditions. Simultaneously, the photometric compensation vectors are used to adjust the shadow distribution in the scene to conform to the characteristics of the actual environment. Furthermore, dynamic uncertainty parameter values ​​are used to calculate dynamic object uncertainty data to identify the impact of lighting changes on moving objects in the scene. After data processing, the optimized training data can be used for sputtering model training, enabling the autonomous driving system to maintain stable positioning even in low-light environments.

[0048] Sub-step 2013: Based on all the obtained output images, pixel color prediction images, and dynamic object uncertainty data, construct a sputtering training loss function for the three-dimensional Gaussian sputtering model, and train the three-dimensional Gaussian sputtering model using the sputtering training loss function.

[0049] In some embodiments of this application, a sputtering training loss function is constructed using image data obtained from the sputtering image training set to optimize the adaptability of 3DGS in cross-domain visual scenes. Specifically, a loss function for training the sputtering model can be constructed by integrating all output images, pixel color prediction images, and dynamic object uncertainty data, enabling the model to simultaneously optimize its robustness to illumination, tone changes, and dynamic scene effects. The loss function includes a pixel color prediction error term to correct color deviations in the model's output images, and a dynamic object uncertainty constraint term to reduce visual errors caused by moving object interference. In this process, the loss function serves as an optimization objective, guiding model training to make it more generalizable in cross-domain image generation tasks and ensuring that rendered images meet high accuracy requirements in different scenes.

[0050] In a specific example, data training for an augmented reality navigation system was used to improve the stability of the visual positioning system under various environmental conditions. First, image data under different lighting conditions and dynamic scene changes were collected, and corresponding rendered images were generated using a 3D Gaussian sputtering model. Then, a sputtering training loss function was constructed, incorporating color errors in pixel color prediction images and uncertainty data of dynamic objects to optimize the model's adaptability to changes in time and viewpoint. After training, the optimized 3D Gaussian sputtering model can more accurately render high-precision images, enabling the augmented reality system to maintain stable visual recognition capabilities under complex lighting and dynamic environmental conditions.

[0051] Optionally, in order to construct a sputtering training loss function for the 3D Gaussian sputtering model based on all the obtained output images, pixel color prediction images, and dynamic object uncertainty data, sub-step 2013 includes the following sub-steps:

[0052] Sub-step 20131: Construct a systematic loss function for the 3D Gaussian sputtering model based on all output images, pixel color prediction images, and dynamic object uncertainty data; and construct a structural similarity loss function for the 3D Gaussian sputtering model based on all output images and pixel color prediction images.

[0053] In some embodiments of this application, a systematic loss function and a structural similarity loss function for 3DGS are constructed to optimize the rendering accuracy and stability of the model in cross-domain visual scenes. Specifically, a systematic loss function is established by analyzing all output images, pixel color prediction images, and dynamic object uncertainty data to measure the overall matching degree between the rendered image and the real scene. Simultaneously, a structural similarity loss function is constructed based on the output image and the pixel color prediction image to ensure consistency in color distribution of the rendered image. The systematic loss function is used to optimize the lighting and geometric consistency of cross-domain images, while the structural similarity loss function is used to ensure the color and texture stability of the image. In this process, the joint optimization of the two types of loss functions enables the 3D Gaussian sputtering model to adapt to different lighting conditions, viewpoint changes, and dynamic object interference, thereby improving the robustness of the visual positioning system.

[0054] In a specific example, rendering optimization is applied to an augmented reality navigation system to ensure the system can generate high-precision visual data under varying lighting and dynamic scene changes. First, multiple viewpoint images of the scene are collected, and this data is used to construct pixel color prediction images and dynamic object uncertainty data. Then, a systematic loss function is constructed based on the output images to optimize lighting consistency in the augmented reality environment, while a structural similarity loss function is constructed based on the pixel color prediction images to improve the visual stability of the rendered images. After training, the optimized 3D Gaussian sputtering model can generate high-quality visual data under different time and environmental conditions, thereby enhancing the interactive experience of the augmented reality system.

[0055] Sub-step 20132: The regularized dual mean of the systematic loss function and the structural similarity loss function is determined as the sputtering training loss function.

[0056] In some embodiments of this application, a sputtering training loss function is constructed using the regularized dual mean of a systematic loss function and a structural similarity loss function to optimize the adaptability of 3DGS in cross-domain vision tasks. Specifically, a systematic loss function is first determined, which measures the overall error of the model-rendered image in terms of illumination consistency, geometric matching, and control of dynamic object influences. Then, a structural similarity loss function is determined, which focuses on optimizing the tonal stability between the pixel color prediction image and the output image. To ensure the balance of the loss functions, this step uses the regularized dual mean method, dynamically adjusting the weight ratio of the two loss functions during the optimization process to ensure a highly stable model training objective under various ambient lighting and dynamic scenes. Finally, this sputtering training loss function serves as the optimization basis, enabling the 3D Gaussian sputtering model to have stronger adaptability and generalization ability in visual localization tasks.

[0057] In a specific example, visual optimization is applied to an intelligent surveillance system to ensure accurate scene feature recognition under varying lighting conditions. First, images from multiple time periods within the monitored environment are collected, and these data are used to calculate a systematic loss function and a structural similarity loss function. Then, a regularized dual mean method is employed to optimize the weight ratio of the two loss functions, allowing the model to dynamically adjust its lighting adaptability and color stability during training. After training, the optimized intelligent surveillance system maintains stable target recognition capabilities under complex environments such as day-night cycles and dynamic interference, while reducing visual errors caused by uneven lighting.

[0058] In some embodiments of this application, cross-domain 3DGS is constructed based on anchor points to distribute local 3D Gaussian (Scaffold-GS). Therefore, two learnable photometric embeddings are first constructed for each image in the training set. The photometric histogram of the encoded image is displayed using photometric embedding, and the spatial location of the photometric embedding is consistent with the spatial location of the training data. Two decoupled multilayer perceptrons are constructed. It can embed light intensity and anchor features from Scaffold-GS Together, they are input into a multilayer perceptron to predict the color vector of 3DGS. Photometric compensation vector and dynamic uncertain parameter values :

[0059] ,

[0060] The pixel color prediction map can then be calculated based on the 3DGS rasterization steps. Uncertainty data of dynamic objects :

[0061] ,

[0062] Where i and j are representative elements in a dataset of size N. This is the opacity channel parameter.

[0063] Therefore, we can use the main domain data to train a cross-domain 3DGS, learn the precise geometry of the scene and the appearance of the main domain, and then use the training loss L as:

[0064] ,

[0065] Where C is the true value of pixel color (output image), and M = M(x), The Structural Similarity Index Measure (SSIM) function is used for loss. This is the regularization parameter.

[0066] Step 202: Generate a prior prediction image corresponding to the second image in the second image domain using an image editing model, so as to obtain a prior dataset of the second image dataset in the second image domain.

[0067] In some embodiments of this application, prior prediction images are generated based on the main domain (second image domain) image to construct a prior dataset for subsequent visual localization model training. Specifically, the main domain image can be processed by an image editing model to generate a prior prediction image consistent with the features of the target scene. The image editing model utilizes learned patterns of illumination, tone, and dynamic feature changes to ensure that the generated prediction image maintains consistency with the actual main domain scene in terms of visual feature distribution. In this process, the prior prediction image is not only used for data expansion but also serves as a reference for further optimizing the image editing model to improve its ability to model the features of the main domain environment. Ultimately, the generated prior dataset effectively enriches the visual information in the main domain scene, providing more stable and accurate data support for model training in subsequent steps.

[0068] In a specific example, visual localization optimization is applied to urban road scenarios to improve the consistency of image features within the main domain environment. First, main domain road images are acquired at different time intervals, and corresponding prior prediction images are generated using an image editing model. These prior prediction images include adjustments for factors such as changes in illumination and the influence of dynamic objects. Then, the generated prior prediction images are combined with the original main domain images to construct a prior dataset for further optimization of the visual localization model. During the optimization process, the model can more accurately learn the visual features of the main domain environment, thereby improving the localization accuracy and stability of the autonomous driving system under different lighting conditions.

[0069] Step 203: Fine-tune the image editing model based on the corresponding second image and the prior prediction image in the prior dataset.

[0070] Among them, the dynamic denoising performance index of the fine-tuned image editing model meets the preset anti-dynamic interference threshold requirements.

[0071] In some embodiments of this application, a prior dataset is used to optimize the image editing model, enabling it to more stably handle dynamic interference in the primary domain (second image domain) image and ensure high visual consistency in the generated secondary domain (first image domain) image. Specifically, the parameters of the image editing model can be adjusted by comparing the feature differences between the primary domain image and the prior predicted image in the prior dataset to improve its denoising capability under different lighting conditions. During fine-tuning, the image editing model learns patterns of lighting, tone changes, and dynamic object interference, enabling the fine-tuned model to generate more accurate secondary domain images in complex scenes. Ultimately, the optimized image editing model, when performing image conversion tasks, can reduce visual noise caused by dynamic objects and meet the preset anti-dynamic interference threshold requirements, thereby improving the stability and accuracy of the visual positioning system.

[0072] In a specific example, data preprocessing is applied to an autonomous driving system to improve visual recognition capabilities in nighttime scenes (sub-domain). First, daytime road images are acquired in the main domain, and corresponding nighttime images are generated using an image editing model. Then, the original nighttime scene images are compared with the generated nighttime images to identify visual errors caused by dynamic object interference. The parameters of the image editing model are adjusted to reduce the influence of unstable light sources and moving objects during subsequent image conversion. After fine-tuning, the optimized image editing model generates nighttime images that not only conform to the characteristics of the actual scene but also effectively reduce dynamic object interference, improving the robustness and reliability of the autonomous driving system in complex environments.

[0073] Step 204: Generate a second image in a second image domain corresponding to the first image in the first image domain using an image editing model, so as to obtain an extended dataset of the first image dataset in the first image domain.

[0074] The method shown in this step has been explained in step 101 and will not be repeated here.

[0075] Optionally, step 204 includes the following sub-steps:

[0076] Sub-step 2041 involves inputting the pose channel data of each first image in the first image dataset into the three-dimensional Gaussian sputtering model as input data to obtain a rendered image in the first image neighborhood corresponding to each first image.

[0077] In some embodiments of this application, rendered images are generated using subdomain (first image domain) image data to ensure that the generated images visually conform to the feature distribution of the subdomain and provide high-quality input for subsequent cross-domain image conversion. Specifically, the pose channel data of the subdomain images can be extracted and input into 3DGS. This model is then used for 3D reconstruction and rendering to generate rendered images corresponding to each subdomain image. The 3D Gaussian sputtering model uses point cloud representation, which includes color, photometric embedding, and dynamic confidence parameters to ensure that the output rendered images maintain stability in terms of lighting consistency and structure restoration. In this process, the generated rendered images can accurately capture the photometric features in the subdomain scene and provide a complete spatial representation, providing high-quality cross-domain data support for subsequent image conversion steps.

[0078] In a specific example, cross-environment visual modeling is applied to intelligent monitoring systems to improve their scene adaptability under low-light conditions. First, multiple nighttime (sub-domain) indoor environmental images are acquired, and the corresponding pose channel data is extracted. Then, this data is input into a 3D Gaussian sputtering model to generate corresponding rendered images, ensuring that the rendered images maintain visual characteristics consistent with the actual sub-domain scene. Furthermore, the illumination characteristics of the generated images are adjusted by photometric embedding parameters to accurately reflect the illumination changes in the nighttime scene. After processing, the generated rendered images can be used for subsequent image transformation tasks, enabling the intelligent monitoring system to more effectively adapt to scene changes across illumination conditions and improve its recognition stability in low-light environments.

[0079] Sub-step 2042 involves inputting the transformation instructions for the second image domain and each rendered image into the image editing model to obtain a second image corresponding to each first image in the first image dataset.

[0080] In some embodiments of this application, rendered images and conversion instructions are used to convert images from a secondary domain (first image domain) to a primary domain (second image domain) to construct domain-consistent visual data. Specifically, the generated rendered image and conversion instructions for the primary domain can be input into an image editing model to generate a primary domain image corresponding to the secondary domain image. The image editing model utilizes a pre-trained visual feature conversion mechanism, combined with illumination adjustment, tone matching, and dynamic detail restoration methods, to ensure that the output primary domain image not only conforms to the primary domain features in terms of color and illumination characteristics but also reduces visual errors caused by scene transitions. In this process, the converted primary domain image further enriches the cross-domain training data, enabling the visual localization model to more accurately adapt to different environmental scenes and improve the generalization ability of cross-domain visual localization.

[0081] In a specific example, image transformation is applied to an autonomous driving system to optimize the vehicle's visual recognition capabilities under day-night changing conditions. First, nighttime (secondary domain) road scenes are acquired, and a corresponding rendered image is generated using a 3D Gaussian sputtering model. Then, a day-night transformation command is input, and the rendered image is fed into an image editing model to generate a transformed image matching the daytime (primary domain) scene. After transformation, the image's illumination consistency and environmental feature matching are verified, and the transformed image data is used to train a visual localization model to improve the vehicle's localization accuracy in different time-related scenarios.

[0082] Step 205: Fine-tune the sputtering model of the three-dimensional Gaussian sputtering model adapted to the feature distribution of the second image domain based on the first and second images in the extended dataset.

[0083] Among them, the finely tuned three-dimensional Gaussian sputtering model is adapted to the feature distribution of the second image domain, and also to the feature distribution of the first image domain.

[0084] The method shown in this step has been explained in step 102 and will not be repeated here.

[0085] Step 206: Based on the three-dimensional Gaussian sputtering model that has been fine-tuned by the sputtering model, generate a training set of positioning images for the image positioning model, and train the image positioning model based on the training set of positioning images to obtain the trained image positioning model.

[0086] The method shown in this step has been explained in step 103 and will not be repeated here.

[0087] Optionally, the 3D Gaussian sputtering model includes a photometric embedding data layer and a pose embedding data layer. Step 206 includes the following sub-steps:

[0088] Sub-step 2061: Using each photometric embedding in the photometric embedding data layer as an anchor point, construct a three-dimensional virtual sphere according to a preset sampling interval, and randomly obtain a preset number of image generation poses in each three-dimensional virtual sphere, so that each image generation pose and the photometric embedding corresponding to the image generation pose are associated by a neural network according to a preset association function, while forming a statistical representation according to a preset distribution function.

[0089] In some embodiments of this application, a three-dimensional virtual sphere is constructed using photometric embeddings in the photometric embedding data layer, and image generation poses are generated for image localization training to optimize the adaptability of 3DGS in cross-domain visual scenes. Specifically, a three-dimensional virtual sphere can be constructed using photometric embeddings as anchor points and according to a preset sampling interval to ensure uniform distribution of illumination features in space. Subsequently, a preset number of image generation poses are obtained in each three-dimensional virtual sphere through random sampling, enabling it to effectively cover the range of illumination changes in the scene. During this process, each image generation pose is associated with the photometric embedding through a neural network, and a statistical representation is formed according to a preset distribution function to ensure that illumination change patterns can be accurately modeled. Ultimately, this improves the adaptability of the visual localization model to illumination changes, resulting in higher illumination consistency in the generated training data, which helps to improve the stability of cross-domain visual localization.

[0090] In a specific example, cross-illumination visual localization optimization is applied to urban road scenarios to improve the adaptability of autonomous driving systems under different time and illumination conditions. First, photometric embedding data is collected in daytime (primary domain) and nighttime (secondary domain) road scenes, and multiple 3D virtual spheres are constructed according to a preset sampling interval to ensure that illumination information covers the entire scene. Then, several images are randomly sampled within each sphere to generate poses, and a neural network is used to establish the correlation between photometric embeddings and image-generated poses to ensure that illumination features in the scene are accurately captured. After data processing, the generated illumination statistical representation is used to optimize the visual localization model of the autonomous driving system, enabling it to maintain stable localization performance during day-night transitions.

[0091] Sub-step 2062 involves inputting the image generation pose and photometric embedding corresponding to each original image from the original image dataset containing the first original image in the first image domain and / or the second original image in the second image domain into the three-dimensional Gaussian sputtering model that has been fine-tuned by the sputtering model, so as to form a localization image training set by using all the rendered images corresponding to each original image.

[0092] In some embodiments of this application, a finely tuned 3DGS is used to render the original image dataset to construct a training set of localization images for optimizing the image localization model. Specifically, the image generation pose and photometric embedding, which contain a subdomain (first image domain) and a main domain (second image domain) corresponding to the original image, are respectively input into a 3D Gaussian sputtering model, which generates corresponding rendered images based on the model's photometric embedding and pose embedding. This model uses a photometric embedding data layer to adjust the lighting characteristics of the image, ensuring that the generated rendered image maintains visual consistency under different lighting conditions, while the pose embedding data layer ensures that the spatial relationships from different viewpoints remain stable. In this process, the generated rendered image not only covers the feature distributions of the main and subdomains but also provides high-quality data support for the subsequent training of the image localization model, thereby optimizing the accuracy and robustness of cross-domain visual localization.

[0093] In a specific example, this is applied to visual localization training for an autonomous driving system to improve the vehicle's recognition capabilities under different lighting conditions. First, daytime (primary domain) and nighttime (secondary domain) road scene images are acquired. The images contained in these images are used to generate pose and photometric embeddings, which are then input into a finely tuned 3D Gaussian sputtering model for rendering. The model uses a photometric embedding data layer to adjust scene lighting, ensuring that the nighttime rendered images match the actual environment in terms of lighting consistency. Simultaneously, a pose embedding data layer ensures spatial consistency across different viewpoints. After processing, all rendered images are integrated into a localization image training set and used to train the visual localization model of the autonomous driving system, thereby enhancing the system's localization stability in complex lighting environments.

[0094] Sub-step 2063 generates a theoretical positioning pose corresponding to each rendered image based on each rendered image in the positioning image training set, and inputs each rendered image in the positioning image training set into the image positioning model to obtain an output positioning pose corresponding to each rendered image.

[0095] In some embodiments of this application, the theoretical positioning poses corresponding to the rendered images in the positioning image training set are generated and input into the image positioning model to obtain the corresponding output positioning poses, thereby optimizing the training process of the visual positioning system. Specifically, based on the pose information of the rendered images, a theoretical positioning pose corresponding to each rendered image can be generated through theoretical modeling methods to ensure that the model has accurate reference data during training. Subsequently, all rendered images are input into the image positioning model to learn the visual features of different viewpoints and lighting conditions in the scene, and output positioning results matching each rendered image. In this process, the theoretical positioning pose, as an idealized positioning reference, helps guide the model to optimize prediction accuracy, while the output positioning pose of the image positioning model is used to evaluate the model's learning ability, thereby improving the stability and robustness of the visual positioning task.

[0096] In a specific example, visual localization training is applied to a robot navigation system to ensure accurate positioning under varying lighting and viewing angles. First, the theoretical localization pose is calculated based on rendered images from the training dataset, and a pose calculation model is used to ensure the match between the theoretical pose and the actual environment. Then, the rendered images are input into the image localization model, allowing it to learn different scene features and generate corresponding output localization poses. After processing, the differences between the theoretical and output localization poses are compared, and this data is used to optimize the parameters of the localization model, thereby improving the robot's autonomous localization capabilities in complex environments.

[0097] Optionally, in order to generate a theoretical localization pose corresponding to each rendered image based on each rendered image in the localization image training set, sub-step 2063 includes the following sub-steps:

[0098] Sub-step 20631 transforms the panoramic image component in the rendered image to linear coordinate space to obtain the spatial coordinate component of the theoretical positioning pose, and transforms the image coordinate component in the rendered image to the global image coordinate space to obtain the position attention component of the theoretical positioning pose.

[0099] In some embodiments of this application, the panoramic image components in the rendered image are transformed to a linear coordinate space to ensure that the theoretical positioning pose maintains accurate spatial consistency across different viewpoints. Furthermore, the image coordinate components are transformed to a global image coordinate space to optimize positioning accuracy. Specifically, a linear coordinate transformation is first performed on the panoramic image components of the rendered image, enabling the spatial coordinate components to accurately represent the three-dimensional geometric relationships of the image in a standardized coordinate system. Subsequently, the image coordinate components are mapped to the global image coordinate space to construct the position attention component of the theoretical positioning pose, ensuring that image features from different viewpoints are correctly correlated. During this process, the spatial coordinate components are used to optimize the geometric consistency of the positioning system, while the position attention component improves the model's robustness to dynamically changing scenes through cross-domain visual feature alignment, ultimately enhancing the stability and accuracy of the visual positioning task. In some embodiments of this application, since the currently generated image is only a 360-degree format image, we need to map it to obtain cross-camera domain images. To save memory, we use an online mapping method when training the visual positioning method: i.e., cross-domain images (spatial coordinate components). It can be done through formula Find out, where It is a panoramic image component. The back projection function of a panoramic camera is represented. It is a camera model The projection function of (fisheye camera, pinhole camera, etc.); and in order to construct a positional attention mechanism to guide the alignment of appearance features: a global image coordinate space can first be constructed based on the panoramic image. This space can cover the field of view of any camera, and then the image coordinates from images from different cameras are... Transform to global image coordinate space As positional attention: ,in The projection function of a panoramic camera is represented. It is a camera model Back projection function of (fisheye camera, pinhole camera, etc.).

[0100] In a specific example, visual scene alignment is applied to a robot navigation system to enhance the robot's localization consistency across different scenarios. First, multi-view panoramic images of the indoor environment are acquired, and corresponding rendered images are generated using a 3D Gaussian sputtering model. Then, linear coordinate transformations are performed on the panoramic image components to ensure that the spatial coordinate components accurately reflect the environmental geometry. Simultaneously, the image coordinate components are transformed to the global image coordinate space to construct position attention components. After processing, the optimized theoretical localization pose is used to train the robot navigation system, enabling it to maintain stable visual recognition capabilities under different viewpoints and ambient lighting conditions, thereby improving the accuracy and robustness of navigation tasks.

[0101] Sub-step 20632 involves concatenating the spatial coordinate components with the position attention components to obtain the theoretical positioning pose.

[0102] In some embodiments of this application, spatial coordinate components and position attention components are concatenated to generate a theoretical localization pose, thereby ensuring that the visual localization system can accurately learn spatial relationships and visual features in cross-domain scenes. Specifically, spatial coordinate components can first be used to describe the spatial geometric relationships of the rendered image in a standardized coordinate system to ensure that image features remain consistent across different viewpoints. Subsequently, the position attention components are used to model the pixel weight distribution in the global image coordinate space, enabling the localization system to optimize and adjust for different camera viewpoints. In this process, spatial coordinate components ensure the geometric consistency of the visual localization system, while position attention components optimize the stability of visual features under different device or lighting conditions. Ultimately, the concatenated theoretical localization pose not only maintains high accuracy in geometric structure but also improves the model's adaptability to changes in visual features, thereby enhancing the robustness of the visual localization task.

[0103] In a specific example, scene matching is applied to an indoor navigation system to optimize the robot's localization accuracy under different viewing conditions. First, multiple viewing angle images of the indoor environment are acquired, and corresponding rendered images are generated using a 3D Gaussian sputtering model. Then, spatial coordinate components are extracted to ensure consistent geometric relationships across different viewpoints, and position attention components are calculated from the global image coordinate space to optimize the alignment process of visual features. Finally, these two components are concatenated to generate a theoretical localization pose, which is used to train the navigation system, enabling it to maintain stable localization capabilities under different locations and device conditions, thus improving its adaptability to complex environments.

[0104] Sub-step 2064: Based on the one-to-one correspondence between the theoretical positioning pose and the output positioning pose, train the image positioning model to obtain the trained image positioning model.

[0105] In some embodiments of this application, the image localization model is trained using both the theoretical localization pose and the output localization pose to optimize its localization accuracy and stability in cross-domain visual scenes. Specifically, the theoretical localization pose can be used as a standard reference, compared with the output localization pose, and the error between the two can be used to construct an optimization target to adjust the model parameters. During this process, the localization model learns the influence of different lighting conditions, viewpoint changes, and dynamic environmental factors on visual features, and gradually optimizes its adaptability in cross-domain scenes. This step also incorporates a loss calculation method based on a positional attention mechanism, enabling the model to adaptively correct visual deviations caused by camera distortion or lighting changes in cross-domain data, thereby further improving the accuracy and robustness of visual localization.

[0106] In a specific example, this is applied to visual localization optimization in autonomous driving systems to enhance vehicle stability in complex environments. First, a training set of localization images is constructed based on nighttime (secondary domain) and daytime (primary domain) images, generating corresponding theoretical localization poses. Then, all rendered images are input into the image localization model, and the output localization images are recorded. By calculating the error between the theoretical and output localization poses, the parameters of the localization model are adjusted to better adapt to changes in visual features under different lighting conditions and dynamic environments. After training, the optimized visual localization model maintains high localization accuracy under day-night transitions or environmental changes, improving the robustness of the autonomous driving system.

[0107] In summary, in this embodiment, an image editing model is used to convert images from the first image domain to images from the second image domain to expand the dataset, enrich the training data, and improve the adaptability of the visual positioning model to different environmental features. This reduces the robustness problem caused by the long-tail distribution of the dataset in traditional methods. Furthermore, the paired first and second images in the expanded dataset are used to fine-tune the 3D Gaussian sputtering model, making it adaptable to the feature distribution of both the first and second image domains. This optimizes the model's adaptability to changes in features such as illumination and hue in cross-domain images, enabling it to effectively capture visual information in multiple environments and improve the stability and accuracy of cross-domain visual positioning. The fine-tuned sputtering model is used to generate a positioning image training set for training the image positioning model, ensuring that the training data has consistent cross-domain feature distribution and providing good constraints on the model training process. This allows the model to better adapt to cross-domain scenarios in practical applications, improve visual positioning accuracy, and reduce positioning errors caused by device distortion or environmental changes. Finally, the visual positioning model is trained using the positioning image training set, enabling it to achieve stable positioning under various device and scenario conditions, reducing the impact of data scarcity on the model's generalization ability. Therefore, the method based on the embodiments of this application, through data expansion, model fine-tuning, and training optimization, enables the visual positioning system to possess stronger adaptability and generalization ability, while reducing the positioning accuracy decline caused by environmental differences in traditional methods, thereby meeting the practical requirements of stability and high accuracy for cross-domain visual positioning applications. It plays a crucial role in solving the problem of insufficient robustness of existing visual positioning methods due to long-tailed dataset distribution, device distortion, and cross-environmental lighting variations.

[0108] Table 1 shows a comparison of the errors between this application and the prior art:

[0109] Specifically, in the comparison, this application set up a testing environment on Ubuntu 20.04, equipped with an Intel Gold 6330 series CPU with a processing frequency of 2.00 GHz, and an NVIDIA GTX 4090 graphics processor with a core frequency of 2235 MHz and a video memory capacity of 24 GB. The errors were then compared with those of the panoramic visual positioning dataset (360Loc). For example, 6.5 / 39.0 in the table means that in this positioning result, the distance error is 6.5 meters and the angle error is 39.0°.

[0110] Table 1 - Comparison of Errors Between This Application and Prior Art

[0111]

[0112] As can be seen from Table 1, the positioning results obtained by the method according to the embodiments of this application show significant improvement in both distance error and angle error.

[0113] like Figure 3 As shown in the embodiments, this application also discloses a cross-domain three-dimensional Gaussian sputtering enhanced visual positioning system 30, including:

[0114] The extended dataset module 301 is used to generate a second image in a second image domain corresponding to the first image in the first image domain through an image editing model, so as to obtain an extended dataset of the first image dataset in the first image domain;

[0115] The sputtering fine-tuning module 302 is used to fine-tune the sputtering model of the three-dimensional Gaussian sputtering model adapted to the feature distribution of the second image domain based on the first and second images in the extended dataset, so that the fine-tuned three-dimensional Gaussian sputtering model is adapted to the feature distribution of the second image domain as well as the feature distribution of the first image domain.

[0116] The positioning training module 303 is used to generate a positioning image training set for the image positioning model based on the three-dimensional Gaussian sputtering model that has been fine-tuned by the sputtering model, and to train the image positioning model based on the positioning image training set to obtain the trained image positioning model.

[0117] Optionally, the cross-domain 3D Gaussian sputtering-enhanced visual positioning system 30 also includes:

[0118] The sputtering training module is used to train the initialized 3D Gaussian sputtering model based on the sputtering image training set containing the second image in the second image neighborhood, so that the image fidelity of the rendered image output by the trained 3D Gaussian sputtering model in the second image neighborhood is higher than the preset image fidelity threshold requirement.

[0119] Optionally, a color multilayer perceptron and a photometric compensation multilayer perceptron are established in the 3D Gaussian sputtering model. The sputtering training module includes:

[0120] The data preparation submodule is used to input the second images in the sputtering image training set into the three-dimensional Gaussian sputtering model respectively, so as to obtain the color vector of each second image in the sputtering image training set from the color multilayer perceptron, and obtain the photometric compensation vector and dynamic uncertainty parameter value of each second image in the sputtering image training set from the photometric compensation multilayer perceptron, and obtain the output image of each second image in the sputtering image training set to the three-dimensional Gaussian sputtering model.

[0121] The calculation submodule is used to determine the pixel color prediction image of each second image in the sputtered image training set based on the color vector and photometric compensation vector of each second image in the sputtered image training set, and to determine the dynamic object uncertainty data of each second image in the sputtered image training set based on the dynamic uncertainty parameter value of each second image in the sputtered image training set.

[0122] The training submodule is used to construct a sputtering training loss function for the 3D Gaussian sputtering model based on all the obtained output images, pixel color prediction images, and dynamic object uncertainty data, and to train the 3D Gaussian sputtering model using the sputtering training loss function.

[0123] Optionally, the training submodule includes:

[0124] The loss component unit is used to construct a systematic loss function for the 3D Gaussian sputtering model based on all output images, pixel color prediction images, and dynamic object uncertainty data, and to construct a structural similarity loss function for the 3D Gaussian sputtering model based on all output images and pixel color prediction images.

[0125] The loss unit is used to determine the sputtering training loss function by the regularized dual mean of the systematic loss function and the structural similarity loss function.

[0126] Optionally, the 3D Gaussian sputtering model is trained on a training set of sputtered images that include the second image in the second image's neighborhood. The extended dataset module 301 includes:

[0127] An extended rendering submodule is used to input the pose channel data of each first image in the first image dataset into the three-dimensional Gaussian sputtering model as input data to obtain a rendered image in the first image neighborhood corresponding to each first image.

[0128] An extended transformation submodule is used to input transformation instructions for the second image domain, as well as each rendered image, into the image editing model to obtain a second image corresponding to each first image in the first image dataset.

[0129] Optionally, the cross-domain 3D Gaussian sputtering-enhanced visual positioning system 30 also includes:

[0130] The prior dataset module is used to generate prior prediction images corresponding to the second images in the second image domain through the image editing model, so as to obtain the prior dataset of the second image dataset in the second image domain;

[0131] The interference fine-tuning module is used to fine-tune the image editing model based on the corresponding second image and the prior prediction image in the prior dataset, so that the dynamic denoising performance index of the fine-tuned image editing model meets the preset anti-dynamic interference threshold requirements.

[0132] Optionally, the 3D Gaussian sputtering model includes a photometric embedding data layer and a pose embedding data layer, and the localization training module 303 includes:

[0133] The rendering deployment submodule is used to construct a three-dimensional virtual sphere with each photometric embedding in the photometric embedding data layer as an anchor point according to a preset sampling interval, so as to randomly obtain a preset number of image generation poses in each three-dimensional virtual sphere, so that each image generation pose and the photometric embedding corresponding to the image generation pose are associated by a neural network according to a preset association function, while forming a statistical representation according to a preset distribution function.

[0134] The original training set submodule is used to input the image generation pose and photometric embedding corresponding to each original image from the original image dataset containing the first original image in the first image domain and / or the second original image in the second image domain into the three-dimensional Gaussian sputtering model that has been fine-tuned by the sputtering model, so as to form a localization image training set by all the rendered images corresponding to each original image.

[0135] The data collection submodule is used to generate a theoretical localization pose corresponding to each rendered image based on each rendered image in the localization image training set, and input each rendered image in the localization image training set into the image localization model to obtain the output localization pose corresponding to each rendered image.

[0136] The localization training submodule is used to train the image localization model based on the one-to-one corresponding theoretical localization pose and output localization pose, so as to obtain the trained image localization model.

[0137] Optionally, the data collection submodule includes:

[0138] The transformation unit is used to transform the panoramic image components in the rendered image to the linear coordinate space to obtain the spatial coordinate components of the theoretical positioning pose, and to transform the image coordinate components in the rendered image to the global image coordinate space to obtain the position attention components of the theoretical positioning pose.

[0139] The stitching unit is used to stitch together spatial coordinate components and position attention components to obtain the theoretical positioning pose.

[0140] In summary, in this embodiment, an image editing model is used to convert images from the first image domain to images from the second image domain to expand the dataset, enrich the training data, and improve the adaptability of the visual positioning model to different environmental features. This reduces the robustness problem caused by the long-tail distribution of the dataset in traditional methods. Furthermore, the paired first and second images in the expanded dataset are used to fine-tune the 3D Gaussian sputtering model, making it adaptable to the feature distribution of both the first and second image domains. This optimizes the model's adaptability to changes in features such as illumination and hue in cross-domain images, enabling it to effectively capture visual information in multiple environments and improve the stability and accuracy of cross-domain visual positioning. The fine-tuned sputtering model is used to generate a positioning image training set for training the image positioning model, ensuring that the training data has consistent cross-domain feature distribution and providing good constraints on the model training process. This allows the model to better adapt to cross-domain scenarios in practical applications, improve visual positioning accuracy, and reduce positioning errors caused by device distortion or environmental changes. Finally, the visual positioning model is trained using the positioning image training set, enabling it to achieve stable positioning under various device and scenario conditions, reducing the impact of data scarcity on the model's generalization ability. Therefore, the method based on the embodiments of this application, through data expansion, model fine-tuning, and training optimization, enables the visual positioning system to possess stronger adaptability and generalization ability, while reducing the positioning accuracy decline caused by environmental differences in traditional methods, thereby meeting the practical requirements of stability and high accuracy for cross-domain visual positioning applications. It plays a crucial role in solving the problem of insufficient robustness of existing visual positioning methods due to long-tailed dataset distribution, device distortion, and cross-environmental lighting variations.

[0141] This application also provides a computer-readable storage medium storing a computer program. When executed by a processor, this computer program implements the various processes described in the above embodiments of the cross-domain 3D Gaussian sputtering-enhanced visual positioning method, achieving the same technical effects. To avoid repetition, it will not be described again here. The computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.

[0142] Figure 4 This is a block diagram of an electronic device 700 provided in an embodiment of this application. For example, the electronic device 700 may be a mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, medical device, fitness equipment, personal digital assistant, etc.

[0143] Reference Figure 4The electronic device 700 may include one or more of the following components: a processing component 702, a memory 704, a power supply component 706, a multimedia component 708, an audio component 710, an input / output (I / O) interface 712, a sensor component 714, and a communication component 716.

[0144] Processing component 702 typically controls the overall operation of electronic device 700, such as operations associated with display, telephone calls, data communication, camera operation, and recording operations. Processing component 702 may include one or more processors 720 to execute instructions to complete all or part of the steps described above in the cross-domain 3D Gaussian sputtering-enhanced visual positioning method. Furthermore, processing component 702 may include one or more modules to facilitate interaction between processing component 702 and other components. For example, processing component 702 may include a multimedia module to facilitate interaction between multimedia component 708 and processing component 702.

[0145] Memory 704 is used to store various types of data to support the operation of electronic device 700. Examples of this data include instructions for any application or method operating on electronic device 700, contact data, phonebook data, messages, pictures, multimedia, etc. Memory 704 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0146] Power supply component 706 provides power to various components of electronic device 700. Power supply component 706 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to electronic device 700.

[0147] Multimedia component 708 includes a screen that provides an output interface between electronic device 700 and user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may not only sense the boundaries of touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 708 includes a front-facing camera and / or a rear-facing camera. When electronic device 700 is in an operating mode, such as shooting mode or multimedia mode, the front-facing camera and / or rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.

[0148] Audio component 710 is used to output and / or input audio signals. For example, audio component 710 includes a microphone (MIC) used to receive external audio signals when electronic device 700 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 704 or transmitted via communication component 716. In some embodiments, audio component 710 also includes a speaker for outputting audio signals.

[0149] I / O interface 712 provides an interface between processing component 702 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.

[0150] Sensor assembly 714 includes one or more sensors for providing state assessments of various aspects of electronic device 700. For example, sensor assembly 714 may detect the on / off state of electronic device 700, the relative positioning of components such as the display and keypad of electronic device 700, changes in position of electronic device 700 or a component of electronic device 700, the presence or absence of user contact with electronic device 700, orientation or acceleration / deceleration of electronic device 700, and temperature changes of electronic device 700. Sensor assembly 714 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 714 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 714 may also include an accelerometer, gyroscope, magnetometer, pressure sensor, or temperature sensor.

[0151] Communication component 716 facilitates wired or wireless communication between electronic device 700 and other devices. Electronic device 700 can access wireless networks based on communication standards, such as WiFi, carrier networks (such as 2G, 3G, 4G, or 7G), or combinations thereof. In one exemplary embodiment, communication component 716 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 716 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0152] In an exemplary embodiment, the electronic device 700 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to implement the cross-domain three-dimensional Gaussian sputtering-enhanced visual positioning method provided in the embodiments of this application.

[0153] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 704 including instructions, which can be executed by a processor 720 of an electronic device 700 to complete the above-described cross-domain three-dimensional Gaussian sputtering-enhanced visual positioning method. For example, the non-transitory storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0154] In an exemplary embodiment, the electronic device 700 may also be provided as a server. The electronic device 700 includes a processing component 702, which further includes one or more processors 720, and memory resources represented by memory 704 for storing instructions executable by the processing component 702, such as application programs. The application programs stored in memory 704 may include one or more modules, each corresponding to a set of instructions. Furthermore, the processing component 702 is configured to execute instructions to perform the cross-domain 3D Gaussian sputtering-enhanced visual positioning method provided in this application embodiment.

[0155] Electronic device 700 may also include a power supply component 706 configured to perform power management of electronic device 700, a wired or wireless communication component 716 configured to connect electronic device 700 to a network, and an input / output (I / O) interface 712. Electronic device 700 may operate on an operating system stored in memory 704, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, or similar.

[0156] This application also provides a computer program product, including a computer program that implements a cross-domain 3D Gaussian sputtering-enhanced visual positioning method when executed by a processor.

[0157] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the application disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the claims below.

[0158] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the claimed rights.

[0159] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0160] It will be readily apparent to those skilled in the art that any combination of the above embodiments is feasible. Therefore, any combination of the above embodiments is an implementation scheme of this application. However, due to space limitations, this specification will not describe them in detail here.

[0161] The cross-domain 3D Gaussian sputtering-enhanced visual positioning method provided herein is not inherently related to any particular computer, virtual system, or other device. Various general-purpose systems can also be used in conjunction with the teachings herein. Based on the above description, the required structure for constructing a system with the scheme of this application is obvious. Furthermore, this application is not directed to any particular programming language. It should be understood that the content of this application described herein can be implemented using various programming languages, and the above description of specific languages ​​is for the purpose of disclosing the best mode of implementation of this application.

[0162] Numerous specific details are set forth in the specification provided herein. However, it will be understood that embodiments of this application may be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this specification.

[0163] Similarly, it should be understood that, for the sake of brevity and to aid in understanding one or more of the various aspects of this application, in the above description of exemplary embodiments of this application, various features of this application are sometimes grouped together into a single embodiment, figure, or description thereof. However, this method of disclosure should not be construed as reflecting an intention that the claimed application requires more features than expressly recited in each claimed claim. Rather, as reflected in the claimed claims, the application aspects comprise fewer than all features of a single foregoing disclosed embodiment. Therefore, the specific implementation followed by the claimed claims is thus expressly incorporated into that specific implementation, wherein the content recited in each claimed claim itself serves as a separate embodiment of this application.

[0164] Those skilled in the art will understand that modules in the device of the embodiments can be adaptively changed and placed in one or more devices different from that embodiment. Modules, units, or components in the embodiments can be combined into a single module, unit, or component, and further, they can be divided into multiple sub-modules, sub-units, or sub-components. Except where at least some of such features and / or processes or units are mutually exclusive, any combination of all features disclosed in the application documents of this application and all processes or units of any method or device so disclosed can be employed. Unless expressly stated otherwise, each feature disclosed in the application documents of this application may be replaced by an alternative feature that serves the same, equivalent, or similar purpose.

[0165] Furthermore, those skilled in the art will understand that although some embodiments described herein include certain features but not others included in other embodiments, combinations of features from different embodiments are intended to be within the scope of this application and form different embodiments. For example, in each of the claimed claims, any one of the claimed embodiments can be used in any combination.

[0166] The various component embodiments of this application can be implemented in hardware, or as software modules running on one or more processors, or a combination thereof. Those skilled in the art will understand that microprocessors or digital signal processors (DSPs) can be used in practice to implement some or all of the functions of some or all of the components in the cross-domain 3D Gaussian sputtering-enhanced visual positioning method according to the embodiments of this application. This application can also be implemented as a device or apparatus program (e.g., a computer program and computer program product) for performing part or all of the methods described herein. Such an implementation of this application can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, provided on a carrier signal, or provided in any other form.

[0167] In another embodiment of the present invention, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to execute the cross-domain three-dimensional Gaussian sputtering-enhanced visual positioning method of the present application embodiments.

[0168] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)).

[0169] It should be noted that the above embodiments are illustrative of this application and not restrictive, and that those skilled in the art can devise alternative embodiments without departing from the scope of the claims. In the claims, any reference signs placed between parentheses should not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. This application can be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In the unit claims listing several means, several of these means may be embodied by the same hardware item. The use of the words first, second, and third, etc., does not indicate any order. These words can be interpreted as names.

[0170] It should be noted that, for the sake of simplicity, the method embodiments of this application are all described as a series of actions. However, those skilled in the art should understand that the embodiments of this application are not limited to the described order of actions, because according to the embodiments of this application, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential to the embodiments of this application.

[0171] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, for embodiments of systems or devices, since they are basically similar to the method embodiments, the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0172] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.

Claims

1. A visual positioning method based on cross-domain 3D Gaussian sputtering enhancement, characterized in that, include: A second image, corresponding to a first image in a first image domain, is generated using an image editing model to obtain an extended dataset of the first image dataset in the first image domain. Based on the first and second images corresponding to the extended dataset, the three-dimensional Gaussian sputtering model adapted to the feature distribution of the second image domain is fine-tuned so that the fine-tuned three-dimensional Gaussian sputtering model adapts to the feature distribution of the second image domain as well as the feature distribution of the first image domain. Based on the three-dimensional Gaussian sputtering model that has been fine-tuned by the sputtering model, a training set of positioning images for the image positioning model is generated, and the image positioning model is trained based on the training set of positioning images to obtain the trained image positioning model. The three-dimensional Gaussian sputtering model includes a photometric embedding data layer and a pose embedding data layer. The step of generating a localization image training set for the image localization model based on the fine-tuned three-dimensional Gaussian sputtering model, and training the image localization model using the localization image training set to obtain the trained image localization model, includes: Using each photometric embedding in the photometric embedding data layer as an anchor point, a three-dimensional virtual sphere is constructed according to a preset sampling interval. A preset number of image generation poses are randomly obtained in each three-dimensional virtual sphere. This allows each image generation pose and the photometric embedding corresponding to the image generation pose to be associated by a neural network according to a preset association function, while simultaneously forming a statistical representation according to a preset distribution function. The image generation pose and photometric embedding corresponding to each original image in the original image dataset containing the first original image in the first image domain and / or the second original image in the second image domain are respectively input into the three-dimensional Gaussian sputtering model that has been fine-tuned by the sputtering model, so as to form the localization image training set by all the rendered images corresponding to each original image. Generate a theoretical positioning pose corresponding to each rendered image based on each rendered image in the positioning image training set, and input each rendered image in the positioning image training set into the image positioning model to obtain an output positioning pose corresponding to each rendered image; Based on the one-to-one correspondence between the theoretical positioning pose and the output positioning pose, the image positioning model is trained to obtain the trained image positioning model.

2. The method for enhancing visual localization based on cross-domain 3D Gaussian sputtering as described in claim 1, characterized in that, The cross-domain 3D Gaussian sputtering-enhanced visual localization method also includes: The initial 3D Gaussian sputtering model is trained using a sputtering image training set containing the second image in the second image domain, so that the image fidelity of the rendered image output by the trained 3D Gaussian sputtering model in the second image domain is higher than the preset image fidelity threshold requirement.

3. The method for enhancing visual positioning based on cross-domain 3D Gaussian sputtering as described in claim 2, characterized in that, The three-dimensional Gaussian sputtering model incorporates a color multilayer perceptron and a photometric compensation multilayer perceptron. The training of the initialized three-dimensional Gaussian sputtering model using a training set of sputtered images containing the second image in the second image's neighborhood includes: The second images in the sputtering image training set are respectively input into the three-dimensional Gaussian sputtering model to obtain the color vector of each second image in the sputtering image training set from the color multilayer perceptron, and to obtain the photometric compensation vector and dynamic uncertainty parameter value of each second image in the sputtering image training set from the photometric compensation multilayer perceptron, and to obtain the output image of each second image in the sputtering image training set for the three-dimensional Gaussian sputtering model. The pixel color prediction image of each second image in the sputtered image training set is determined based on the color vector and photometric compensation vector of each second image in the sputtered image training set, and the dynamic object uncertainty data of each second image in the sputtered image training set is determined based on the dynamic uncertainty parameter value of each second image in the sputtered image training set. Based on all the obtained output images, the pixel color prediction images, and the dynamic object uncertainty data, a sputtering training loss function is constructed for the three-dimensional Gaussian sputtering model, and the sputtering model is trained using the sputtering training loss function.

4. The method for enhancing visual positioning based on cross-domain 3D Gaussian sputtering as described in claim 3, characterized in that, The step of constructing a sputtering training loss function for the three-dimensional Gaussian sputtering model based on all the obtained output images, the pixel color prediction images, and the dynamic object uncertainty data includes: A systematic loss function for the three-dimensional Gaussian sputtering model is constructed based on all the output images, the pixel color prediction images, and the dynamic object uncertainty data; and a structural similarity loss function for the three-dimensional Gaussian sputtering model is constructed based on all the output images and the pixel color prediction images. The regularized dual mean of the systematic loss function and the structural similarity loss function is determined as the sputtering training loss function.

5. The method for enhancing visual localization based on cross-domain 3D Gaussian sputtering as described in claim 1, characterized in that, The three-dimensional Gaussian sputtering model is trained on a training set of sputtered images that includes a second image in the second image domain. The step of generating a second image in the second image domain, corresponding to the first image in the first image domain, through an image editing model includes: The pose channel data of each first image in the first image dataset is respectively input into the three-dimensional Gaussian sputtering model to obtain a rendered image in the domain of the first image corresponding to each first image. The transformation instructions for the second image domain, and each of the rendered images, are input into the image editing model to obtain a second image corresponding to each of the first images in the first image dataset.

6. The method for enhancing visual localization based on cross-domain 3D Gaussian sputtering as described in claim 1, characterized in that, Before generating a second image in a second image domain corresponding to the first image in the first image domain using an image editing model, the cross-domain 3D Gaussian sputtering-enhanced visual localization method further includes: The image editing model generates a prior prediction image corresponding to the second image in the second image domain, thereby obtaining a prior dataset of the second image dataset in the second image domain. The image editing model is fine-tuned based on the second image and the prior prediction image in the prior dataset, so that the dynamic denoising performance index of the fine-tuned image editing model meets the preset anti-dynamic interference threshold requirement.

7. A cross-domain 3D Gaussian sputtering-enhanced visual positioning system, characterized in that, include: An extended dataset module is used to generate a second image in a second image domain corresponding to a first image in a first image domain, through an image editing model, so as to obtain an extended dataset of the first image dataset in the first image domain; The sputtering fine-tuning module is used to fine-tune the sputtering model of the three-dimensional Gaussian sputtering model adapted to the feature distribution of the second image domain according to the first image and the second image corresponding to the extended dataset, so that the fine-tuned three-dimensional Gaussian sputtering model is adapted to the feature distribution of the second image domain as well as the feature distribution of the first image domain. The positioning training module is used to generate a positioning image training set for the image positioning model based on the three-dimensional Gaussian sputtering model that has been fine-tuned by the sputtering model, and to train the image positioning model based on the positioning image training set to obtain the trained image positioning model. The three-dimensional Gaussian sputtering model includes a photometric embedding data layer and a pose embedding data layer, and the localization training module includes: The rendering deployment submodule is used to construct a three-dimensional virtual sphere with each photometric embedding in the photometric embedding data layer as an anchor point according to a preset sampling interval, so as to randomly obtain a preset number of image generation poses in each three-dimensional virtual sphere, so that each image generation pose and the photometric embedding corresponding to the image generation pose are associated by a neural network according to a preset association function, while forming a statistical representation according to a preset distribution function. The original training set submodule is used to input the image generation pose and photometric embedding corresponding to each original image from the original image dataset containing the first original image in the first image domain and / or the second original image in the second image domain into the three-dimensional Gaussian sputtering model that has been fine-tuned by the sputtering model, so as to form the localization image training set by all the rendered images corresponding to each original image. The data collection submodule is used to generate a theoretical positioning pose corresponding to each rendered image based on each rendered image in the positioning image training set, and input each rendered image in the positioning image training set into the image positioning model to obtain an output positioning pose corresponding to each rendered image. The localization training submodule is used to train the image localization model based on the one-to-one correspondence between the theoretical localization pose and the output localization pose, so as to obtain the trained image localization model.

8. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, which, when executed by a processor, implements the cross-domain three-dimensional Gaussian sputtering-enhanced visual positioning method as described in any one of claims 1 to 6.

9. An electronic device, characterized in that, It includes a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein when executed by the processor, the computer program implements the steps of the cross-domain three-dimensional Gaussian sputtering enhanced visual positioning method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • New view angle synthesis method and system based on generative adversarial strategy and Gaussian sputtering

    CN119379548A

  • Surface modeling method for porcelain bushing power transformation equipment based on unmanned aerial vehicle and light three-dimensional Gaussian

    CN119832154A