Method and system for enhancing visual positioning based on cross-domain three-dimensional Gaussian sputtering
Through the combination of image editing model and three-dimensional Gaussian sputtering model, the extended data set is generated and the model is fine-tuned, which solves the problem of insufficient robustness in cross-domain visual positioning and improves the stability and accuracy of the visual positioning system in different environments.
Patent Information
- Application Number
- CN202510873576.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-06-27
AI Technical Summary
The existing visual positioning methods are not robust in cross-domain scenarios due to the long-tail distribution of the data set, equipment distortion and cross-ambient lighting changes, making it difficult to adapt to different environmental characteristics, resulting in a decrease in positioning accuracy.
The image editing model generates an extended data set, converting the images of the first image field into images of the second image field to enrich the training data, and fine-tuning the three-dimensional Gaussian sputtering model to adapt to the feature distribution of the two image fields, generating a positioning image training set, and optimizing the cross-domain adaptability of the visual positioning model.
It improves the stability and accuracy of the visual positioning system in multiple environments, reduces positioning errors caused by equipment distortion and environmental changes, enhances the adaptability and generalization capabilities of the model, and meets the stability and high-precision requirements of cross-domain visual positioning.
Smart Images

Figure CN120388074A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of image positioning, and specifically relates to a cross-domain three-dimensional Gaussian sputtering enhanced visual positioning method, system, storage medium, device, and computer program product. Background Art
[0002] In modern image positioning tasks, high-precision visual positioning technology is widely used in fields such as autonomous driving, robot navigation, intelligent monitoring, and augmented reality. These applications usually require stable and accurate positioning on image data collected by different devices.
[0003] Currently, visual positioning methods based on absolute pose regression have received extensive attention in cross-domain visual positioning tasks due to their simple architecture, easy deployment, and high computational efficiency. Traditional methods mainly rely on a single pinhole camera model for training to achieve image positioning within a single image domain.
[0004] However, due to the long-tailed distribution problem of the training dataset, existing methods are difficult to adapt to complex lighting and tone changes in sub-domains (such as night images). Summary of the Invention
[0005] This application aims to provide a cross-domain three-dimensional Gaussian sputtering enhanced visual positioning method, system, storage medium, device, and computer program product, which at least solves the problem of insufficient adaptability of image positioning to sub-domains.
[0006] In a first aspect, an embodiment of this application discloses a cross-domain three-dimensional Gaussian sputtering enhanced visual positioning method, including: Generating a second image in a second image domain corresponding to a first image in a first image domain through an image editing model to obtain an extended dataset of the first image dataset in the first image domain; Fine-tuning a three-dimensional Gaussian sputtering model adapted to the feature distribution of the second image domain according to the corresponding first image and second image in the extended dataset, so that the fine-tuned three-dimensional Gaussian sputtering model not only adapts to the feature distribution of the second image domain but also adapts to the feature distribution of the first image domain; Generating a positioning image training set for an image positioning model according to the three-dimensional Gaussian sputtering model after the sputtering model fine-tuning is completed, and training the image positioning model according to the positioning image training set to obtain the trained image positioning model.
[0007] In a second aspect, an embodiment of this application also discloses a cross-domain three-dimensional Gaussian sputtering enhanced visual positioning system, including: An extended dataset module, configured to generate, through an image editing model, a second image in a second image domain corresponding to a first image in a first image domain, so as to obtain an extended dataset of the first image dataset in the first image domain; A sputtering fine-tuning module, configured to fine-tune a three-dimensional Gaussian sputtering model adapted to the feature distribution of the second image domain according to the corresponding first image and second image in the extended dataset, so that the fine-tuned three-dimensional Gaussian sputtering model is adapted to the feature distribution of the second image domain while also being adapted to the feature distribution of the first image domain; A positioning training module, configured to generate a positioning image training set for an image positioning model according to the three-dimensional Gaussian sputtering model after the sputtering model fine-tuning is completed, and perform positioning model training on the image positioning model according to the positioning image training set, so as to obtain the trained image positioning model.
[0008] In a third aspect, an embodiment of the present application further discloses a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps described in the first aspect or the second aspect are implemented.
[0009] In a fourth aspect, an embodiment of the present application further discloses an electronic device, including a processor, a memory, and a computer program stored on the memory and executable on the processor, and when the computer program is executed by the processor, the steps described in the first aspect or the second aspect are implemented.
[0010] In a fifth aspect, an embodiment of the present application further discloses a computer program product, on which a computer program is stored, and when the computer program is executed by a processor, the steps described in the first aspect or the second aspect are implemented.
[0011] In summary, in the embodiments of the present application, the image in the first image domain is converted into the image in the second image domain through an image editing model to expand the data set, enrich the training data, improve the adaptability of the visual positioning model to different environmental features, thereby reducing the robustness problem caused by the long-tailed distribution of the data set in the traditional method. Furthermore, the paired first image and second image in the expanded data set are used to fine-tune the three-dimensional Gaussian sputtering model, so that it can adapt to the feature distributions in both the first image domain and the second image domain at the same time. In this way, the adaptability of the model to feature changes such as illumination and hue in cross-domain images is optimized, enabling it to effectively capture visual information in multiple environments, improving the stability and accuracy of cross-domain visual positioning. The fine-tuned sputtering model is used to generate a positioning image training set for training the image positioning model, ensuring that the training data has consistent cross-domain feature distributions and forming a good constraint on the model training process, enabling it to better adapt to cross-domain scenarios in practical applications, improving the visual positioning accuracy, and reducing the positioning error caused by device distortion or environmental changes. Finally, the visual positioning model is trained through the positioning image training set, enabling it to achieve stable positioning under various device and scene conditions, reducing the impact of data scarcity on the model generalization ability. Thus, based on the method of the embodiments of the present application, through data expansion, model fine-tuning, and training optimization, the visual positioning system has stronger adaptability and generalization ability, while reducing the phenomenon of decreased positioning accuracy caused by environmental differences in the traditional method, thereby meeting the actual requirements of cross-domain visual positioning applications for stability and high precision. It plays a key role in solving the problem of insufficient robustness of existing visual positioning methods caused by the long-tailed distribution of the data set, device distortion, and cross-environmental illumination changes. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of illustrating the preferred embodiments and are not considered to be a limitation of the present application. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings: Figure 1 is a flowchart of the steps of a cross-domain three-dimensional Gaussian sputtering enhanced visual positioning method provided by an embodiment of the present application; Figure 2 is a flowchart of the steps of another cross-domain three-dimensional Gaussian sputtering enhanced visual positioning method provided by an embodiment of the present application; Figure 3 is a schematic structural diagram of a cross-domain three-dimensional Gaussian sputtering enhanced visual positioning system provided by an embodiment of the present application; Figure 4 is a block diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0013] Exemplary embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present application can be more thoroughly understood and the scope of the present application can be fully conveyed to those skilled in the art.
[0014] Figure 1 This is a cross-domain three-dimensional Gaussian sputtering enhanced visual positioning method provided by this embodiment, which specifically includes the following steps: Step 101, generate a second image in the second image domain corresponding to the first image in the first image domain through an image editing model to obtain an extended dataset of the first image dataset in the first image domain.
[0015] In some embodiments of the present application, an image in the main domain (second image domain) is generated based on an image in the sub-domain (first image domain) to construct an extended dataset for visual positioning optimization. Specifically, the sub-domain image can be transformed through an image editing model to generate image data that conforms to the feature distribution of the main domain. The image editing model utilizes a pre-trained visual feature transformation mechanism, combined with methods such as illumination adjustment, hue matching, and dynamic detail restoration, to make the transformed main-domain image match the real-scene features in terms of visual consistency. During this process, the generated extended dataset can not only enhance the adaptability of the visual positioning model to different environmental features but also reduce the robustness problems caused by the long-tail distribution of the dataset in traditional methods, thereby improving the accuracy and stability of the cross-domain visual positioning system.
[0016] In a specific example, it is applied to the environmental adaptability optimization of an autonomous driving system to improve the visual positioning ability of the vehicle under different lighting and environmental conditions. First, collect night (sub-domain) road scene images and use an image editing model to transform them into day (main-domain) road scene images, while adjusting the lighting intensity, road reflectivity, and dynamic elements to match the feature distribution of the main domain. Subsequently, combine the transformed main-domain images with the original day scene images to construct an extended dataset for optimizing the visual positioning model. During subsequent training, the model can more accurately learn the visual features of the main-domain environment, thereby improving the robustness and positioning accuracy of the autonomous driving system in complex environments.
[0017] Step 102, fine-tune the three-dimensional Gaussian sputtering model adapted to the feature distribution of the second image domain according to the corresponding first image and second image in the extended dataset.
[0018] Among them, the fine-tuned three-dimensional Gaussian sputtering model adapts to the feature distribution of the second image domain while also adapting to the feature distribution of the first image domain.
[0019] In some embodiments of the present application, an extended dataset will be used to optimize the 3D Gaussian Splatting (3DGS) model, enabling it to be adaptable in both the primary domain (the second image domain) and the secondary domain (the first image domain). Specifically, by adjusting the model parameters, it can stably render high-precision 3D scene images under different scene feature distribution conditions. The extended dataset contains paired images of the primary domain and the secondary domain. Using these images to fine-tune the splatting model enables it not only to adapt to the feature distribution of the primary domain but also to effectively adjust to the secondary domain environment. During the fine-tuning process of the splatting model, the model learns the lighting changes, hue shifts, and dynamic scene features between cross-domain images, enabling the fine-tuned model to generate stable and domain-consistent rendering results in cross-domain visual localization tasks, reducing cross-environment visual errors and improving localization accuracy.
[0020] In a specific example, it is applied to the cross-time visual localization optimization of an indoor navigation system to enhance the localization stability of the system under different lighting conditions. Nighttime (secondary domain) environment images have been collected, and corresponding daytime (primary domain) images have been generated. These paired images are used to fine-tune the 3D Gaussian Splatting model. During the fine-tuning process, the model learns the lighting change characteristics at different time periods, enabling the finally rendered nighttime scene to not only have the geometric accuracy of the primary domain but also reasonably reflect the lighting adjustment characteristics of the secondary domain. The optimized model can maintain high-precision visual localization in the nighttime environment, making the navigation system still highly robust in complex lighting change scenarios.
[0021] Step 103: Generate a localization image training set for the image localization model based on the 3D Gaussian Splatting model that has completed fine-tuning, and train the image localization model according to the localization image training set to obtain the trained image localization model.
[0022] In some embodiments of the present application, the fine-tuned 3DGS will be used to generate a localization image training set, and the image localization model will be optimized based on this training set to enhance its cross-domain visual localization ability. Specifically, the fine-tuned 3D Gaussian Splatting model can be used to render a series of images for the localization task, ensuring that the generated images cover the visual feature distributions of the primary domain (the second image domain) and the secondary domain (the first image domain). During this process, the model learns the lighting changes, hue matching, and dynamic interference control between cross-domain images, enabling the finally constructed localization image training set to effectively reflect the visual features in the real scene under different times and different device conditions. Subsequently, this training set is used to optimize the image localization model, enabling it to maintain stability and high precision when facing complex environments, thereby enhancing the robustness and generalization ability of cross-domain visual localization.
[0023] In a specific example, it is applied to the visual positioning optimization of an intelligent monitoring system to ensure that the system can accurately identify the target position even under the conditions of day-night conversion or device switching. First, a positioning image training set is generated based on the fine-tuned three-dimensional Gaussian sputtering model, which includes indoor environment images under different lighting conditions. Subsequently, the visual positioning model is trained using this training set so that it can accurately learn the effects of lighting changes and device distortions in the scene. After training, the optimized visual positioning model can stably locate the target in the monitoring scene at different time periods, while reducing the errors caused by lighting or device switching, thereby improving the stability and reliability of the system in cross-domain scenarios.
[0024] In summary, in the embodiment of the present application, the image in the first image domain is converted into the image in the second image domain through the image editing model to expand the data set, enrich the training data, and improve the adaptability of the visual positioning model to different environmental features, thereby reducing the robustness problem caused by the long-tail distribution of the data set in the traditional method. Furthermore, the paired first image and second image in the expanded data set are used to fine-tune the three-dimensional Gaussian sputtering model so that it can adapt to the feature distributions of both the first image domain and the second image domain at the same time. In this way, the adaptability of the model to the feature changes such as lighting and hue in cross-domain images is optimized, enabling it to effectively capture visual information in multiple environments, improving the stability and accuracy of cross-domain visual positioning. The fine-tuned sputtering model is used to generate a positioning image training set for training the image positioning model, ensuring that the training data has the consistency of cross-domain feature distribution and forming a good constraint on the model training process, enabling it to better adapt to the cross-domain scenarios in actual applications, improving the visual positioning accuracy, and reducing the positioning errors caused by device distortions or environmental changes. Finally, the visual positioning model is trained through the positioning image training set so that it can achieve stable positioning under various device and scene conditions, reducing the impact of data scarcity on the model generalization ability. Therefore, based on the method of the embodiment of the present application, through data expansion, model fine-tuning, and training optimization, the visual positioning system has stronger adaptability and generalization ability, while reducing the phenomenon of decreased positioning accuracy caused by environmental differences in the traditional method, thus meeting the actual requirements of cross-domain visual positioning applications for stability and high precision. It plays a key role in solving the problem of insufficient robustness of existing visual positioning methods caused by the long-tail distribution of the data set, device distortions, and cross-environment lighting changes.
[0025] Figure 2 This is another cross-domain three-dimensional Gaussian sputtering enhanced visual positioning method provided by the embodiment of the present application, which specifically includes the following steps: Step 201, perform sputtering model training on the initialized three-dimensional Gaussian sputtering model according to the sputtering image training set including the second image in the second image domain.
[0026] Among them, the trained three-dimensional Gaussian splatting model outputs a rendered image with an image fidelity higher than the preset image fidelity threshold requirement in the second image domain.
[0027] In some embodiments of the present application, the initialized three-dimensional Gaussian splatting model (3D Gaussian Splatting, 3DGS) will be trained to ensure that its feature adaptation ability in the main domain (the second image domain) meets the preset requirements. Specifically, based on the splatting image training set containing main domain images, by adjusting the model parameters, it can accurately render high-precision three-dimensional scene images. The splatting image training set contains image data from multiple perspectives in the scene, and combines Photometric Embedding technology to maintain lighting consistency and improve rendering quality. During this process, the model gradually learns the lighting changes corresponding to each perspective, so that after training, the generated rendered image meets the preset image fidelity threshold requirement in terms of image fidelity, thereby optimizing the model's adaptation ability in cross-domain scenarios.
[0028] In a specific example, it is applied to data preprocessing in the autonomous driving scenario to improve the accuracy of visual localization in the main domain (such as the daytime scenario). First, road environment images under different lighting conditions during the day are collected, and these images are used to construct a splatting image training set, which includes data from multiple perspectives under different weather conditions. Subsequently, the initialized three-dimensional Gaussian splatting model is used to train this training set, enabling it to learn the photometric change rules under different lighting conditions. After training, the rendered image generated by the model can meet the image fidelity of the experimental environment and maintain lighting consistency.
[0029] Optionally, a color multi-layer perceptron and a photometric compensation multi-layer perceptron are established in the three-dimensional Gaussian splatting model. Step 201 includes the following sub-steps: Sub-step 2011, input the second images in the splatting image training set into the three-dimensional Gaussian splatting model respectively, so as to obtain the color vectors of each second image in the splatting image training set from the color multi-layer perceptron, and obtain the photometric compensation vectors and dynamic uncertainty parameter values of each second image in the splatting image training set from the photometric compensation multi-layer perceptron, and obtain the output images of each second image in the splatting image training set for the three-dimensional Gaussian splatting model.
[0030] In some embodiments of the present application, color vectors, photometric compensation vectors, and dynamic uncertainty parameter values of the main domain (second image domain) images in the sputtering image training set are obtained to support the subsequent optimization training of the sputtering model. Specifically, each main domain image in the sputtering image training set can be input into 3DGS, and the color vectors can be extracted from the Color Multi-Layer Perceptron (Color-MLP) respectively, the photometric compensation vectors and dynamic uncertainty parameter values can be extracted from the Photometric Compensation Multi-Layer Perceptron (PC-MLP), and the model output images corresponding to each main domain image in the sputtering image training set can be obtained. In this process, the color multi-layer perceptron is used to encode the color features of the image, so that the generated output image maintains a high accuracy in terms of hue consistency, while the photometric compensation multi-layer perceptron is used to model the lighting changes and dynamic interference factors of the scene, so that the subsequent sputtering model training can more accurately simulate different lighting environments. Finally, the obtained color vectors, photometric compensation vectors, and dynamic uncertainty parameter values provide crucial support for the construction of the sputtering training loss function, ensuring that the three-dimensional Gaussian sputtering model meets the preset goals in terms of cross-domain image adaptation ability.
[0031] In a specific example, it is applied to the environmental light adaptation optimization of the indoor positioning scenario to improve the accuracy of the visual positioning system under different lighting conditions. First, indoor environment images under different lighting conditions during the day (main domain) are collected, and these images are processed using the three-dimensional Gaussian sputtering model. During the processing, the color vectors of each main domain image are extracted from the color multi-layer perceptron to ensure the color consistency of the subsequent rendered images, and at the same time, the photometric compensation vectors are obtained from the photometric compensation multi-layer perceptron to correct the light offset of the images under different lighting conditions. In addition, the dynamic uncertainty parameter values are obtained to adjust the interference response ability of the model to moving objects. After the processing is completed, the optimized data can be used to construct the sputtering training loss function, ensuring that the visual positioning model can maintain a stable positioning accuracy under various indoor lighting conditions.
[0032] Sub-step 2012: Determine the pixel color prediction image of each second image in the sputtering image training set according to the color vector and photometric compensation vector of each second image in the sputtering image training set, and determine the dynamic object uncertainty data of each second image in the sputtering image training set according to the dynamic uncertainty parameter value of each second image in the sputtering image training set.
[0033] In some embodiments of the present application, a pixel color prediction image will be generated using the color vectors and photometric compensation vectors of the images in the main domain (the second image domain) of the sputtering image training set, and the dynamic object uncertainty data will be calculated based on the dynamic uncertainty parameter values to enhance the adaptability of the model to cross-domain illumination and dynamic scene interference. Specifically, the pixel-level color distribution can be predicted by analyzing the color vectors of each main domain image to ensure that the generated image conforms to the main domain characteristics in terms of hue consistency. At the same time, the photometric compensation vector is used to adjust the image brightness and shadow distribution to optimize the illumination adaptation of the main domain image. Subsequently, the dynamic object uncertainty data in the scene is calculated based on the dynamic uncertainty parameter values, and this data is used to estimate the degree of influence of moving objects in the scene on the visual features of the image. During this process, the pixel color prediction image ensures the consistency of the generated training data in terms of visual feature distribution, while the dynamic object uncertainty data is used to optimize the robustness of the model to dynamic scene changes, providing effective support for subsequent sputtering model training.
[0034] In a specific example, it is applied to the visual optimization of an autonomous driving system to improve the visual recognition accuracy in a night driving scene (the secondary domain). First, road scene images under different lighting conditions during the day (the main domain) are collected, and the color vectors and photometric compensation vectors of each image are extracted. Subsequently, the pixel color prediction image is calculated using the color vectors to ensure that the generated image can still maintain visual consistency under different lighting conditions. At the same time, the photometric compensation vector is used to adjust the shadow distribution in the scene to conform to the actual environmental characteristics. In addition, the dynamic object uncertainty data is calculated using the dynamic uncertainty parameter values to identify the influence of the lighting changes of moving objects in the scene. After the data processing is completed, the optimized training data can be used for sputtering model training, enabling the autonomous driving system to maintain stable positioning in low-light environments.
[0035] Sub-step 2013, construct a sputtering training loss function for the three-dimensional Gaussian sputtering model according to the obtained all output images, pixel color prediction images, and dynamic object uncertainty data, and perform sputtering model training on the three-dimensional Gaussian sputtering model through the sputtering training loss function.
[0036] In some embodiments of the present application, image data obtained from the sputtering image training set will be used to construct a sputtering training loss function to optimize the adaptability of 3DGS in cross-domain visual scenarios. Specifically, the loss function for sputtering model training can be constructed by integrating all output images, pixel color prediction images, and dynamic object uncertainty data, enabling the model to simultaneously optimize the robustness against illumination, hue changes, and dynamic scene impacts. The loss function includes a pixel color prediction error term for correcting the color deviation of the model output image, and also includes a dynamic object uncertainty constraint term to reduce the visual error caused by the interference of moving objects. During this process, the loss function serves as the optimization objective to guide the model training, making it more generalizable in cross-domain image generation tasks and ensuring that the rendered images meet the high-precision requirements in different scenarios.
[0037] In a specific example, it is applied to the data training of an augmented reality navigation system to improve the stability of the visual positioning system under multiple environmental conditions. First, image data under different illumination and dynamic scene changes are collected, and the corresponding rendered images are generated using a three-dimensional Gaussian sputtering model. Subsequently, a sputtering training loss function is constructed to combine the color error of the pixel color prediction image and the dynamic object uncertainty data to optimize the adaptability of the model to different time and perspective changes. After training, the optimized three-dimensional Gaussian sputtering model can render high-precision images more accurately, enabling the augmented reality system to maintain stable visual recognition capabilities under complex illumination and dynamic environmental conditions.
[0038] Optionally, in order to construct a sputtering training loss function for the three-dimensional Gaussian sputtering model based on all the obtained output images, pixel color prediction images, and dynamic object uncertainty data, sub-step 2013 includes the following sub-steps: Sub-step 20131, construct a systematic loss function for the three-dimensional Gaussian sputtering model based on all the output images, pixel color prediction images, and dynamic object uncertainty data, and construct a structural similarity loss function for the three-dimensional Gaussian sputtering model based on all the output images and pixel color prediction images.
[0039] In some embodiments of the present application, a systematic loss function and a structural similarity loss function for constructing 3DGS will be formulated to optimize the rendering accuracy and stability of the model in cross-domain visual scenarios. Specifically, a systematic loss function can be established by analyzing all output images, pixel color prediction images, and dynamic object uncertainty data to measure the overall matching degree between the rendered image and the real scene. Meanwhile, a structural similarity loss function is constructed based on the output image and the pixel color prediction image to ensure the consistency of the color distribution of the rendered image. The systematic loss function is used to optimize the lighting and geometric consistency of cross-domain images, while the structural similarity loss function is used to ensure the color and texture stability of the images. During this process, the joint optimization of the two types of loss functions enables the three-dimensional Gaussian sputtering model to adapt to different lighting, perspective changes, and dynamic object interferences, thereby enhancing the robustness of the visual positioning system.
[0040] In a specific example, it is applied to the rendering optimization of an augmented reality navigation system to ensure that the system can still generate high-precision visual data under different lighting and dynamic scene changes. First, multiple perspective images in the scene are collected, and pixel color prediction images and dynamic object uncertainty data are constructed using this data. Subsequently, a systematic loss function is constructed based on the output image to optimize the lighting consistency in the augmented reality environment, and a structural similarity loss function is constructed based on the pixel color prediction image to improve the visual stability of the rendered image. After training, the optimized three-dimensional Gaussian sputtering model can generate high-quality visual data under different times and environmental conditions, thereby enhancing the interaction experience of the augmented reality system.
[0041] Sub-step 20132: Determine the regularized dual average value of the systematic loss function and the structural similarity loss function as the sputtering training loss function.
[0042] In some embodiments of the present application, the regularized dual average value of the systematic loss function and the structural similarity loss function will be used to construct a sputtering training loss function to optimize the adaptability of 3DGS in cross-domain visual tasks. Specifically, the systematic loss function can be determined first, which measures the overall error of the model's rendered image in terms of lighting consistency, geometric structure matching, and dynamic object influence control; then the structural similarity loss function can be determined, which focuses on optimizing the hue stability between the pixel color prediction image and the output image. To ensure the balance of the loss function, the regularized dual average value method is adopted in this step to dynamically adjust the weight ratio of the two loss functions during the optimization process to ensure that a high-stability model training target can be obtained under various environmental lighting and dynamic scenarios. Finally, this sputtering training loss function serves as the optimization basis, enabling the three-dimensional Gaussian sputtering model to have stronger adaptability and generalization ability in visual positioning tasks.
[0043] In a specific example, it is applied to the visual optimization of an intelligent monitoring system to ensure that the system can accurately identify scene features under different lighting conditions. First, images at multiple time periods in the monitoring environment are collected, and these data are used to calculate the systematic loss function and the structural similarity loss function. Subsequently, the regularized dual average method is adopted to optimize the weight ratio of the two loss functions, enabling the model to dynamically adjust the lighting adaptation ability and color stability during the training process. After training, the optimized intelligent monitoring system can maintain stable target recognition ability in complex environments such as day-night changes and dynamic interference, and reduce visual errors caused by uneven lighting.
[0044] In some embodiments of the present application, the cross-domain 3DGS is constructed based on the local three-dimensional Gaussian (Scaffold-GS) distributed by the anchor points. Then, two learnable photometric embeddings are constructed for each image in the training set. And the photometric histogram of the encoded image is displayed using the photometric embedding, and the spatial position of the photometric embedding is consistent with the spatial position of the training data. Two decoupled multi-layer perceptrons are constructed. can input the photometric embedding and the anchor point features from the Scaffold-GS together into the multi-layer perceptron to predict the color vector of the 3DGS, the photometric compensation vector and the dynamic uncertainty parameter value : , Then, according to the steps of rasterizing the 3DGS, the pixel color prediction map and the dynamic object uncertainty data can be calculated: , where i and j are a certain representative element in the dataset with a data capacity of N, is the opacity channel parameter.
[0045] Then, the main domain cross-domain 3DGS can be trained using the main domain data to learn the precise geometry and the appearance of the main domain of the scene. Furthermore, the training loss L used is: , where C is the pixel color ground truth (output image), M = M(x), is the Structural Similarity Index Measure (SSIM) function, is the regularization parameter.
[0046] Step 202: Generate a prior prediction image corresponding to the second image in the second image domain through an image editing model to obtain a prior dataset of the second image dataset in the second image domain.
[0047] In some embodiments of the present application, a prior prediction image will be generated based on the main domain (second image domain) image to construct a prior dataset for subsequent visual localization model training. Specifically, the image editing model can be used to process the main domain image so that it generates a prior prediction image consistent with the target scene features. The image editing model utilizes the learned laws of illumination, hue, and dynamic feature changes to ensure that the generated prediction image maintains consistency with the actual main domain scene in terms of visual feature distribution. During this process, the prior prediction image is not only used for data expansion but also serves as a reference for further optimizing the image editing model to improve its ability to model the main domain environmental features. Finally, the generated prior dataset effectively enriches the visual information in the main domain scene, providing more stable and accurate data support for model training in subsequent steps.
[0048] In a specific example, it is applied to the optimization of visual localization in an urban road scene to improve the image feature consistency in the main domain environment. First, the main domain road images at different time periods are collected, and the corresponding prior prediction images are generated using the image editing model, including adjustments for factors such as illumination changes and the influence of dynamic objects. Subsequently, the generated prior prediction images are combined with the original main domain images to construct a prior dataset for further optimizing the visual localization model. During the optimization process, the model can more accurately learn the visual features of the main domain environment, thereby improving the positioning accuracy and stability of the autonomous driving system under different illumination conditions.
[0049] Step 203: Fine-tune the image editing model according to the second image and the prior prediction image in the prior dataset.
[0050] Among them, the dynamic denoising performance index of the fine-tuned image editing model meets the requirements of a preset anti-dynamic interference threshold.
[0051] In some embodiments of the present application, a prior dataset will be used to optimize the image editing model, enabling it to more stably handle dynamic interference in the main domain (second image domain) images and ensuring that the generated sub-domain (first image domain) images have high visual consistency. Specifically, by comparing the feature differences between the main domain images and the prior prediction images in the prior dataset, the parameters of the image editing model can be adjusted to improve its denoising ability under different lighting conditions. The image editing model learns the patterns of lighting, hue changes, and dynamic object interference during the fine-tuning process, enabling the fine-tuned model to generate more accurate sub-domain images in complex scenarios. Finally, when the optimized image editing model performs the image conversion task, it can reduce the visual noise caused by dynamic objects and meet the preset anti-dynamic interference threshold requirements, thereby enhancing the stability and accuracy of the visual positioning system.
[0052] In a specific example, it is applied to the data preprocessing of an autonomous driving system to improve the visual recognition ability in night scenes (sub-domain). First, road images in the main domain (daytime) are collected, and the corresponding night images are generated using the image editing model. Subsequently, the original night scene images are compared with the generated night images to identify the visual errors caused by dynamic object interference. The parameters of the image editing model are adjusted to reduce the influence of unstable light sources and moving objects in subsequent image conversion processes. After fine-tuning, the night images generated by the optimized image editing model not only conform to the actual scene characteristics but also effectively reduce dynamic object interference, improving the robustness and reliability of the autonomous driving system in complex environments.
[0053] Step 204, generate a second image in the second image domain corresponding to the first image in the first image domain through the image editing model, so as to obtain an extended dataset of the first image dataset in the first image domain.
[0054] The method shown in this step has been described in step 101 and will not be elaborated here.
[0055] Optionally, step 204 includes the following sub-steps: Sub-step 2041, respectively use the pose channel data of each first image in the first image dataset as input data to input into a three-dimensional Gaussian sputtering model to obtain a rendered image in the first image domain corresponding to each first image.
[0056] In some embodiments of the present application, rendering images will be generated using sub-domain (first image domain) image data to ensure that the generated images conform to the feature distribution of the sub-domain in terms of visual features and provide high-quality input for subsequent cross-domain image conversion. Specifically, the pose channel data of the sub-domain images can be extracted and input into the 3DGS, and this model is used for three-dimensional reconstruction and rendering to generate rendering images corresponding to each sub-domain image. The three-dimensional Gaussian sputtering model adopts a point cloud representation, which includes color, photometric embedding, and dynamic confidence parameters inside, ensuring the stability of the output rendering images in terms of lighting consistency and structural restoration. During this process, the generated rendering images can accurately capture the photometric features in the sub-domain scene and provide a complete spatial representation, providing high-quality cross-domain data support for subsequent image conversion steps.
[0057] In a specific example, it is applied to cross-environment visual modeling of an intelligent monitoring system to improve the scene adaptability of the monitoring system under low-light conditions. First, multiple night (sub-domain) indoor environment images are collected, and the corresponding pose channel data are extracted. Subsequently, these data are input into the three-dimensional Gaussian sputtering model to generate corresponding rendering images, ensuring that the rendered images are consistent with the actual sub-domain scene in terms of visual features. In addition, the lighting characteristics of the generated images are adjusted through photometric embedding parameters so that they can correctly reflect the lighting changes in the night scene. After processing, the generated rendering images can be used for subsequent image conversion tasks, enabling the intelligent monitoring system to more effectively adapt to scene changes across lighting conditions and improving its recognition stability in low-light environments.
[0058] Sub-step 2042: The conversion instruction for the second image domain and each rendering image are respectively input into the image editing model to obtain the second images respectively corresponding to each first image in the first image dataset.
[0059] In some embodiments of the present application, rendering images and conversion instructions are utilized to achieve image conversion from a sub-domain (the first image domain) to a main domain (the second image domain), so as to construct visually consistent data. Specifically, the generated rendering image and the conversion instructions for the main domain can be input into an image editing model together, enabling it to generate a main domain image corresponding to the sub-domain image. The image editing model utilizes a pre-trained visual feature conversion mechanism, combined with methods such as lighting adjustment, tone matching, and dynamic detail restoration, so that the output main domain image not only conforms to the main domain characteristics in terms of color and lighting properties, but also can reduce the visual errors caused by scene conversion. During this process, the converted main domain image further enriches the cross-domain training data, enabling the visual localization model to more accurately adapt to different environmental scenarios and improving the generalization ability of cross-domain visual localization.
[0060] In a specific example, image conversion is applied to an autonomous driving system to optimize the vehicle's visual recognition ability under day-night change conditions. First, a road scene at night (sub-domain) is collected, and a corresponding rendering image is generated using a three-dimensional Gaussian sputtering model. Subsequently, day-night conversion instructions are input, and the rendering image is input into an image editing model to generate a converted image that matches the day (main domain) scene. After the conversion is completed, the lighting consistency and environmental feature matching degree of the image are verified, and the converted image data is used to train the visual localization model to improve the vehicle's positioning accuracy in different time scenarios.
[0061] Step 205: Fine-tune the three-dimensional Gaussian sputtering model adapted to the feature distribution of the second image domain according to the first image and the second image pair in the extended dataset.
[0062] Among them, the fine-tuned three-dimensional Gaussian sputtering model adapts to the feature distribution of the second image domain while also adapting to the feature distribution of the first image domain.
[0063] The method shown in this step has been described in step 102 and will not be elaborated here.
[0064] Step 206: Generate a positioning image training set for the image positioning model according to the three-dimensional Gaussian sputtering model that has completed fine-tuning of the sputtering model, and train the image positioning model according to the positioning image training set to obtain a trained image positioning model.
[0065] The method shown in this step has been described in step 103 and will not be elaborated here.
[0066] Optionally, the three-dimensional Gaussian sputtering model includes a photometric embedding data layer and a pose embedding data layer. Step 206 includes the following sub-steps: Sub-step 2061: Using each photometric embedding in the photometric embedding data layer as an anchor point, construct three-dimensional virtual spheres at a preset sampling interval to randomly obtain a preset number of image generation poses in each three-dimensional virtual sphere, so that while the neural network associates each image generation pose with the photometric embedding corresponding to the image generation pose according to a preset association function, a statistical representation is formed according to a preset distribution function.
[0067] In some embodiments of the present application, three-dimensional virtual spheres will be constructed using the photometric embeddings in the photometric embedding data layer, and image generation poses for image localization training will be generated to optimize the adaptation ability of 3DGS in cross-domain visual scenarios. Specifically, the photometric embeddings can be used as anchor points, and three-dimensional virtual spheres can be constructed at a preset sampling interval to ensure the uniform distribution of lighting features in space. Subsequently, through random sampling, a preset number of image generation poses are obtained in each three-dimensional virtual sphere, enabling it to effectively cover the lighting change range in the scene. During this process, each image generation pose is associated with the photometric embedding through a neural network, and at the same time, a statistical representation is formed according to a preset distribution function to ensure that the lighting change pattern can be accurately modeled. Ultimately, the adaptation ability of the visual localization model to lighting changes is improved, making the generated training data have higher lighting consistency, which helps to enhance the stability of cross-domain visual localization.
[0068] In a specific example, it is applied to the cross-lighting visual localization optimization of urban road scenarios to improve the adaptation ability of the autonomous driving system under different times and lighting conditions. First, photometric embedding data is collected in daytime (primary domain) and nighttime (secondary domain) road scenarios, and multiple three-dimensional virtual spheres are constructed according to a preset sampling interval to ensure that the lighting information can cover the entire scene. Subsequently, several image generation poses are randomly sampled within each sphere, and the association between the photometric embedding and the image generation pose is established through a neural network to ensure that the lighting features in the scene are accurately captured. After the data processing is completed, the generated lighting statistical representation is used to optimize the visual localization model of the autonomous driving system, enabling it to maintain stable localization performance during the day-night transition.
[0069] Sub-step 2062: Input the image generation poses and photometric embeddings corresponding to each original image in the original image dataset containing the first original image of the first image domain and / or the second original image of the second image domain into the three-dimensional Gaussian sputtering model that has been fine-tuned through the sputtering model, so as to form a localization image training set composed of all the rendered images corresponding to each original image obtained.
[0070] In some embodiments of the present application, a fine-tuned 3DGS will be used to render the original image dataset to construct a localization image training set for optimizing the image localization model. Specifically, the image generation poses and photometric embeddings corresponding to the original images, which include the sub-domain (the first image domain) and the main domain (the second image domain), can be respectively input into the three-dimensional Gaussian sputtering model, so that the corresponding rendered images are generated according to the photometric embedding (Photometric Embedding) and pose embedding (PoseEmbedding) of the model. The model uses the photometric embedding data layer to adjust the lighting characteristics of the image, so that the generated rendered images can still maintain visual consistency under different lighting environments. At the same time, the pose embedding data layer is used to ensure that the spatial relationships of different perspectives remain stable. In this process, the generated rendered images not only cover the feature distributions of the main domain and the sub-domain, but also provide high-quality data support for the subsequent training of the image localization model to optimize the accuracy and robustness of cross-domain visual localization.
[0071] In a specific example, it is applied to the visual localization training of an autonomous driving system to improve the vehicle's recognition ability under different lighting conditions. First, road scene images during the day (main domain) and at night (sub-domain) are collected, and the image generation poses and photometric embeddings contained in these images are input into the fine-tuned three-dimensional Gaussian sputtering model for rendering. The model uses the photometric embedding data layer to adjust the scene lighting, so that the rendered images at night match the actual environment in terms of lighting consistency. At the same time, the pose embedding data layer is used to ensure the spatial consistency between different perspectives. After the processing is completed, all the rendered images are integrated into a localization image training set and used to train the visual localization model of the autonomous driving system to enhance the localization stability of the system in complex lighting environments.
[0072] Sub-step 2063: Generate a theoretical localization pose corresponding to each rendered image in the localization image training set, and input each rendered image in the localization image training set into the image localization model respectively to obtain an output localization pose corresponding to each rendered image.
[0073] In some embodiments of the present application, the rendered images in the positioning image training set are used to generate corresponding theoretical positioning poses, which are then input into the image positioning model to obtain corresponding output positioning poses, thereby optimizing the training process of the visual positioning system. Specifically, first, based on the pose information of the rendered images, the theoretical positioning poses corresponding to each rendered image are generated through theoretical modeling methods to ensure that the model has accurate reference data during the training process. Subsequently, all the rendered images are respectively input into the image positioning model, enabling it to learn the visual features under different perspectives and lighting conditions in the scene and output positioning results that match each rendered image. During this process, the theoretical positioning pose serves as an idealized positioning reference, helping to guide the model to optimize the prediction accuracy, while the output positioning pose of the image positioning model is used to evaluate the learning ability of the model, thereby improving the stability and robustness of the visual positioning task.
[0074] In a specific example, it is applied to the visual positioning training of a robot navigation system to ensure that the robot can still maintain accurate positioning under different lighting and perspective conditions. First, the theoretical positioning pose is calculated based on the rendered images in the training dataset, and a pose calculation model is used to ensure the matching degree between the theoretical positioning pose and the actual environment. Subsequently, the rendered images are input into the image positioning model, enabling it to learn different scene features and generate corresponding output positioning poses. After the processing is completed, the differences between the theoretical positioning pose and the output positioning pose are compared, and these data are used to optimize the parameters of the positioning model to improve the robot's autonomous positioning ability in complex environments.
[0075] Optionally, in order to generate the theoretical positioning pose corresponding to each rendered image in the positioning image training set, sub-step 2063 includes the following sub-steps: Sub-step 20631, convert the panoramic image component in the rendered image to the linear coordinate space to obtain the spatial coordinate component of the theoretical positioning pose, and convert the image coordinate component in the rendered image to the global image coordinate space to obtain the position attention component of the theoretical positioning pose.
[0076] In some embodiments of the present application, the panoramic image component in the rendered image is converted to a linear coordinate space to ensure that the theoretical positioning pose can maintain accurate spatial consistency under different perspectives, and is based on the image coordinate component to be converted to the global image coordinate space to optimize the positioning accuracy. Specifically, a linear coordinate transformation can be first performed on the panoramic image component of the rendered image, so that the spatial coordinate component can accurately express the three-dimensional geometric relationship of the image in the standardized coordinate system. Subsequently, the image coordinate component is mapped to the global image coordinate space to construct the position attention component of the theoretical positioning pose and ensure that the image features under different perspectives can be correctly associated. During this process, the spatial coordinate component is used to optimize the geometric consistency of the positioning system, while the position attention component improves the robustness of the model to dynamically changing scenes through cross-domain visual feature alignment, ultimately enhancing the stability and accuracy of the visual positioning task. In some embodiments of the present application, since the currently generated image is only a 360-format image, we need to map to obtain an image across the camera domain. To save memory, we adopt an online mapping method when training the visual positioning method: that is, the cross-domain image (spatial coordinate component) can be obtained through the formula , where is the panoramic image component, represents the back-projection function of the panoramic camera, is the camera model (such as a fisheye camera, a pinhole camera, etc.)'s projection function; and to construct a position attention mechanism to guide the alignment of appearance features: a global image coordinate space can be first constructed based on the panoramic image, and this space can cover the field of view of any camera, and then the image coordinates from images of different cameras are converted to the global image coordinate space as the position attention: , where represents the projection function of the panoramic camera, is the camera model (such as a fisheye camera, a pinhole camera, etc.)'s back-projection function.
[0077] In a specific example, it is applied to visual scene alignment of a robot navigation system to enhance the positioning consistency of the robot in different scenarios. First, multi-view panoramic images of the indoor environment are collected, and corresponding rendered images are generated using a three-dimensional Gaussian sputtering model. Subsequently, a linear coordinate transformation is performed on the panoramic image components to ensure that the spatial coordinate components can accurately reflect the environmental geometric structure, and at the same time, the image coordinate components are transformed into the global image coordinate space to construct the position attention component. After the processing is completed, the optimized theoretical positioning pose is used to train the robot navigation system, enabling it to maintain stable visual recognition capabilities under different perspectives and environmental lighting conditions, and improving the accuracy and robustness of the navigation task.
[0078] Sub-step 20632: Concatenate the spatial coordinate component and the position attention component to obtain the theoretical positioning pose.
[0079] In some embodiments of the present application, the spatial coordinate component and the position attention component are concatenated to generate the theoretical positioning pose, thereby ensuring that the visual positioning system can accurately learn the spatial relationships and visual features in cross-domain scenarios. Specifically, the spatial coordinate component can be first used to describe the spatial geometric relationship of the rendered image in the standardized coordinate system to ensure the consistency of image features under different perspectives. Subsequently, the position attention component is used to model the pixel weight distribution in the global image coordinate space, enabling the positioning system to be optimized and adjusted for different camera perspectives. In this process, the spatial coordinate component ensures the geometric consistency of the visual positioning system, while the position attention component optimizes the stability of visual features under different devices or lighting conditions. Finally, the concatenated theoretical positioning pose not only maintains high precision in geometric structure but also improves the adaptability of the model to visual feature changes, thereby enhancing the robustness of the visual positioning task.
[0080] In a specific example, it is applied to scene matching of an indoor navigation system to optimize the positioning accuracy of the robot under different perspective conditions. First, multiple perspective images of the indoor environment are collected, and corresponding rendered images are generated using a three-dimensional Gaussian sputtering model. Subsequently, the spatial coordinate component is extracted to ensure the consistency of geometric relationships under different perspectives, and the position attention component is calculated from the global image coordinate space to optimize the alignment process of visual features. Finally, these two components are concatenated to generate the theoretical positioning pose, which is used to train the navigation system to maintain stable positioning capabilities under different positions and device conditions and improve the adaptability to complex environments.
[0081] Sub-step 2064: According to the one-to-one corresponding theoretical positioning pose and output positioning pose, perform positioning model training on the image positioning model to obtain the trained image positioning model.
[0082] In some embodiments of the present application, the image localization model will be trained using the theoretical localization pose and the output localization pose to optimize its localization accuracy and stability in cross - domain visual scenarios. Specifically, the theoretical localization pose can be used as a standard reference, compared with the output localization pose, and an optimization objective can be constructed using the error between the two to adjust the model parameters. During this process, the localization model learns the effects of different lighting conditions, perspective changes, and dynamic environmental factors on visual features and gradually optimizes its adaptability in cross - domain scenarios. This step also combines a loss calculation method based on the position attention mechanism, enabling the model to adaptively correct visual deviations caused by camera distortion or lighting changes in cross - domain data, thereby further improving the accuracy and robustness of visual localization.
[0083] In a specific example, it is applied to the optimization of visual localization in an autonomous driving system to enhance the stability of the vehicle in complex environments. First, a localization image training set is constructed based on night (sub - domain) and day (main - domain) images, and the corresponding theoretical localization poses are generated. Subsequently, all the rendered images are input into the image localization model, and the output localization images are recorded. By calculating the error between the theoretical localization pose and the output localization pose, the parameters of the localization model are adjusted to make it more adaptable to the visual feature changes under different lighting and dynamic environments. After training, the optimized visual localization model can maintain a high localization accuracy under conditions such as day - night transition or environmental changes, improving the robustness of the autonomous driving system.
[0084] In summary, in the embodiments of the present application, the image in the first image domain is converted into the image in the second image domain through the image editing model to expand the data set, enrich the training data, improve the adaptability of the visual positioning model to different environmental features, thereby reducing the robustness problem caused by the long-tailed distribution of the data set in the traditional method. Furthermore, the paired first image and second image in the expanded data set are used to fine-tune the three-dimensional Gaussian sputtering model so that it can adapt to the feature distributions in both the first image domain and the second image domain at the same time. In this way, the adaptability of the model to feature changes such as illumination and hue in cross-domain images is optimized, enabling it to effectively capture visual information in multiple environments, improving the stability and accuracy of cross-domain visual positioning. The fine-tuned sputtering model is used to generate a positioning image training set for training the image positioning model, ensuring that the training data has cross-domain feature distribution consistency and forming a good constraint on the model training process, enabling it to better adapt to cross-domain scenarios in practical applications, improving the visual positioning accuracy, and reducing the positioning error caused by device distortion or environmental changes. Finally, the visual positioning model is trained through the positioning image training set so that it can achieve stable positioning under various device and scenario conditions, reducing the impact of data scarcity on the model generalization ability. Thus, based on the method of the embodiments of the present application, through data expansion, model fine-tuning, and training optimization, the visual positioning system has stronger adaptability and generalization ability, while reducing the phenomenon of decreased positioning accuracy caused by environmental differences in the traditional method, thereby meeting the actual requirements of cross-domain visual positioning applications for stability and high precision. It plays a key role in solving the problem of insufficient robustness caused by the long-tailed distribution of the data set, device distortion, and cross-environmental illumination changes in the existing visual positioning methods.
[0085] As shown in Table 1, it is an error comparison table between the present application and the prior art: Specifically, when making the comparison, the present application set up the relevant environment under the test environment of Ubuntu 20.04, equipped with an Intel Gold 6330 series central processing unit with a processing frequency of 2.00 GHz, and additionally equipped with an NVIDIA GTX 4090 graphics processing unit with a core frequency of 2235 MHz and a video memory capacity of 24 GB, and then made a data comparison with the error of the panoramic visual positioning data set (360Loc). For example, the meaning of 6.5 / 39.0 recorded in the table indicates that in this positioning result, the distance error is 6.5 meters and the angle error is 39.0°.
[0086] Table 1 - Error Comparison Table between the Present Application and the Prior Art
[0087] As can be seen from Table 1, the positioning results obtained by the method according to the embodiments of the present application have obvious improvements in both distance error and angle error.
[0088] As shown Figure 3 in the figure, the embodiment of the present application also discloses a cross-domain three-dimensional Gaussian sputtering enhanced visual positioning system 30, including: An extended dataset module 301, configured to generate a second image in a second image domain corresponding to a first image in a first image domain through an image editing model, so as to obtain an extended dataset of the first image dataset in the first image domain; A sputtering fine-tuning module 302, configured to fine-tune a three-dimensional Gaussian sputtering model adapted to the feature distribution of the second image domain according to the corresponding first image and second image in the extended dataset, so that the fine-tuned three-dimensional Gaussian sputtering model not only adapts to the feature distribution of the second image domain, but also adapts to the feature distribution of the first image domain; A positioning training module 303, configured to generate a positioning image training set for the image positioning model according to the three-dimensional Gaussian sputtering model after the sputtering model fine-tuning is completed, and perform positioning model training on the image positioning model according to the positioning image training set to obtain a trained image positioning model.
[0089] Optionally, the cross-domain three-dimensional Gaussian sputtering enhanced visual positioning system 30 further includes: A sputtering training module, configured to perform sputtering model training on an initialized three-dimensional Gaussian sputtering model according to a sputtering image training set including a second image in a second image domain, so that the output rendering image of the trained three-dimensional Gaussian sputtering model in the second image domain has an image fidelity higher than a preset image fidelity threshold requirement.
[0090] Optionally, a color multi-layer perceptron and a photometric compensation multi-layer perceptron are established in the three-dimensional Gaussian sputtering model, and the sputtering training module includes: A data preparation sub-module, configured to input the second images in the sputtering image training set into the three-dimensional Gaussian sputtering model respectively, so as to obtain the color vector of each second image in the sputtering image training set from the color multi-layer perceptron, and obtain the photometric compensation vector and dynamic uncertainty parameter value of each second image in the sputtering image training set from the photometric compensation multi-layer perceptron, and obtain the output image of each second image in the sputtering image training set for the three-dimensional Gaussian sputtering model; A calculation sub-module, configured to determine a pixel color prediction image of each second image in the sputtering image training set according to the color vector and photometric compensation vector of each second image in the sputtering image training set, and determine the dynamic object uncertainty data of each second image in the sputtering image training set according to the dynamic uncertainty parameter value of each second image in the sputtering image training set; A training sub-module, configured to construct a sputtering training loss function for a three-dimensional Gaussian sputtering model based on all the obtained output images, pixel color prediction images, and dynamic object uncertainty data, and perform sputtering model training on the three-dimensional Gaussian sputtering model through the sputtering training loss function.
[0091] Optionally, the training sub-module includes: A loss component unit, configured to construct a systematic loss function for the three-dimensional Gaussian sputtering model based on all the output images, pixel color prediction images, and dynamic object uncertainty data, and construct a structural similarity loss function for the three-dimensional Gaussian sputtering model based on all the output images and pixel color prediction images; A loss unit, configured to determine the regularized dual average value of the systematic loss function and the structural similarity loss function as the sputtering training loss function.
[0092] Optionally, the three-dimensional Gaussian sputtering model is trained based on a sputtering image training set including second images in a second image domain. The extended dataset module 301 includes: An extended rendering sub-module, configured to respectively use the pose channel data of each first image in the first image dataset as input data to input into the three-dimensional Gaussian sputtering model to obtain a rendered image in the first image domain corresponding to each first image; An extended conversion sub-module, configured to input the conversion instruction for the second image domain and each rendered image into an image editing model respectively to obtain second images corresponding to each first image in the first image dataset respectively.
[0093] Optionally, the cross-domain three-dimensional Gaussian sputtering enhanced visual positioning system 30 further includes: A prior dataset module, configured to generate a prior prediction image corresponding to the second image in the second image domain through the image editing model to obtain a prior dataset of the second image dataset in the second image domain; An interference fine-tuning module, configured to perform editing model fine-tuning on the image editing model according to the corresponding second image and prior prediction image in the prior dataset, so that the dynamic denoising performance index of the fine-tuned image editing model meets the requirements of a preset anti-dynamic interference threshold.
[0094] Optionally, the three-dimensional Gaussian sputtering model includes a photometric embedding data layer and a pose embedding data layer. The positioning training module 303 includes: A rendering deployment sub-module, which uses each photometric embedding in the photometric embedding data layer as an anchor to construct three-dimensional virtual spheres at a preset sampling interval, randomly obtains a preset number of image generation poses in each three-dimensional virtual sphere, so that while the neural network association is performed between each image generation pose and the photometric embedding corresponding to the image generation pose according to a preset association function, a statistical representation is formed according to a preset distribution function; An original training set sub-module, which inputs the image generation poses and photometric embeddings corresponding to each original image in the original image dataset containing the first original images in the first image domain and / or the second original images in the second image domain into a three-dimensional Gaussian sputtering model that has been fine-tuned by a sputtering model, so as to form a positioning image training set composed of all the rendered images corresponding to each original image; A data collection sub-module, which generates a theoretical positioning pose corresponding to each rendered image in the positioning image training set, and inputs each rendered image in the positioning image training set into an image positioning model to obtain an output positioning pose corresponding to each rendered image; A positioning training sub-module, which trains the image positioning model according to the one-to-one corresponding theoretical positioning pose and output positioning pose to obtain a trained image positioning model.
[0095] Optionally, the data collection sub-module includes: A conversion unit, which converts the panoramic image component in the rendered image to a linear coordinate space to obtain the spatial coordinate component of the theoretical positioning pose, and converts the image coordinate component in the rendered image to a global image coordinate space to obtain the position attention component of the theoretical positioning pose; A splicing unit, which splices the spatial coordinate component and the position attention component to obtain the theoretical positioning pose.
[0096] In summary, in the embodiments of the present application, the image in the first image domain is converted into the image in the second image domain through the image editing model to expand the data set, enrich the training data, improve the adaptability of the visual positioning model to different environmental features, and thus reduce the robustness problem caused by the long-tail distribution of the data set in the traditional method. Furthermore, the paired first image and second image in the expanded data set are used to fine-tune the three-dimensional Gaussian sputtering model to make it adapt to the feature distributions in both the first image domain and the second image domain at the same time. In this way, the adaptability of the model to the feature changes such as illumination and hue in cross-domain images is optimized, enabling it to effectively capture visual information in multiple environments, improving the stability and accuracy of cross-domain visual positioning. The fine-tuned sputtering model is used to generate a positioning image training set for training the image positioning model, ensuring that the training data has cross-domain feature distribution consistency and forming a good constraint on the model training process, enabling it to better adapt to the cross-domain scenarios in actual applications, improving the visual positioning accuracy, and reducing the positioning error caused by device distortion or environmental changes. Finally, the visual positioning model is trained through the positioning image training set, enabling it to achieve stable positioning under various device and scenario conditions, reducing the impact of data scarcity on the model generalization ability. Therefore, based on the method of the embodiments of the present application, through data expansion, model fine-tuning, and training optimization, the visual positioning system has stronger adaptability and generalization ability, while reducing the phenomenon of decreased positioning accuracy caused by environmental differences in the traditional method, thus meeting the actual requirements of cross-domain visual positioning applications for stability and high precision. It plays a key role in solving the problem of insufficient robustness caused by the long-tail distribution of the data set, device distortion, and cross-environment illumination changes in the existing visual positioning methods.
[0097] The embodiments of the present application further provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements each process of the above-mentioned embodiment of the method for enhancing visual positioning based on cross-domain three-dimensional Gaussian sputtering and can achieve the same technical effect. To avoid repetition, it will not be elaborated here. Among them, the computer-readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.
[0098] Figure 4 It is a block diagram of an electronic device 700 provided by the embodiments of the present application. For example, the electronic device 700 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0099] Refer to Figure 4, the electronic device 700 may include one or more of the following components: a processing component 702, a memory 704, a power component 706, a multimedia component 708, an audio component 710, an input / output (I / O) interface 712, a sensor component 714, and a communication component 716.
[0100] The processing component 702 generally controls the overall operation of the electronic device 700, such as operations associated with display, telephone calls, data communications, camera operations, and recording operations. The processing component 702 may include one or more processors 720 to execute instructions to complete all or part of the steps of the above-mentioned cross-domain three-dimensional Gaussian sputtering enhanced vision positioning method. In addition, the processing component 702 may include one or more modules to facilitate the interaction between the processing component 702 and other components. For example, the processing component 702 may include a multimedia module to facilitate the interaction between the multimedia component 708 and the processing component 702.
[0101] The memory 704 is used to store various types of data to support the operation of the electronic device 700. Examples of such data include instructions for any application or method operating on the electronic device 700, contact data, phone book data, messages, pictures, multimedia, etc. The memory 704 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.
[0102] The power component 706 provides power to various components of the electronic device 700. The power component 706 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 700.
[0103] The multimedia component 708 includes a screen that provides an output interface between the electronic device 700 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operations. In some embodiments, the multimedia component 708 includes a front camera and / or a rear camera. When the electronic device 700 is in an operating mode, such as a shooting mode or a multimedia mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.
[0104] The audio component 710 is used to output and / or input audio signals. For example, the audio component 710 includes a microphone (MIC) that is used to receive external audio signals when the electronic device 700 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 704 or sent via the communication component 716. In some embodiments, the audio component 710 further includes a speaker for outputting audio signals.
[0105] The I / O interface 712 provides an interface between the processing component 702 and a peripheral interface module, and the peripheral interface module can be a keyboard, a click wheel, buttons, etc. These buttons can include but are not limited to: a home button, a volume button, a power button, and a lock button.
[0106] The sensor component 714 includes one or more sensors for providing a status assessment of various aspects of the electronic device 700. For example, the sensor component 714 can detect the on / off state of the electronic device 700, the relative positioning of components, such as the display and the keypad of the electronic device 700. The sensor component 714 can also detect a change in the position of the electronic device 700 or a component of the electronic device 700, the presence or absence of user contact with the electronic device 700, the orientation or acceleration / deceleration of the electronic device 700, and the temperature change of the electronic device 700. The sensor component 714 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor component 714 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 714 can further include an acceleration sensor, a gyro sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0107] The communication component 716 is used to facilitate communication between the electronic device 700 and other devices in a wired or wireless manner. The electronic device 700 can access a communication standard-based wireless network, such as WiFi, a carrier network (such as 2G, 3G, 4G, or 7G), or a combination thereof. In an exemplary embodiment, the communication component 716 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 716 further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra-Wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0108] In an exemplary embodiment, the electronic device 700 can be implemented by one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for implementing the cross-domain three-dimensional Gaussian sputtering enhanced visual positioning method provided in the embodiments of the present application.
[0109] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 704 including instructions, and the above instructions can be executed by a processor 720 of the electronic device 700 to complete the above-mentioned cross-domain three-dimensional Gaussian sputtering enhanced visual positioning method. For example, the non-transitory storage medium can be a ROM, Random Access Memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0110] In an exemplary embodiment, the electronic device 700 can also be provided as a server. The electronic device 700 includes a processing component 702, which further includes one or more processors 720, and memory resources represented by the memory 704 for storing instructions executable by the processing component 702, such as application programs. The application programs stored in the memory 704 can include one or more modules each corresponding to a set of instructions. In addition, the processing component 702 is configured to execute instructions to perform the cross-domain three-dimensional Gaussian sputtering enhanced visual positioning method provided in the embodiments of the present application.
[0111] The electronic device 700 may further include a power supply component 706 configured to perform power management of the electronic device 700, a wired or wireless communication component 716 configured to connect the electronic device 700 to a network, and an input / output (I / O) interface 712. The electronic device 700 may operate based on an operating system stored in the memory 704, such as WindowsServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM or the like.
[0112] The embodiments of the present application also provide a computer program product, including a computer program, which when executed by a processor, implements a cross-domain three-dimensional Gaussian sputtering enhanced vision positioning method.
[0113] Those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the application disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common knowledge or conventional technical means in the technical field not disclosed herein. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present application are pointed out by the following claims.
[0114] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the claims required.
[0115] Each embodiment in this specification is described in a progressive manner, with the key point of each embodiment being the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other.
[0116] It is easy for those skilled in the art to think that any combination application of the above-mentioned various embodiments is feasible. Therefore, any combination of the above-mentioned various embodiments is an embodiment of the present application. However, due to space limitations, this specification will not elaborate on them one by one here.
[0117] The cross-domain three-dimensional Gaussian sputtering enhanced vision positioning method provided herein is not inherently related to any specific computer, virtual system, or other device. Various general-purpose systems can also be used in conjunction with the teachings herein. Based on the above description, the structure required to construct a system with the solution of the present application is obvious. In addition, the present application is not directed to any specific programming language. It should be understood that the content of the present application described herein can be implemented using various programming languages, and the description of the specific language above is to disclose the best implementation mode of the present application.
[0118] In the specification provided herein, a number of specific details are set forth. However, it will be understood that embodiments of the present application may be practiced without these specific details. In some instances, well-known methods, structures and techniques have not been shown in detail so as not to obscure an understanding of the present specification.
[0119] Similarly, it should be understood that, in order to streamline the present application and assist in understanding one or more of the various aspects thereof, in the above description of the exemplary embodiments of the present application, the various features of the present application are sometimes grouped together in a single embodiment, figure, or description thereof. However, the disclosed method should not be construed as reflecting an intention that the claimed present application requires more features than are expressly recited in each claimed right. Rather, as reflected in the claimed rights, the aspects of the application lie in less than all the features of the single embodiment disclosed previously. Thus, the specific embodiments followed by the claimed rights are hereby expressly incorporated into the specific embodiments, where the content recited in each claimed right itself serves as a separate embodiment of the present application.
[0120] Those skilled in the art can understand that the modules in the devices in the embodiments can be adaptively changed and arranged in one or more devices different from the embodiments. The modules or units or components in the embodiments can be combined into one module or unit or component, and in addition, they can be divided into multiple sub-modules or sub-units or sub-components. Except that at least some of such features and / or processes or units are mutually exclusive, any combination can be used to combine all the features disclosed in the application documents of the present application and all the processes or units of any method or device thus disclosed. Unless otherwise expressly stated, each feature disclosed in the application documents of the present application can be replaced by an alternative feature providing the same, equivalent or similar purpose.
[0121] In addition, those skilled in the art can understand that, although some of the embodiments described herein include certain features included in other embodiments rather than other features, the combination of the features of different embodiments means that it is within the scope of the present application and forms different embodiments. For example, among the various claimed rights, any one of the claimed embodiments can be used in any combination.
[0122] Each component embodiment of the present application can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. Those skilled in the art should understand that a microprocessor or a digital signal processor (DSP) can be used in practice to implement some or all of the functions of some or all of the components in the cross-domain three-dimensional Gaussian sputtering enhanced vision positioning method according to the embodiments of the present application. The present application can also be implemented as a device or apparatus program (such as a computer program and a computer program product) for executing part or all of the methods described herein. Such a program for implementing the present application can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.
[0123] In another embodiment provided by the present invention, there is also provided a computer program product containing instructions, which when run on a computer, causes the computer to execute the cross-domain three-dimensional Gaussian sputtering enhanced vision positioning method according to the embodiments of the present application.
[0124] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described according to the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).
[0125] It should be noted that the above embodiments illustrate the present application rather than limit the present application, and those skilled in the art can design alternative embodiments without departing from the scope of the claimed rights. In the claimed rights, any reference signs placed between parentheses shall not be construed as limiting the claimed rights. The word "comprising" does not exclude the presence of elements or steps not listed in the claimed rights. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present application can be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In the unit claims listing several devices, several of these devices may be embodied by the same item of hardware. The use of the words first, second, and third, etc. does not denote any order. These words can be interpreted as names.
[0126] It should be noted that, for the method embodiments of the present application, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of the present application are not limited by the described order of actions, because according to the embodiments of the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential for the embodiments of the present application.
[0127] Each embodiment in this specification is described in a related manner. The same or similar parts among the embodiments can be referred to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the embodiments of the system or device, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.
[0128] The above is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.
Claims
1. A cross-domain three-dimensional Gaussian sputtering enhanced visual positioning method, characterized in that, Including: Generating a second image in a second image domain corresponding to a first image in a first image domain through an image editing model to obtain an extended dataset of a first image dataset in the first image domain; Fine-tuning a three-dimensional Gaussian sputtering model adapted to the feature distribution of the second image domain according to the corresponding first and second images in the extended dataset, so that the fine-tuned three-dimensional Gaussian sputtering model not only adapts to the feature distribution of the second image domain but also adapts to the feature distribution of the first image domain; Generating a positioning image training set for an image positioning model according to the three-dimensional Gaussian sputtering model after the sputtering model fine-tuning is completed, and training the image positioning model according to the positioning image training set to obtain the trained image positioning model.
2. The cross-domain three-dimensional Gaussian sputtering enhanced visual positioning method according to claim 1, wherein The cross-domain three-dimensional Gaussian sputtering enhanced visual positioning method further includes: Training a sputtering model of an initialized three-dimensional Gaussian sputtering model according to a sputtering image training set including a second image in a second image domain, so that the rendered image output by the trained three-dimensional Gaussian sputtering model in the second image domain has an image fidelity higher than a preset image fidelity threshold requirement.
3. The cross-domain three-dimensional Gaussian sputtering enhanced visual positioning method according to claim 2, wherein A color multi-layer perceptron and a photometric compensation multi-layer perceptron are established in the three-dimensional Gaussian sputtering model. The training of the sputtering model of the initialized three-dimensional Gaussian sputtering model according to the sputtering image training set including the second image in the second image domain includes: Inputting the second images in the sputtering image training set into the three-dimensional Gaussian sputtering model respectively to obtain the color vectors of each second image in the sputtering image training set from the color multi-layer perceptron, and obtaining the photometric compensation vectors and dynamic uncertainty parameter values of each second image in the sputtering image training set from the photometric compensation multi-layer perceptron, and obtaining the output images of each second image in the sputtering image training set for the three-dimensional Gaussian sputtering model; Determining a pixel color prediction image of each second image in the sputtering image training set according to the color vectors and photometric compensation vectors of each second image in the sputtering image training set, and determining the dynamic object uncertainty data of each second image in the sputtering image training set according to the dynamic uncertainty parameter values of each second image in the sputtering image training set; Constructing a sputtering training loss function for the three-dimensional Gaussian sputtering model according to all the obtained output images, the pixel color prediction images, and the dynamic object uncertainty data, and training the sputtering model of the three-dimensional Gaussian sputtering model through the sputtering training loss function.
4. The cross-domain three-dimensional Gaussian sputtering enhanced visual positioning method according to claim 3, wherein The constructing of the sputtering training loss function for the three-dimensional Gaussian sputtering model according to all the obtained output images, the pixel color prediction images, and the dynamic object uncertainty data includes: Construct a systematic loss function for the three-dimensional Gaussian sputtering model based on all of the output images, the pixel color prediction images, and the dynamic object uncertainty data, and construct a structural similarity loss function for the three-dimensional Gaussian sputtering model based on all of the output images and the pixel color prediction images; Determine the regularized dual average value of the systematic loss function and the structural similarity loss function as the sputtering training loss function.
5. The cross-domain three-dimensional Gaussian sputtering enhanced visual positioning method according to claim 1, wherein The three-dimensional Gaussian sputtering model is trained based on a sputtering image training set including second images in a second image domain. The generation of the second image in the second image domain corresponding to the first image in the first image domain through the image editing model includes: Respectively use the pose channel data of each first image in the first image dataset as input data and input them into the three-dimensional Gaussian sputtering model to obtain rendered images corresponding to each first image in the first image domain; Respectively input the conversion instruction for the second image domain and each of the rendered images into the image editing model to obtain second images corresponding to each first image in the first image dataset.
6. The cross-domain three-dimensional Gaussian sputtering enhanced visual positioning method according to claim 1, wherein Before generating the second image in the second image domain corresponding to the first image in the first image domain through the image editing model, the cross-domain three-dimensional Gaussian sputtering enhanced visual positioning method further includes: Generate a prior prediction image corresponding to the second image in the second image domain through the image editing model to obtain a prior dataset of the second image dataset in the second image domain; Fine-tune the image editing model according to the corresponding second images and prior prediction images in the prior dataset so that the dynamic denoising performance index of the fine-tuned image editing model meets the requirements of a preset anti-dynamic interference threshold.
7. The cross-domain three-dimensional Gaussian sputtering enhanced visual positioning method according to claim 1, wherein The three-dimensional Gaussian sputtering model includes a photometric embedding data layer and a pose embedding data layer. Generating a positioning image training set for the image positioning model according to the three-dimensional Gaussian sputtering model after the sputtering model fine-tuning is completed, and training the image positioning model according to the positioning image training set to obtain the trained image positioning model includes: Using each photometric embedding in the photometric embedding data layer as an anchor point, constructing a three-dimensional virtual sphere at a preset sampling interval to randomly obtain a preset number of image generation poses in each three-dimensional virtual sphere, so that while each image generation pose and the photometric embedding corresponding to the image generation pose are associated through a neural network according to a preset association function, a statistical representation is formed according to a preset distribution function; Input the image generation poses and photometric embeddings corresponding to each original image in the original image dataset containing the first original image of the first image domain and / or the second original image of the second image domain into the three-dimensional Gaussian sputtering model that has been fine-tuned by the sputtering model, so as to form the positioning image training set by all the rendered images obtained corresponding to each original image; Generate the theoretical positioning poses corresponding to each rendered image according to each rendered image in the positioning image training set, and input each rendered image in the positioning image training set into the image positioning model respectively to obtain the output positioning poses corresponding to each rendered image; Perform positioning model training on the image positioning model according to the corresponding theoretical positioning poses and output positioning poses one by one to obtain the trained image positioning model.
8. A cross-domain three-dimensional Gaussian sputtering enhanced visual positioning system, characterized in that, Comprising: An extended dataset module, configured to generate a second image corresponding to a first image in a second image domain through an image editing model, so as to obtain an extended dataset of the first image dataset in the first image domain; A sputtering fine-tuning module, configured to perform sputtering model fine-tuning on a three-dimensional Gaussian sputtering model adapted to the feature distribution of the second image domain according to the corresponding first image and second image in the extended dataset, so that the fine-tuned three-dimensional Gaussian sputtering model is adapted to the feature distribution of the second image domain while also being adapted to the feature distribution of the first image domain; A positioning training module, configured to generate a positioning image training set for the image positioning model according to the three-dimensional Gaussian sputtering model that has been fine-tuned by the sputtering model, and perform positioning model training on the image positioning model according to the positioning image training set to obtain the trained image positioning model.
9. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, it implements the cross-domain three-dimensional Gaussian sputtering enhanced visual positioning method according to any one of claims 1 to 7.
10. An electronic device, characterized in that, Comprising a processor, a memory, and a computer program stored on the memory and executable on the processor, and when the computer program is executed by the processor, it implements the steps of the cross-domain three-dimensional Gaussian sputtering enhanced visual positioning method according to any one of claims 1 to 7.
Citation Information
Patent Citations
New view angle synthesis method and system based on generative adversarial strategy and Gaussian sputtering
CN119379548A
Surface modeling method for porcelain bushing power transformation equipment based on unmanned aerial vehicle and light three-dimensional Gaussian
CN119832154A
Multi-view image reconstruction method and device based on three-dimensional Gaussian sputtering representation
CN120070799A
Four-dimensional object and scene model synthesis using generative models
US20250182404A1
Cited By
Three-dimensional reconstruction model training method and device, equipment and medium
CN121564459A