Visual positioning method, device and equipment and storage medium
Through the multimodal generation model, the target image is transferred in style, sensitive information is hidden, and the static image features are precisely positioned, which solves the accuracy and privacy leakage of visual positioning in complex outdoor monitoring scenes, and achieves efficient visual positioning.
Patent Information
- Application Number
- CN202411788234.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-06
- Publication Date
- 2025-05-16
AI Technical Summary
In complex and changeable outdoor surveillance scenarios, existing visual positioning technologies are difficult to achieve accurate positioning, and there is a problem of data privacy leakage.
The multimodal generation model is used for style transfer, hiding sensitive information in the target image, and coarsely positioning the image through the style transfer image, and accurately refine the positioning is combined with the static image features of the target image.
While maintaining positioning accuracy, it effectively hides scene sensitive information, improves visual positioning effect, and avoids data privacy leakage.
Smart Images

Figure CN120014220A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the field of image processing technology, and in particular to a visual positioning method, device, equipment and storage medium. Background Art
[0002] With the development of my country's low-altitude economy, the importance of low-altitude image acquisition equipment in the process of urban digitization is increasing. In monitoring scenarios, Global Navigation Satellite System (GNSS) and wireless network (Wi-Fi) signals are often unavailable, and only discrete video data is used as a reference. In order to solve this problem, scholars began to study camera positioning technology based on pure vision.
[0003] Visual positioning aims to analyze visual clues in images to determine the camera pose, and is a hot research topic in the field of computer vision. The input of the visual positioning algorithm in the prior art is the image to be tested and a pre-prepared reference database, and the output is the six-degree-of-freedom pose of the camera, which includes three-dimensional rotation angles and position components. At present, a common problem with visual positioning algorithms is data privacy leakage. For complex and changeable outdoor monitoring scenes, visual positioning is still difficult to be accurate. On the one hand, when the front-end user transmits the image to be positioned to the back-end positioning server, the unprocessed image is prone to leak sensitive information such as people and vehicles; on the other hand, when the back-end server presents the visual positioning results, it is easy to leak three-dimensional model information, causing damage to the rights and interests of digital assets. Summary of the invention
[0004] The purpose of the present invention is to at least provide a visual positioning method, device, equipment and storage medium, which can at least solve the problem of difficulty in visual positioning in complex and changeable outdoor monitoring scenes and the resulting privacy leakage, at least be able to perform visual positioning on complex monitoring images, and be able to hide scene-sensitive information while maintaining positioning accuracy, thereby improving the visual positioning effect.
[0005] To solve the above technical problems, at least one embodiment of the present application provides a visual positioning method, including a front-end processing stage and a back-end positioning stage;
[0006] The front-end processing stage includes:
[0007] Get the target image;
[0008] Performing feature extraction on the target image to obtain static image features of the target image;
[0009] Inputting the target image into a pre-trained multimodal generative model for style transfer processing to obtain a style-transferred image after hiding sensitive information of the target image;
[0010] The backend positioning stage includes:
[0011] Retrieving and roughly locating the style transfer image in a pre-built reference database, and determining a candidate reference image that is most similar to the style transfer image as a rough positioning result;
[0012] Feature matching is performed based on the coarse positioning result and the static image features of the target image to obtain a feature matching result, and the position and posture information corresponding to the target image is determined based on the feature matching result.
[0013] At least one embodiment of the present application further provides a visual positioning device, comprising: a front-end processing module and a back-end positioning module; the front-end processing module comprises: a data acquisition unit, used to acquire a target image, and perform feature extraction on the target image to obtain static image features of the target image; a style transfer unit, used to input the target image into a pre-trained multimodal generation model for style transfer processing, and obtain a style transfer image after hiding sensitive information of the target image;
[0014] The back-end positioning module includes: a retrieval coarse positioning unit, which is used to perform retrieval coarse positioning in a pre-constructed reference database based on the style migration image, and determine a candidate reference image that is most similar to the style migration image as a coarse positioning result; a precise positioning unit, which is used to perform feature matching based on the coarse positioning result and the static image features of the target image to obtain a feature matching result, and determine the position and posture information corresponding to the target image based on the feature matching result.
[0015] At least one embodiment of the present application also provides an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the above-mentioned visual positioning method.
[0016] At least one embodiment of the present application further provides a computer-readable storage medium storing a computer program, wherein the computer program implements the above-mentioned visual positioning method when executed by a processor.
[0017] The visual positioning method, apparatus, device and storage medium provided in the embodiments of the present application can perform style transfer on the target image through a multimodal generation model, thereby hiding sensitive information in the target image and avoiding leakage after visual positioning, and extracting features for visual positioning. By using the style transfer image for coarse positioning and then combining it with the target image for precise and detailed positioning, the accuracy of posture information determination can be improved, thereby improving the visual positioning effect.
[0018] In some optional embodiments, the front-end processing stage further includes: obtaining style transfer prompts and sensitive information hiding prompts;
[0019] The step of inputting the target image into a pre-trained multimodal generative model to perform style transfer processing on the target image to obtain a style transfer image after hiding sensitive information from the target image comprises:
[0020] Decomposing and vectorizing the acquisition style transfer prompt and the sensitive information hiding prompt into a text encoding of the multimodal generation model;
[0021] The text encoding of the converted multimodal generative model and the input of the target image into the pre-trained multimodal generative model are subjected to style transfer processing to obtain a style transferred image after hiding sensitive information of the target image.
[0022] In some optional embodiments, the multimodal generation model includes a Stable Diffusion model, into which a pre-trained ControlNet plug-in network is inserted for providing conditional constraint control on the multimodal generation model.
[0023] In some optional embodiments, each reference image in the reference database is generated by traversing and rendering the target space using a renderer using a three-dimensional reconstruction model.
[0024] In some optional embodiments, the retrieving and coarse positioning based on the style transfer image in a pre-built reference database to determine a candidate reference image most similar to the style transfer image as a coarse positioning result includes:
[0025] Based on the style transfer image, a search is performed in a pre-built reference database to obtain a plurality of candidate reference images that are most similar to the style transfer image as search recommendation results;
[0026] The retrieval recommendation results are sorted from high to low by similarity using a cosine similarity calculation formula to determine the rough positioning result.
[0027] In some optional embodiments, performing feature matching based on the coarse positioning result and the static image features of the target image to obtain a feature matching result, and determining the position and posture information corresponding to the target image based on the feature matching result includes:
[0028] Rendering is performed using a renderer based on the coarse positioning result to obtain a rendered reference image;
[0029] Extracting features from the rendered reference image to obtain reference image features, and fusing the reference image features with the static image features to obtain a fused feature image;
[0030] The feature image is input into a preset visual positioning algorithm for calculation to obtain the posture information output by the visual positioning algorithm.
[0031] In some optional embodiments, inputting the feature image into a preset visual positioning algorithm for calculation to obtain the posture information output by the visual positioning algorithm includes:
[0032] Determining whether the pose information output by the visual positioning algorithm successfully locates the pose information corresponding to the target image based on a preset backtracking condition;
[0033] If not, re-enter the step according to the backtracking condition to perform a search and rough positioning based on the style transfer image in a pre-built reference database, determine a candidate reference image most similar to the style transfer image as a rough positioning result, and update the rough positioning result.
[0034] In some optional embodiments, judging whether the pose information output by the visual positioning algorithm successfully locates the pose information corresponding to the target image based on a preset backtracking condition includes:
[0035] Determine a reprojection error value between a feature point of the rendered reference image and a feature point of the pose information according to a feature matching result of the feature image and pose information corresponding to a current feature matching result;
[0036] Comparing the reprojection error value with a preset threshold range to obtain a comparison result;
[0037] Whether the position and posture information corresponding to the target image is successfully located is determined according to the comparison result.
[0038] In some optional embodiments, determining whether the position and posture information corresponding to the target image is successfully located according to the comparison result includes:
[0039] When the comparison result is that the reprojection error value is less than the minimum value of the preset threshold range, it is determined that the positioning is successful, and the posture information is output;
[0040] When the comparison result shows that the reprojection error value is within the preset threshold range, it is recorded as a backtracking point, and the visual optimization algorithm is iteratively optimized with the currently determined posture information until the reprojection error value is less than the minimum value of the preset threshold range;
[0041] When the comparison result is that the reprojection error value is greater than the maximum value of the preset threshold range, it is further determined whether the current point is a recorded backtracking point. If not, the positioning is determined to have failed and the iteration is terminated, and the process is re-circulated to the step of retrieving rough positioning based on the style transfer image in a pre-built reference database. If so, the visual optimization algorithm is iteratively optimized with the corresponding backtracking point until the reprojection error value is less than the minimum value of the preset threshold range. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] One or more embodiments are exemplarily described by the pictures in the corresponding drawings, and these exemplary descriptions do not constitute limitations on the embodiments.
[0043] Figure 1 This is the process of the visual positioning method provided by an embodiment of the present application Figure 1 ;
[0044] Figure 2 This is the process of the visual positioning method provided by an embodiment of the present application Figure 2 ;
[0045] Figure 3 This is the process of the visual positioning method provided by an embodiment of the present application Figure 3 ;
[0046] Figure 4 This is the process of the visual positioning method provided by an embodiment of the present application Figure 4 ;
[0047] Figure 5 This is the process of the visual positioning method provided by an embodiment of the present application Figure 5 ;
[0048] Figure 6 This is the process of the visual positioning method provided by an embodiment of the present application Figure 6 ;
[0049] Figure 7 is a detailed flow chart of a visual positioning method provided by another embodiment of the present application;
[0050] Figure 8 is a schematic diagram of an iterative optimization process of a visual positioning method provided in an embodiment of the present application;
[0051] Fig. 9 is an exemplary block diagram of a visual positioning device provided by another embodiment of the present application;
[0052] Fig.10 is a structural schematic diagram of an electronic device provided by another embodiment of the present application;
[0053] Fig.11It is a schematic diagram of the structure of a computer-readable storage medium provided by another embodiment of the present application. DETAILED DESCRIPTION
[0054] To make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the embodiments of the present application will be described in detail below in conjunction with the accompanying drawings. However, it will be appreciated by those skilled in the art that in the present application, many technical details are proposed in order to enable the reader to better understand the present application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical scheme claimed in the present application can also be implemented. The division of the following embodiments is for the convenience of description, and the specific implementation of the present application should not be construed as any limitation, and the various embodiments can be combined and referenced with each other under the premise of no contradiction.
[0055] In order to facilitate understanding of the embodiments of the present application, relevant content about visual positioning is first introduced here.
[0056] Visual localization technology aims to analyze visual clues in images to determine the camera pose and is a hot research topic in the field of computer vision. The input of the visual localization algorithm is the image to be tested and a pre-prepared reference database. The output is the six-degree-of-freedom pose of the camera, which includes three-dimensional rotation angle and position components. Factors such as season, weather, lighting and human activities cause differences in image appearance, which often affect the success rate of localization [1]. In recent years, scholars have mainly studied the problem of long-term visual localization. According to the different ways of using reference data, mainstream visual localization algorithms can be divided into three categories [2]: methods based on two-dimensional images, methods based on three-dimensional point cloud registration, and structured localization methods based on the combination of two and three dimensions. Methods based on two-dimensional image matching directly match the image to be tested with database images of known poses to solve the pose. The DenseVLAD algorithm proposed by Torii et al. [3] uses the normalized RootSIFT descriptor to compress the image and construct a visual bag-of-words vector for retrieval and localization. It has achieved good results on the Google Street View dataset. Arandjelovic et al. [4] proposed the NetVLAD model, which uses a deep convolutional neural network to improve the robustness of the descriptor, generalization and retrieval success rate. On this basis, the Patch-NetVLAD[5] algorithm improved the network structure and focused on local details of the image. PoseNet[6] used an end-to-end convolutional neural network to directly regress pose parameters to replace the retrieval step. Sattler et al.[7] analyzed the limitations of image-based visual localization methods in positioning accuracy and pointed out that the fundamental reason is that two-dimensional images cannot cover the geometric structure of three-dimensional scenes, resulting in the inability of deep learning networks to understand perspective. In recent years, scholars have begun to study the integration of three-dimensional scene information into visual localization algorithms to improve accuracy. The visual localization method based on three-dimensional point cloud first estimates the depth map in the image, constructs a local point cloud, and then matches it with the existing point cloud in the database. Representative methods include PointNet[8] and PointNet++[9]. This type of method has high positioning accuracy and fast positioning speed in indoor scenes, so it is widely used in embedded device environments such as robots. However, this type of method lacks understanding of image texture and semantics and is easily disturbed by moving and occluded objects. In recent years, structured localization algorithms based on the combination of two and three dimensions have become mainstream. Hloc
[10] proposed a hierarchical structure to process two-dimensional and three-dimensional data separately and align the results in the positioning backend. PixLoc
[11] further introduced a pixel coordinate loss function to optimize the positioning results. MeshLoc
[12] and VirtualLoc
[13] dynamically generate reference images by using a 3D scene model containing textures to replace discrete point cloud data. Currently, a common problem with visual positioning algorithms is data privacy leakage.On the one hand, when the front-end user transmits the image to be located to the back-end positioning server, the unprocessed image is prone to leak sensitive information such as people and vehicles; on the other hand, when the back-end server presents the visual positioning results, it is easy to leak the three-dimensional model information, causing damage to the rights and interests of digital assets. To solve this problem, Speciale et al.
[14] used straight line features and point cloud matching in the image to reduce privacy leakage. SegLoc
[15] avoids privacy leakage by compressing images and three-dimensional scenes into semantic segmentation forms. However, these methods lose the visual clues in the image and lose the positioning accuracy. Therefore, how to hide sensitive information without losing image details and meeting the positioning accuracy is a basic problem.
[0057] In recent years, scholars have begun to study style transfer methods based on large models. This method can be used for texture synthesis
[16] and color information hiding
[17] to avoid privacy leakage
[18] . Among them, the Stable Difussion (SD) large model proposed by Rombach et al.
[19] generates images by first adding noise to the image and then using a deep convolutional neural network to remove the noise. This method introduces a Contrastive Language-Image Pre-training (CLIP)
[20] model when training the image generation layer encoder, uses user text prompts to guide the network to generate images that meet user preferences, and supports text-guided fine-tuning. However, the images generated by this method are often too different from the original images, resulting in feature loss, and are not suitable for visual localization tasks. Zhang et al.
[21] proposed an improved version of the ControlNet (CN), which superimposes a conditional constraint branch model on the original diffusion model. By adding conditional constraints during the training process, the generated image is closer to the original image in terms of layout structure. Large visual models lack spatial location information
[22] and cannot directly answer visual localization problems. To this end, Ai Haojun et al.
[23] used feature clustering technology to generate data fingerprints for images generated by diffusion models, and achieved certain success in indoor visual positioning. However, how to guide large models to generate images for visual positioning in complex and changeable outdoor monitoring scenarios remains a difficult problem.
[0058] In order to solve the above-mentioned technical problems of difficulty in visual positioning in complex and changeable outdoor monitoring scenarios and privacy leakage, the present invention proposes a visual positioning method. The implementation details of the visual positioning method of this embodiment are described in detail below. The following content is only the implementation details provided for easy understanding and is not necessary for the implementation of this solution.
[0059] Embodiment 1:
[0060] The visual positioning method of this embodiment can be applied to electronic devices with communication, computing and data storage capabilities. The specific process can be as follows: Figure 1 As shown, it includes the front-end processing stage and the back-end positioning stage, specifically:
[0061] The front-end processing stage includes:
[0062] Step 101, acquiring a target image;
[0063] Specifically, the target image is captured by a camera.
[0064] In some examples, aerial views of cities can be captured by cameras mounted on drones.
[0065] Step 102, extracting features of the target image to obtain static image features of the target image;
[0066] Specifically, feature extraction is performed on the target image, specifically, gradient features in the target image are extracted to obtain static image features of the target image.
[0067] Furthermore, after step 102, the front-end processing stage also uses a preprocessing algorithm to extract the main constraint information in the image and generate a constraint image as a condition for controlling the style transfer model. Specifically, the constraint image is inserted into the multi-modal generation model in step 103 through a plug-in network to provide conditional constraint control and improve the style transfer accuracy.
[0068] Step 103: Input the target image into a pre-trained multimodal generative model to perform style transfer processing, and obtain a style transferred image after hiding sensitive information of the target image.
[0069] Specifically, before executing step 103, the front-end processing stage also includes: obtaining style transfer prompts and sensitive information hiding prompts.
[0070] like Figure 2 As shown, step 103 specifically includes:
[0071] Step 1031: Decompose and vectorize the acquired style transfer prompt and the sensitive information hiding prompt into a text encoding of the multimodal generation model;
[0072] Step 1032: The text encoding of the converted multimodal generative model and the input of the target image into the pre-trained multimodal generative model are processed by style transfer to obtain a style transferred image after hiding sensitive information of the target image.
[0073] In this embodiment, the multimodal generation model includes a Stable Diffusion model, in which a pre-trained ControlNet plug-in network is inserted to provide conditional constraint control for the multimodal generation model. Specifically, the style transfer prompt is decomposed into positive prompt words, and the sensitive hidden prompt is decomposed into reverse prompt words. The positive prompt words and reverse prompt words are vectorized into text encoding and input into the Stable Diffusion large model. The system also vectorizes the time step information into position encoding and passes it to the Stable Diffusion model to control the noise and denoising process of the diffusion model. In order to avoid the loss of the main content of the image due to excessive style transfer, the number of iteration steps is set to 20. This paper uses the pre-trained Stable Diffusion 1.5 version large model as the basic model, and is used for conditional constraint control with the pre-trained ControlNet1.1 version plug-in. When in use, ControlNet is used as a plug-in network of Stable Diffusion, and has the same network structure as shared in the noise adding part, but ControlNet also receives the edge constraints of the pre-processed image at the same time, and uses a zero convolution layer to connect the denoising network of Stable Diffusion for fine-tuning. The outputs of the two are finally superimposed to output a style transfer image with conditional constraints.
[0074] Furthermore, the back-end positioning stage is a continuation of the front-end processing stage. After being processed in the front-end processing stage, the style transfer image and the static image features of the target image are transmitted to the back-end positioning stage for processing.
[0075] Specifically, the backend positioning stage includes:
[0076] Step 201: performing a search and rough positioning in a pre-built reference database based on the style transfer image, and determining a candidate reference image that is most similar to the style transfer image as a rough positioning result;
[0077] Specifically, each reference image in the reference database is rendered by traversing the target space using a 3D reconstruction model through a renderer. In this embodiment, a 3D reconstruction model of a city scene is used to traverse the city space for rendering, and a virtual image is generated to construct a reference database for covering the target space.
[0078] Furthermore, the reference database includes two types of data: offline data and online data, both of which are generated by rendering the three-dimensional scene model by the rendering engine. Offline data can be prepared in advance based on the reconstruction of the three-dimensional scene, and the low-altitude range of the target city scene is traversed and rendered at fixed intervals by the renderer. The traversal interval is different depending on the size of the city scene. Among them, the online data is the data information stored in the reference database after executing the method of this application.
[0079] In some examples, step 201, such as Figure 3 As shown, specifically including:
[0080] Step 2011: searching a pre-built reference database based on the style transfer image to obtain a plurality of candidate reference images that are most similar to the style transfer image as search recommendation results;
[0081] Step 2012: sort the search recommendation results from high to low by similarity using a cosine similarity calculation formula to determine the rough positioning result.
[0082] Specifically, after the offline data in the reference database is prepared, the NetVLAD algorithm is used to compress the image into a 4096-bit retrieval descriptor. When the front-end transmits the relevant image to be located to the back-end, the visual positioning module calls the NetVLAD algorithm to compress the front-end image into a 4096-bit descriptor, retrieves and recommends the most similar candidate reference image in the offline database, and the specific cosine similarity calculation formula is:
[0083]
[0084] Where: p and q are descriptor vectors. Due to the environmental changes in outdoor scenes, there are still cases of recommendation errors even without style transfer.
[0085] In some examples, when step 2012 is sorted, the top 10 recommended results with the highest similarity are selected for sorting, and then the top 10 recommended results are used as rough positioning results for feature matching in the first pose calculation, and are adjusted according to the judgment result of the output module. If the judgment result is positioning failure, the recommended results with 10 times the current candidate number are selected for retrieval backtracking. If the judgment result is successful, the online data preparation mode is entered, and the reference image under the current pose is rendered using the previous positioning result, and the parallax is reduced by iteration to further improve the positioning accuracy.
[0086] Step 202: perform feature matching based on the coarse positioning result and the static image features of the target image to obtain a feature matching result, and determine the position and posture information corresponding to the target image based on the feature matching result.
[0087] In some embodiments, step 202, such as Figure 4 As shown, steps 2021-2023 are included, specifically including:
[0088] Step 2021: Rendering is performed using a renderer based on the coarse positioning result to obtain a rendered reference image;
[0089] Specifically, the rough positioning results are programmed using shaders under the open source OpenGL rendering pipeline, which can achieve a high-resolution real-time rendering speed of 3 milliseconds per frame on ordinary graphics card devices.
[0090] Step 2022: extracting features from the rendered reference image to obtain reference image features, and fusing the reference image features with the static image features to obtain a fused feature image;
[0091] Step 2023: input the feature image into a preset visual positioning algorithm for calculation to obtain the posture information output by the visual positioning algorithm.
[0092] Specifically, the rendered reference image uses the East North Up (ENU) coordinate system to map the XYZ coordinate axes of the world coordinate system, whose origin is the coordinate center of the three-dimensional model, and contains accurate latitude, longitude and altitude information. The purpose of using this coordinate system is to easily convert the three-dimensional coordinates into real geophysical coordinates. For the target image, the camera coordinate system with the camera optical center as the origin is used, where the X axis points to the right side of the image, the Y axis points to the top of the image, and the Z axis is opposite to the camera line of sight.
[0093] To facilitate matrix penalty, the present embodiment adopts homogeneous coordinate representation. Suppose the three-dimensional point X in the world coordinate system is [x w ,y w ,z w ,1] T , its two-dimensional point in the image coordinate system is x=[x c, y c ,-f,1] T , where f is the equivalent focal length of the camera. Then the homogeneous matrix M∈R 4 Represents the pose matrix, which satisfies the following relationship:
[0094]
[0095] Among them: M can be decomposed into two parts, R and t, R∈R 3 It is a unit orthogonal matrix used to represent rotation, t is a three-dimensional vector used to represent translation, and d is the depth, which is obtained from the depth buffer of the rendering pipeline.
[0096] In some embodiments, Figure 5As shown, step 2023 includes:
[0097] Step 301: judging whether the pose information output by the visual positioning algorithm successfully locates the pose information corresponding to the target image based on a preset backtracking condition;
[0098] If yes, output the posture information;
[0099] If not, the process re-enters step 2011 to update the coarse positioning result and perform backtracking according to the backtracking condition.
[0100] Specifically, based on the preset backtracking conditions, it is determined whether the pose information output by the visual positioning algorithm successfully locates the pose information corresponding to the target image, such as Figure 6 As shown, including:
[0101] Step 3011, determining a reprojection error value between a feature point of the rendered reference image and a feature point of the pose information according to a feature matching result of the feature image and pose information corresponding to a current feature matching result;
[0102] Step 3012: Compare the reprojection error value with a preset threshold range to obtain a comparison result;
[0103] Step 3013: Determine whether the position and posture information corresponding to the target image is successfully located based on the comparison result.
[0104] The goal of the visual positioning algorithm is to optimize the pose matrix M so that the error between the feature points on the real image and the feature points on the rendered image is minimized. The reprojection error is calculated using the following formula:
[0105]
[0106] Where: X i is the i-th 3D point, x i is the observed feature point on the corresponding real image, d i is the depth of the point. The main function of this formula is to transform the distant 3D world point into the camera space through the pose, and then compress it to the imaging plane through the depth, and measure the error with the 2D features on the image.
[0107] Specifically, step 3013 includes:
[0108] When the comparison result is that the reprojection error value is less than the minimum value of the preset threshold range, it is determined that the positioning is successful, and the posture information is output;
[0109] When the comparison result shows that the reprojection error value is within the preset threshold range, it is recorded as a backtracking point, and the visual optimization algorithm is iteratively optimized with the currently determined posture information until the reprojection error value is less than the minimum value of the preset threshold range;
[0110] When the comparison result is that the reprojection error value is greater than the maximum value of the preset threshold range, further determining whether the current point is a recorded backtracking point;
[0111] If not, it is determined that the positioning has failed and the iteration ends, and the process loops back to step 2011;
[0112] If so, the visual optimization algorithm is iteratively optimized with the corresponding backtracking point until the reprojection error value is less than the minimum value of the preset threshold range.
[0113] In some examples, such as Figure 8 As shown, the minimum reprojection error threshold is E min , the positioning results less than the threshold will be judged as high confidence results and used as system output to end the iteration. The maximum reprojection error threshold is E max , positioning results greater than this threshold need to be backtracked. If there is no backtracking point recorded before, it is determined that the positioning has failed and the iteration ends, returning to the retrieval recommendation step. Positioning results between the two will be recorded as backtracking points and used for iterative optimization. If the number of system iterations reaches the upper limit T max If a low reprojection error is still not obtained, the iteration is stopped and the result with the lowest confidence is the backtracking point with the smallest reprojection error during the iteration. In actual production tasks, the confidence and number of iterations can be dynamically adjusted according to the real-time requirements of the task and the computing speed of the equipment. The default parameter settings in this system are: E min =1×10 -1 、E min =1×10 -5 and T max =1.
[0114] By combining iterative optimization and feature screening with 3D scene rendering, Robust feature matching is provided for visual positioning to accurately solve the camera pose when the image is taken.
[0115] In this embodiment, the style of the target image is transferred through a multimodal generative model, which can hide sensitive information in the target image and extract features for visual positioning. After using the style transfer image for rough positioning, it is combined with the target image for precise and detailed positioning, which can improve the accuracy of posture information determination and improve the visual positioning effect. At the same time, the edge and grayscale gradient feature matching are fused, and iterative optimization and feature screening are combined with three-dimensional scene rendering to provide Robin feature matching for visual positioning, which is used to accurately solve the camera posture when the image is taken, and improve the visual positioning effect.
[0116] Embodiment 2:
[0117] The visual positioning method based on large model privacy protection of this embodiment can be applied to electronic devices with communication, computing and data storage capabilities. The content of the visual positioning method specifically includes a front-end module for style transfer and a back-end module for visual positioning: in the front-end module, the image is style transferred, sensitive information is hidden, and transmitted to the back-end; in the back-end module, the three-dimensional scene model is rendered into a reference image for feature matching and posture optimization solution, and finally the posture information and three-dimensional visualization results are output. The overall architecture design is as follows: Figure 7 The specific process includes:
[0118] (1) Large Model Style Transfer
[0119] The front-end module is responsible for processing input and performing style transfer on the image. The input data of the system includes the image to be located and the style transfer description prompt, as well as the description of the sensitive information content to be hidden. In order to prevent information leakage, this method does not pass the original image directly to the back-end for feature matching or image comparison, but directly extracts static features at the front end and passes them to the back end. However, the performance of the SIFT descriptor will be reduced in the case of perspective transformation with large environmental changes and parallax. The mainstream algorithm expresses the image by training more complex deep learning descriptors, which reduces the accuracy of feature points and also causes the leakage of front-end data. This paper reduces the environmental impact by adding edge feature descriptions to the back end, and uses rendering iterations to gradually reduce the perspective parallax. Input processing also includes a preprocessing process, which uses a preprocessing algorithm to extract the main constraint information in the image and generate a constraint image as a condition for controlling the large style transfer model.
[0120] The large model style transfer submodule first uses the CLIP large model to process the input text, decomposes the style transfer prompt into positive prompt words, and decomposes the sensitive hidden prompt into negative prompt words. The positive prompt words and negative prompt words are vectorized into text encoding and input into the Stable Diffusion large model. The system also vectorizes the time step information into position encoding and passes it to the Stable Diffusion model to control the denoising and denoising process of the diffusion model. In order to avoid the loss of the main content of the image due to excessive style transfer, the number of iterations is set to 20. This method uses the pre-trained Stable Diffusion 1.5 version large model as the basic model, and uses the pre-trained ControlNet 1.1 version plug-in for conditional constraint control. When used, ControlNet is used as a plug-in network of Stable Diffusion. It has the same network structure as the noise adding part, but ControlNet also receives the edge constraints of the pre-processed image at the same time, and uses the zero convolution layer to connect the denoising network of Stable Diffusion for fine-tuning. The outputs of the two are finally superimposed to output a style transfer image with conditional constraints. In this process, the positive cue words guide the diffusion model to transfer the overall style of the image, hiding the visual cues related to the season and weather in the image. The negative cue words guide the diffusion model to further erase local details such as vehicles and pedestrians to avoid the leakage of sensitive information. The control conditions guide the large model to retain visual cues related to visual positioning accuracy, reducing the loss of positioning accuracy. The front-end module only passes the gradient features extracted by the input processing part and the image after style transfer to the back-end, avoiding the leakage of sensitive information of the input image.
[0121] (2) Visual localization for style transfer images
[0122] The back-end module is mainly used for visual positioning services. After receiving the static image features and style-transferred images transmitted by the front-end module, it estimates the camera pose when the image was taken and outputs it to typical downstream applications such as digital twins. Visual positioning is the main core submodule of the back-end. Its main function is to query the database for similar reference images for a given image and solve the camera pose through feature matching. In the problem of visual positioning, the image to be positioned often undergoes feature changes due to factors such as lighting, season, and human activities, and there are apparent differences with the existing reference images in the database. The apparent difference between the image after privacy-preserving style transfer and the database image will be further amplified, affecting the positioning success rate and positioning accuracy. In view of the large difference in the appearance of the front-end and back-end images, this embodiment proposes a visual positioning algorithm for style-transferred images, which adopts strategies such as retrieval, rendering, iteration, and backtracking to solve the image camera pose from coarse to fine.
[0123] 21) Retrieval of coarse positioning
[0124] In the back-end visual positioning architecture of the method in this paper, the image transmitted by the front end must first go through the retrieval coarse positioning process to determine its initial position for the subsequent precise positioning process. The main process of retrieval positioning is to compress and vectorize the front-end image through a convolutional neural network, and retrieve the most similar candidate image from the back-end reference database as the initial positioning result. The mainstream algorithm uses real images captured by the camera to prepare the reference database. This kind of data often cannot cover the vast urban space and lacks accurate pose annotations. This paper uses a 3D reconstruction model of the urban scene to traverse the urban space for rendering through a renderer, generate virtual images to build a reference database, and use it to cover the target space.
[0125] In this method, the reference database includes two types of data: offline data and online data, both of which are generated by rendering the 3D scene model by the rendering engine. Offline data can be prepared in advance based on the reconstruction of the 3D scene, and the low-altitude range of the target city scene is traversed and rendered by the renderer at fixed intervals. The traversal interval varies according to the size of the city scene. In this paper, the traversal intervals set for different scene environments are shown in Table 1.
[0126] Tab.1Interval for offline data generation in different outdoor scenes
[0127] Scene Type Main Target Horizontal interval (m) Vertical separation (m) Yaw angle interval (°) Pitch angle interval (°) Urban area High-rise 100 30 60 30 village Houses 60 20 30 15 Community the way 30 10 15 10
[0128] In Table 1, for large-scale monitoring scenarios dominated by high-rise buildings such as city centers, the offline data collection interval is large; for densely populated areas such as communities, the offline data collection interval is small; and the collection density of scenes such as villages and urban-rural fringe areas is between the two. This setting is mainly distinguished by the different layout methods of monitoring cameras in different scenes.
[0129] After the offline data preparation is completed, the NetVLAD algorithm is used to compress the image into a 4096-bit retrieval descriptor. When the front-end transmits the image to be positioned to the back-end, the visual positioning module calls the NetVLAD algorithm to compress the front-end image into a 4096-bit descriptor, and retrieves and recommends the most similar candidate reference image in the offline database. The recommendation results are sorted according to the cosine similarity, and the calculation formula is:
[0130]
[0131] Where: p and q are descriptor vectors. Due to environmental changes in outdoor scenes, there are still cases of recommendation errors even without style transfer. Therefore, the visual positioning module selects the top 10 recommended results for feature matching in the first pose calculation and adjusts them according to the judgment result of the output module. If the judgment result is positioning failure, the recommended results 10 times the current number of candidates are selected for retrieval backtracking. If the judgment result is successful, it enters the online data preparation mode, uses the previous positioning result to render the reference image in the current pose, and reduces the parallax by iteration to further improve the positioning accuracy.
[0132] 22) Feature matching and precise positioning
[0133] Performing feature matching positioning on the results of retrieval of coarse positioning can improve the visual positioning accuracy to the sub-pixel level. The process first needs to call the renderer to render the scene graph according to the recommended results of retrieval of coarse positioning, and use the scene graph information as a reference image to match the front-end image to obtain feature matching information for feature positioning. Since the renderer can call the vertex and patch information of the three-dimensional scene model, it can generate accurate color and depth information pixel by pixel. This embodiment uses shader programming under the open source OpenGL rendering pipeline, and can achieve a high-resolution real-time rendering speed of 3 milliseconds per frame on an ordinary graphics card device.
[0134] In this embodiment, the virtual scene uses the East North Up (ENU) coordinate system to map the XYZ coordinate axis of the world coordinate system. The origin is the coordinate center of the three-dimensional model, which contains accurate longitude, latitude and altitude information. The purpose of using this coordinate system is to easily convert the three-dimensional coordinates into real geophysical coordinates. For surveillance video images, a camera coordinate system with the optical center of the camera as the origin is used, in which the X-axis points to the right side of the image, the Y-axis points to the top of the image, and the Z-axis is opposite to the camera's line of sight. To facilitate matrix penalty, this article uses homogeneous coordinate representation. Suppose the three-dimensional point X in the world coordinate system is [x w ,y w ,z w ,1] T , its two-dimensional point in the image coordinate system is x=[x c ,y c ,-f,1] T , where f is the equivalent focal length of the camera. Then the homogeneous matrix M∈R 4 Represents the pose matrix, which satisfies the following relationship
[0135]
[0136] Among them: M can be decomposed into two parts, R and t, R∈R 3is a unit orthogonal matrix used to represent rotation, and t is a three-dimensional vector used to represent translation. d is the depth, which is obtained from the depth buffer of the rendering pipeline. The goal of the visual positioning algorithm is to optimize the pose matrix M so that the error between the feature points on the real image and the feature points on the rendering image is minimized. The reprojection error is calculated using the following formula:
[0137]
[0138] Where: X i is the i-th 3D point, x i is the observed feature point on the corresponding real image, d i is the depth of the point. The main function of this formula is to transform the distant 3D world point into the camera space through the pose, and then compress it to the imaging plane through the depth, and measure the error with the 2D features on the image. In this paper, in order to eliminate the impact of image size differences on positioning evaluation, the pixel coordinates will be normalized, and half of the image height will be used as the unit 1 for normalization. In the feature extraction part, this paper uses the scale-invariant feature transform (SIFT) to extract grayscale gradient features. The main reason is that the SIFT descriptor will not change due to image scaling and rotation, and the feature extraction and feature matching processes of the SIFT algorithm are separate. In order to avoid privacy leakage, the system can only pass the compressed feature descriptor to the backend after extracting static features at the front end. At the backend, during the feature fusion process of each iteration step, the static features of the front end remain unchanged, and the dynamic features of the back end change with the different results of each iteration, and are fused and matched with the front end features. In the actual detection process, there will always be bad pixels in the feature detection results. In order to exclude features with large errors, this paper uses the RANSAC-PNP algorithm to filter features by excluding bad pixels and optimize the pose calculation results.
[0139] 23) Iterative optimization and backtracking
[0140] Different from the mainstream visual positioning algorithm, the positioning framework proposed in this embodiment includes iteration and backtracking, that is, based on each positioning result, the positioning is judged whether it is successful according to the reprojection error. If successful, the positioning result is further optimized through iteration, otherwise it is backtracked to the previous result. Figure 8 As shown in Figure 1, the visual positioning algorithm uses the initial pose to render the reference image for feature matching positioning. The initial pose can be guessed based on the task context, or the rough positioning result can be retrieved. The feature matching algorithm outputs the matching result and the estimated pose, and calculates the average reprojection error according to formula (3). The system sets different iteration and backtracking strategies based on the reprojection error. Among them, the minimum reprojection error threshold is E min, the positioning results less than the threshold will be judged as high confidence results and used as system output to end the iteration. The maximum reprojection error threshold is E max , positioning results greater than this threshold need to be backtracked. If there is no backtracking point recorded before, it is determined that the positioning has failed and the iteration ends, returning to the retrieval recommendation step. Positioning results between the two will be recorded as backtracking points and used for iterative optimization. If the number of system iterations reaches the upper limit T max If a low reprojection error is still not obtained, the iteration is stopped and the result with the lowest confidence is the backtracking point with the smallest reprojection error during the iteration. In actual production tasks, the confidence and number of iterations can be dynamically adjusted according to the real-time requirements of the task and the computing speed of the equipment. The default parameter settings in this system are: E min =1×10 -1 、E min =1×10 -5 and T max =10.
[0141] This embodiment performs visual positioning by combining a multi-feature fusion feature matching algorithm. This method is suitable for visual positioning problems in complex surveillance video images with large environmental changes such as indoors and outdoors. The stable diffusion model combined with edge perception can retain some visual features for positioning while protecting privacy. This method can hide scene-sensitive information while maintaining positioning accuracy.
[0142] Embodiment three:
[0143] Another embodiment of the present application relates to a visual positioning device. The implementation details of the visual positioning device of this embodiment are specifically described below. The following content is only for the convenience of understanding the implementation details provided, and is not necessary for the implementation of this solution. The schematic diagram of the visual positioning device of this embodiment can be as follows Fig. 9 As shown, it includes a front-end processing module 801 and a back-end positioning module 802. The front-end processing module 801 includes:
[0144] The data acquisition unit 8011 is used to acquire a target image and perform feature extraction on the target image to obtain static image features of the target image;
[0145] The style transfer unit 8012 is used to input the target image into the pre-trained multimodal generation model for style transfer processing to obtain a style transferred image after hiding the sensitive information of the target image;
[0146] The back-end positioning module 802 includes:
[0147] A retrieval coarse positioning unit 8021 is used to perform retrieval coarse positioning in a pre-built reference database based on the style migration image, and determine a candidate reference image that is most similar to the style migration image as a coarse positioning result;
[0148] The precise positioning unit 8022 is used to perform feature matching based on the rough positioning result and the static image features of the target image to obtain a feature matching result, and determine the position and posture information corresponding to the target image based on the feature matching result.
[0149] It is worth mentioning that all modules involved in this embodiment are logic modules. In practical applications, a logic unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. In addition, in order to highlight the innovative part of this application, this embodiment does not introduce units that are not closely related to solving the technical problems proposed by this application, but this does not mean that there are no other units in this embodiment.
[0150] Embodiment 4:
[0151] Another embodiment of the present application relates to an electronic device, such as Fig.10 As shown, it includes: at least one processor 901; and a memory 902 that is communicatively connected to the at least one processor 901; wherein the memory 902 stores instructions that can be executed by the at least one processor 901, and the instructions are executed by the at least one processor 901 so that the at least one processor 901 can execute the visual positioning method in the above-mentioned embodiments.
[0152] Among them, the memory and the processor are connected in a bus manner, and the bus may include any number of interconnected buses and bridges, and the bus connects various circuits of one or more processors and memories together. The bus can also connect various other circuits such as peripherals, voltage regulators, and power management circuits, which are well known in the art and are therefore not further described herein. The bus interface provides an interface between the bus and the transceiver. The transceiver can be one element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices on a transmission medium. The data processed by the processor is transmitted on a wireless medium via an antenna, and further, the antenna also receives data and transmits the data to the processor.
[0153] The processor is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. Memory can be used to store data used by the processor when performing operations.
[0154] Embodiment five:
[0155] Another embodiment of the present application relates to a computer-readable storage medium, such as Fig.11 As shown, a computer program 31 is stored. When the computer program 31 is executed by a processor, the above method embodiment is implemented.
[0156] That is, those skilled in the art can understand that all or part of the steps in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a program, and the program is stored in a storage medium, including a number of instructions to enable a device (which can be a single-chip microcomputer, chip, etc.) or a processor to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, referred to as: ROM), random access memory (Random Access Memory, referred to as: RAM), disk or optical disk and other media that can store program codes.
[0157] Those skilled in the art will appreciate that the above embodiments are specific embodiments for implementing the present application, and in actual applications, various changes may be made thereto in form and detail without departing from the spirit and scope of the present application.
Claims
1. A visual positioning method, characterized in that: It includes the front-end processing stage and the back-end positioning stage; The front-end processing stage includes: Get the target image; Performing feature extraction on the target image to obtain static image features of the target image; Inputting the target image into a pre-trained multimodal generative model for style transfer processing to obtain a style-transferred image after hiding sensitive information of the target image; The backend positioning stage includes: Retrieving and roughly locating the style transfer image in a pre-built reference database, and determining a candidate reference image that is most similar to the style transfer image as a rough positioning result; Feature matching is performed based on the coarse positioning result and the static image features of the target image to obtain a feature matching result, and the position and posture information corresponding to the target image is determined based on the feature matching result.
2. A visual positioning method according to claim 1, characterized in that: The front-end processing stage also includes: obtaining style transfer prompts and sensitive information hiding prompts; The step of inputting the target image into a pre-trained multimodal generative model to perform style transfer processing on the target image to obtain a style transfer image after hiding sensitive information from the target image comprises: Decomposing and vectorizing the acquisition style transfer prompt and the sensitive information hiding prompt into a text encoding of the multimodal generation model; The text encoding of the converted multimodal generative model and the input of the target image into the pre-trained multimodal generative model are subjected to style transfer processing to obtain a style transferred image after hiding sensitive information of the target image.
3. A visual positioning method according to claim 2, characterized in that: The multimodal generation model includes a Stable Diffusion model, in which a pre-trained ControlNet plug-in network is inserted to provide conditional constraint control for the multimodal generation model.
4. A visual positioning method according to claim 1, characterized in that: Each reference image in the reference database is generated by traversing and rendering the target space using a renderer using a three-dimensional reconstruction model.
5. A visual positioning method according to claim 1, characterized in that: The method of performing a search and rough positioning based on the style transfer image in a pre-built reference database and determining a candidate reference image most similar to the style transfer image as a rough positioning result includes: Based on the style transfer image, a search is performed in a pre-built reference database to obtain a plurality of candidate reference images that are most similar to the style transfer image as search recommendation results; The retrieval recommendation results are sorted from high to low by similarity using a cosine similarity calculation formula to determine the rough positioning result.
6. A visual positioning method according to claim 1, characterized in that: The step of performing feature matching based on the coarse positioning result and the static image feature of the target image to obtain a feature matching result, and determining the position and posture information corresponding to the target image based on the feature matching result includes: Rendering is performed using a renderer based on the coarse positioning result to obtain a rendered reference image; Extracting features from the rendered reference image to obtain reference image features, and fusing the reference image features with the static image features to obtain a fused feature image; The feature image is input into a preset visual positioning algorithm for calculation to obtain the posture information output by the visual positioning algorithm.
7. A visual positioning method according to claim 6, characterized in that: The step of inputting the feature image into a preset visual positioning algorithm for calculation to obtain the position and posture information output by the visual positioning algorithm includes: Determining whether the pose information output by the visual positioning algorithm successfully locates the pose information corresponding to the target image based on a preset backtracking condition; If not, re-enter the step according to the backtracking condition to perform a search and rough positioning based on the style transfer image in a pre-built reference database, determine a candidate reference image most similar to the style transfer image as a rough positioning result, and update the rough positioning result.
8. A visual positioning method according to claim 7, characterized in that: The determining, based on a preset backtracking condition, whether the pose information output by the visual positioning algorithm successfully locates the pose information corresponding to the target image includes: Determine a reprojection error value between a feature point of the rendered reference image and a feature point of the pose information according to a feature matching result of the feature image and pose information corresponding to a current feature matching result; Comparing the reprojection error value with a preset threshold range to obtain a comparison result; Whether the position and posture information corresponding to the target image is successfully located is determined according to the comparison result.
9. A visual positioning method according to claim 8, characterized in that: The determining, according to the comparison result, whether the position and posture information corresponding to the target image is successfully located includes: When the comparison result is that the reprojection error value is less than the minimum value of the preset threshold range, it is determined that the positioning is successful, and the posture information is output; When the comparison result shows that the reprojection error value is within the preset threshold range, it is recorded as a backtracking point, and the visual optimization algorithm is iteratively optimized with the currently determined posture information until the reprojection error value is less than the minimum value of the preset threshold range; When the comparison result is that the reprojection error value is greater than the maximum value of the preset threshold range, it is further determined whether the current point is a recorded backtracking point. If not, the positioning is determined to have failed and the iteration is terminated, and the process is re-circulated to the step of retrieving rough positioning based on the style transfer image in a pre-built reference database. If so, the visual optimization algorithm is iteratively optimized with the corresponding backtracking point until the reprojection error value is less than the minimum value of the preset threshold range.
10. A visual positioning device, characterized in that: include: Front-end processing module and back-end positioning module; The front-end processing module comprises: A data acquisition unit, used to acquire a target image and perform feature extraction on the target image to obtain static image features of the target image; A style transfer unit, used to input a target image into a pre-trained multimodal generative model for style transfer processing, so as to obtain a style-transferred image after hiding sensitive information of the target image; The back-end positioning module includes: A retrieval coarse positioning unit, used to perform retrieval coarse positioning in a pre-built reference database based on the style migration image, and determine a candidate reference image that is most similar to the style migration image as a coarse positioning result; The precise positioning unit is used to perform feature matching based on the rough positioning result and the static image features of the target image to obtain a feature matching result, and determine the position and posture information corresponding to the target image based on the feature matching result.
11. An electronic device, characterized in that: include: at least one processor; as well as, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the visual positioning method according to any one of claims 1 to 8.
12. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the visual positioning method according to any one of claims 1 to 8 is implemented.