An indoor visual positioning method and positioning system based on an occlusion removal algorithm
By combining object detection, image repair and feature matching technologies in the indoor visual positioning system, the problem of occlusion objects affecting positioning accuracy is solved, and higher positioning accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202411638966.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-18
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2044-11-18
AI Technical Summary
In indoor positioning technology, the impact of blocking objects on the camera's field of view leads to incomplete image information, reduces positioning accuracy, and may even lead to the loss of key feature points.
The indoor visual positioning method based on the occlusion and clearing algorithm is adopted to clear the occlusion objects, restore the complete environmental information through object detection, image repair and feature matching technology, and image repair is performed using the image repair model MAT.
Improve the accuracy and robustness of positioning, ensure accurate indoor positioning can be achieved in complex occlusion environments, and improve the overall performance of the visual positioning system.
Smart Images

Figure CN119149765B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of occluded image processing, image restoration, and indoor positioning, and particularly relates to an indoor visual positioning method and positioning system based on an occlusion removal algorithm. Background Art
[0002] Indoor positioning technology has important applications in modern smart homes, security monitoring, robot navigation, and other fields. With the development of technology, vision-based positioning systems have gradually become a research hotspot due to their advantages such as low cost, convenient installation, and rich information. These methods obtain environmental images through cameras and use image features for positioning and navigation. Compared with traditional methods based on sensors such as odometers and lidar, vision-based positioning systems can not only provide more abundant environmental information but also achieve higher positioning accuracy.
[0003] However, due to the presence of a large number of occluding objects in the indoor environment, such as furniture and decorations, these occluding objects will affect the field of view of the camera, resulting in incomplete image information being obtained, which will cause great damage to the positioning accuracy and may even cause the loss of key feature points. In this case, traditional vision-based positioning algorithms cannot effectively process occlusion information, resulting in an increase in positioning error or even inability to position. In addition, occluding objects may also cause misjudgment of vision algorithms, misidentifying occluding objects as environmental features, further exacerbating the difficulty of positioning. Therefore, how to effectively process occluded images and restore complete environmental information has become an important research direction for improving indoor positioning accuracy. Summary of the Invention
[0004] The technical problem to be solved by the present invention is: to overcome the deficiencies of the prior art and provide an indoor visual positioning method and positioning system based on an occlusion removal algorithm that combines object detection, image restoration, and feature matching technologies, improving positioning accuracy and robustness.
[0005] The technical solution adopted by the present invention to solve its technical problems is: the indoor visual positioning method based on the occlusion removal algorithm is characterized by including the following steps:
[0006] Step 1, query the images in the database, extract the feature information of the database images, and save the feature information;
[0007] Step 2, perform occluder detection and remove all occluders in the query images;
[0008] Step 3, use an image restoration model to perform image restoration and restore the query images;
[0009] Step 4, perform indoor positioning.
[0010] Preferably, when performing step 2, it includes: constructing an occluded image dataset; performing occluder detection based on the occluded image dataset; and resetting pixels to eliminate occluders in the query image.
[0011] Preferably, when constructing the occluded image dataset: in the coco128 dataset, obtain several objects by means of image matting. The image matting process involves separating the object of interest from the image to obtain an object image with a transparent background; then, generate the mask images of these objects by retaining the pixels of the objects and setting the pixel values of the remaining positions to (255, 255, 255).
[0012] Load the images of the 7Scenes dataset; then, load the images of all occluding objects and their corresponding masks; randomly select an occluding object and its mask for each query and database image, add random types and degrees of blur information to it, and then place the occluding object with added blur information at a random position on the image at a random scale and rotation angle to generate occluded images. Finally, the occluded images generated one-to-one for all images constitute the occluded image dataset.
[0013] Preferably, the steps of performing occluder detection based on the occluded image dataset are: first, read the bounding box information in the detection results, extract the pixel region of the occluding object, load the YOLOv5 model pre-trained on the COCO dataset, detect the query image, and generate the bounding box and class information containing the occluding object to complete the detection of the occluder.
[0014] Preferably, the process of resetting pixels to eliminate occluders in the query image is: reset the pixels in the occluding object region to pure black pixels (0, 0, 0) to represent the position of the occluding object; at the same time, generate a mask image corresponding to the query image, in which the pixels in the occluding object region are set to pure black pixels and the pixels in other regions are set to pure white pixels (255, 255, 255).
[0015] Preferably, step 4 includes the following steps:
[0016] Step 4-1, extract and save the feature information of the repaired query image in step 3;
[0017] Step 4-2, perform image retrieval; according to the database image feature information and the query image feature information, use the image retrieval algorithm to perform nearest neighbor retrieval of the query image to obtain the nearest neighbor retrieval pairs of the query image;
[0018] Step 4-3, according to the database image feature information, the query image feature information and the nearest neighbor retrieval pairs, use the LightGlue algorithm to perform feature matching and establish the 2D-2D matching relationship between the query image and the database image;
[0019] Step 4-4: Use the indoor visual positioning algorithm based on the essential matrix; based on the 2D-2D matching information, use the five-point method in the RANSAC loop to calculate the essential matrix of the image pair; subsequently, through the epipolar constraint condition, adopt the singular value decomposition and RANSAC algorithm, combine the known absolute pose and essential matrix of the database image, and finally determine the positioning result.
[0020] Preferably, when performing Step 1, first set reference points in the indoor environment and collect image data at the reference points; secondly, use the SuperPoint algorithm and NetVLAD algorithm to extract features from the database images; finally, store the extracted feature information and the corresponding coordinate information to form a feature file.
[0021] Preferably, in Step 3, use the image inpainting model MAT for image inpainting.
[0022] An indoor visual positioning system based on an occlusion removal algorithm, characterized in that it includes an offline unit and an online unit, wherein the offline unit includes:
[0023] An offline image acquisition module, used to query images in the database;
[0024] An offline feature extraction module, connected to the offline image acquisition module, used to extract the feature information of the database images and save the feature information;
[0025] The online unit includes:
[0026] An online image acquisition module, used to obtain query images;
[0027] An occlusion removal module, connected to the online image acquisition module, used to detect occlusions and remove all occlusions in the query images;
[0028] An indoor positioning module, connected to the occlusion removal module and the offline feature extraction module, used to perform indoor positioning.
[0029] Preferably, the indoor positioning module includes:
[0030] A feature extraction sub-module, used to obtain the feature information of the repaired query images;
[0031] An image retrieval sub-module, connected to the feature extraction sub-module, used to perform a nearest neighbor search of the query images using an image retrieval algorithm to obtain the nearest neighbor search pairs of the query images;
[0032] A feature matching sub-module, connected to the image retrieval sub-module, used to establish a 2D-2D matching relationship between the query images and the database images;
[0033] And a positioning sub-module, based on the positioning algorithm of the essential matrix, to obtain the positioning result.
[0034] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0035] In the present indoor visual positioning method and positioning system based on the occlusion removal algorithm, by combining object detection, image inpainting, and feature matching technologies, the positioning accuracy and robustness are improved.
[0036] By detecting and removing occlusions in all query images, it ensures that the pixels of occluded objects do not interfere with subsequent processing, and at the same time provides a clear target area for the image restoration process. In this way, occluded objects can be effectively isolated and strong support is provided for subsequent image inpainting algorithms.
[0037] When performing image inpainting, the image inpainting model MAT (Mask-Aware Transformer for Large Hole Image Inpainting) is adopted. This model combines the advantages of Transformer and convolution. By introducing conditional long-range interactions, it can effectively reconstruct image details in the presence of occlusions, thereby improving the accuracy and robustness of the visual positioning system.
[0038] Through the application of the MAT algorithm, occluded images are effectively inpainted, which is not only more visually coherent but also provides high-quality input data for subsequent visual positioning tasks. The MAT algorithm performs excellently in dealing with complex occlusions and high-resolution images, providing important support for the high precision and high robustness of the visual positioning system.
[0039] In the present indoor visual positioning method based on the occlusion removal algorithm, the SuperPoint feature extraction algorithm and the LightGlue feature matching algorithm are combined. This combination greatly improves the robustness and accuracy of feature detection and matching. After matching feature points, the five-point method and the RANSAC algorithm are used to estimate the essential matrix. Among them, the five-point method is an effective non-linear method that derives the specific form of the essential matrix by solving the geometric constraints of five pairs of matching points. The RANSAC algorithm estimates the essential matrix by iteratively randomly selecting subsets from the feature matching set, and uses consistency checks to eliminate incorrect matches, improving the robustness and accuracy of the estimation. After obtaining the essential matrix, the relative rotation and translation are obtained by performing singular value decomposition on it, and the decomposed rotation matrix and translation vector will be further used to calculate the absolute pose of the query image. When calculating the absolute pose, the epipolar constraint is introduced to ensure the accuracy of the estimation result.
[0040] The essential matrix can effectively fuse multi-view information, improving the accuracy of 3D reconstruction and positioning. Indoor environments usually have many occlusions and unstructured features, and single-view methods are difficult to provide sufficient constraint conditions. In contrast, multi-view methods can provide more constraint conditions through multi-angle view constraints, thereby significantly reducing the positioning error and improving the accuracy and robustness of positioning. In an indoor visual positioning system, this algorithm complements the image inpainting method to jointly improve the positioning accuracy and robustness of the system, ensuring accurate indoor positioning can still be achieved in complex occlusion environments. Description of the Drawings
[0041] Figure 1 It is a flowchart of an indoor visual positioning method based on an occlusion removal algorithm.
[0042] Figure 2 It is a schematic block diagram of the principle of an indoor visual positioning system based on an occlusion removal algorithm.
[0043] Figure 3 It is a curve graph of the cumulative distribution function (CDF) of the positioning error in the chess scene of the 7Scenes public dataset.
[0044] Figure 4 It is a curve graph of the cumulative distribution function (CDF) in the corridor1 scene.
[0045] Figure 5 It is a curve graph of the cumulative distribution function (CDF) in the corridor2 scene.
[0046] Figure 6 It is a curve graph of the cumulative distribution function (CDF) in the corridor3 scene. Detailed Implementation Manner
[0047] Figures 1 to 6 This is the best embodiment of the present invention. The following further describes the present invention with reference to the attached Figures 1 to 6 drawings.
[0048] As Figure 1 shown, an indoor visual positioning method based on an occlusion removal algorithm includes the following steps:
[0049] Step 1, extracting database image features;
[0050] Use the SuperPoint algorithm and the NetVLAD algorithm to extract the feature information of the database images, and save the feature information for use in image matching and positioning.
[0051] First, set reference points in the indoor environment and collect image data at the reference points. Secondly, extract features from the database images. Since the feature information required by the image retrieval module and the pose estimation module is different, two different feature extraction algorithms are used for processing (SuperPoint and NetVLAD). Finally, the extracted feature information and the corresponding coordinate information are stored to form a feature file for subsequent image matching and positioning.
[0052] Step 2, occluded object detection;
[0053] To study the impact of image occlusion on the visual positioning system, an occluded image dataset was constructed in this indoor visual positioning system based on the occlusion removal algorithm. First, 15 types of occluding objects, a total of 363 objects, were selected. They were obtained by cropping in the coco128 dataset. The cropping process involved separating the object of interest from the image to obtain an object image with a transparent background. Then, the mask image of these objects was generated by retaining the pixels of the occluding object and setting the pixel values of the remaining positions to (255, 255, 255).
[0054] When generating the occluded image dataset, first load the images of the 7Scenes dataset. Then, load the images of all occluding objects and their corresponding masks. Next, randomly select an occluding object and its mask for each query and database image, and add random types and degrees of blur information, including no blur, box blur, Gaussian blur, and motion blur. Then, place the occluding object with added blur information at a random position on the image at a random scale and rotation angle to generate occluded images. Finally, the occluded images generated one-to-one for all images constitute the occluded image dataset.
[0055] During the occluded object detection process, the following steps are included:
[0056] Step 2-1, read the bounding box information in the detection result and extract the pixel region of the occluding object.
[0057] Load the YOLOv5 model pre-trained on the COCO dataset and further fine-tune it on a specific occluding object dataset. Then use the fine-tuned model to detect the query image and generate the bounding box and class information containing the occluding object. After detecting the occluding object, the next step is to reset the pixels in the occluding object region for subsequent image restoration and visual positioning processing.
[0058] Step 2-2: Reset the pixels in this area to pure black pixels (0, 0, 0) to represent the position of the occluding object. At the same time, generate a mask image corresponding to the query image. In this mask image, the pixels in the occluding object area are set to pure black pixels, and the pixels in other areas are set to pure white pixels (255, 255, 255).
[0059] Through this step, it is ensured that the pixels of the occluding object will not interfere with subsequent processing, and at the same time, a clear target area is provided for the image restoration process. In this way, the occluding object can be effectively isolated, and strong support is provided for subsequent image inpainting algorithms.
[0060] Step 3: Image inpainting
[0061] To solve the problem of large-area pixel loss and support diverse generation, the image inpainting model MAT (Mask-Aware Transformer for Large Hole Image Inpainting) is adopted. This model combines the advantages of Transformer and convolution. By introducing conditional long-range interaction, it can effectively reconstruct image details in the presence of occlusion, thereby improving the accuracy and robustness of the visual positioning system.
[0062] Through the application of the MAT algorithm, the occluded image is effectively restored, not only being more visually coherent but also providing high-quality input data for subsequent visual positioning tasks. The MAT algorithm performs excellently in dealing with complex occlusions and high-resolution images, providing important support for the high precision and high robustness of the visual positioning system.
[0063] Step 4: Indoor positioning
[0064] In the process of performing this step for indoor positioning, the following steps are included:
[0065] Step 4-1: Extract the feature information of the query image. Perform feature extraction on the query image after occlusion removal. Use the SuperPoint algorithm and NetVLAD algorithm again to extract and save the feature information of the restored query image for use in image matching and positioning.
[0066] Step 4-2: Conduct image retrieval. According to the feature information of the database images and the query image, use the image retrieval algorithm to perform a nearest neighbor search for the query image to obtain the nearest neighbor search pairs of the query image.
[0067] In the process of image retrieval, global feature vectors of the images are generated with the help of algorithms such as CNN, and the similarity between each image in the database and them is calculated using a similarity metric function. Based on these similarities, the pairing relationships between all query images and database images are obtained.
[0068] Using the Patch-NetVLAD algorithm, by calculating the similarity scores between image pairs, to measure their spatial and appearance consistency, and finally obtain the image pairs composed of the query image and its top-k nearest neighbor images. This algorithm reduces the computational cost brought by the cross-matching of block features to a great extent while ensuring the retrieval accuracy.
[0069] Step 4-3, perform feature matching;
[0070] According to the database image feature information, query image feature information and nearest neighbor retrieval pairs, use the LightGlue algorithm to perform feature matching and establish the 2D-2D matching relationship between the query image and the database image.
[0071] Step 4-4, perform indoor positioning;
[0072] Utilize the indoor visual positioning algorithm based on the essential matrix, which uses the geometric relationship between multiple viewpoints in the image to achieve accurate 3D reconstruction and positioning. Based on the 2D-2D matching information, use the five-point method in the RANSAC loop to calculate the essential matrix of the image pair; subsequently, through the epipolar constraint condition, adopt the singular value decomposition (SVD) and RANSAC algorithms, combined with the known absolute pose and essential matrix of the database image, and finally determine the positioning result of the query camera.
[0073] In this indoor visual positioning method based on the occlusion removal algorithm, the SuperPoint feature extraction algorithm and the LightGlue feature matching algorithm are combined, and this combination greatly improves the robustness and accuracy of feature detection and matching. After matching the feature points, the five-point method and the RANSAC algorithm are used to estimate the essential matrix. Among them, the five-point method is an effective non-linear method, which derives the specific form of the essential matrix by solving the geometric constraints of five pairs of matching points. The RANSAC algorithm estimates the essential matrix by iteratively randomly selecting subsets from the feature matching set, and uses the consistency check to eliminate the wrong matches, improving the robustness and accuracy of the estimation. After obtaining the essential matrix, perform singular value decomposition on it to obtain the relative rotation and translation, and the decomposed rotation matrix and translation vector will be further used to calculate the absolute pose of the query image. When calculating the absolute pose, introduce the epipolar constraint to ensure the accuracy of the estimation result.
[0074] The essential matrix can effectively fuse multi-view information, improving the accuracy of 3D reconstruction and positioning. Indoor environments usually have a lot of occlusions and unstructured features. Single-view methods are difficult to provide sufficient constraints, while multi-view methods can provide more constraints through multi-angle view constraints, thus significantly reducing the positioning error and improving the accuracy and robustness of positioning. In an indoor vision positioning system, this algorithm complements the image inpainting method to jointly improve the positioning accuracy and robustness of the system, ensuring accurate indoor positioning can still be achieved in complex occlusion environments.
[0075] As Figure 2 shown, a positioning system for implementing the above indoor vision positioning method based on the occlusion removal algorithm includes an offline unit and an online unit. In the offline unit, there are an offline image acquisition module and an offline feature extraction module. The output of the offline image acquisition module is connected to the input of the offline feature extraction module, and the above-mentioned step 1 is executed through the offline image acquisition module and the offline feature extraction module.
[0076] In the online unit, there are an online image acquisition module, an occlusion removal module, an image inpainting module, and an indoor positioning module. The online image acquisition module is connected to the occlusion removal module, the occlusion removal module is connected to the image inpainting module, the image inpainting module is connected to the indoor positioning module, and the offline feature extraction module is also connected to the indoor positioning module.
[0077] The above-mentioned step 4-1 in step 4 is executed through the online image acquisition module. An occlusion object detection sub-module and a pixel reset sub-module are set in the occlusion removal module, which are respectively used to implement the above-mentioned steps 2-1 to 2-2. The image inpainting module is used to implement the above step 3, and the pixel reset sub-module and the image inpainting module are both connected to the indoor positioning module.
[0078] In the indoor positioning module, there are a feature extraction sub-module, an image retrieval sub-module, a feature matching sub-module, and a positioning sub-module. The feature extraction sub-module, the image retrieval sub-module, the feature matching sub-module, and the positioning sub-module respectively execute the above steps 4-1 to 4-4. Finally, the positioning algorithm based on the essential matrix is executed through the positioning sub-module to obtain the positioning result.
[0079] Through an example, the above indoor vision positioning method and positioning system based on the occlusion removal algorithm are further verified:
[0080] Using the datasets of 7Scenes_occ, 7Scenes_disocc, and 7Scenes, the localization results of the system include the median absolute position error (unit: m) and median absolute rotation error (unit: °) for each scene, as well as the average median localization error for all scenes (including the chess, fire, heads, office, pumpkin, redkitchen, and stairs scenes). The error thresholds (0.25 m, 2°), (0.5 m, 5°), and (5 m, 10°) set by predecessors are adopted, and the passing rates of the query images under each threshold are output.
[0081] Table 1 shows the ablation experiments of the occlusion removal module. The experiments are divided into three groups: The first group uses the 7Scenes_occ dataset, turns off the occlusion removal module, and directly performs localization; the second group uses the 7Scenes_occ dataset, turns on the occlusion removal module, and performs localization after occlusion removal processing; the third group uses the 7Scenes dataset. Although there are no occluded objects in the dataset, the occlusion removal module is still turned on to ensure the fairness of the experiment. In addition, the image retrieval algorithm used in all three groups of experiments is DenseVLAD. As shown in Table 1, the results of the second group and the third group are similar, while the first group has poor localization accuracy due to occlusion. Therefore, directly using occluded images will affect the localization accuracy. The 7Scenes_occ dataset after occlusion removal shows a similar localization accuracy to the 7Scenes dataset, proving the effectiveness of the occlusion removal module. Figure 3 It shows the cumulative distribution function (CDF) curve of the localization error in the chess scene, proving that the occlusion removal module significantly improves the localization accuracy.
[0082] Table 1 Ablation Experiments of the Occlusion Removal Module
[0083]
[0084] Note: In the first (), the first value is the median absolute position error, unit: m; the second value is the median absolute rotation error, unit: (°). In the second (), the three percentages are the passing rates of the query images under each threshold, unit: (%).
[0085] To verify the improved indoor visual positioning algorithm proposed by the present invention, self-made LH data is used to conduct experimental comparisons with a variety of advanced algorithms. The algorithms participating in the comparison are as follows: 1) The five-point algorithm based on SIFT (abbreviated as sift_5pt), which has achieved excellent positioning results on the 7Scenes dataset; 2) The algorithm based on efficient image retrieval (abbreviated as retrieval), which designs a fast retrieval algorithm and has achieved good positioning results on 7Scenes by combining deep learning features; 3) The proposed algorithm based on SuperPoint and SuperGlue (abbreviated as sp+sg_5pt), which has more advantages in positioning accuracy and has become an advanced method in the current indoor positioning field.
[0086] On this basis, the indoor visual positioning method based on the occlusion removal algorithm proposes an improved positioning algorithm based on SuperPoint and LightGlue (abbreviated as sp+lg_5pt). Attached Figures 4 to 6 The cumulative distribution function (CDF) curves of these algorithms in three scenarios of the LH dataset, namely corridor 1 (corridor1), corridor 2 (corridor2), and corridor 3 (corridor3), are respectively shown. In addition, the quantitative evaluation results of the four algorithms are shown in Table 2. According to Figures 4 to 6 and the results in Table 2, it can be clearly seen that the positioning algorithm based on SuperPoint+LightGlue proposed by the indoor visual positioning method based on the occlusion removal algorithm shows higher positioning accuracy in all three scenarios of the LH dataset.
[0087] Table 2 Performance comparison of positioning algorithms under the LH dataset
[0088]
[0089] As can be seen from the above, the indoor visual positioning system based on the occlusion removal algorithm proposed by the indoor visual positioning method based on the occlusion removal algorithm provides an effective solution to the problem of the influence of occluding objects on positioning accuracy. By combining object detection, image restoration, and feature matching technologies, the system not only improves the accuracy and robustness of positioning, but also provides new ideas for visual positioning problems in complex indoor environments.
[0090] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention in other forms. Any person skilled in the relevant art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes. However, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention without departing from the technical content of the present invention still fall within the protection scope of the technical solution of the present invention.
Claims
1. An indoor visual positioning method based on an occlusion removal algorithm, characterized in that: The steps include: Step 1, query the image in the database, extract the feature information of the database image, and save the feature information; Step 2: Detect occluders and remove all occluders in the query image. Step 3: Use the image restoration model to perform image restoration and restore the query image; Step 4: Perform indoor positioning; When executing step 2, it includes: constructing an occlusion image dataset; performing occlusion detection based on the occlusion image dataset; resetting pixels to eliminate occlusions in the query image; When constructing the occluded image dataset: several occluded objects are obtained in the coco128 dataset by matting. The matting process involves separating the object of interest from the image to obtain an object image with a transparent background; then, mask images of these objects are generated by retaining the pixels of the occluded objects and setting the pixel values of the remaining positions to (255, 255, 255); Load the images of the 7Scenes dataset; then, load all the images of the occluded objects and their corresponding masks; randomly select an occluded object and its mask for each query and database image, and add random types and degrees of blur information to it, then place the occluded object with blurred information at random positions on the image with random scales and rotation angles to generate occluded images. Finally, the occluded images generated one-to-one for all images constitute the occluded image dataset; Step 4 includes the following steps: Step 4-1, extracting and saving feature information of the query image restored in step 3; Step 4-2, performing image retrieval; performing a neighbor retrieval of the query image using an image retrieval algorithm based on the database image feature information and the query image feature information to obtain a neighbor retrieval pair of the query image; Step 4-3, based on the database image feature information, the query image feature information and the nearest neighbor retrieval pair, use the LightGlue algorithm to perform feature matching and establish a 2D-2D matching relationship between the query image and the database image; Step 4-4, using the indoor visual positioning algorithm based on the essential matrix; based on the 2D-2D matching information, the five-point method in the RANSAC cycle is used to calculate the essential matrix of the image pair; then, through the epipolar constraint condition, the singular value decomposition and RANSAC algorithm are used, combined with the known absolute pose and essential matrix of the database image, to finally determine the positioning result.
2. The indoor visual positioning method based on the occlusion removal algorithm according to claim 1 is characterized in that: The steps for occluded object detection based on the occluded image dataset are: first read the bounding box information in the detection result, extract the pixel area of the occluded object, load the YOLOv5 model pre-trained on the COCO dataset, and further fine-tune it on a specific occluded object dataset, then use the fine-tuned model to detect the query image, generate the bounding box and category information containing the occluded object, and complete the occluded object detection.
3. The indoor visual positioning method based on the occlusion removal algorithm according to claim 2 is characterized in that: The process of pixel resetting to eliminate occluders in the query image is as follows: the pixels in the occluded object area are reset to pure black pixels (0, 0, 0) to indicate the position of the occluded object; at the same time, a mask image corresponding to the query image is generated, in which the pixels in the occluded object area are set to pure black pixels, and the pixels in other areas are set to pure white pixels (255, 255, 255).
4. The indoor visual positioning method based on the occlusion removal algorithm according to claim 1, characterized in that: When executing step 1, first set the reference point in the indoor environment and collect image data at the reference point; secondly, use the SuperPoint algorithm and NetVLAD algorithm to extract features from the database image; finally, store the extracted feature information and the corresponding coordinate information to form a feature file.
5. The indoor visual positioning method based on the occlusion removal algorithm according to claim 1, characterized in that: In step 3, the image restoration model MAT is used to perform image restoration.
6. A positioning system for implementing the indoor visual positioning method based on the occlusion removal algorithm according to any one of claims 1 to 5, characterized in that: It includes offline units and online units, where the offline units include: Offline image acquisition module, used to query images in the database; The offline feature extraction module is connected to the offline image acquisition module and is used to extract feature information of the database image and save the feature information; Online units include: An online image acquisition module, used to obtain query images; The occlusion removal module is connected to the online image acquisition module and is used to detect occlusions and remove occlusions in all query images; The indoor positioning module is connected to the occlusion removal module and the offline feature extraction module and is used for indoor positioning.
7. The positioning system according to claim 6, characterized in that: The indoor positioning module includes: A feature extraction submodule is used to obtain feature information of the restored query image; The image retrieval submodule is connected to the feature extraction submodule and is used to perform a neighbor retrieval of the query image using an image retrieval algorithm to obtain a neighbor retrieval pair of the query image; The feature matching submodule is connected to the image retrieval submodule and is used to establish a 2D-2D matching relationship between the query image and the database image; And the positioning submodule obtains the positioning result based on the positioning algorithm of the essential matrix.
Citation Information
Patent Citations
Method and system for identifying facial features based on a security video
CN109492614A
Video processing method and device, electronic equipment and storage medium
CN116132732A
Indoor visual positioning system and method based on blurred image
CN118379476A
Method for generating occlusion data set based on specified occlusion rate
CN118691924A