Building indoor positioning method and system based on AI image retrieval
By building an environmental prior knowledge base and real-time image quality assessment, combined with spatial discreteness and mismatch diagnostic index, the problem of visual positioning errors of AI models in large buildings was solved, and the accurate interpretation and robustness of positioning results were achieved.
Patent Information
- Application Number
- CN202511005830.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-22
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-07-22
AI Technical Summary
When existing AI models perform visual positioning in large buildings, there are areas of similar visual features that lead to positioning errors. They lack the ability to diagnose the causes of low confidence, and the positioning system lacks reliability and robustness, making it impossible to effectively apply it in complex scenarios.
Build an environmental prior knowledge base, quantify the scene visual ambiguity index, combine real-time image quality assessment and candidate set spatial discreteness to generate a mismatch diagnosis index, and achieve accurate interpretation and robustness judgment of positioning results through multi-dimensional diagnostic logic.
It achieves accurate interpretation and reliability judgment of positioning results, improves the application value and user trust of the system in complex scenarios, and provides intelligent decision-making support.
Smart Images

Figure CN120510522B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of indoor positioning technology, and in particular to a method and system for indoor building positioning based on AI image retrieval. Background Art
[0002] In indoor visual positioning technology based on AI models, existing technologies face the following technical bottlenecks:
[0003] There is a mismatch between the raw confidence output by the AI model and the actual positioning difficulty of the scene. Large buildings often contain a large number of areas with highly similar visual features, such as corridors with repetitive structures, symmetrical atriums, or office areas with a unified style. The inherent visual ambiguity of this scene will cause the AI model to produce similar responses at different physical locations, leading to positioning errors. Traditional positioning systems lack a quantitative assessment mechanism for the inherent difficulty of this scene. The model's own confidence score alone cannot accurately reflect the true reliability of the positioning results.
[0004] The system lacks the ability to diagnose the root cause of low-confidence positioning results. When the model outputs a low-confidence result, the reason may be poor input image quality, such as motion blur, dark lighting, or overexposure; or the user is in an area of high visual ambiguity. Existing methods are usually unable to distinguish between these two situations, resulting in the inability to provide users with targeted and executable corrective instructions.
[0005] This technical limitation leads to insufficient reliability and robustness of the positioning system. In actual deployment, the system cannot accurately interpret the validity of the positioning results, nor can it provide intelligent decision support when positioning fails, limiting its application value and user trust in complex scenarios. Summary of the Invention
[0006] The purpose of the present invention is to provide a building indoor positioning method based on AI image retrieval, which solves the problems existing in the background technology.
[0007] To solve the above technical problems, the present invention provides a method for indoor building positioning based on AI image retrieval, comprising: constructing an environmental priori knowledge base, the environmental priori knowledge base storing a plurality of logical areas divided according to structured data of the building, and a scene visual ambiguity index calculated for each logical area to characterize the difficulty of positioning the logical area;
[0008] Acquire real-time images for positioning, perform multi-dimensional quality assessment on the real-time images, and generate comprehensive image quality scores;
[0009] Process the real-time image to obtain a candidate positioning result set containing a preset number k positioning results, and use the positioning result with the highest confidence in the candidate positioning result set as the initial positioning result, and use the confidence corresponding to the initial positioning result as the original confidence score;
[0010] Calculate the spatial dispersion of the candidate positioning result set in the physical space according to the candidate positioning result set;
[0011] According to the initial positioning result, the target logical area to which the initial positioning result belongs is determined, and the scene visual ambiguity index corresponding to the target logical area is extracted from the environmental prior knowledge base;
[0012] Based on the original confidence score and the scene visual ambiguity index, a mismatch diagnostic index is generated for quantifying the degree of matching between the original confidence score and the scene visual ambiguity index;
[0013] The final positioning state is generated according to the mismatch diagnostic index, the comprehensive image quality score and the spatial dispersion. The final positioning state includes the positioning dispersion failure state and the positioning mismatch state.
[0014] When the final positioning state is a positioning mismatch state, the calibration positioning result is output;
[0015] When the final positioning state is a positioning dispersion failure state, the user command is output.
[0016] Preferably, an environmental prior knowledge base is constructed, including:
[0017] Divide the building space into multiple logical areas based on the building structured data;
[0018] Collect and annotate image sequences with three-dimensional coordinate labels within each logical area to form a historical image set;
[0019] By analyzing the average feature similarity between images within each logical area and the average feature similarity between images in different logical areas in the historical image collection, the scene visual ambiguity index of each logical area is calculated.
[0020] Preferably, generating a comprehensive image quality score includes:
[0021] By applying a reference-free image evaluation algorithm, the blurriness, brightness, contrast, and occlusion ratio indicators of the real-time image are calculated in parallel and weighted fusion is performed to generate a comprehensive image quality score.
[0022] Preferably, obtaining a candidate positioning result set includes:
[0023] Extract visual feature descriptors of real-time images;
[0024] Retrieve and match visual feature descriptors with historical image sets labeled with 3D coordinates;
[0025] The probability outputs of the top k confidence rankings generated during the matching process are used as the candidate positioning result set.
[0026] Preferably, the step of generating the final positioning state includes:
[0027] When the spatial discreteness is greater than a preset discreteness threshold, the final positioning state is determined to be a positioning dispersion failure state.
[0028] Preferably, generating a mismatch diagnostic index comprises:
[0029] Based on the original confidence score, calculate the model's confidence level, which reflects the degree of mismatch in the current positioning.
[0030] The model uncertainty is divided by the scene visual ambiguity index extracted from the environment prior knowledge base to generate a mismatch diagnosis index.
[0031] Preferably, outputting a user instruction corresponding to the final positioning state includes:
[0032] When the final positioning state is a positioning dispersion failure state, generating and outputting a user instruction for guiding the user to move out of the current target logical area;
[0033] When the final positioning state is determined to be a positioning mismatch state, the initial positioning result is output as a calibration positioning result, and this positioning event is recorded for subsequent optimization of the environmental prior knowledge base.
[0034] We also provide building indoor positioning systems based on AI image retrieval, including:
[0035] a knowledge base construction unit for constructing an environmental priori knowledge base, the environmental priori knowledge base storing a plurality of logical areas divided according to pre-collected building environment data, and a scene visual ambiguity index calculated for each logical area and used to characterize the difficulty of locating the logical area;
[0036] An image quality assessment unit, configured to perform a multi-dimensional quality assessment on the acquired real-time image to be positioned, and generate a comprehensive image quality score;
[0037] A real-time positioning unit is configured to process the real-time image to obtain a candidate positioning result set including a preset number k of positioning results, use the positioning result with the highest confidence in the candidate positioning result set as the initial positioning result, and use the confidence corresponding to the initial positioning result as the original confidence score;
[0038] A dispersion calculation unit, configured to calculate the spatial dispersion of the candidate positioning result set in the physical space based on the candidate positioning result set;
[0039] A priori extraction unit is used to determine the target logical area to which the initial positioning result belongs based on the initial positioning result, and to extract the scene visual ambiguity index corresponding to the target logical area from the environmental prior knowledge base;
[0040] a mismatch diagnosis unit for generating a mismatch diagnosis index for quantifying a degree of matching between the original confidence score and the positioning difficulty based on the obtained original confidence score and the extracted scene visual ambiguity index;
[0041] The decision output unit is used to generate a final positioning state according to the mismatch diagnosis index, the comprehensive image quality score, and the spatial discreteness, and output the corresponding calibration positioning result or user instruction according to the final positioning state.
[0042] Compared with the prior art, the present invention has the following beneficial effects:
[0043] 1. By building an environmental prior knowledge base to quantify the scene visual ambiguity index, combined with real-time image quality assessment, candidate set spatial discreteness calculation, and mismatch diagnosis index generation, accurate interpretation and reliability judgment of positioning results are achieved, avoiding the mismatch between the original confidence and the scene positioning difficulty.
[0044] 2. Through dual-factor attribution diagnostic logic, the positioning confidence problem is decomposed into two dimensions: input quality and scene-inherent attributes, so as to accurately determine the root cause of low confidence and have the ability to diagnose the root cause of low-confidence positioning results.
[0045] 3. By building a closed-loop decision-making process, it is possible to output system states with clear semantics and provide reasonable behavioral recommendations, thereby improving robustness and expanding its application value in complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention, and those skilled in the art can derive other drawings based on these drawings without inventive effort.
[0047] Figure 1 Flowchart of the method of the present invention. DETAILED DESCRIPTION
[0048] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0049] Example 1:
[0050] See also Figure 1 , the present invention provides a building indoor positioning method based on AI image retrieval, comprising: constructing an environmental priori knowledge base, the environmental priori knowledge base storing a plurality of logical areas divided according to structured data of the building, and a scene visual ambiguity index calculated for each logical area and used to characterize the difficulty of positioning the logical area;
[0051] Acquire real-time images for positioning; perform multi-dimensional quality assessment on real-time images to generate comprehensive image quality scores;
[0052] Process the real-time image to obtain a candidate positioning result set containing a preset number k positioning results, and use the positioning result with the highest confidence in the candidate positioning result set as the initial positioning result, and use the confidence corresponding to the initial positioning result as the original confidence score;
[0053] Calculate the spatial dispersion of the candidate positioning result set in the physical space according to the candidate positioning result set;
[0054] According to the initial positioning result, the target logical area to which the initial positioning result belongs is determined, and the scene visual ambiguity index corresponding to the target logical area is extracted from the environmental prior knowledge base;
[0055] Based on the original confidence score and the scene visual ambiguity index, a mismatch diagnostic index is generated for quantifying the degree of matching between the original confidence score and the scene visual ambiguity index;
[0056] The final positioning state is generated according to the mismatch diagnostic index, the comprehensive image quality score and the spatial dispersion. The final positioning state includes the positioning dispersion failure state and the positioning mismatch state.
[0057] When the final positioning state is a positioning mismatch state, the calibration positioning result is output;
[0058] When the final positioning state is a positioning dispersion failure state, output the user command;
[0059] This embodiment provides a building indoor positioning method based on AI image retrieval. The implementation of this method is designed as a complete adaptive confidence calibration framework. By constructing an environmental prior knowledge base, the positioning difficulty of different logical areas in the building is pre-quantified. This difficulty is characterized by the scene visual ambiguity index. In the real-time positioning stage, the acquired real-time image is first subjected to a multi-dimensional quality assessment to obtain an image comprehensive quality score. This can prevent poor quality image input from interfering with subsequent processing. By processing the high-quality real-time image, a set of k candidate results is generated, among which the one with the highest confidence is selected as the initial positioning result, and its confidence is used as the original confidence score. The spatial discreteness of the candidate result set is calculated in parallel, the scene visual ambiguity index of the corresponding area is extracted from the knowledge base, and the mismatch diagnostic index is generated in combination with the original confidence. The final positioning status of the system is jointly determined by the three dimensions of image quality, spatial discreteness and mismatch diagnostic index.
[0060] This method deconstructs the potential causes of positioning dispersion failure and conducts a comprehensive diagnosis from three levels: image quality, cohesion of model judgment, and the match between model confidence and scene difficulty. It achieves accurate interpretation and reliability judgment of positioning results. It can provide calibrated positioning results when positioning mismatches occur, or provide clear user guidance when positioning dispersion fails, significantly improving the robustness of the positioning system.
[0061] Example 2:
[0062] Build an environment prior knowledge base, including:
[0063] Divide the building space into multiple logical areas based on the building structured data;
[0064] Collect and annotate image sequences with three-dimensional coordinate labels within each logical area to form a historical image set;
[0065] By analyzing the average feature similarity between images within each logical region and the average feature similarity between images in different logical regions in the historical image collection, the scene visual ambiguity index of each logical region is calculated;
[0066] Get the candidate positioning result set, including:
[0067] Extract visual feature descriptors of real-time images;
[0068] Retrieve and match visual feature descriptors with historical image sets labeled with 3D coordinates;
[0069] The probability outputs of the top k confidence rankings generated during the matching process are used as the candidate positioning result set;
[0070] In this embodiment, the construction of the environmental prior knowledge base is first based on the building structured data, such as CAD floor plans, to discretize the building space into multiple logical areas with functional or visual consistency, such as corridors, atriums or meeting room groups, and assign a unique identifier to each area. Subsequently, system deployment engineers brought professional equipment to collect image sequences in each logical area and accurately marked the three-dimensional coordinates of the image frames. , forming a labeled historical image set ; Scene visual ambiguity index The calculation of is performed, and the index is defined as the ratio of the average feature similarity between images within the region to the average feature similarity between the region and other regions, which is expressed by the following formula:
[0071] ,
[0072] in: Indicates area The scene visual ambiguity index is used to measure the "inherent difficulty" of positioning in a specific logical area; Indicates area Average feature similarity between internal images; Indicates area Average feature similarity with images in other regions;
[0073] During real-time positioning, a visual feature descriptor is first extracted from the real-time image captured by the user. ; This descriptor and historical image collections Perform high-speed retrieval and matching; the purpose of this process is to find the reference image that is most similar to the current scene in the historical data. The k results with the highest confidence output by the matching algorithm are adopted to form a candidate positioning result set ; By pre-calculating and storing the scene visual ambiguity index of each logical area This method converts the fuzzy experience of the scene's "positioning difficulty" into computable prior data; during positioning, through the retrieval and matching of feature descriptors and historical image sets, a compact set of candidate positioning results can be quickly generated, providing a reliable data basis for subsequent diagnosis and decision-making links, enabling the system to predict and effectively respond to positioning challenges caused by scene self-similarity or similarity between regions.
[0074] Example 3:
[0075] Generates a comprehensive image quality score, including:
[0076] By applying a reference-free image evaluation algorithm, the blur, brightness, contrast, and occlusion ratio indicators of the real-time image are calculated in parallel and weighted fusion is performed to generate a comprehensive image quality score;
[0077] In this embodiment, the generation of the comprehensive image quality score begins by applying multiple reference-free image evaluation algorithms to the real-time image; these algorithms are executed in parallel to calculate four core quality indicators of the image: blurriness, ,brightness , contrast And the occlusion ratio ; These indicators are then weighted and fused to generate a normalized image quality score ; The weighted fusion process is modeled as:
[0078] ,
[0079] in: Represents the comprehensive image quality score, quantifying the quality of the input image; Indicates the blur of the image; Indicates the brightness of the image; Indicates the contrast of the image; Indicates the occlusion ratio of the image; 、 、 、 Represent the weights of blur, brightness, contrast, and occlusion ratio indicators respectively;
[0080] Each weight is pre-calibrated through experiments based on the degree of influence of each indicator on the performance of the visual positioning model. This multi-dimensional quality assessment mechanism provides the system with a reliable input "filter" by quantifying image clarity, lighting conditions, and information integrity. It eliminates low-quality input caused by factors such as motion blur, over-exposure, or severe occlusion at the beginning of the positioning process, effectively preventing the introduction of uncertainty from the input, thereby ensuring the stability and accuracy of subsequent positioning solutions, and is the first line of defense for achieving robust positioning.
[0081] Example 4:
[0082] The steps to generate the final positioning state include:
[0083] When the spatial dispersion is greater than the preset dispersion threshold, the final positioning state is determined to be a positioning dispersion failure state;
[0084] Generates mismatch diagnostic indices including:
[0085] Based on the original confidence score, calculate the model's confidence level, which reflects the degree of mismatch in the current positioning.
[0086] The model uncertainty is divided by the scene visual ambiguity index extracted from the environment prior knowledge base to generate a mismatch diagnosis index;
[0087] In this embodiment, the candidate positioning result set The spatial dispersion of Calculations are performed to verify the stability of the AI model’s judgment; the calculation formula is:
[0088] ,
[0089] in: Indicates the spatial dispersion of the candidate positioning result set, which is used to measure the degree of dispersion of the candidate positioning results of the AI model in the physical space; Indicates the number of candidate positioning results; Represents the three-dimensional coordinates of the i-th candidate positioning result; Represents the geometric centroid of all candidate positioning result coordinates;
[0090] If the calculated spatial dispersion Exceeds a preset discreteness threshold , the final positioning state will be directly judged as a dispersion failure; at the same time, a core mismatch diagnostic index is generated to quantify how well the model's real-time performance matches the inherent challenges of the scene; first, from the raw confidence score Calculate the model's uncertainty, i.e. ; Then, divide this uncertainty by the scene visual ambiguity index corresponding to the current target logical area queried from the environmental prior knowledge base ; Mismatch diagnostic index The generating formula is defined as:
[0091] ,
[0092] in: represents the mismatch diagnostic index, which quantifies the gap between the model performance and the objective scenario challenge; Indicates the original confidence score output by the AI model, reflecting the model's initial confidence in the positioning result; The scene visual ambiguity index of the logical area to which the initial positioning result belongs represents the inherent positioning difficulty of the scene;
[0093] Spatial discreteness The introduction of the mismatch diagnostic index enables the system to quickly identify catastrophic positioning dispersion failures where the model output is extremely inconsistent in physical space; The construction of provides a deeper diagnostic capability. It no longer looks at the confidence level in isolation, but innovatively combines it with the positioning difficulty of the scene. By making associations, it is possible to accurately distinguish between the two fundamentally different situations of "the model is not confident even in simple scenarios" and "the model shows reasonably low confidence in difficult scenarios", thus achieving a leap from measurement to diagnosis.
[0094] Example 5:
[0095] Output user instructions corresponding to the final positioning state, including:
[0096] When the final positioning state is a positioning dispersion failure state, generating and outputting a user instruction for guiding the user to move out of the current target logical area;
[0097] When the final positioning state is determined to be a positioning mismatch state, the initial positioning result is output as a calibration positioning result, and the positioning event is recorded for subsequent optimization of the environmental prior knowledge base;
[0098] In this embodiment, the final decision and output of the system is based on a multi-dimensional state decision mechanism; the basis for the decision is the comprehensive image quality score calculated in real time. , candidate set space dispersion and mismatch diagnostic index , comparing them with a set of preset thresholds 、 and Compare and generate the final positioning state ;according to The system outputs the corresponding user command or positioning result; if the final positioning state is determined to be a positioning failure, for example due to or , the system will generate and output clear guidance instructions to the user, such as "Positioning dispersion failed, please try to move to an area with more obvious features and try again"; if the final positioning state is judged to be a positioning mismatch state, this usually means that the image quality is good and the candidate set converges, but the confidence of the model does not fully match the scene difficulty. At this time, the system will use the initial positioning result with the highest confidence in the initial positioning result set as the final calibration positioning result Output to the user and display on the map; at the same time, the event of "successful positioning in ambiguous scenes" includes the image, location, and Such information will be fully recorded for future long-term analysis and iterative optimization of the environmental prior knowledge base; this state-based output strategy upgrades the system from a simple positioning tool to an intelligent diagnostic partner; it not only provides location information, but also explains the reliability of positioning, and provides feasible solutions and behavioral suggestions for positioning difficulties caused by subjective and objective reasons. This closed-loop "perception-diagnosis-decision-making" process greatly enhances the robustness of the system and user trust, and provides data support for the system's continuous self-optimization.
[0099] Example 6:
[0100] Building indoor positioning system based on AI image retrieval, including:
[0101] The knowledge base construction unit is used to construct an environmental prior knowledge base, which stores a plurality of logical areas divided according to pre-collected building environment data, and a scene visual ambiguity index calculated for each logical area and used to characterize the difficulty of locating the logical area;
[0102] The image quality assessment unit is used to perform multi-dimensional quality assessment on the acquired real-time image to be positioned and generate a comprehensive image quality score;
[0103] The real-time positioning unit is used to process the real-time image to obtain a candidate positioning result set including a preset number k positioning results, take the positioning result with the highest confidence in the candidate positioning result set as the initial positioning result, and take the confidence corresponding to the initial positioning result as the original confidence score;
[0104] The dispersion calculation unit is used to calculate the spatial dispersion of the candidate positioning result set in the physical space based on the candidate positioning result set;
[0105] The prior extraction unit is used to determine the target logical area to which it belongs based on the initial positioning result, and extract the scene visual ambiguity index corresponding to the target logical area from the environmental prior knowledge base;
[0106] The mismatch diagnosis unit is used to generate a mismatch diagnosis index for quantifying the degree of matching between the original confidence score and the positioning difficulty based on the obtained original confidence score and the extracted scene visual ambiguity index;
[0107] The decision output unit is used to generate a final positioning state according to the mismatch diagnosis index, the image comprehensive quality score, and the spatial discreteness, and output a corresponding calibration positioning result or user instruction according to the final positioning state;
[0108] The AI image retrieval building indoor positioning system described in this embodiment is constructed as a set of highly coordinated modular units; the knowledge base construction unit works offline before the system is deployed, and is responsible for dividing the physical space of the building into logical areas, and calculating the scene visual ambiguity index for each area to form an environmental prior knowledge base; when the system is running, the image quality assessment unit first receives the real-time image and outputs the comprehensive image quality score, which plays a pre-filtering role; the high-quality image is passed to the real-time positioning unit, which generates a set of k candidate positioning results and their corresponding original confidence scores through image retrieval and matching; then, the discreteness calculation unit, the prior extraction unit and the mismatch diagnosis unit work in parallel: the discreteness calculation unit quantifies the consistency of the candidate result set in the physical space; the prior extraction unit accurately extracts the current location from the knowledge base based on the initial positioning result. The mismatch diagnosis unit combines the original confidence and the scene ambiguity index to generate the core mismatch diagnosis index. Finally, all these real-time and a priori information are summarized in the decision output unit, which makes a comprehensive judgment on the image quality, result discreteness and mismatch index according to a set of preset rules to generate the final system state and decide whether to output a confirmed calibration positioning result or issue an actionable guidance instruction to the user. This system design, which combines serial and parallel units with clear responsibilities, decomposes complex positioning problems into a series of clear steps including data preparation, quality control, preliminary solution, multi-dimensional diagnosis and intelligent decision-making. Each unit performs its duties and cooperates closely to ensure that every link from input to output is processed in a refined manner, ultimately achieving reliable, explainable and interactive intelligent indoor positioning.
[0109] The above are merely preferred embodiments of the present invention and do not limit the present invention in any other form. Any technician familiar with the profession may use the technical content disclosed above to change or modify it into an equivalent embodiment with equivalent changes and apply it to other fields. However, any simple modification, equivalent change and modification made to the above embodiment based on the technical essence of the present invention without departing from the content of the technical solution of the present invention shall still fall within the scope of protection of the technical solution of the present invention.
Claims
1. A building indoor positioning method based on AI image retrieval, characterized in that: include: Constructing an environmental prior knowledge base, which stores multiple logical areas divided according to the structured data of the building, and a scene visual ambiguity index calculated for each logical area to characterize the difficulty of locating the logical area. The scene visual ambiguity index is defined as the ratio of the average feature similarity between images within the area to the average feature similarity between images of the area and other areas. Acquire real-time images for positioning, perform multi-dimensional quality assessment on the real-time images, and generate comprehensive image quality scores; Process the real-time image to obtain a candidate positioning result set containing a preset number k positioning results, and use the positioning result with the highest confidence in the candidate positioning result set as the initial positioning result, and use the confidence corresponding to the initial positioning result as the original confidence score; Calculate the spatial dispersion of the candidate positioning result set in the physical space according to the candidate positioning result set; According to the initial positioning result, the target logical area to which the initial positioning result belongs is determined, and the scene visual ambiguity index corresponding to the target logical area is extracted from the environmental prior knowledge base; Based on the original confidence score and the scene visual ambiguity index, a mismatch diagnostic index is generated for quantifying the degree of matching between the original confidence score and the scene visual ambiguity index; The final positioning state is generated according to the mismatch diagnostic index, the comprehensive image quality score and the spatial dispersion. The final positioning state includes the positioning dispersion failure state and the positioning mismatch state. When the final positioning state is a positioning mismatch state, the calibration positioning result is output; When the final positioning state is a positioning dispersion failure state, the user command is output.
2. The building indoor positioning method based on AI image retrieval according to claim 1 is characterized in that: Build an environment prior knowledge base, including: Divide the building space into multiple logical areas based on the building structured data; Collect and annotate image sequences with three-dimensional coordinate labels within each logical area to form a historical image set; By analyzing the average feature similarity between images within each logical area and the average feature similarity between images in different logical areas in the historical image collection, the scene visual ambiguity index of each logical area is calculated.
3. The building indoor positioning method based on AI image retrieval according to claim 1 is characterized in that: Generates a comprehensive image quality score, including: By applying a reference-free image evaluation algorithm, the blurriness, brightness, contrast, and occlusion ratio indicators of the real-time image are calculated in parallel and weighted fusion is performed to generate a comprehensive image quality score.
4. The building indoor positioning method based on AI image retrieval according to claim 2 is characterized in that: Get the candidate positioning result set, including: Extract visual feature descriptors of real-time images; Retrieve and match visual feature descriptors with historical image sets labeled with 3D coordinates; The probability outputs of the top k confidence rankings generated during the matching process are used as the candidate positioning result set.
5. The building indoor positioning method based on AI image retrieval according to claim 1 is characterized in that: The steps to generate the final positioning state include: When the spatial discreteness is greater than a preset discreteness threshold, the final positioning state is determined to be a positioning dispersion failure state.
6. The building indoor positioning method based on AI image retrieval according to claim 1, characterized in that: Generates mismatch diagnostic indices including: Based on the original confidence score, calculate the model's confidence level, which reflects the degree of mismatch in the current positioning. The model uncertainty is divided by the scene visual ambiguity index extracted from the environment prior knowledge base to generate a mismatch diagnosis index.
7. The building indoor positioning method based on AI image retrieval according to claim 1 is characterized in that: Output user instructions corresponding to the final positioning state, including: When the final positioning state is a positioning dispersion failure state, generating and outputting a user instruction for guiding the user to move out of the current target logical area; When the final positioning state is determined to be a positioning mismatch state, the initial positioning result is output as a calibration positioning result, and this positioning event is recorded for subsequent optimization of the environmental prior knowledge base.
8. A building indoor positioning system based on AI image retrieval, applied to the building indoor positioning method based on AI image retrieval according to any one of claims 1 to 7, characterized in that: include: a knowledge base construction unit for constructing an environmental priori knowledge base, the environmental priori knowledge base storing a plurality of logical areas divided according to pre-collected building environment data, and a scene visual ambiguity index calculated for each logical area and used to characterize the difficulty of locating the logical area; An image quality assessment unit, configured to perform a multi-dimensional quality assessment on the acquired real-time image to be positioned, and generate a comprehensive image quality score; A real-time positioning unit is configured to process the real-time image to obtain a candidate positioning result set including a preset number k of positioning results, use the positioning result with the highest confidence in the candidate positioning result set as the initial positioning result, and use the confidence corresponding to the initial positioning result as the original confidence score; A dispersion calculation unit, configured to calculate the spatial dispersion of the candidate positioning result set in the physical space based on the candidate positioning result set; A priori extraction unit is used to determine the target logical area to which the initial positioning result belongs based on the initial positioning result, and to extract the scene visual ambiguity index corresponding to the target logical area from the environmental prior knowledge base; a mismatch diagnosis unit for generating a mismatch diagnosis index for quantifying a degree of matching between the original confidence score and the positioning difficulty based on the obtained original confidence score and the extracted scene visual ambiguity index; The decision output unit is used to generate a final positioning state according to the mismatch diagnosis index, the comprehensive image quality score, and the spatial discreteness, and output the corresponding calibration positioning result or user instruction according to the final positioning state.
Citation Information
Patent Citations
Indoor vision rapid matching and positioning method and system based on space optimization strategy
CN112614162A
Intelligent agent, indoor navigation method and equipment thereof, medium and product
CN119443287A