Visual positioning method based on indoor fine three-dimensional model
By constructing and correcting feature datasets, generating feature sequences in real time, and combining them with position estimation data for matching and local updates, the problem of balancing positioning accuracy and real-time performance in indoor navigation is solved, and efficient indoor positioning and navigation are achieved.
Patent Information
- Application Number
- CN202510649838.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-10-10
AI Technical Summary
Existing indoor navigation technologies have difficulty balancing positioning accuracy and real-time performance in complex environments, especially in large-scale, multi-level indoor spaces. How to efficiently match dynamic images and static models has become a technical barrier that needs to be overcome urgently.
The image data obtained by the camera is combined with the attitude angular velocity to construct a feature data set and perform corrections. The feature sequence is generated in real time and matched with the position estimation data. Local updates and path planning are performed in a dynamic environment, and the occlusion area ratio weight is introduced to optimize the navigation path.
It significantly improves the accuracy, stability and efficiency of indoor positioning, adapts to dynamic environmental changes, meets real-time requirements, and is suitable for robot navigation, augmented reality, virtual reality and intelligent building management.
Smart Images

Figure CN120765731A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of three-dimensional model visual positioning, and in particular to a method for visual positioning based on indoor fine three-dimensional models. Background Art
[0002] Indoor navigation technology, a key branch of intelligent positioning, plays a crucial role in modern urbanization. With the growing demand for indoor space utilization, the need for navigation in complex architectural environments such as large shopping malls, airports, and hospitals is becoming increasingly prominent. The accuracy and practicality of indoor navigation technology directly impact user experience and space management efficiency. Unlike outdoor navigation, which relies on satellite signals, indoor navigation, due to environmental complexity and signal obstruction, has become a challenging area in positioning technology that urgently requires breakthroughs.
[0003] Current mainstream indoor navigation methods rely on technologies such as Bluetooth beacons, Wi-Fi signals, or inertial navigation. However, these solutions have significant limitations in practical applications. Bluetooth and Wi-Fi positioning are limited by uneven signal coverage and high equipment deployment costs, often reaching only meter-level accuracy, making them difficult to meet the needs of refined navigation. While inertial navigation requires no external signal support, it is susceptible to cumulative errors, leading to significant positioning deviations after prolonged use. These shortcomings leave users struggling with positioning drift and inaccurate navigation in complex indoor environments.
[0004] Against this backdrop, the core challenges facing indoor navigation are becoming increasingly clear, particularly how to leverage visual information to achieve high-precision, real-time positioning. Existing technologies face bottlenecks in building detailed indoor models and real-time image matching. These bottlenecks are manifested in the difficulty of integrating multi-source data, the computational complexity of matching model construction with actual scenes, and the insufficient image processing capabilities of mobile devices in dynamic environments. These unresolved technical factors lead to the unique challenge of balancing positioning accuracy and real-time performance. This is particularly true in large-scale, multi-layered indoor spaces, where efficiently matching dynamic images with static models becomes a critical technical barrier that needs to be overcome.
[0005] Therefore, building a detailed 3D model of the indoor scene based on multi-source data and achieving precise spatial matching through real-time mobile imagery and back-end visual computing technology has become a key issue in improving the practicality and accuracy of indoor navigation solutions. This issue focuses on the coordinated optimization of data integration, model publishing, and image matching, aiming to overcome the accuracy and efficiency limitations of existing technologies and provide a more reliable path for indoor navigation. Summary of the Invention
[0006] In view of this, the purpose of the present invention is to propose a method for visual positioning based on a fine indoor three-dimensional model, which can solve at least one technical problem mentioned in the background technology.
[0007] According to one aspect of the present invention, a method for visual positioning based on a fine three-dimensional indoor model is provided, the method comprising:
[0008] The camera acquires the indoor original image and synchronously collects the camera's attitude angular velocity to construct a first feature data set, and performs correction to generate a second feature data set;
[0009] Aligning the second feature data set with a pre-built indoor real-scene three-dimensional model to generate first position estimation data under spatial boundary constraints; locating feature point regions from the real-time acquired dynamic image data based on the first position estimation data, and tracking point motion trajectories in the movement trajectory offset to generate a first feature sequence;
[0010] Matching the first feature sequence with the first position estimation data and fusing them to generate a first posture dataset; when the distance deviation between the first posture dataset and the three-dimensional model position exceeds a preset threshold, performing a local update on the indoor real scene three-dimensional model to generate a first three-dimensional model;
[0011] Matching the first three-dimensional model with the first feature sequence to generate a first positioning data set;
[0012] According to the first positioning data set combined with the spatial boundary constraints in the dynamic environment, a first navigation instruction sequence is generated by calculating a navigation path change trend taking into account the weight of the occluded area ratio.
[0013] In the aforementioned technical solution, this approach uses multi-source data fusion technology to combine image data collected by the camera with attitude angular velocity to construct and correct a feature dataset. This design significantly improves positioning accuracy and robustness. The introduction of attitude angular velocity dynamically compensates for high-frequency jitter or systematic errors in the image data, ensuring the stability of the feature data and significantly improving the reliability of positioning results.
[0014] In dynamic environments, this solution uses feature point area positioning and trajectory tracking technology to generate feature sequences in real time, combining them with position estimation data for matching and posture fusion. This design effectively adapts to environmental changes, ensuring the system's stability and adaptability in dynamic scenarios. Through feature point tracking and trajectory offset processing, the system updates positioning data in real time, avoiding positioning deviations caused by environmental changes and ensuring the real-time and accurate positioning results.
[0015] When the positional deviation between the pose dataset and the 3D model exceeds a preset threshold, this solution uses a local update mechanism to dynamically adjust the 3D model structure. This design avoids the high computational cost of global updates while ensuring the real-time and accuracy of the model, making it particularly suitable for efficient positioning in dynamic environments.
[0016] In terms of path planning, the scheme introduces an occlusion area proportion weight, dynamically calculates the trend of the navigation path, and generates a navigation instruction sequence. This path planning method can effectively avoid occlusion areas, optimize path selection, reduce blind spots in path planning, and significantly improve navigation efficiency and success rate.
[0017] In the feature data and three-dimensional model alignment process, the scheme introduces a spatial boundary constraint to generate position estimation data. This constraint can limit the range of positioning results and avoid drift of positioning results, significantly improving the stability of positioning results, especially in complex or dynamic environments.
[0018] The scheme emphasizes real-time performance, and through optimization algorithms and local update mechanisms, ensures the rapid response capability of feature point area positioning, trajectory tracking, and model updating. This design significantly improves the overall efficiency of the system in dynamic environments, meeting real-time requirements.
[0019] The scheme is suitable for various indoor scenarios, including robot navigation, augmented reality (AR), virtual reality (VR), and intelligent building management. The wide range of application scenarios significantly improves the market value of the scheme. At the same time, the scheme introduces feature point tracking, local updating, and path optimization based on existing technology, solving the limitations of existing technology in dynamic environments, and making significant technological progress.
[0020] In summary, the scheme significantly improves the accuracy, stability, and efficiency of indoor visual positioning and navigation through multi-source data fusion, dynamic environment adaptability, local update mechanisms, and path planning optimization, providing an efficient and reliable solution for positioning and navigation in dynamic environments.
[0021] In some embodiments, the indoor original image is obtained through a camera, and the attitude angular velocity of the camera is synchronously collected to construct a first feature data set, including:
[0022] The original image data is obtained from the indoor camera, and the attitude angular velocity recorded by the sensor is synchronously collected to obtain an initial data set. If the image data and angular velocity data timestamps in the initial data set are aligned, the time axis alignment algorithm is used to correct the deviation to obtain a synchronous data set.
[0023] According to the synchronous data set, the Kalman filter algorithm is used to fuse the mobile trajectory offset to obtain trajectory optimization data. For the image sequence in the trajectory optimization data, the inter-frame illumination change feature is calculated to obtain an illumination feature sequence. From the illumination feature sequence and the trajectory optimization data, a multi-dimensional description vector is extracted to obtain a vector representation set.
[0024] If the dimension of the vector representation set meets the preset threshold, the principal component analysis algorithm is used for dimension reduction processing to obtain the first feature data set.
[0025] In the aforementioned technical solution, this approach ensures temporal consistency between image data and attitude and angular velocity data through timestamp alignment and timeline alignment algorithms. This synchronization eliminates errors caused by time offsets during data acquisition, provides a reliable foundation for multi-source data fusion, and ensures the accuracy and robustness of subsequent processing.
[0026] The Kalman filter algorithm is used to fuse trajectory deviations, effectively filtering out noise and deviations in the data and generating more stable trajectory optimization data. As a classic dynamic system state estimation algorithm, the Kalman filter can process random noise and systematic deviations in data in real time, making it particularly suitable for data fusion in dynamic environments.
[0027] To address the impact of illumination changes on visual localization, this solution calculates inter-frame illumination variation characteristics and generates an illumination feature sequence. This process significantly enhances feature robustness, reduces the interference of illumination changes on feature extraction, and ensures feature stability under varying lighting conditions. The extraction of illumination features provides more reliable visual information for subsequent processing.
[0028] A multidimensional description vector is extracted from the illumination feature sequence and trajectory optimization data to generate a vector representation set. The multidimensional description vector can comprehensively capture image features, enhance their distinguishability and expressiveness, and provide richer information for subsequent feature matching and positioning.
[0029] The principal component analysis (PCA) algorithm is used to reduce the dimensionality of the vector representation set to generate the first feature data set. The PCA algorithm can effectively reduce the data dimension and remove redundant information while retaining key features, significantly improving data processing efficiency.
[0030] This solution organically combines multiple steps, including data synchronization, filtering and fusion, illumination feature extraction, multidimensional feature description, and dimensionality reduction, to form a complete feature data construction process. This integrated design not only improves the overall performance of the system, but also ensures the accuracy, stability, and efficiency of the feature data.
[0031] This solution is particularly well-suited for a variety of indoor scenarios in dynamic environments and complex lighting conditions, including robotic navigation, augmented reality (AR), virtual reality (VR), and intelligent building management. Through techniques such as data synchronization and correction, Kalman filter fusion, illumination feature extraction, multidimensional feature description, and principal component analysis dimensionality reduction, this solution constructs a high-quality first feature dataset, providing a solid foundation for subsequent positioning and navigation. These technical elements work together to significantly improve the accuracy, stability, and processing efficiency of feature data, addressing the limitations of existing technologies in dynamic environments. This represents a significant technological advancement and promises broad market application value.
[0032] In some embodiments, performing correction to generate a second feature dataset includes:
[0033] A deep learning network is used to extract dynamic blur features from the first feature dataset, and the view rotation matrix is calculated based on the attitude angular velocity to generate a corrected image feature set. If the timestamp of the corrected image feature set is aligned with the sensor data, a time synchronization algorithm is used to adjust the deviation to obtain a synchronized feature set.
[0034] Based on the synchronized feature set, texture density features are extracted to generate an intermediate feature set containing texture information. Target area detection is performed on the intermediate feature set through a convolutional neural network to obtain a target distance deviation vector. If the dimension of the target distance deviation vector exceeds a preset threshold, a dimensionality reduction algorithm is used to generate an optimized feature vector. Based on the optimized feature vector, the texture density features and the target distance deviation are fused to generate a second feature data set.
[0035] In the above technical solution, this solution uses a deep learning network to accurately extract dynamic blur features from the first feature data set, and combines this with the attitude angular velocity to calculate the view rotation matrix, generating a corrected image feature set. This process effectively corrects dynamic blur in the image, significantly improving feature clarity and accuracy. Dynamic blur is a key factor affecting visual positioning accuracy. By extracting blurred features through a deep learning network and correcting them with attitude angular velocity, this fundamentally addresses the negative impact of dynamic blur on positioning accuracy, ensuring optimal image feature quality.
[0036] If the timestamps of the corrected image feature sets deviate from those of the sensor data, a high-precision time synchronization algorithm is used to correct them, resulting in a strictly aligned synchronized feature set. This time synchronization ensures data consistency in the temporal dimension, completely eliminating positioning errors caused by time deviations. This provides a solid foundation for multi-source data fusion and significantly improves the reliability and accuracy of subsequent processing.
[0037] Based on the synchronized feature set, this solution further extracts texture density features to generate an intermediate feature set containing rich texture information. Texture density features significantly enhance feature discrimination, making them more stable and robust in complex environments. Texture information, a core feature in visual localization, is refined through texture density feature extraction, ensuring its stability and adaptability in diverse environmental conditions.
[0038] A convolutional neural network efficiently detects target regions from the intermediate feature set, generating precise target distance deviation vectors. This processing not only accurately locates the target region but also calculates its distance deviation in real time, providing critical quantitative information for subsequent positioning and navigation. Convolutional neural networks, with their high precision and efficiency in target detection, ensure accurate and real-time target region recognition.
[0039] If the dimension of the target distance deviation vector exceeds a preset threshold, an advanced dimensionality reduction algorithm is used to generate a compact and complete optimized feature vector. This dimensionality reduction process not only reduces data dimensionality and improves computational efficiency, but also preserves key feature information, effectively eliminating data redundancy and further improving data processing efficiency and accuracy.
[0040] Finally, based on the optimized feature vectors, texture density features and target distance deviation are fused to generate a comprehensive second feature dataset. This feature fusion strategy comprehensively integrates multiple feature information, significantly improving the expressiveness of the feature dataset and positioning accuracy. Feature fusion, a core step in improving positioning accuracy, generates a more comprehensive and accurate feature dataset through the synergy of multi-dimensional features, providing strong support for high-precision positioning and navigation.
[0041] This solution seamlessly integrates multiple key technical links, including dynamic blur feature extraction, time synchronization correction, texture density feature extraction, target area detection, dimensionality reduction, and feature fusion, to create an efficient and accurate feature data generation process. This highly integrated design not only significantly improves the overall system performance but also ensures the real-time and accuracy of data processing, making it particularly suitable for real-time positioning and navigation in dynamic environments.
[0042] This solution generates a high-quality second feature dataset through innovative techniques such as dynamic blur feature extraction, time synchronization correction, texture density feature extraction, target area detection, dimensionality reduction, and feature fusion. These techniques work together to significantly improve the accuracy, stability, and processing efficiency of the feature dataset from multiple dimensions, laying a solid foundation for subsequent high-precision positioning and intelligent navigation.
[0043] In some embodiments, aligning the second feature dataset with a pre-built indoor real-scene 3D model to generate first position estimation data constrained by a spatial boundary includes:
[0044] The second feature data set is roughly aligned with the pre-established indoor real-scene 3D model. The sampling interval of feature points is optimized through acquisition frequency analysis to generate a uniformly distributed feature point set. If the acquisition frequency fluctuation exceeds a preset threshold, the sampling interval is optimized through an adaptive adjustment algorithm to obtain a uniformly distributed feature point set.
[0045] Extract key points of the second feature data set based on the evenly distributed feature point set to generate an initial feature descriptor; use a key point matching algorithm to compare the initial feature descriptor with a reference point set of the indoor real-scene 3D model to obtain a roughly aligned transformation matrix;
[0046] Using a transformation matrix, the second feature data set is preliminarily aligned with the indoor real-scene 3D model to generate an aligned feature point cloud. If the deviation between the aligned feature point cloud and the 3D model exceeds a preset threshold, the transformation matrix is adjusted using an iterative closest point algorithm to obtain an optimized feature point cloud.
[0047] Based on the optimized feature point cloud and combined with the spatial boundary constraints, the first position estimation data is generated; through the spatial boundary constraint check, if the first position estimation data exceeds the boundary range, the estimation parameters are adjusted to obtain the estimation data that meets the constraints.
[0048] In the aforementioned technical solution, this approach combines rough alignment with optimized feature point sampling intervals to generate a uniformly distributed set of feature points, ensuring their spatial uniformity within the 3D model, thereby significantly improving alignment accuracy and stability. Uniformly distributed feature points effectively avoid the problem of overly dense or sparse local features, ensuring global consistency during the alignment process.
[0049] When the acquisition frequency fluctuates beyond a preset threshold, the system dynamically optimizes the feature point sampling interval through an adaptive adjustment algorithm to ensure uniform distribution of feature points under varying environmental conditions. This adaptive mechanism significantly improves the robustness of alignment and enhances the system's adaptability to dynamic environmental changes.
[0050] Using a keypoint matching algorithm, the initial feature descriptor is efficiently matched with the reference point set of the indoor real-world 3D model, generating a roughly aligned initial transformation matrix. The transformation matrix is further refined using the Iterative Closest Point (ICP) algorithm to generate an optimized feature point cloud. This dual optimization mechanism significantly improves alignment accuracy while reducing computational resource consumption.
[0051] Based on the optimized feature point cloud, spatial boundary constraints are incorporated to generate the first position estimate. This constraint ensures that the positioning result remains within a reasonable range, effectively preventing drift. This constraint mechanism significantly improves the stability and reliability of positioning results, especially in complex or dynamic environments.
[0052] This step organically combines multiple algorithms, including feature point sampling optimization, key point matching, iterative closest point algorithm, and spatial boundary constraints, to form a complete alignment process. This multi-algorithm collaborative optimization design not only significantly improves alignment accuracy but also enables real-time responsiveness to dynamic environmental changes through adaptive adjustment algorithms and iterative optimization mechanisms.
[0053] By optimizing feature point sampling, key point matching, transformation matrix optimization, and spatial boundary constraints, this solution significantly improves alignment accuracy and stability, providing a reliable foundation for subsequent positioning and navigation. This solution is particularly suitable for complex scenarios such as robot navigation, augmented reality (AR), virtual reality (VR), and intelligent building management, demonstrating significant technological advancement and broad market application value.
[0054] In some embodiments, locating a feature point region from dynamic image data collected in real time based on the first position estimation data and tracking the point motion trajectory in the movement trajectory offset to generate a first feature sequence includes:
[0055] The real-time collected dynamic image data is acquired, the occlusion area ratio is determined, and a first image sequence is obtained. The first image sequence is processed by a feature point positioning algorithm to locate the feature point area and obtain a first feature point set. The points in the first feature point set are tracked using an optical flow method, and the motion trajectory offset is calculated to obtain a first trajectory sequence.
[0056] In the aforementioned technical solution, this solution effectively avoids occlusions in the image by accurately identifying the proportion of occlusion areas, ensuring the accuracy of feature point positioning. Occlusion area identification is crucial for feature point positioning in dynamic environments, significantly reducing positioning errors caused by occlusion and improving system robustness.
[0057] An advanced feature point localization algorithm is used to process the first image sequence, precisely locating the feature point region and generating the first feature point set. This process ensures high accuracy and stability of the feature points, providing a solid foundation for subsequent tracking and positioning. As one of the core technologies of visual positioning, the feature point localization algorithm efficiently extracts key feature points from images and is crucial for achieving high-precision positioning.
[0058] The optical flow method is used to track points in the first feature point set, calculate the trajectory offset, and generate the first trajectory sequence. As a classic and efficient feature point tracking algorithm, the optical flow method can capture the trajectory of feature points in real time, providing the system with continuous motion information. This processing is particularly suitable for tracking feature points in dynamic environments, ensuring the stability of feature point positioning and tracking.
[0059] This step, through real-time acquisition of dynamic image data and optical flow tracking, can rapidly respond to environmental changes and significantly improve the system's real-time performance. Real-time performance is a key requirement for positioning and navigation in dynamic environments. This solution ensures both efficiency and real-time performance through optimized algorithms and efficient processing.
[0060] This solution organically combines multiple steps, including occlusion region identification, feature point location, and optical flow tracking, to form a complete feature point location and tracking process. This integrated design not only improves overall system performance but also significantly enhances the accuracy and stability of feature point location and tracking through real-time processing and dynamic adaptation.
[0061] By leveraging technologies such as occluded area recognition, feature point localization algorithms, and optical flow tracking, this solution significantly improves feature point localization and tracking capabilities in dynamic environments, providing a reliable foundation for subsequent positioning and navigation. This solution is particularly suitable for complex scenarios such as robot navigation, augmented reality (AR), virtual reality (VR), and intelligent building management, demonstrating significant technological advancement and broad market application value.
[0062] In some embodiments, matching the first feature sequence with the first position estimation data and fusing them to generate a first posture dataset includes:
[0063] Based on the first feature sequence and the first position estimation data, a refined match is performed, and the spatial geometric constraints of the inter-frame illumination change and the static background ratio are fused through the Kalman filter algorithm to generate a first posture data set.
[0064] In the aforementioned technical solution, this solution uses refined matching technology to accurately match feature sequences with position estimation data, significantly improving the accuracy and reliability of pose estimation. Refined matching ensures accurate correspondence between feature points, effectively reducing matching errors and providing a solid foundation for pose estimation.
[0065] The Kalman filter algorithm incorporates spatial geometric constraints such as inter-frame illumination variations and static background proportions to effectively handle noise and bias in the data, generating a stable and accurate pose dataset. As a classic state estimation algorithm, the Kalman filter can handle data fluctuations in dynamic environments in real time, ensuring the continuity and stability of pose estimation.
[0066] By integrating spatial geometric constraints such as inter-frame illumination variations and static background proportions, this solution comprehensively considers the impact of environmental factors on pose estimation, significantly improving the accuracy of the estimation results. The synergistic effect of multiple constraints enhances the robustness of pose estimation and effectively reduces estimation errors caused by environmental changes.
[0067] This step processes feature sequences and position estimation data in real time in a dynamic environment, adapting to environmental changes and ensuring the real-time and accuracy of posture data. Real-time performance is a key requirement for positioning and navigation in dynamic environments. This solution ensures system efficiency and real-time responsiveness through optimized algorithms and efficient processing mechanisms.
[0068] This solution organically combines multiple steps, including refined matching, the Kalman filter algorithm, and various constraints, to form a complete posture data generation process. This integrated design not only improves the overall performance of the system but also significantly enhances the accuracy and stability of posture estimation through real-time processing and dynamic adaptation.
[0069] By integrating multiple constraints through refined matching and a Kalman filter algorithm, this solution significantly improves pose estimation capabilities in dynamic environments, providing reliable data support for subsequent positioning and navigation. This solution is particularly suitable for complex scenarios such as robot navigation, augmented reality (AR), virtual reality (VR), and intelligent building management, demonstrating significant technological advancement and broad market application value.
[0070] In some embodiments, when a distance deviation between the first posture dataset and the three-dimensional model position exceeds a preset threshold, performing a local update on the indoor real scene three-dimensional model to generate a first three-dimensional model includes:
[0071] By comparing the first posture data set with the predicted position of the three-dimensional model, the distance deviation is calculated, and it is determined whether it exceeds a preset threshold to obtain a deviation detection result; if the deviation detection result exceeds the preset threshold, multi-view views are extracted from the dynamic image data, and a multi-view geometry algorithm is used to perform view registration to obtain a registered view set;
[0072] Based on the registered view set, the density distribution of feature points is analyzed, the high-density feature point areas are determined, and feature point optimization data is generated. The feature point optimization data is matched with the local area of the indoor real scene model, and the data is fused using the weighted average method to obtain the locally updated model fragment.
[0073] The three-dimensional geometric information is extracted from the locally updated model fragments, and the topological consistency with the first three-dimensional model is determined to obtain a topological verification result. If the topological verification result is consistent, the locally updated model fragments are integrated into the indoor real-scene model to generate a first three-dimensional model. Based on the geometric structure of the first three-dimensional model, the surface texture is optimized using a stereoscopic microscopy algorithm to obtain the final three-dimensional model output.
[0074] In the above technical solution, this approach accurately calculates the distance deviation by comparing the first pose dataset with the 3D model's predicted position and determines whether it exceeds a preset threshold. This deviation detection mechanism promptly detects deviations between the model and actual position, providing a reliable basis for subsequent model updates. As a key step in ensuring model accuracy, deviation detection can trigger timely model updates, significantly improving the model's real-time performance and accuracy.
[0075] Extract multi-viewpoint views from dynamic image data and efficiently register them using a multi-view geometry algorithm to generate a set of registered views. As a core technology for 3D reconstruction, multi-view geometry effectively processes images from different viewpoints, significantly improving the accuracy of view registration and making it suitable for updating 3D models in dynamic environments.
[0076] Based on the registered view set, the density distribution of feature points is analyzed, high-density feature point areas are identified, and optimized feature point data is generated. This process highlights key features in the scene and provides more reliable data support for model updates. Feature point optimization not only enhances the model's discrimination and stability but also provides a high-quality data foundation for subsequent model matching and fusion.
[0077] By matching the optimized feature point data with local areas of the indoor real-world model and fusing the data using a weighted average method, we generate locally updated model fragments. This local update mechanism dynamically adjusts the model to adapt to environmental changes while avoiding the high computational cost of global updates, ensuring the real-time and accuracy of the model.
[0078] The 3D geometric information is extracted from the locally updated model fragments, and their topological consistency with the original 3D model is determined to obtain a topological verification result. If the topological verification result is consistent, the locally updated model fragments are seamlessly integrated into the indoor real-world model. This verification mechanism ensures the accuracy and completeness of model updates, avoids model structural errors caused by local updates, and ensures the overall consistency and reliability of the model.
[0079] Finally, based on the geometry of the first 3D model, a stereoscopic microscopy algorithm is used to optimize the surface texture and generate the final 3D model output. This process significantly improves the detail and quality of the model, enhancing the visual effect of the model and making it suitable for application scenarios requiring high-precision models.
[0080] This step organically combines multiple steps, including deviation detection, multi-view geometry algorithms, feature point optimization, local updates, topology verification, and surface texture optimization, to form a complete model update process. This integrated design not only improves the overall performance of the system but also significantly enhances the accuracy and efficiency of updating indoor real-world 3D models through real-time processing and dynamic adaptation. This solution is particularly suitable for complex scenarios such as robot navigation, augmented reality (AR), virtual reality (VR), and intelligent building management, offering significant technological advancements and broad market application value.
[0081] In some embodiments, matching the first three-dimensional model with the first feature sequence to generate a first positioning dataset includes:
[0082] The geometric structure and texture density are obtained from the first three-dimensional model, and key point descriptions are extracted in combination with the first feature sequence to generate initial matching parameters. The initial matching parameters are optimized using a bundle adjustment algorithm to correct for texture density deviations to obtain a first optimized model. If the view rotation matrix consistency of the first optimized model is lower than a preset threshold, the view rotation parameters are iteratively adjusted to obtain a second optimized model.
[0083] A lightweight network is used to compress the feature sequence of the second optimization model to generate a compressed feature set; based on the compressed feature set and the second optimization model, a fast matching calculation is performed to obtain a low-latency matching result; based on the low-latency matching result, a first positioning data set is generated; for the first positioning data set, the texture density and perspective rotation consistency are verified to obtain the final positioning data set.
[0084] In the aforementioned technical solution, this approach obtains geometric structure and texture density information, combines it with the first feature sequence to accurately extract key point descriptions and generate initial matching parameters. This multi-dimensional parameter generation method significantly improves positioning accuracy and provides more comprehensive positioning information for subsequent processing.
[0085] The bundle adjustment algorithm is used to optimize the initial matching parameters, correct for texture density deviations, and generate the first optimized model. As a classic method for optimizing camera parameters and 3D point positions, the bundle adjustment algorithm can significantly improve model accuracy and ensure the reliability of positioning and modeling.
[0086] If the consistency of the view rotation matrix of the first optimized model falls below a preset threshold, the view rotation parameters are iteratively adjusted to generate a second optimized model. This consistency verification mechanism effectively avoids positioning errors caused by view changes and ensures the stability and reliability of the model.
[0087] The lightweight network efficiently compresses the feature sequence of the second optimization model to generate a compact compressed feature set. This compression process reduces the amount of data while retaining key feature information, significantly improving data processing efficiency and laying the foundation for subsequent calculations.
[0088] Based on the compressed feature set and the second optimization model, a fast matching calculation is performed to generate low-latency matching results. The fast matching calculation significantly improves the real-time performance of the system, meeting the critical requirements of real-time positioning and navigation in dynamic environments.
[0089] The final positioning dataset is generated by verifying the texture density and view rotation consistency of the low-latency matching results. This verification mechanism ensures the accuracy and reliability of the positioning data, avoids positioning errors caused by data inconsistencies, and improves the credibility of the positioning results.
[0090] This step organically combines multiple steps, including geometric structure acquisition, feature extraction, bundle adjustment optimization, view rotation adjustment, feature compression, and fast matching, to form a complete positioning data generation process. This integrated design not only improves the overall performance of the system but also significantly enhances positioning accuracy and efficiency through real-time processing and dynamic adaptation.
[0091] This solution significantly improves positioning accuracy and efficiency through the integration of geometric structure and texture density, bundle adjustment algorithm optimization, view rotation matrix consistency verification, feature sequence compression, and fast matching calculation. It provides reliable technical support for real-time positioning and navigation in dynamic environments. This solution is particularly suitable for complex scenarios such as robot navigation, augmented reality (AR), virtual reality (VR), and intelligent building management, demonstrating significant technological advancement and broad market application value.
[0092] In some embodiments, a first navigation instruction sequence is generated based on the first positioning data set in combination with spatial boundary constraints in a dynamic environment by calculating a navigation path change trend taking into account a weight of a proportion of an occluded area, including:
[0093] Obtain the first positioning data set, combine it with the spatial boundaries in the dynamic environment, and extract environmental constraint features through preprocessing to obtain an initial positioning distribution. If there are occluded areas in the initial positioning distribution, use a region segmentation method to determine the range of the occluded areas and obtain the occluded area distribution. Based on the occluded area distribution, calculate the proportional weight of each area, generate a weight matrix through the weight allocation model, and determine the path optimization parameters.
[0094] The weight matrix is fused through the A-star algorithm, and the navigation path is calculated based on the spatial boundary constraints to obtain the path change trend. If there is a deviation in the path change trend, the path is recalculated by dynamically adjusting the weight matrix to obtain the optimized navigation path. Based on the optimized navigation path, the first navigation instruction sequence is generated to determine the final instruction output.
[0095] In the aforementioned technical solution, this solution extracts environmental constraint features through preprocessing and generates an initial positioning distribution, providing a solid foundation for path planning. The extraction of environmental constraint features effectively identifies boundaries and obstacles in the environment, ensuring the rationality and feasibility of path planning.
[0096] When occluded regions exist in the initial positioning distribution, a region segmentation method is used to precisely determine the extent of these regions and generate an occluded region distribution. This process effectively identifies occluded regions, avoids blind spots in path planning, and improves the robustness and accessibility of path planning.
[0097] Based on the distribution of blocked areas, the weight of each area is calculated, and a weight matrix is generated through a weight distribution model to determine the path optimization parameters. This weight distribution optimizes path planning, balancing path length, safety, and efficiency, ensuring optimal path planning.
[0098] By integrating the weight matrix with the A-star algorithm, the navigation path is calculated based on spatial boundary constraints and path change trends are generated. As a classic path planning algorithm, the A-star algorithm can effectively handle complex environmental constraints, improving the efficiency and accuracy of path planning.
[0099] If there are deviations in the path change trend, the path is recalculated by dynamically adjusting the weight matrix to generate an optimized navigation path. This dynamic adjustment ensures the real-time and adaptability of path planning and improves the reliability of navigation.
[0100] Finally, a first sequence of navigation instructions is generated based on the optimized navigation path, and the final instruction output is determined. This process converts the path planning results into specific navigation instructions, guiding the movement of the robot or device and ensuring that the device operates efficiently according to the planned path.
[0101] This step organically combines multiple steps, including environmental constraint feature extraction, occlusion area processing, weight matrix generation, A-star algorithm integration, path optimization and adjustment, and navigation command generation, to form a complete path planning process. This integrated design significantly improves the overall performance and real-time performance of the system, making it particularly suitable for real-time path planning and navigation in dynamic environments.
[0102] This solution significantly improves the accuracy and efficiency of path planning through technical means such as environmental constraint feature extraction, occlusion area processing, weight matrix generation, A-star algorithm fusion, and path optimization adjustment. It provides reliable technical support for complex scenarios such as robot navigation, augmented reality (AR), virtual reality (VR), and intelligent building management.
[0103] According to another aspect of the present invention, a device for visual positioning based on a fine indoor three-dimensional model is provided, based on the above method, comprising:
[0104] A feature data set generation module is used to acquire the indoor original image through the camera and synchronously collect the camera's attitude angular velocity, construct a first feature data set, and perform correction to generate a second feature data set;
[0105] A feature sequence generation module is configured to align the second feature data set with a pre-built indoor real-scene three-dimensional model to generate first position estimation data under spatial boundary constraints; based on the first position estimation data, locate feature point regions from the real-time acquired dynamic image data, and track point motion trajectories in the movement trajectory offset to generate a first feature sequence;
[0106] a posture data generation and update module, configured to match the first position estimation data with the first feature sequence and fuse them to generate a first posture data set; and when the distance deviation between the first posture data set and the three-dimensional model position exceeds a preset threshold, perform a local update on the indoor real scene three-dimensional model to generate a first three-dimensional model;
[0107] a positioning data optimization module, configured to match the first three-dimensional model with the first feature sequence to generate a first positioning data set;
[0108] The navigation path generation module is used to generate a first navigation instruction sequence by calculating a navigation path change trend taking into account a weight of a proportion of an occluded area based on the first positioning data set in combination with spatial boundary constraints in a dynamic environment.
[0109] In the above technical solution, in order to better use the above method, this application proposes a testing device based on dynamic tracking of code links. Each module corresponds to each step of the above method. The specific principles have been described above and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0110] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0111] Figure 1 This is a flow chart of an embodiment of a method for visual positioning based on a fine indoor three-dimensional model according to the present invention;
[0112] Figure 2 It is a structural diagram of an embodiment of a device for visual positioning based on a fine indoor three-dimensional model of the present invention. DETAILED DESCRIPTION
[0113] The present invention will be described in further detail below with reference to the accompanying drawings and examples. It is particularly noted that the following examples are intended only to illustrate the present invention and are not intended to limit the scope of the present invention. Similarly, the following examples are only some embodiments of the present invention and are not intended to be exhaustive. All other embodiments obtained by those of ordinary skill in the art without creative effort are intended to fall within the scope of protection of the present invention.
[0114] Example 1
[0115] See also Figure 1 , a method for visual positioning based on a fine indoor three-dimensional model, the method comprising:
[0116] S1. Acquire the indoor original image through the camera and synchronously collect the camera's attitude angular velocity to construct a first feature data set, and perform correction to generate a second feature data set;
[0117] S2. Aligning the second feature data set with the pre-built indoor real-scene three-dimensional model to generate first position estimation data under spatial boundary constraints; locating the feature point area from the real-time acquired dynamic image data based on the first position estimation data, and tracking the point motion trajectory in the movement trajectory offset to generate a first feature sequence;
[0118] S3, matching the first feature sequence with the first position estimation data, and fusing them to generate a first posture dataset; when the distance deviation between the first posture dataset and the three-dimensional model position exceeds a preset threshold, performing a local update on the indoor real scene three-dimensional model to generate a first three-dimensional model;
[0119] S4. Matching the first three-dimensional model with the first feature sequence to generate a first positioning data set;
[0120] S5. Generate a first navigation instruction sequence by calculating a navigation path change trend taking into account a weight of a proportion of an occluded area based on the first positioning data set and the spatial boundary constraints in a dynamic environment.
[0121] In the aforementioned technical solution, this approach uses multi-source data fusion technology to combine image data collected by the camera with attitude angular velocity to construct and correct a feature dataset. This design significantly improves positioning accuracy and robustness. The introduction of attitude angular velocity dynamically compensates for high-frequency jitter or systematic errors in the image data, ensuring the stability of the feature data and significantly improving the reliability of positioning results.
[0122] In dynamic environments, this solution uses feature point area positioning and trajectory tracking technology to generate feature sequences in real time, combining them with position estimation data for matching and posture fusion. This design effectively adapts to environmental changes, ensuring the system's stability and adaptability in dynamic scenarios. Through feature point tracking and trajectory offset processing, the system updates positioning data in real time, avoiding positioning deviations caused by environmental changes and ensuring the real-time and accurate positioning results.
[0123] When the positional deviation between the pose dataset and the 3D model exceeds a preset threshold, this solution uses a local update mechanism to dynamically adjust the 3D model structure. This design avoids the high computational cost of global updates while ensuring the real-time and accuracy of the model, making it particularly suitable for efficient positioning in dynamic environments.
[0124] In terms of path planning, this solution introduces weights for the proportion of occluded areas, dynamically calculates the changing trends of the navigation path, and generates a sequence of navigation instructions. This path planning method effectively avoids occluded areas, optimizes path selection, reduces blind spots in path planning, and significantly improves navigation efficiency and success rate.
[0125] During the alignment process between feature data and the 3D model, this solution introduces spatial boundary constraints to generate position estimates. These constraints limit the range of positioning results, prevent drift, and significantly improve the stability of positioning results, especially in complex or dynamic environments.
[0126] This solution emphasizes real-time performance, ensuring rapid response in steps such as feature point area positioning, trajectory tracking, and model updates through optimized algorithms and local update mechanisms. This design significantly improves the overall efficiency of the system in dynamic environments and meets real-time requirements.
[0127] This solution is applicable to a variety of indoor scenarios, including robot navigation, augmented reality (AR), virtual reality (VR), and intelligent building management. This wide range of applications significantly enhances the solution's market value. Furthermore, building on existing technologies, this solution introduces innovative features such as feature point tracking, local updates, and path optimization, addressing the limitations of existing technologies in dynamic environments and representing a significant technological advancement.
[0128] In summary, this solution significantly improves the accuracy, stability, and efficiency of indoor visual positioning and navigation through technical means such as multi-source data fusion, dynamic environment adaptability, local update mechanism, and path planning optimization, providing an efficient and reliable solution for positioning and navigation in dynamic environments.
[0129] In this embodiment, the camera acquires the indoor original image and simultaneously collects the camera's attitude angular velocity to construct a first feature data set, including:
[0130] Acquire raw image data from indoor cameras and synchronously collect the attitude and angular velocity recorded by the sensor to obtain an initial data set. If the timestamps of the image data and angular velocity data in the initial data set are aligned, correct the deviation using a time axis alignment algorithm to obtain a synchronized data set.
[0131] Based on the synchronized data set, the Kalman filter algorithm is used to fuse the movement trajectory offset to obtain trajectory optimization data. For the image sequence in the trajectory optimization data, the inter-frame illumination change characteristics are calculated to obtain an illumination feature sequence. Multidimensional description vectors are extracted from the illumination feature sequence and the trajectory optimization data to obtain a vector representation set.
[0132] If the dimension of the vector representation set meets the preset threshold, the principal component analysis algorithm is used for dimensionality reduction to obtain the first feature data set.
[0133] For example, in an indoor environment, an RGB camera with a resolution of 1920×1080 is deployed to collect raw image data at a frame rate of 30, while an MPU6050 sensor is used to synchronously record the three-axis angular velocity at a frequency of 100Hz (range ±2000° / s, accuracy 0.1°). The image data is detected in real time by the YOLOv5 algorithm using the ORB feature points in the scene, extracting 500-800 feature descriptors (256-dimensional vectors) per frame. At the same time, the LK optical flow method is used to calculate the displacement of feature points between adjacent frames, with an average of 12.3 pixels (standard deviation ±3.5). The sensor data is converted from angular velocity to attitude angle using the fourth-order Runge-Kutta integral algorithm, and its pitch angle error is controlled within 0.5°. To fuse multi-source data, an extended Kalman filter (EKF) was constructed. The state vector contained 6-DOF pose (x, y, z displacement accuracy ±2 mm, roll / pitch / yaw angle error ±0.3°) and illumination parameters (HSV spatial brightness rate of change 0.15 / s). The process noise covariance matrix Q was set to diag(0.1, 0.1, 0.1, 0.05, 0.05, 0.05, 0.01). The feature dataset was compressed using PCA to reduce the ORB descriptor to 64 dimensions, which was then combined with the pose data to form a 720-dimensional feature vector.
[0134] In this embodiment, performing correction to generate the second feature data set includes:
[0135] A deep learning network is used to extract dynamic blur features from the first feature dataset, and the view rotation matrix is calculated based on the attitude angular velocity to generate a corrected image feature set. If the timestamp of the corrected image feature set is aligned with the sensor data, a time synchronization algorithm is used to adjust the deviation to obtain a synchronized feature set.
[0136] Based on the synchronized feature set, texture density features are extracted to generate an intermediate feature set containing texture information. Target area detection is performed on the intermediate feature set through a convolutional neural network to obtain a target distance deviation vector. If the dimension of the target distance deviation vector exceeds a preset threshold, a dimensionality reduction algorithm is used to generate an optimized feature vector. Based on the optimized feature vector, the texture density features and the target distance deviation are fused to generate a second feature data set.
[0137] For example, ResNet50 was first used as the base network architecture, with an RGB image with a resolution of 256×256 as input. When extracting the initial feature map through the convolutional layer, the convolution kernel size was set to 3×3, the stride was 1, the padding was 1, and the ReLU activation function was used. During the dynamic blur feature extraction stage, a temporal module consisting of three LSTM layers was designed, with 128 hidden units in each layer. When processing a sequence of five consecutive image frames, the 2048-dimensional features output by the average pooling of the last layer of ResNet50 were used as the LSTM input. A time-expansion calculation was performed to obtain a blur score for each frame. The value was normalized to a range of 0-1, with values above 0.7 considered severe blur. For attitude and angular velocity data, an IMU sensor acquires three-axis angular velocity at a frequency of 100Hz. The rotation matrix is calculated using the fourth-order Runge-Kutta method. The correction process is triggered when the root mean square value of the angular velocity exceeds 15° / s. SVD decomposition is used to solve the optimal rotation matrix. The texture density of the corrected image is calculated using the local binary pattern (LBP) algorithm. Eight sampling points with a neighborhood radius of 2 pixels are set. The LBP histogram entropy of each 16×16 image block is calculated as a texture density indicator. Areas with an entropy value below 2.3 are identified as texture-missing. The resulting second feature dataset contains the motion blur score (e.g., 0.68), the corrected texture density (e.g., 2.8), and the target distance deviation (disparity calculated using a stereo matching algorithm and converted to actual distance, exemplified by 1.2 meters) for each frame. These features are fused through a fully connected layer and input into the subsequent decision module.
[0138] In this embodiment, aligning the second feature dataset with the pre-built indoor real scene 3D model to generate first position estimation data under spatial boundary constraints includes:
[0139] The second feature data set is roughly aligned with the pre-established indoor real-scene 3D model. The sampling interval of feature points is optimized through acquisition frequency analysis to generate a uniformly distributed feature point set. If the acquisition frequency fluctuation exceeds a preset threshold, the sampling interval is optimized through an adaptive adjustment algorithm to obtain a uniformly distributed feature point set.
[0140] Extract key points of the second feature data set based on the evenly distributed feature point set to generate an initial feature descriptor; use a key point matching algorithm to compare the initial feature descriptor with a reference point set of the indoor real-scene 3D model to obtain a roughly aligned transformation matrix;
[0141] Using a transformation matrix, the second feature data set is preliminarily aligned with the indoor real-scene 3D model to generate an aligned feature point cloud. If the deviation between the aligned feature point cloud and the 3D model exceeds a preset threshold, the transformation matrix is adjusted using an iterative closest point algorithm to obtain an optimized feature point cloud.
[0142] Based on the optimized feature point cloud and combined with the spatial boundary constraints, the first position estimation data is generated; through the spatial boundary constraint check, if the first position estimation data exceeds the boundary range, the estimation parameters are adjusted to obtain the estimation data that meets the constraints.
[0143] For example, when roughly aligning the second feature data set with the indoor real-scene three-dimensional model, a fast matching algorithm based on feature descriptors can be used. For example, the SIFT algorithm is used to extract key points and calculate a 128-dimensional feature vector. The nearest neighbor search is performed through the FLANN matcher, and the matching threshold ratio is set to 0.7 to filter out false matches. For the optimization of acquisition frequency fluctuations, an adaptive sampling strategy can be used. When the density of feature points is detected to be lower than 0.2 per square meter, the lidar scanning frequency is increased from 10Hz to 15Hz. At the same time, a Gaussian mixture model is used to analyze the spatial distribution of feature points. If the standard deviation of a certain area exceeds 1.5 meters, supplementary sampling is triggered. When generating spatial boundary constraints, the RANSAC algorithm is used to fit the plane boundary, the number of iterations is set to 500 times, the internal point threshold is set to 0.05 meters, and when the deviation between the wall point cloud and the model is detected to be greater than 0.1 meters, local optimization is performed using the ICP algorithm, and the iteration error is set to 1e-5. For the first position estimation, a graph optimization framework is used to construct a pose graph, and feature point observations are used as edge constraints. The diagonal elements of the information matrix are set to [1, 1, 1, 0.5, 0.5, 0.5] corresponding to the translation and rotation weights. The Levenberg-Marquardt algorithm is used to solve the problem, and the final pose estimate is output when the median reprojection error is less than 2 pixels.
[0144] In this embodiment, based on the first position estimation data, the feature point area is located from the dynamic image data collected in real time, and the motion trajectory of the points in the moving trajectory offset is tracked to generate a first feature sequence, including: obtaining the dynamic image data collected in real time, determining the proportion of the occluded area, and obtaining a first image sequence; processing the first image sequence through a feature point positioning algorithm, locating the feature point area, and obtaining a first feature point set; using the optical flow method to track the points in the first feature point set, calculating the motion trajectory offset, and obtaining a first trajectory sequence.
[0145] For example, specifically, when a mobile camera collects dynamic image data in real time, the proportion of the occluded area is first calculated through an image processing algorithm. For example, a threshold segmentation method based on pixel grayscale value is used, the threshold is set to 128, and the image is divided into an occluded area and a non-occluded area. The pixel ratio of the occluded area is statistically calculated to be 15%. Next, the SIFT algorithm is used to locate the feature point area and extract the key points in the image. For example, 200 feature points are detected in an image with a resolution of 640×480. Then, the motion trajectory of these feature points is tracked using the optical flow method. The Lucas-Kanade optical flow algorithm is used, the window size is set to 15×15 pixels, and the displacement of each feature point in consecutive frames is calculated. For example, the displacement of a feature point between two frames is (5, 3) pixels. In order to eliminate unstable feature points, a check is performed based on geometric consistency, and the relative position changes between feature points are calculated. The error threshold is set to 2 pixels, and feature points with errors greater than the threshold are eliminated. Finally, 150 stable feature points are retained. Finally, the motion trajectories of these stable feature points are arranged in a time series to generate the first feature sequence. For example, the sequence length is 30 frames, with each frame containing the coordinate information of 150 feature points. Through these steps, the complete processing flow from image acquisition to feature sequence generation is completed, providing a reliable data foundation for subsequent analysis.
[0146] In this embodiment, matching the first feature sequence with the first position estimation data and fusing them to generate a first posture dataset includes:
[0147] Based on the first feature sequence and the first position estimation data, a refined match is performed, and the spatial geometric constraints of the inter-frame illumination change and the static background ratio are integrated through the Kalman filter algorithm to generate a first posture data set.
[0148] For example, in the fine matching process, first, key points are extracted by the first feature sequence, such as detecting feature points in the image using the SIFT algorithm, and generating 128-dimensional descriptors. Assuming that the input image resolution is 1920x1080, 500 feature points are detected, the descriptors of each feature point are matched with the first position estimation data, the nearest neighbor distance ratio is calculated using the KNN algorithm (k=2), the threshold is set to 0.7, and 300 matching points are selected. Next, the Kalman filter algorithm is used to model the inter-frame illumination change, assuming that the illumination change follows a Gaussian distribution, the initial state covariance matrix is 0.1, the process noise covariance is 0.01, and the observation noise covariance is 0.05. Through the prediction and update steps, the illumination parameters are gradually optimized. At the same time, combined with the spatial geometric constraint of the static background ratio, assuming that the static background ratio is 70%, the RANSAC algorithm is used to estimate the fundamental matrix, and the false matching points of the dynamic background are removed, finally 200 accurate matching points are retained. Based on these matching points, the camera pose is solved by the PnP algorithm, assuming that the camera intrinsic matrix is [800, 0, 960; 0, 800, 540; 0, 0, 1], the first pose data set is generated, including the position and rotation matrix of the camera, for example, the position is [1.2, 0.8, 3.5], and the rotation matrix is [0.99, -0.05, 0.12; 0.06, 0.99, -0.03; -0.12, 0.04, 0.99], thereby completing the entire fine matching and pose estimation process.
[0149] In this embodiment, when the distance deviation of the first pose data set and the three-dimensional model position exceeds the preset threshold, the indoor real scene three-dimensional model is locally updated to generate the first three-dimensional model, including:
[0150] By comparing the first pose data set with the three-dimensional model predicted position, the distance deviation is calculated to determine whether it exceeds the preset threshold, and the deviation detection result is obtained; if the deviation detection result exceeds the preset threshold, multi-view views are extracted from the dynamic image data, and multi-view geometric algorithm is used for view registration to obtain the registered view set;
[0151] According to the registered view set, the feature point density distribution is analyzed to determine the high-density feature point region, and the feature point optimization data is generated; the feature point optimization data is matched with the local area of the indoor real scene model, the weighted average method is used to fuse the data, and the locally updated model segment is obtained;
[0152] The three-dimensional geometric information is extracted from the locally updated model segment, and the topological consistency with the first three-dimensional model is judged to obtain the topological verification result; if the topological verification result is consistent, the locally updated model segment is integrated into the indoor real scene model to generate the first three-dimensional model; according to the geometric structure of the first three-dimensional model, the surface texture is optimized by the stereomicroscopic algorithm to obtain the final three-dimensional model output.
[0153] For example, when the distance deviation between the first pose dataset and the predicted position of the 3D model exceeds a preset threshold, the system first processes the dynamic image data using a multi-view geometry algorithm to extract key feature points. For example, when a deviation exceeding 5 cm is detected in a certain area, the system uses the SIFT algorithm to extract at least 100 feature points from images from multiple viewpoints and calculates their positions in 3D space. Next, the system performs local optimization on high-density areas based on the density distribution of feature points. For example, if the density of feature points in a certain area reaches 50 per square meter, the system uses the Bundle Adjustment algorithm to optimize these feature points to ensure the accuracy of their 3D coordinates. Finally, the system fuses the optimized feature points with the original 3D model to generate a first 3D model. For example, the optimized feature points are registered with the model using the ICP algorithm, ensuring that the accuracy of local updates is within 2 mm. This entire process is automated, ensuring real-time updates and high-precision matching of the 3D model.
[0154] In this embodiment, matching a first three-dimensional model with a first feature sequence to generate a first positioning dataset includes: obtaining a geometric structure and texture density from the first three-dimensional model, extracting key point descriptions in combination with the first feature sequence, and generating initial matching parameters; optimizing the initial matching parameters using a bundle adjustment algorithm, correcting for texture density deviations, and obtaining a first optimized model; and if the view rotation matrix consistency of the first optimized model is lower than a preset threshold, iteratively adjusting the view rotation parameters to obtain a second optimized model.
[0155] A lightweight network is used to compress the feature sequence of the second optimization model to generate a compressed feature set; based on the compressed feature set and the second optimization model, a fast matching calculation is performed to obtain a low-latency matching result; based on the low-latency matching result, a first positioning data set is generated; for the first positioning data set, the texture density and perspective rotation consistency are verified to obtain the final positioning data set.
[0156] For example, during the 3D model optimization process, a bundle adjustment algorithm is used to optimize the model's texture density deviation based on the first 3D model and the first feature sequence. By calculating the local density distribution of the model's surface texture and setting a target density of 100 texture points per square centimeter, the texture mapping parameters are iteratively adjusted using the bundle adjustment algorithm to keep the deviation between the actual density and the target density within ±5%. Simultaneously, to optimize the consistency of the view rotation matrix, a quaternion-based rotation matrix optimization method is used to minimize the difference in rotation matrices between adjacent viewpoints, ensuring that the error angle of the rotation matrix does not exceed 0.1 degrees. During the lightweight network compression stage, a lightweight network structure based on depthwise separable convolution is used to reduce the matching calculation latency from the original 50 milliseconds to 10 milliseconds, significantly improving computational efficiency. Ultimately, through the above optimization and compression processes, a first positioning dataset is generated. This dataset contains 1,000 samples, each of which contains 3D model coordinates, texture density distribution, and view rotation matrix information, providing high-quality data support for subsequent positioning and navigation tasks.
[0157] In this embodiment, a first navigation instruction sequence is generated by calculating a navigation path change trend that takes into account the weight of the proportion of occluded areas based on the first positioning data set in combination with spatial boundary constraints in a dynamic environment. The method includes: obtaining the first positioning data set, extracting environmental constraint features through preprocessing in combination with the spatial boundaries in the dynamic environment, and obtaining an initial positioning distribution; if there are occluded areas in the initial positioning distribution, determining the range of the occluded areas using a region segmentation method to obtain an occluded area distribution; calculating the proportion weight of each area based on the occluded area distribution, generating a weight matrix through a weight allocation model, and determining path optimization parameters;
[0158] The weight matrix is fused through the A-star algorithm, and the navigation path is calculated based on the spatial boundary constraints to obtain the path change trend. If there is a deviation in the path change trend, the path is recalculated by dynamically adjusting the weight matrix to obtain the optimized navigation path. Based on the optimized navigation path, the first navigation instruction sequence is generated to determine the final instruction output.
[0159] For example, in a dynamic environment, first, a first positioning data set is collected by a laser radar, for example, point cloud data in a range of 10 m x 10 m is obtained, the sampling frequency is 20 Hz, and the coordinate accuracy is ± 2 cm. After spatial indexing of the point cloud using a KD tree, the wall surface boundary is fitted based on the RANSAC algorithm, and three linear boundary equations are extracted: y = 0, x = 10, y = 0.5x + 5, and the fitting error threshold is set to 0.03 m. Then, a dynamic cost map is constructed, and the trajectory prediction result of the moving obstacle (for example, a pedestrian moving at a speed of 1.2 m / s in a direction of 30 degrees) is converted into a time sequence grid, each grid cell is 0.1 m x 0.1 m, and the time resolution is 0.5 s. In the A-star algorithm, an occlusion proportion weight factor a is introduced, the value of a is determined by calculating the line-of-sight flux, for example, when the obstacle coverage rate in the field of view angle of 60 degrees of the sensor reaches 40%, a is taken as 0.6. The Manhattan distance is used as the heuristic function in path searching, the node expansion step is 0.2 m, and the steering constraint (maximum turning angle 0.5 radian) is considered. The final generated navigation instruction sequence contains 5 waypoints: (2.1, 3.4), (4.7, 5.2), (6.8, 6.0), (8.3, 4.9), (9.6, 2.1), each waypoint is accompanied by a speed instruction in the interval of 0.3-0.8 m / s, and the path curvature is ensured to be continuous through cubic spline interpolation. In the path optimization stage, the gradient descent method is used to locally adjust the path segment in the occlusion area, and after 10 iterations, the occlusion exposure time is reduced from 3.2 s to 1.8 s.
[0160] Embodiment two
[0161] Please refer to Figure 2 An apparatus for visual positioning based on an indoor fine three-dimensional model, based on the above method, comprising:
[0162] A feature data set generation module for acquiring indoor raw images through a camera and synchronously collecting the attitude angular velocity of the camera, constructing a first feature data set, and generating a second feature data set after correction;
[0163] A feature sequence generation module for aligning the second feature data set with a pre-constructed indoor real scene three-dimensional model to generate a first position estimation data under spatial boundary constraints; positioning a feature point region from real-time collected dynamic image data according to the first position estimation data, and tracking the motion trajectory of the points in the moving trajectory deviation to generate a first feature sequence;
[0164] A pose data generation and update module for matching the first feature sequence with the first position estimation data, and fusing to generate a first pose data set; when the distance deviation of the first pose data set from the position of the three-dimensional model exceeds a preset threshold, performing local update on the indoor real scene three-dimensional model to generate a first three-dimensional model;
[0165] a positioning data optimization module, configured to match the first three-dimensional model with the first feature sequence to generate a first positioning data set;
[0166] The navigation path generation module is used to generate a first navigation instruction sequence by calculating a navigation path change trend taking into account a weight of a proportion of an occluded area based on the first positioning data set in combination with spatial boundary constraints in a dynamic environment.
[0167] In this embodiment, in the above technical solution, in order to better use the above method, the present application proposes a testing device based on dynamic tracking of code links. Each module corresponds to each step of the above method. The specific principles have been described above and will not be repeated here.
[0168] The above descriptions are only some embodiments of the present invention and do not limit the scope of protection of the present invention. Any equivalent device or equivalent process transformation made by using the contents of the description and drawings of the present invention, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A method for visual positioning based on indoor fine three-dimensional models, characterized in that: The method comprises: The camera acquires the indoor original image and synchronously collects the camera's attitude angular velocity to construct a first feature data set, and performs correction to generate a second feature data set; Aligning the second feature data set with a pre-built indoor real-scene three-dimensional model to generate first position estimation data under spatial boundary constraints; locating feature point regions from the real-time acquired dynamic image data based on the first position estimation data, and tracking point motion trajectories in the movement trajectory offset to generate a first feature sequence; Matching the first feature sequence with the first position estimation data and fusing them to generate a first posture dataset; when the distance deviation between the first posture dataset and the three-dimensional model position exceeds a preset threshold, performing a local update on the indoor real scene three-dimensional model to generate a first three-dimensional model; Matching the first three-dimensional model with the first feature sequence to generate a first positioning data set; According to the first positioning data set combined with the spatial boundary constraints in the dynamic environment, a first navigation instruction sequence is generated by calculating a navigation path change trend taking into account the weight of the occluded area ratio.
2. The method for visual positioning based on an indoor fine three-dimensional model according to claim 1, characterized in that: The camera acquires the original indoor image and simultaneously collects the camera's attitude and angular velocity to construct the first feature dataset, including: Acquire raw image data from indoor cameras and synchronously collect the attitude and angular velocity recorded by the sensor to obtain an initial data set. If the timestamps of the image data and angular velocity data in the initial data set are aligned, correct the deviation using a time axis alignment algorithm to obtain a synchronized data set. Based on the synchronized data set, the Kalman filter algorithm is used to fuse the movement trajectory offset to obtain trajectory optimization data. For the image sequence in the trajectory optimization data, the inter-frame illumination change characteristics are calculated to obtain an illumination feature sequence. Multidimensional description vectors are extracted from the illumination feature sequence and the trajectory optimization data to obtain a vector representation set. If the dimension of the vector representation set meets the preset threshold, the principal component analysis algorithm is used for dimensionality reduction to obtain the first feature data set.
3. The method for visual positioning based on indoor fine three-dimensional model according to claim 1, characterized in that: Correction is performed to generate a second feature data set, including: A deep learning network is used to extract dynamic blur features from the first feature dataset, and the view rotation matrix is calculated based on the attitude angular velocity to generate a corrected image feature set. If the timestamp of the corrected image feature set is aligned with the sensor data, a time synchronization algorithm is used to adjust the deviation to obtain a synchronized feature set. Based on the synchronized feature set, texture density features are extracted to generate an intermediate feature set containing texture information. Target area detection is performed on the intermediate feature set through a convolutional neural network to obtain a target distance deviation vector. If the dimension of the target distance deviation vector exceeds a preset threshold, a dimensionality reduction algorithm is used to generate an optimized feature vector. Based on the optimized feature vector, the texture density features and the target distance deviation are fused to generate a second feature data set.
4. The method for visual positioning based on indoor fine three-dimensional model according to claim 1, characterized in that: Aligning the second feature dataset with the pre-built indoor real scene 3D model to generate first position estimation data under spatial boundary constraints, including: The second feature data set is roughly aligned with the pre-established indoor real-scene 3D model. The sampling interval of feature points is optimized through acquisition frequency analysis to generate a uniformly distributed feature point set. If the acquisition frequency fluctuation exceeds a preset threshold, the sampling interval is optimized through an adaptive adjustment algorithm to obtain a uniformly distributed feature point set. Extract key points of the second feature data set based on the evenly distributed feature point set to generate an initial feature descriptor; use a key point matching algorithm to compare the initial feature descriptor with a reference point set of the indoor real-scene 3D model to obtain a roughly aligned transformation matrix; Using a transformation matrix, the second feature data set is preliminarily aligned with the indoor real-scene 3D model to generate an aligned feature point cloud. If the deviation between the aligned feature point cloud and the 3D model exceeds a preset threshold, the transformation matrix is adjusted using an iterative closest point algorithm to obtain an optimized feature point cloud. Based on the optimized feature point cloud and combined with the spatial boundary constraints, the first position estimation data is generated; through the spatial boundary constraint check, if the first position estimation data exceeds the boundary range, the estimation parameters are adjusted to obtain the estimation data that meets the constraints.
5. The method for visual positioning based on indoor fine three-dimensional model according to claim 1, characterized in that: Locating a feature point region from the real-time collected dynamic image data according to the first position estimation data, and tracking the point motion trajectory in the movement trajectory offset to generate a first feature sequence, including: The real-time collected dynamic image data is acquired, the occlusion area ratio is determined, and a first image sequence is obtained. The first image sequence is processed by a feature point positioning algorithm to locate the feature point area and obtain a first feature point set. The points in the first feature point set are tracked using an optical flow method, and the motion trajectory offset is calculated to obtain a first trajectory sequence.
6. The method for visual positioning based on indoor fine three-dimensional model according to claim 1, characterized in that: Matching the first feature sequence with the first position estimation data and fusing them to generate a first posture dataset includes: Based on the first feature sequence and the first position estimation data, a refined match is performed, and the spatial geometric constraints of the inter-frame illumination change and the static background ratio are integrated through the Kalman filter algorithm to generate a first posture data set.
7. The method for visual positioning based on indoor fine three-dimensional model according to claim 1, characterized in that: When the distance deviation between the first posture dataset and the three-dimensional model position exceeds a preset threshold, performing a local update on the indoor real scene three-dimensional model to generate a first three-dimensional model, including: By comparing the first posture data set with the predicted position of the three-dimensional model, the distance deviation is calculated, and it is determined whether it exceeds a preset threshold to obtain a deviation detection result; if the deviation detection result exceeds the preset threshold, multi-view views are extracted from the dynamic image data, and a multi-view geometry algorithm is used to perform view registration to obtain a registered view set; Based on the registered view set, the density distribution of feature points is analyzed, the high-density feature point areas are determined, and feature point optimization data is generated. The feature point optimization data is matched with the local area of the indoor real scene model, and the data is fused using the weighted average method to obtain the locally updated model fragment. The three-dimensional geometric information is extracted from the locally updated model fragments, and the topological consistency with the first three-dimensional model is determined to obtain a topological verification result. If the topological verification result is consistent, the locally updated model fragments are integrated into the indoor real-scene model to generate a first three-dimensional model. Based on the geometric structure of the first three-dimensional model, the surface texture is optimized using a stereoscopic microscopy algorithm to obtain the final three-dimensional model output.
8. The method for visual positioning based on indoor fine three-dimensional model according to claim 1, characterized in that: Matching the first three-dimensional model with the first feature sequence to generate a first positioning data set includes: The geometric structure and texture density are obtained from the first three-dimensional model, and key point descriptions are extracted in combination with the first feature sequence to generate initial matching parameters. The initial matching parameters are optimized using a bundle adjustment algorithm to correct for texture density deviations to obtain a first optimized model. If the view rotation matrix consistency of the first optimized model is lower than a preset threshold, the view rotation parameters are iteratively adjusted to obtain a second optimized model. A lightweight network is used to compress the feature sequence of the second optimization model to generate a compressed feature set; based on the compressed feature set and the second optimization model, a fast matching calculation is performed to obtain a low-latency matching result; based on the low-latency matching result, a first positioning data set is generated; for the first positioning data set, the texture density and perspective rotation consistency are verified to obtain the final positioning data set.
9. The method for visual positioning based on indoor fine three-dimensional model according to claim 1, characterized in that: Based on the first positioning data set and in combination with the spatial boundary constraints in the dynamic environment, a first navigation instruction sequence is generated by calculating a navigation path change trend taking into account a weight of a proportion of an occluded area, including: Obtain the first positioning data set, combine it with the spatial boundaries in the dynamic environment, and extract environmental constraint features through preprocessing to obtain an initial positioning distribution. If there are occluded areas in the initial positioning distribution, use a region segmentation method to determine the range of the occluded areas and obtain the occluded area distribution. Based on the occluded area distribution, calculate the proportional weight of each area, generate a weight matrix through the weight allocation model, and determine the path optimization parameters. The weight matrix is fused through the A-star algorithm, and the navigation path is calculated based on the spatial boundary constraints to obtain the path change trend. If there is a deviation in the path change trend, the path is recalculated by dynamically adjusting the weight matrix to obtain the optimized navigation path. Based on the optimized navigation path, the first navigation instruction sequence is generated to determine the final instruction output.
10. A device for visual positioning based on a fine indoor three-dimensional model, characterized in that: The method according to any one of claims 1 to 9, comprising: A feature data set generation module is used to acquire the indoor original image through the camera and synchronously collect the camera's attitude angular velocity, construct a first feature data set, and perform correction to generate a second feature data set; A feature sequence generation module is configured to align the second feature data set with a pre-built indoor real-scene three-dimensional model to generate first position estimation data under spatial boundary constraints; based on the first position estimation data, locate feature point regions from the real-time acquired dynamic image data, and track point motion trajectories in the movement trajectory offset to generate a first feature sequence; a posture data generation and update module, configured to match the first position estimation data with the first feature sequence and fuse them to generate a first posture data set; and when the distance deviation between the first posture data set and the three-dimensional model position exceeds a preset threshold, perform a local update on the indoor real scene three-dimensional model to generate a first three-dimensional model; a positioning data optimization module, configured to match the first three-dimensional model with the first feature sequence to generate a first positioning data set; The navigation path generation module is used to generate a first navigation instruction sequence by calculating a navigation path change trend taking into account a weight of a proportion of an occluded area based on the first positioning data set in combination with spatial boundary constraints in a dynamic environment.
Citation Information
Cited By
Electric power inspection robot detection method and system based on machine vision
CN121074026A
A machine vision-based power inspection robot detection method and system
CN121074026B
Robot vision mode matching system in dynamic scene
CN121447653A
Control method and system for binocular visual catheter with posture correction function
CN121465492A
A control method and system for a binocular visual catheter with attitude correction
CN121465492B