Inertial vision fusion positioning method and system based on graph optimization in indoor cross-floor environment

By fusing inertial visual information through a factor graph model, the problems of cumulative drift and visual matching failure in indoor cross-floor positioning were solved, achieving high-precision, high-continuity, and high-robust indoor positioning.

CN121346780APending Publication Date: 2026-01-16HUAIYIN TEACHERS COLLEGE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511565816.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-30
Publication Date
2026-01-16

AI Technical Summary

Technical Problem

Existing indoor positioning technologies struggle to simultaneously guarantee high accuracy, high continuity, and high robustness in complex, multi-story environments. Inertial navigation suffers from cumulative drift, while visual positioning is prone to matching failures.

Method used

An inertial-visual fusion localization method based on graph optimization is adopted. MEMS-IMU and visual data are fused through a factor graph model. The global optimal 3D position and attitude of the pedestrian are obtained by nonlinear optimization. By combining inertial pre-integration and visual absolute pose information, inertial drift is suppressed and localization continuity is maintained.

Benefits of technology

It significantly improves the positioning accuracy and stability across floors, achieving decimeter-level position accuracy and 3.59-degree attitude accuracy, with a 100% positioning success rate, and maintains stable positioning performance under dynamic scenes and changes in lighting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121346780A_ABST
    Figure CN121346780A_ABST
Patent Text Reader

Abstract

The invention relates to an inertial vision fusion positioning method and system based on graph optimization in an indoor cross-floor environment, and the method comprises the steps: obtaining the short-distance relative pose estimation data of a pedestrian based on the multi-source motion parameters of the pedestrian; based on the priori three-dimensional feature map and the image of the current scene of the pedestrian, acquiring visual absolute pose data; and based on a pre-constructed factor graph model, fusing the relative pose estimation data and the visual absolute pose data, and obtaining the global optimal three-dimensional position and pose of the pedestrian through nonlinear optimization solution. According to the method, continuity of deep coupling inertial navigation and absolute precision of visual positioning are optimized through the factor graph, inertial navigation accumulated drift is effectively inhibited, the problem of visual matching failure is solved, decimeter-level positioning precision (the average error is 0.24 m) and 100% continuous positioning are realized in an indoor cross-floor complex scene, and positioning robustness is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of indoor positioning technology, in particular to an inertial-vision fusion positioning method and system based on graph optimization in an indoor cross-floor environment. BACKGROUND

[0002] Indoor positioning technology is essential for location-based services. Global Navigation Satellite System (GNSS) performs well outdoors, but its signal is severely attenuated or even lost indoors due to building obstruction. Existing indoor positioning technologies include Wi-Fi, Bluetooth, Ultra-Wideband (UWB), Inertial Navigation (INS), geomagnetic positioning, and vision positioning, etc. The popularity of smartphones provides convenience for pedestrian positioning using their built-in sensors (such as MEMS-IMU, camera).

[0003] Inertial navigation (such as Pedestrian Dead Reckoning PDR) has small computation and high frequency, and can provide relative motion information, but has cumulative error, which will lead to positioning drift after a long time, especially in long-distance scenarios such as cross-floor, the accuracy will decrease sharply. Vision positioning based on prior three-dimensional map can provide absolute pose (6 degrees of freedom), and has high accuracy, but it is easy to fail in matching in the case of texture missing, light change or dynamic scene, resulting in discontinuous positioning.

[0004] Although existing fusion methods (such as Kalman filter) can improve performance to some extent, they still have deficiencies in suppressing long-term drift of inertial navigation, overcoming matching failure of vision, and achieving global optimization and error suppression of multi-source data in complex indoor cross-floor scenarios. It is difficult to ensure high accuracy, high continuity and strong robustness of positioning at the same time.

[0005] Therefore, there is an urgent need for a method that can effectively fuse inertial and vision information in complex indoor cross-floor environments, and achieve continuous, stable and high-precision three-dimensional positioning. SUMMARY

[0006] The purpose of the present application is to provide an inertial-vision fusion positioning method and system based on graph optimization in an indoor cross-floor environment, to solve the problems of cumulative drift and matching failure of existing single sensor positioning methods, and the problem of insufficient global optimization of existing fusion methods in complex scenarios, and to realize high-precision, high-continuity and high-robustness estimation of pedestrian position in indoor three-dimensional space.

[0007] To achieve the above purpose, the present application provides the following solutions:

[0008] An inertial-vision fusion positioning method based on graph optimization in an indoor cross-floor environment, comprising:

[0009] obtaining relative pose estimation data of a pedestrian in a short distance based on multi-source motion parameters of the pedestrian;

[0010] Based on the prior three-dimensional feature map and the image of the current scene of the pedestrian, visual absolute pose data is obtained;

[0011] Based on the pre-constructed factor graph model, the relative pose estimation data and the visual absolute pose data are fused, and the global optimal three-dimensional position and attitude of the pedestrian are obtained by nonlinear optimization.

[0012] Optionally, based on the multi-source motion parameters of the pedestrian, the relative pose estimation data of the pedestrian in a short distance is obtained, including:

[0013] A mobile device with a MEMS-IMU is used to collect the multi-source motion parameters of the pedestrian; wherein the MEMS-IMU includes an accelerometer, a gyroscope, and a magnetometer;

[0014] Based on the multi-source motion parameters, gait detection is performed to obtain the relative pose estimation data of the pedestrian in a short distance.

[0015] Optionally, based on the multi-source motion parameters, gait detection includes:

[0016] Based on the multi-source motion parameters, a pedestrian dead reckoning algorithm is used to perform relative position calculation of the pedestrian in a short distance, combined with the decoupling strategy of attitude calculation and position calculation, and relative pose information is output.

[0017] Optionally, obtaining visual absolute pose data includes:

[0018] A prior three-dimensional feature map of the positioning area is constructed; wherein the prior three-dimensional feature map includes 3D point cloud features;

[0019] An image acquisition device in the mobile device is used to obtain the image of the current scene of the pedestrian;

[0020] 2D visual features of the current scene image are extracted;

[0021] The 2D visual features are matched with the 3D point cloud features to establish the conversion relationship between pixel coordinates and world coordinates;

[0022] Based on the conversion relationship between pixel coordinates and world coordinates, a random sample consensus algorithm is used to screen reliable feature matching pairs;

[0023] A perspective N-point algorithm is used to solve the inliers of the feature matching pairs to estimate the six-degree-of-freedom absolute pose of the image acquisition device in the current world coordinate system.

[0024] Optionally, the factor graph model comprises: variable nodes representing system states, and factor nodes representing constraints; wherein the variable nodes comprise position, attitude, velocity, IMU bias, and the factor nodes comprise inertial pre-integration constraint, visual absolute pose constraint, state transition constraint, and IMU bias prior constraint.

[0025] Optionally, based on the pre-constructed factor graph model, the fusion of the relative pose estimation data and the visual absolute pose data comprises:

[0026] converting the relative pose estimation data into an inertial pre-integration factor and adding the inertial pre-integration factor into the factor graph model;

[0027] converting the visual absolute pose data into a visual positioning factor and adding the visual positioning factor into the factor graph model;

[0028] and introducing zero velocity update and zero angular rate update constraint factors into the factor graph model to suppress drift caused by IMU cumulative error.

[0029] Optionally, the state transition constraint is a model describing the evolution of the state over time.

[0030] The IMU bias prior constraint is a smooth constraint on the bias of the accelerometer and the gyroscope.

[0031] Optionally, the obtaining of the globally optimal three-dimensional position and attitude of the pedestrian through nonlinear optimization solving comprises:

[0032] solving the entire factor graph model by using a nonlinear optimization library, minimizing the sum of squares of all factor error terms, and obtaining optimal estimation values of all state variables, i.e., the globally optimal three-dimensional position and attitude of the pedestrian.

[0033] An inertial visual fusion positioning system based on graph optimization in an indoor cross-floor environment, the system comprising:

[0034] a relative pose estimation module configured to obtain relative pose estimation data of a pedestrian for a short distance based on multi-source motion parameters of the pedestrian;

[0035] a visual pose solving module configured to obtain visual absolute pose data based on an a priori three-dimensional feature map and an image of a current scene of the pedestrian;

[0036] a factor graph fusion positioning module configured to fuse the relative pose estimation data and the visual absolute pose data based on a pre-constructed factor graph model, and to obtain globally optimal three-dimensional position and attitude of the pedestrian through nonlinear optimization solving.

[0037] The present application has the following advantages:

[0038] The application provides an inertial vision fusion positioning method and system based on graph optimization in an indoor cross-floor environment. BRIEF DESCRIPTION OF DRAWINGS

[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed in the embodiments. Obviously, the drawings described below only show some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without any creative effort based on these drawings.

[0040] Figure 1 A flowchart of an inertial vision fusion positioning method based on graph optimization in an indoor cross-floor environment is provided for the embodiments of the present application.

[0041] Figure 2 A pose solving flowchart based on a MEMS inertial sensor is provided for the embodiments of the present application.

[0042] Figure 3 A factor graph structure diagram of inertial / vision fusion is provided for the embodiments of the present application.

[0043] Figure 4 An experimental test scene and walking path diagram is provided for the embodiments of the present application.

[0044] Figure 5Fig. 1 is a positioning trajectory effect comparison diagram of a test route B-C and C-D according to an embodiment of the present application; wherein (a) is a positioning trajectory effect comparison diagram of the test route B-C, and (b) is a positioning trajectory effect comparison diagram of the test route C-D. DETAILED DESCRIPTION

[0045] The technical solutions in the embodiments of the present application will be apparently and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative labor fall within the protection scope of the present application.

[0046] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, the present application will be further described in detail below with reference to the drawings and specific embodiments.

[0047] As shown in Fig. 1, the embodiment proposes an inertial vision fusion positioning method based on graph optimization in an indoor cross-floor environment, which comprises the following steps. Figure 1

[0048] Obtaining relative pose estimation data of a pedestrian in a short distance based on multi-source motion parameters of the pedestrian;

[0049] Obtaining visual absolute pose data based on a prior three-dimensional feature map and an image of a current scene of the pedestrian;

[0050] Fusing the relative pose estimation data and the visual absolute pose data based on a pre-constructed factor graph model, and obtaining a global optimal three-dimensional position and attitude of the pedestrian through nonlinear optimization solution.

[0051] Further, the step of obtaining relative pose estimation data of a pedestrian in a short distance based on multi-source motion parameters of the pedestrian comprises the following steps.

[0052] Collecting multi-source motion parameters of the pedestrian by using a mobile device with a MEMS-IMU; wherein the MEMS-IMU comprises an accelerometer, a gyroscope and a magnetometer; and the multi-source motion parameters comprise speed data, angular velocity data and heading data;

[0053] The MEMS-IMU is a micro-electromechanical system inertial measurement unit. In the embodiment, it is a combination of sensors built-in in a smart phone, including an accelerometer, a gyroscope and a magnetometer, which are used to obtain multi-source motion parameters of the pedestrian. The acceleration data is obtained from the accelerometer, the angular velocity data is obtained from the gyroscope, and the heading data is calculated by fusing the magnetometer and the accelerometer information (for example, used for heading angle estimation in a pedestrian dead reckoning PDR algorithm).

[0054] ​Based on the multi-source motion parameters, gait detection is performed to obtain relative pose estimation data of the pedestrian in a short distance.

[0055] Specifically, in the embodiment, first, data acquisition and preprocessing are performed:

[0056] First, multi-source motion parameters are obtained by using the built-in accelerometer, gyroscope and magnetometer of the smart phone; for example, the pedestrian holds the smart phone and walks in an indoor cross-floor environment (such as a multi-story building including a lobby, stairs and a corridor). The built-in MEMS-IMU (accelerometer, gyroscope and magnetometer) of the smart phone collects inertial data at a frequency of 50 Hz. The built-in camera of the smart phone captures scene images at a frequency of 5 Hz (which can be uniformly down-sampled to 640x480 resolution for processing). An accurate prior three-dimensional feature map of the experimental area is obtained or pre-constructed (which can be constructed by an SFM method, including a 3D point cloud and its feature descriptor), and the smart phone obtains measurement values by using a barometer.

[0057] Then, a robust pedestrian dead reckoning (PDR) algorithm is used to perform relative position calculation of the pedestrian in a short distance in combination with a decoupling strategy of attitude calculation and position calculation, and relative pose information (including three-dimensional position increment and attitude change) is output.

[0058] The absolute pose result output by visual positioning is used to correct the barometer data to obtain relatively accurate three-dimensional elevation information. In order to obtain accurate elevation information, the barometer measurement values are corrected by using the pose output result of visual positioning, so as to improve the accuracy of three-dimensional position calculation.

[0059] The three-dimensional elevation information is used to correct the elevation value in the relative pose estimation, and is used as part of the state variable in the factor graph fusion positioning. Specifically, in the relative pose estimation, the elevation information obtained by correcting the barometer data and the visual positioning is integrated into the position calculation formula to provide the Z-axis component in the global coordinate system. In the factor graph optimization, the elevation information is part of the position state variable, which is solved by nonlinear optimization, and finally the global optimal three-dimensional position and attitude of the pedestrian are output.

[0060] Specifically, in the embodiment, the relative pose estimation is performed as follows: Figure 2 As shown in the figure, the MEMS-IMU data is processed:

[0061] Gait detection (such as peak detection and zero-crossing detection) is performed.

[0062] Attitude calculation: the attitude is calculated based on the gyroscope integration and the accelerometer / magnetometer respectively, and the gradient descent method (or other fusion filter) is used for fusion to obtain the optimal quaternion attitude.

[0063] Step length estimation: the step length is estimated according to the characteristics (such as variance and frequency spectrum) of the acceleration data.

[0064] heading estimate: fusion of gyro angular velocity integration and magnetometer corrected heading.

[0065] height correction using barometer data and visual positioning feedback.

[0066] Based on the PDR principle, the relative position increment and attitude change in the local coordinate system at the current time are calculated based on the position at the previous time, the current step length and the heading, to form a relative pose estimate.

[0067] Further, acquiring the visual absolute pose data comprises:

[0068] constructing or acquiring a prior three-dimensional feature map (containing 3D point cloud and feature descriptor) of the positioning area;

[0069] acquiring an image of the current scene of the pedestrian using the built-in camera of the smart phone;

[0070] extracting 2D visual features of the current image;

[0071] matching the 2D visual features with the 3D point cloud features in the prior three-dimensional feature map to establish a conversion relationship between pixel coordinates and world coordinates;

[0072] screening reliable feature matching pairs using a random sample consensus (RANSAC) algorithm; for example, using the RANSAC algorithm (such as 500 iterations) to remove false matches to obtain reliable 2D-3D feature matching pairs (inliers);

[0073] solving the matching inliers using a perspective N-point (PnP) algorithm to estimate the six-degree-of-freedom (6DoF) absolute pose of the camera in the current world coordinate system. Specifically, using the camera intrinsic parameters (obtained from EXIF or calibrated) and the matching inliers, the PnP problem is solved to calculate the absolute position and attitude of the current camera in the world coordinate system.

[0074] Further, the factor graph model comprises variable nodes representing system states, and factor nodes representing constraints; wherein the variable nodes comprise position, attitude, velocity, IMU bias, and the factor nodes comprise inertial pre-integration constraints, visual absolute pose constraints, state transition constraints, and IMU bias prior constraints.

[0075] Further, based on the pre-constructed factor graph model, fusing the relative pose estimate data and the visual absolute pose data comprises:

[0076] converting the relative pose estimate data into an inertial pre-integration factor and adding it to the factor graph model;

[0077] convert the visual absolute pose data into visual positioning factors, and add the visual positioning factors into a factor graph model;

[0078] and introduce zero velocity update (ZUPT) and zero angular rate update (ZARU) constraints into the factor graph model to suppress the drift caused by the accumulated errors of the IMU.

[0079] Further, the global optimal three-dimensional position and posture of the pedestrian are obtained by solving the factor graph model through nonlinear optimization, including:

[0080] The global optimal three-dimensional position and posture of the pedestrian are obtained by solving the entire factor graph model through a nonlinear optimization library to minimize the sum of squares of all factor error terms and obtain optimal estimates of all state variables.

[0081] Specifically, in the embodiment, the fusion positioning based on the factor graph optimization includes: Figure 3 as shown in the following figure, a factor graph model is constructed:

[0082] Variable node: represents the system state (position, posture, velocity, IMU bias).

[0083] Factor node:

[0084] Inertial pre-integration factor (black box): converts the relative motion between adjacent state nodes in step 2 into pre-integrated quantities, and constrains the adjacent states and .

[0085] Visual positioning factor (yellow box): the visual absolute pose calculated in step 3 is taken as an observation to constrain the state at the corresponding time.

[0086] State transition factor (blue box): describes the model of the state evolution over time (usually based on IMU dynamics or constant velocity model).

[0087] IMU bias prior factor (green box): applies a smoothing constraint (such as random walk model) to the zero bias of the IMU accelerometer and gyroscope.

[0088] Zero velocity update (ZUPT) and zero angular rate update (ZARU) factors: when the static moment (zero velocity) of the foot landing is detected, the velocity and angular velocity are constrained to be zero, effectively suppressing the accumulated error drift.

[0089] The above factors are connected to the factor graph according to their corresponding state variables and time stamps.

[0090] Using nonlinear optimization libraries (such as g2o, GTSAM), the entire factor graph is solved, minimizing the sum of squares of all factor error terms (i.e., maximizing the posterior probability) to obtain all state variables. The optimal estimate is the fused, globally consistent 3D pedestrian position. and posture .

[0091] Location result output: The optimal location obtained through the above optimization. and posture As the final 3D positioning result of the pedestrian at the current moment, it can be used for navigation or location services.

[0092] This embodiment also proposes an inertial-visual fusion positioning system based on graph optimization in an indoor multi-floor environment, including:

[0093] Relative pose estimation module: used to obtain relative pose estimation data of pedestrians over short distances based on multi-source motion parameters of pedestrians;

[0094] Visual pose calculation module: used to obtain visual absolute pose data based on prior 3D feature map and the current scene image of pedestrian;

[0095] Factor graph fusion localization module: Based on a pre-built factor graph model, it fuses the relative pose estimation data and the visual absolute pose data, and obtains the pedestrian's global optimal 3D position and pose through nonlinear optimization.

[0096] Furthermore, the factor graph fusion localization module employs the Levenberg-Marquardt algorithm for nonlinear optimization;

[0097] The system is deployed on a client-server architecture, with visual positioning running on the server side and inertial data processing running on the mobile device.

[0098] Experimental results: In such Figure 4 Experiments were conducted in a complex indoor multi-floor scene (including a lobby, staircases, and corridors) to verify the results. The walking paths included AB, BC, and CD. The comparison methods included PDR, INS, 3D-LBMS (INS / PDR fusion), and pure visual localization (IBL).

[0099] Positioning accuracy (as shown in Table 1): The average three-dimensional position error of the method (MVL) in this embodiment is 0.24 meters, and the average angle error is 3.59 degrees. It is significantly better than PDR (2.09m, 25.83°), INS (2.85m, 31.79°) and 3D-LBMS (1.94m, 23.99°), and slightly worse than pure vision positioning IBL (0.19m, 3.53°), but the difference is very small.

[0100] Table 1

[0101]

[0102] Continuity and robustness: the positioning failure rate of the method (MVL) of the embodiment is 0%, while the failure rates of the pure visual positioning IBL on the three paths are 6%, 11% and 7% (average 8%) respectively. This shows that the method of the embodiment effectively overcomes the failure problem of visual positioning, and realizes 100% continuous positioning.

[0103] Cross-floor performance: as shown in Figure 5 the trajectories (MVL trajectories in the figure) of the method (MVL) of the application closely follow the real trajectories (Ground Truth) on the path B-C containing stairs and the long corridor path C-D, with the maximum error controlled at a low level (path 3 maximum error <1.1 meters), far superior to other inertial or loosely coupled methods, proving its excellent performance in complex spaces across floors. Figure 5 (a) is a positioning trajectory effect comparison chart of the test route B-C, Figure 5 (b) is a positioning trajectory effect comparison chart of the test route C-D.

[0104] In summary, the embodiment deeply couples the inertial navigation (MEMS-IMU) and visual positioning capabilities of the smartphone through an innovative factor graph optimization framework, effectively solving the two core problems of inertial navigation cumulative drift and visual matching failure, and realizing high-precision (decimeter level), high-continuity (100% success rate) and high-robustness three-dimensional pedestrian positioning in indoor cross-floor complex environments, which has important practical value.

[0105] The above-described embodiments are only descriptions of the preferred modes of the present application and do not limit the scope of the present application. Without departing from the design spirit of the present application, various modifications and improvements to the technical solutions of the present application made by those of ordinary skill in the art shall fall within the protection scope determined by the claims of the present application.

Claims

1. An inertial visual fusion positioning method based on graph optimization in an indoor cross-floor environment, characterized in that, The method comprises the following steps: obtaining relative pose estimation data of a pedestrian in a short distance based on multi-source motion parameters of the pedestrian; obtaining visual absolute pose data based on a prior three-dimensional feature map and an image of a current scene of the pedestrian; fusing the relative pose estimation data and the visual absolute pose data based on a pre-constructed factor graph model, and obtaining global optimal three-dimensional position and posture of the pedestrian through nonlinear optimization.

2. The method of claim 1, wherein, The step of obtaining relative pose estimation data of a pedestrian in a short distance based on multi-source motion parameters of the pedestrian comprises the following steps: collecting multi-source motion parameters of the pedestrian by using a mobile device with a MEMS-IMU; wherein the MEMS-IMU comprises an accelerometer, a gyroscope and a magnetometer; and the multi-source motion parameters comprise velocity data, angular velocity data and heading data; performing gait detection based on the multi-source motion parameters to obtain relative pose estimation data of the pedestrian in a short distance.

3. The method of claim 2, wherein, The step of performing gait detection based on the multi-source motion parameters comprises the following steps: performing relative position resolution of the pedestrian in a short distance by using a pedestrian dead reckoning algorithm in combination with a decoupling strategy of posture resolution and position resolution based on the multi-source motion parameters, and outputting relative pose information.

4. The method of claim 1, wherein, The step of obtaining visual absolute pose data comprises the following steps: constructing a prior three-dimensional feature map of a positioning area; wherein the prior three-dimensional feature map comprises 3D point cloud features; obtaining an image of a current scene of the pedestrian by using an image acquisition device in the mobile device; extracting 2D visual features of the current scene image; matching the 2D visual features with the 3D point cloud features to establish a conversion relationship between pixel coordinates and world coordinates; screening reliable feature matching pairs by using a random sample consensus algorithm based on the conversion relationship between the pixel coordinates and the world coordinates; estimating a six-degree-of-freedom absolute pose of the image acquisition device in a current world coordinate system by using a perspective N-point algorithm to solve inliers in the feature matching pairs.

5. The method of claim 1, wherein, The factor graph model comprises variable nodes representing system states and factor nodes representing constraints; wherein the variable nodes comprise position, posture, velocity and IMU bias, and the factor nodes comprise inertial pre-integration constraints, visual absolute pose constraints, state transition constraints and IMU bias prior constraints.

6. The method of claim 5, wherein, The step of fusing the relative pose estimation data and the visual absolute pose data based on the pre-constructed factor graph model comprises the following steps: converting the relative pose estimation data into inertial pre-integration factors and adding the inertial pre-integration factors to the factor graph model; converting the visual absolute pose data into visual positioning factors and adding the visual positioning factors to the factor graph model; and introducing zero velocity update and zero angular rate update constraint factors to the factor graph model to suppress drift caused by IMU cumulative error.

7. The method according to claim 5, wherein the state transition constraint is a model describing the evolution of the state over time; and the IMU bias prior constraint is a smoothing constraint on the bias of the accelerometer and the gyroscope. The step of obtaining global optimal three-dimensional position and posture of the pedestrian through nonlinear optimization comprises the following steps: ​ 8. The method of claim 1, wherein, ​ The nonlinear optimization library is used to solve the whole factor graph model, and the sum of squares of all factor error terms is minimized to obtain the optimal estimated value of all state variables, i.e. the global optimal three-dimensional position and attitude of the pedestrian.

9. An inertial visual fusion positioning system based on graph optimization in an indoor cross-floor environment, characterized in that, The system for implementing the indoor floor-crossing environment-based graph optimization inertial vision fusion positioning method according to any one of claims 1-8 comprises: a relative pose estimation module configured to obtain relative pose estimation data of a pedestrian at a short distance based on multi-source motion parameters of the pedestrian; a vision pose solution module configured to obtain vision absolute pose data based on a prior three-dimensional feature map and an image of a current scene of the pedestrian; a factor graph fusion positioning module configured to fuse the relative pose estimation data and the vision absolute pose data based on a pre-constructed factor graph model, and obtain the global optimal three-dimensional position and attitude of the pedestrian through nonlinear optimization solution.