Visual inertial fused localization method and device
By constructing a residual model of a second-order partial derivative matrix and performing elimination, the problem of high computational cost in the MSCKF algorithm is solved, and efficient updates of visual-inertial fusion positioning are achieved.
Patent Information
- Application Number
- PCT/CN2024/089422
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-05-09
- Filing Date
- 2024-04-23
- Publication Date
- 2026-01-15
AI Technical Summary
The traditional MSCKF algorithm involves a large amount of computation during state updates, which affects positioning efficiency.
By stacking and merging individual models with multiple feature points in multiple camera states, a residual model including a second-order partial derivative matrix is constructed, and elimination is performed to avoid left null space calculation and QR decomposition dimensionality reduction calculation, and EKF update is performed directly.
It improves positioning efficiency, reduces computational load, and increases computational speed while ensuring positioning accuracy.
Smart Images

Figure CN2024089422_15012026_PF_FP_ABST
Abstract
Description
Visual-inertial fusion positioning method and equipment
[0001] This application claims priority to Chinese Patent Application No. 2023105176004, filed on May 9, 2023, entitled "Visual-Inertial Fusion Positioning Method and Device", the entire contents of which are incorporated herein by reference. Technical Field
[0002] This disclosure relates to the field of navigation and positioning technology, and in particular to a visual-inertial fusion positioning method and device. Background Technology
[0003] Visual Inertial Odometry (VIO) is an algorithm that integrates data from cameras and Inertial Measurement Units (IMUs) to achieve Simultaneous Localization and Mapping (SLAM). It can be broadly categorized into two types: filtering-based and optimization-based. The Multi-State Constraint Kalman Filter (MSCKF) is a filtering-based VIO algorithm.
[0004] However, the inventors discovered that the traditional MSCKF algorithm has at least the following technical problems: the computational load during the state update process is large, which affects the positioning efficiency.
[0005] Summary of the Invention
[0006] This disclosure provides a visual-inertial fusion positioning method and device to overcome the problem that the MSCKF algorithm has a large computational load during state update, which affects positioning efficiency.
[0007] In a first aspect, embodiments of this disclosure provide a visual-inertial fusion localization method, including:
[0008] Obtain target observation data corresponding to the target state vector; the target observation data includes observation data of at least one feature point; the observation data of each feature point includes observation data of the feature point in at least one camera state;
[0009] A first residual model is constructed based on the target observation data; the first residual model includes a product term of a first-order partial derivative matrix and a system state matrix, and a first observation noise term; the first-order partial derivative matrix includes a first first-order partial derivative matrix obtained by differentiating the target state vector and a second first-order partial derivative matrix obtained by differentiating the three-dimensional coordinate vector of the at least one feature point; the system state matrix includes a target state error vector and a three-dimensional coordinate error vector of the feature point; the target state vector includes an inertial measurement unit (IMU) state vector and a camera state vector;
[0010] The residual model is projected based on the first-order partial derivative matrix to obtain a second residual model; the second residual model includes a product term of the second-order partial derivative matrix and the system state matrix, and a covariance term corresponding to the first observation noise term;
[0011] Eliminate the three-dimensional coordinate error vector of the second residual model to obtain the target residual model, and update the target state vector according to the target residual model using EKF.
[0012] In a second aspect, embodiments of this disclosure provide a visual-inertial fusion positioning device, comprising:
[0013] An acquisition module is used to acquire target observation data corresponding to a target state vector; the target observation data includes observation data of at least one feature point; the observation data of each feature point includes observation data of the feature point in at least one camera state; the target state vector includes multiple camera states;
[0014] A construction module is used to construct a first residual model based on the target observation data. The first residual model includes a product term of a first-order partial derivative matrix and a system state matrix, as well as a first observation noise term. The first-order partial derivative matrix includes a first-order partial derivative matrix obtained by differentiating the target state vector and a second-order partial derivative matrix obtained by differentiating the three-dimensional coordinate vector of the at least one feature point. The system state matrix includes a target state error vector and a three-dimensional coordinate error vector of the feature point. The target state vector includes an inertial measurement unit (IMU) state vector and a camera state vector.
[0015] The projection module is used to project the residual model based on the first-order partial derivative matrix to obtain a second residual model; the second residual model includes a product term of the second-order partial derivative matrix and the system state matrix, and a covariance term corresponding to the first observation noise term;
[0016] The elimination module is used to eliminate the three-dimensional coordinate error vector of the second residual model to obtain the target residual model, and to update the target state vector according to the target residual model.
[0017] Thirdly, embodiments of this disclosure provide an electronic device, including: a processor and a memory;
[0018] The memory stores computer-executed instructions;
[0019] The processor executes computer execution instructions stored in the memory, causing the at least one processor to perform the visual-inertial fusion localization method as described in the first aspect and various possible designs of the first aspect.
[0020] Fourthly, embodiments of this disclosure provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the visual-inertial fusion positioning method described in the first aspect and various possible designs of the first aspect.
[0021] Fifthly, embodiments of this disclosure provide a computer program product, including a computer program that, when executed by a processor, implements the visual-inertial fusion positioning method as described in the first aspect and various possible designs of the first aspect. Attached Figure Description
[0022] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 is a schematic diagram illustrating the principle of a visual-inertial fusion positioning method provided in an embodiment of this disclosure;
[0024] Figure 2 is a schematic flowchart of the visual-inertial fusion positioning method provided in an embodiment of this disclosure;
[0025] Figure 3 is a schematic diagram of the observation relationship between multiple feature points and multiple camera states in the current sliding window corresponding to the target state vector provided in the embodiments of this disclosure;
[0026] Figure 4 is a schematic flowchart of the visual-inertial fusion positioning method provided in this embodiment of the present disclosure.
[0027] Figure 5 is a schematic diagram of the observation relationship between multiple feature points corresponding to the target state vector provided in the embodiments of this disclosure and multiple camera states in the current sliding window and multiple historical camera states in the shared window.
[0028] Figure 6 is a structural block diagram of the visual-inertial fusion positioning device provided in an embodiment of this disclosure;
[0029] Figure 7 is a schematic diagram of the hardware structure of the visual-inertial fusion positioning device provided in an embodiment of this disclosure. Detailed Implementation
[0030] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0031] The visual-inertial fusion positioning method and device provided in this embodiment first acquires target observation data corresponding to the target state vector. The target observation data includes observation data of at least one feature point. The observation data of each feature point includes observation data of the feature point in at least one camera state. A first residual model is constructed based on the target observation data. The first residual model includes a product term of a first-order partial derivative matrix and a system state matrix, as well as a first observation noise term. The first-order partial derivative matrix includes a first first-order partial derivative matrix obtained by differentiating the target state vector and a second first-order partial derivative matrix obtained by differentiating the three-dimensional coordinate vector of at least one feature point. The system state matrix includes a target state error vector and a three-dimensional coordinate error vector of the feature point. The target state vector includes an inertial measurement unit (IMU) state vector and a camera state vector. Based on the first-order partial derivative matrix, the residual model is projected to obtain a second residual model. The second residual model includes a product term of a second-order partial derivative matrix and a system state matrix, as well as a covariance term corresponding to the first observation noise term. The three-dimensional coordinate error vector of the second residual model is eliminated to obtain a target residual model. The target state vector is then updated based on the target residual model. This embodiment obtains a first residual model by stacking and merging individual models of multiple feature points under multiple camera states, then projects it to obtain a second residual model including a square matrix of second-order partial derivatives. Subsequently, the second residual model is eliminated to obtain the target residual model for state update. This calculation process does not involve left null space calculation and projection, as well as QR decomposition dimensionality reduction calculation, thus saving computation and improving positioning efficiency.
[0032] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0033] Visual Inertial Odometry (VIO) is an algorithm that integrates data from cameras and Inertial Measurement Units (IMUs) to achieve Simultaneous Localization and Mapping (SLAM). It can be broadly categorized into filtering-based and optimization-based algorithms. The Multi-State Constraint Kalman Filter (MSCKF) is a filtering-based VIO algorithm. However, MSCKF involves significant computational overhead during state updates, impacting localization efficiency. Specifically, state updates require left null space calculations and projections for each feature point observation. Since the number of rows in the resulting matrix after left null space projection far exceeds the number of columns, QR decomposition is also necessary for dimensionality reduction. This series of complex calculations incurs substantial CPU overhead, affecting both computational and localization efficiency.
[0034] To address the aforementioned technical problems, the inventors of this application have discovered that a first residual model obtained by fully stacking and matrix merging individual models of multiple feature points under multiple camera states can be projected to form a second residual model, including a second-order partial derivative matrix. This second residual model is then eliminated to obtain a target residual model for state updating. This calculation process does not involve left null space calculation and projection, nor QR decomposition dimensionality reduction calculation, thus saving computational resources and improving positioning efficiency. Based on this, embodiments of this disclosure provide a visual-inertial fusion positioning method.
[0035] Referring to Figure 1, which is a schematic diagram of the principle of a visual-inertial fusion localization method provided in an embodiment of this disclosure. As shown in Figure 1, the visual-inertial fusion system includes a localization device 101, an IMU 102, and a camera 103. The localization device 101 employs an Extended Kalman Filter (EKF) framework. It predicts the EKF based on IMU data (IMU readings) perceived by the IMU 102 and updates the EKF based on image frames captured by the camera 103, i.e., the observation data of feature points. Specifically, based on the camera 103, the camera states at different times (e.g., camera 103 position P and attitude quaternion Q) are added to the state vector. Feature points can be seen by multiple cameras 103, thereby forming geometric constraints between multiple camera states (Multi-State). These geometric constraints can then be used to construct a residual model, thereby updating the EKF. The positioning device 101 can be applied to fields such as robotics, drones, augmented reality (AR) and virtual reality (VR).
[0036] In the specific implementation process, the state vector and covariance are first initialized. Then, during the prediction phase, based on each IMU reading, the system's basic state and system covariance are propagated from time k to time k+1 of the image, typically processing multiple IMU readings. During the update phase, after camera 103 acquires an image frame at time k+1, it calculates the current camera state based on the image frame and adds the current camera state to the state vector to obtain the target state vector, while simultaneously expanding the state covariance. A target residual model is constructed based on the observation data corresponding to the target state vector, and the EKF is then updated based on the target residual model.
[0037] Specifically, in constructing the target residual model, target observation data corresponding to the target state vector is obtained; the target observation data includes observation data of at least one feature point; the observation data of each feature point includes observation data of the feature point under at least one camera state; a first residual model is constructed based on the target observation data; the first residual model includes a product term of a first-order partial derivative matrix and a system state matrix, and a first observation noise term; the first-order partial derivative matrix includes a first first-order partial derivative matrix obtained by differentiating the target state vector and a second first-order partial derivative matrix obtained by differentiating the three-dimensional coordinate vector of the at least one feature point. The system state matrix includes a target state error vector and a three-dimensional coordinate error vector of feature points. The target state vector includes an inertial measurement unit (IMU) state vector and a camera state vector. Based on the first-order partial derivative matrix, the residual model is projected to obtain a second residual model. The second residual model includes a product term of the second-order partial derivative matrix and the system state matrix, and a covariance term corresponding to the first observation noise term. The three-dimensional coordinate error vector of the second residual model is eliminated to obtain a target residual model, and the target state vector is updated using EKF based on the target residual model.
[0038] The visual-inertial fusion localization method provided in this embodiment stacks and merges the individual models of multiple feature points under multiple camera states to obtain a first residual model, which is then projected to obtain a second residual model including a square matrix of second-order partial derivatives. Subsequently, the second residual model is eliminated to obtain a target residual model for state update. This calculation process does not involve left null space calculation and projection, as well as QR decomposition dimensionality reduction calculation, thus saving computational load and improving localization efficiency.
[0039] Referring to Figure 2, which is a schematic flowchart of the visual-inertial fusion positioning method provided in this embodiment, the method of this embodiment can be applied to the positioning device shown in Figure 1. The visual-inertial fusion positioning method includes:
[0040] 201. Obtain the target observation data corresponding to the target state vector; the target observation data includes the observation data of at least one feature point; the observation data of each feature point includes the observation data of the feature point in at least one camera state.
[0041] In this embodiment of the disclosure, the target state vector is the system state vector at the update timing determined by the update strategy. The system state vector XI includes a basic state XA (IMU state) and an extended state XB (camera state). The basic state XA may include the IMU's attitude, displacement, velocity, accelerometer deviation, gyroscope deviation, and may also include rotational and / or translational extrinsic parameters from the camera to the IMU. The extended state XB may include the camera's rotation, displacement, and may also include velocity, etc.
[0042] In one embodiment of this disclosure, to ensure that the system state vector does not increase indefinitely and affect computational efficiency, an upper limit can be set for the number of expansions in the sliding window (i.e., the sliding window) of the system state vector. For example, this upper limit can be set to n, then XI = [XA, XB0, XB1, ..., XBs]. When s = n, before further expansion, existing expansion states in the sliding window need to be reduced to ensure that the number of expansion states in the sliding window is less than or equal to the upper limit n. The specific reduction strategy (number of reductions, reduction positions, etc.) can be determined according to the actual situation, and this embodiment of the disclosure does not limit this.
[0043] In this embodiment of the disclosure, the observation data of a feature point in camera mode is the relevant data of an image frame containing the feature point captured by the camera in that camera mode. The content of the target observation data is limited according to the adopted update strategy.
[0044] For example, in one possible implementation, all observation data within the sliding window corresponding to feature points that cannot be observed further in the current sliding window can be identified as target observation data. Then, a residual model is constructed based on the target observation data, and an extended Kalman filter (EKF) is used for state updates.
[0045] In another possible implementation, the upper limit of the number of expanded states in the sliding window can be set to m. When s == m, all observations from the first two or more camera states, the last two or more camera states, or one or more camera states before and after the last two are extracted from the sliding window in chronological order. The extracted observation data is determined as the target observation data, and a residual model is then constructed based on this target observation data, using EKF for state updates. As for which windows to extract for state updates, a comprehensive decision can be made based on whether the current system is in motion or whether the amount of motion (e.g., the distance traveled) has reached a specific limit.
[0046] In another possible implementation, to improve computational accuracy, in addition to the target observation data determined by the methods described above, historical camera state observation data outside the current sliding window (i.e., earlier than the earliest time point in the sliding window) that is associated with at least one feature point in the sliding window can also be determined as target observation data. For details, please refer to the descriptions in subsequent embodiments; they will not be repeated here.
[0047] Referring to Figure 3, Figure 3 is a schematic diagram of the observation relationship between multiple feature points and multiple camera states in the current sliding window corresponding to the target state vector provided in this embodiment of the present disclosure. As shown in Figure 3, the pentagram represents the feature points, the box represents the extended camera state sequence, and the dashed line represents the observation relationship between the feature points and the camera states. The current sliding window corresponding to the target state vector includes feature points f0 to f4, and the camera states include Ca0 to Ca4. Feature point f0 is observed by the camera in five camera states (Ca0 to Ca4), feature point f1 is observed by the camera in three camera states (Ca2 to Ca4), feature point f2 is observed by the camera in four camera states (Ca1 to Ca4), feature point f3 is observed by the camera in three camera states (Ca2 to Ca4), and feature point f4 is observed by the camera in three camera states (Ca2 to Ca4).
[0048] In some disclosed embodiments, before constructing the first residual model based on the target observation data, the method may further include: acquiring multiple IMU data after the previous time step and before the current time step; predicting the system state vector of the previous time step based on the multiple IMU data steps to obtain the system state vector to be augmented; acquiring an image frame at the current time step; calculating the camera state vector corresponding to the image frame based on the image frame; and augmenting the system state vector to be augmented based on the camera state vector to obtain the target state vector.
[0049] Specifically, during the prediction phase, multiple IMU data points from time k to time k+1 can be integrated to propagate the system's base state and covariance from time k to time k+1. After obtaining the image frame at time k+1, camera state calculation and augmentation are performed to obtain the target state vector.
[0050] 202. Construct a first residual model based on the target observation data; the first residual model includes a product term of a first-order partial derivative matrix and a system state matrix, and a first observation noise term; the first-order partial derivative matrix includes a first first-order partial derivative matrix obtained by differentiating the target state vector and a second first-order partial derivative matrix obtained by differentiating the three-dimensional coordinate vector of the at least one feature point; the system state matrix includes a target state error vector and a three-dimensional coordinate error vector of the feature point; the target state vector includes an inertial measurement unit (IMU) state vector and a camera state vector.
[0051] Specifically, feature points can be triangulated (i.e., fitted with the three-dimensional coordinates of feature points) based on the current camera state in the sliding window. Then, based on the preset state update strategy (refer to the example description of step 201 above), the observation data of multiple feature points in multiple camera states can be determined. Thus, based on the three-dimensional coordinates of multiple feature points and the corresponding observation data of multiple camera states, the first residual model can be constructed.
[0052] In the triangulation estimation process, multiple camera states from different viewpoints (different positions) and the corresponding observation data of feature points can be used. A nonlinear optimization algorithm is then used to solve for the 3D coordinates of the feature points based on the camera states. After obtaining the 3D coordinates of the feature points in the camera states through triangulation estimation, the 3D coordinates of the feature points in the world coordinate system can be calculated based on these 3D coordinates and the camera's pose in the world coordinate system. Where W represents the World Coordinate System, P represents the Position, and i represents the i-th feature point f.
[0053] In some disclosed embodiments, constructing a first residual model based on the target observation data may include: for each of the at least one feature point and each of the at least one camera states corresponding to the feature point, constructing a third residual model based on the observation data of the feature point in the target observation data under the camera state; the third residual model includes a product term of a first partial derivative matrix and a target state error vector, a product term of a second partial derivative matrix and a three-dimensional coordinate error vector of the feature point, and a second observation noise term; the first partial derivative matrix is the first-order partial derivative of the residual function with respect to the target state vector, and the second partial derivative matrix is the first-order partial derivative of the residual function with respect to the three-dimensional coordinate vector of the feature point; stacking the third residual models corresponding to multiple feature points respectively to obtain a fourth residual model; and performing matrix merging on the product terms in the fourth residual model to obtain the first residual model.
[0054] In some disclosed embodiments, the step of constructing a third residual model for each of the at least one feature point and each of the at least one camera states corresponding to the feature point, based on the observation data of the feature point in the target observation data under the camera states, may include: constructing a residual function for each of the at least one feature point and each of the at least one camera states corresponding to the feature point, based on the observation data of the feature point in the target observation data under the camera states; the residual function is related to the target state vector and the three-dimensional coordinate vector of the feature point; and linearizing the residual function to obtain the third residual model.
[0055] In some disclosed embodiments, stacking the third residual models corresponding to the plurality of feature points to obtain a fourth residual model may include: stacking the third residual models corresponding to the feature points to the at least one camera state to obtain a fifth residual model; and stacking the fourth residual models corresponding to the at least one feature point to obtain a fourth residual model.
[0056] In this embodiment of the disclosure, the first-order partial derivative matrix can be a Jacobian matrix.
[0057] For example, in vision, the constraint refers to the reprojection error of feature points to the camera, that is, the error between the two-dimensional coordinates of the feature points actually observed by the camera and the estimated three-dimensional coordinates of the feature points projected onto the two-dimensional coordinates of the image.
[0058] First, a residual function r = f(XI, Pf) is constructed based on the reprojection error, where r is the residual, XI is the target state vector, and Pf is the three-dimensional coordinates of the feature point.
[0059] Secondly, based on the observation of the i-th feature point in the j-th camera state, a third residual model is constructed:
[0060] Where, r i,j J represents the observation residual of the i-th feature point in the j-th camera state; x(i,j) This represents the Jacobian matrix obtained by taking the residual with respect to the target state vector, which is also the first partial derivative matrix. J represents the small amount of system state error, which is also the target state error vector. f(i,j) This represents the Jacobian matrix of the residual with respect to the 3D coordinates of the i-th feature point, which is also the second partial derivative matrix. This represents the small amount of 3D coordinate error of the i-th feature point in the world coordinate system, i.e., the 3D coordinate error vector, n. i,j This represents the observation noise of the system for the i-th feature point in the j-th camera state, also known as the second observation noise.
[0061] Furthermore, based on all observation data of the camera state sequence specified in the sliding window for the i-th feature point, a residual model is constructed. This can be achieved by stacking Equation 1 to form the fourth residual model with the following matrix pattern:
[0062] Where, r i J represents the permutation or stacking of the observation residuals of the i-th feature point across all specified camera states; x(i) This represents the permutation or stacking of the Jacobian matrix of the residuals with respect to the system state, indicating a small system state error. This indicates the permutation or stacking of the Jacobian matrix calculated for the 3D coordinates of the residual corresponding to the i-th feature point. This represents the small amount of 3D coordinate error of the i-th feature point in the world coordinate system, n. i This represents the stacked combination of the observation noise of the i-th feature point under all specified camera states.
[0063] Furthermore, based on all observations of the relevant camera states within the sliding window for feature points that can be used for state updates, a fifth residual model is constructed:
[0064] Among them, J x J represents the Jacobian matrix of the observation residuals over all specified feature points within a sliding window, expressed as a function of the system state variables. f This represents the Jacobian matrix calculated using the observation residuals within all sliding windows of all specified feature points against the 3D coordinates of all feature points used to construct the residuals, where n represents the first observation noise.
[0065] Next, the fifth residual model can be matrix-merged to obtain the first residual model:
[0066] It can be seen that the system Jacobian matrix in equation (4) above is J = [J x J f The system status is... This represents the error quantity between the system's basic state and its expanded state. This represents the error in the three-dimensional coordinates of all specified feature points; however, the EKF system state variables themselves do not contain the three-dimensional coordinates of any feature points. Therefore, elimination is required subsequently.
[0067] To better illustrate the meaning of formula (4), an example is provided below with reference to Figure 3. Assume... s=4, and the camera pose sequence within the current sliding window can observe feature points 0, 1, 2, 3, and 4. There are 5 expanded camera states within the current sliding window, and the observation relationship between feature points and camera states is shown in Figure 3. Therefore, formula (4) can be expressed in the following form.
[0068] The corresponding covariance R is:
[0069] 203. Project the residual model onto the first-order partial derivative matrix to obtain a second residual model; the second residual model includes a product term of the second-order partial derivative matrix and the system state matrix, and a covariance term corresponding to the first observation noise term.
[0070] Specifically, to obtain the second-order partial derivative matrix, the projection matrix can be determined based on the dimension of the first-order partial derivative matrix. For example, the transpose of the first-order partial derivative matrix can be used as the projection matrix.
[0071] For example, in order to perform fast elimination, the first residual model shown in formula (4) can first be projected onto [J]. x J f ] T The following second residual model is obtained.
[0072] 204. Eliminate the three-dimensional coordinate error vector of the second residual model to obtain the target residual model, and update the target state vector according to the target residual model using EKF.
[0073] Specifically, after obtaining a second residual model including a square matrix by projecting the first residual model, Gaussian elimination (specifically, Schul elimination) can be used to eliminate the three-dimensional coordinate error vector in the system state matrix of the second residual model.
[0074] For example, in formula (6), the n vector usually has all members set to be the same. Let each element of n be u. Formula (6) can be simplified as follows:
[0075] Right now:
[0076] In formula (8) These are the variables that need to be eliminated, in order to eliminate them. Equation (8) can be projected into the L space. The following residual model was obtained:
[0077] After simplification, we get:
[0078] Since the EKF system state only includes the camera pose and not the three-dimensional coordinates of the feature points, we can obtain:
[0079] As can be seen from formula (11), the target residual model after elimination can be reorganized into the following standard paradigm.
[0080] Therefore, EKF can be updated based on formula (12) and in combination with the following formula.
[0081] The gain calculation formula is: K = PJ T (JPJ T +R) -1 (13)
[0082] Where K represents the Kalman gain, P is the system covariance, and J is the Jacobian matrix.
[0083] The formula for calculating the state correction is: △X=Kr (14)
[0084] The state update formula is: X = X + ΔX (15)
[0085] The covariance update formula is: P = (I - KJ)P(I - KJ) T +KRK T (16)
[0086] In some disclosed embodiments, after the EKF update is completed based on the above formula, the observation data corresponding to the feature points already used for updating will be set, so that the observation data will no longer be used for subsequent EKF updates. After setting, the update result can be output, such as the update result of the target state vector.
[0087] The disclosed embodiments of formulas (1) to (16) above project the residual model (including the noise model) with non-existent state variables to the Jacobian transpose space, and after projecting to the Jacobian transpose space, project it to a specific space (e.g., the space related to the Shure elimination method) to obtain the part of the system state that can be mathematically eliminated directly (i.e., the derivation result contains a 0 matrix multiplied by the part of the system state that does not need to be eliminated), thereby obtaining the projected residual model that does not contain the part of the system state that does not need to be eliminated. By projecting the noise in the residual model through the above Shure elimination matrix, the observation noise model n corresponding to the target state variable is obtained, and then the EKF is updated. This achieves improved positioning efficiency while ensuring positioning accuracy.
[0088] As described above, the first residual model, obtained by stacking and merging individual models of multiple feature points under multiple camera states, is projected to obtain a second residual model including a square matrix of second-order partial derivatives. Then, the second residual model is eliminated to obtain the target residual model for state update. This calculation process does not involve left null space calculation and projection, as well as QR decomposition dimensionality reduction calculation, which saves computation and improves localization efficiency.
[0089] Referring to Figure 4, which is a schematic flowchart of the visual-inertial fusion localization method provided in this embodiment, to improve calculation accuracy, this embodiment adds the application of observation data within a shared viewing window. The visual-inertial fusion localization method includes:
[0090] 401. Obtain first observation data of at least one feature point in the current sliding window corresponding to the target state vector. Obtain second observation data of at least one feature point in at least one historical camera state outside the current sliding window, and determine the first observation data and the second observation data as the target observation data.
[0091] 402. For each of the at least one feature points and each of the at least one historical camera states corresponding to the feature points, a sixth residual model is constructed based on the observation data of the feature points in the target observation data under the historical camera states; the sixth residual model includes a product term of a third partial derivative matrix and the three-dimensional coordinate error vector of the feature points, and a third observation noise term; the third partial derivative matrix is the first-order partial derivative of the residual function constructed with the historical camera states as constants with respect to the three-dimensional coordinate vector of the feature points.
[0092] 403. Stack the third residual model and the sixth residual model corresponding to the multiple feature points respectively to obtain the fourth residual model.
[0093] 404. Perform matrix merging on the product terms in the fourth residual model to obtain the first residual model.
[0094] Specifically, the feature points in the current sliding window may have historical camera states outside the current sliding window that have an observation relationship with them. If these historical camera states are added to the constraints on the three-dimensional coordinates of the feature points, the amount of observation data can be increased, which helps to improve the accuracy of the first-order partial derivatives of the three-dimensional coordinates, thereby improving the accuracy of the residual model and the accuracy of EKF updates.
[0095] For example, referring to Figure 5, Figure 5 is a schematic diagram of the observation relationship between multiple feature points corresponding to the target state vector provided in this embodiment of the present disclosure and multiple camera states in the current sliding window and multiple historical camera states in the shared window. As shown in Figure 5, based on the observation relationship shown in Figure 3, an observation relationship between feature point f0 in the shared window and historical camera states Ca-2 to Ca-1 in the shared window is added. The shared window refers to a window outside the current sliding window that contains historical camera states that have an observation relationship with at least one feature point in the sliding window. Therefore, the constraint between feature point f0 and historical camera states Ca-2 to Ca-1 in the shared window can be added to the residual model to obtain the following residual model:
[0096] By comparing formula (17) with formula (5), it can be found that the observation data based on the common window increases the constraint of the three-dimensional coordinate error vector of the feature points.
[0097] This embodiment incorporates the residuals formed by the observations of landmarks within the current sliding window and those observed by the old window outside the sliding window into the current residual construction, and sets the pose (position and attitude) of the old window as a constant, so that its residuals constrain the current EKF solution. This makes the EKF solution more accurate and the positioning accuracy higher.
[0098] 405. Project the residual model onto the first-order partial derivative matrix to obtain a second residual model; the second residual model includes a product term of the second-order partial derivative matrix and the system state matrix, and a covariance term corresponding to the first observation noise term;
[0099] 406. Eliminate the three-dimensional coordinate error vector of the second residual model to obtain the target residual model, and update the target state vector according to the target residual model using EKF.
[0100] Steps 405 to 406 in this embodiment are similar to steps 203 to 204 in the above embodiment, and will not be repeated here.
[0101] As can be seen from the above description, by adding the observation data of historical camera states that have an observation relationship with at least one feature point in the current sliding window (outside the current sliding window, i.e., within the common window) to the construction of the residual model, the amount of observation data can be increased, which helps to improve the accuracy of the first-order partial derivative of the three-dimensional coordinates, thereby improving the accuracy of the residual model and the accuracy of EKF updates.
[0102] Corresponding to the visual-inertial fusion positioning method in the above embodiments, Figure 6 is a structural block diagram of the visual-inertial fusion positioning device provided in the embodiments of this disclosure. For ease of explanation, only the parts related to the embodiments of this disclosure are shown. Referring to Figure 6, the device 60 includes: an acquisition module 601, a construction module 602, a projection module 603, and a elimination module 604.
[0103] The acquisition module 601 is used to acquire target observation data corresponding to the target state vector; the target observation data includes observation data of at least one feature point; the observation data of each feature point includes observation data of the feature point in at least one camera state; the target state vector includes multiple camera states.
[0104] The construction module 602 is used to construct a first residual model based on the target observation data. The first residual model includes a product term of a first-order partial derivative matrix and a system state matrix, as well as a first observation noise term. The first-order partial derivative matrix includes a first-order partial derivative matrix obtained by differentiating the target state vector and a second-order partial derivative matrix obtained by differentiating the three-dimensional coordinate vector of the at least one feature point. The system state matrix includes a target state error vector and a three-dimensional coordinate error vector of the feature point. The target state vector includes an inertial measurement unit (IMU) state vector and a camera state vector.
[0105] The projection module 603 is used to project the residual model based on the first-order partial derivative matrix to obtain a second residual model; the second residual model includes a product term of the second-order partial derivative matrix and the system state matrix, and a covariance term corresponding to the first observation noise term;
[0106] The elimination module 604 is used to eliminate the three-dimensional coordinate error vector of the second residual model to obtain the target residual model, and to update the target state vector according to the target residual model.
[0107] In one embodiment of this disclosure, the construction module 602 is specifically configured to: construct a third residual model for each of the at least one feature points and each of the at least one camera states corresponding to the feature points, based on the observation data of the feature points in the target observation data under the camera states; the third residual model includes a product term of a first partial derivative matrix and a target state error vector, a product term of a second partial derivative matrix and a three-dimensional coordinate error vector of the feature points, and a second observation noise term; the first partial derivative matrix is the first-order partial derivative of the residual function with respect to the target state vector, and the second partial derivative matrix is the first-order partial derivative of the residual function with respect to the three-dimensional coordinate vector of the feature points; stack the third residual models corresponding to the multiple feature points respectively to obtain a fourth residual model; and perform matrix merging on the product terms in the fourth residual model to obtain a first residual model.
[0108] The construction module 602 is specifically used to: for each of the at least one feature points and each of the at least one camera states corresponding to the feature points, construct a residual function based on the observation data of the feature points in the target observation data under the camera states; the residual function is related to the target state vector and the three-dimensional coordinate vector of the feature points; and linearize the residual function to obtain a third residual model.
[0109] The construction module 602 is specifically used to: stack the third residual models corresponding to the feature points and the at least one camera state to obtain a fifth residual model; and stack the fourth residual models corresponding to the at least one feature point to obtain a fourth residual model.
[0110] The acquisition module 601 is specifically used to: acquire first observation data of at least one feature point in the current sliding window corresponding to the target state vector; acquire second observation data of at least one feature point in at least one historical camera state outside the current sliding window; and determine the first observation data and the second observation data as target observation data.
[0111] The construction module 602 is further specifically configured to: for each of the at least one feature points and each of the at least one historical camera states corresponding to the feature points, construct a sixth residual model based on the observation data of the feature points in the target observation data under the historical camera states; the sixth residual model includes a product term of a third partial derivative matrix and the three-dimensional coordinate error vector of the feature points, and a third observation noise term; the third partial derivative matrix is the first-order partial derivative of the residual function constructed with the historical camera states as constants with respect to the three-dimensional coordinate vector of the feature points; and stack the third residual models and the sixth residual models corresponding to multiple feature points respectively to obtain a fourth residual model.
[0112] The device 60 also includes a prediction and augmentation module for: acquiring multiple IMU data after the previous time and before the current time; predicting the system state vector of the previous time based on the multiple IMU data to obtain the system state vector to be augmented; acquiring an image frame at the current time, calculating the camera state vector corresponding to the image frame based on the image frame, and augmenting the system state vector to be augmented based on the camera state vector to obtain the target state vector.
[0113] The elimination module 604 is specifically used to: eliminate the three-dimensional coordinate error vector of the second residual model based on the Schur elimination method to obtain the target residual model.
[0114] The device provided in this embodiment can be used to execute the technical solutions of the above method embodiments. Its implementation principle and technical effect are similar, and will not be described again here.
[0115] To implement the above embodiments, this disclosure also provides an electronic device.
[0116] Referring to Figure 7, a schematic diagram of the structure of an electronic device 700 suitable for implementing embodiments of the present disclosure is shown. The electronic device 700 can be a terminal device or a server. The terminal device can include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, personal digital assistants (PDAs), portable Android devices (PADs), portable media players (PMPs), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. The electronic device shown in Figure 7 is merely an example and should not be construed as limiting the functionality and scope of the embodiments of the present disclosure.
[0117] As shown in Figure 7, the electronic device 700 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage device 708 into a random access memory (RAM) 703. The RAM 703 also stores various programs and data required for the operation of the electronic device 700. The processing unit 701, ROM 702, and RAM 703 are interconnected via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0118] Typically, the following devices can be connected to I / O interface 705: input devices 706 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 707 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 708 including, for example, magnetic tapes, hard disks, etc.; and communication devices 709. Communication device 709 allows electronic device 700 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 7 shows electronic device 700 with various devices, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0119] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 709, or installed from storage device 708, or installed from ROM 702. When the computer program is executed by processing device 701, it performs the functions defined in the methods of embodiments of this disclosure.
[0120] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0121] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0122] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the methods shown in the above embodiments.
[0123] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0124] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0125] The units described in the embodiments of this disclosure can be implemented in software or in hardware. The name of a unit does not necessarily limit the unit itself; for example, the first acquisition unit can also be described as "a unit that acquires at least two Internet Protocol addresses".
[0126] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0127] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
Claims
1. A visual-inertial fusion localization method, comprising: Obtain the target observation data corresponding to the target state vector; The target observation data includes observation data for at least one feature point; The observation data for each feature point includes the observation data of the feature point in at least one camera state; A first residual model is constructed based on the target observation data; the first residual model includes a product term of a first-order partial derivative matrix and a system state matrix, and a first observation noise term; the first-order partial derivative matrix includes a first first-order partial derivative matrix obtained by differentiating the target state vector and a second first-order partial derivative matrix obtained by differentiating the three-dimensional coordinate vector of the at least one feature point; the system state matrix includes a target state error vector and a three-dimensional coordinate error vector of the feature point; the target state vector includes an inertial measurement unit (IMU) state vector and a camera state vector; The residual model is projected based on the first-order partial derivative matrix to obtain a second residual model; the second residual model includes a product term of the second-order partial derivative matrix and the system state matrix, and a covariance term corresponding to the first observation noise term; Eliminate the three-dimensional coordinate error vector of the second residual model to obtain the target residual model, and update the target state vector according to the target residual model using EKF.
2. The method according to claim 1, wherein, The step of constructing the first residual model based on the target observation data includes: For each of the at least one feature points and each of the at least one camera states corresponding to the feature points, a third residual model is constructed based on the observation data of the feature points in the target observation data under the camera states. The third residual model includes a product term of a first partial derivative matrix and a target state error vector, a product term of a second partial derivative matrix and a three-dimensional coordinate error vector of the feature points, and a second observation noise term. The first partial derivative matrix is the first-order partial derivative of the residual function with respect to the target state vector, and the second partial derivative matrix is the first-order partial derivative of the residual function with respect to the three-dimensional coordinate vector of the feature points. The third residual models corresponding to multiple feature points are stacked to obtain a fourth residual model; The product terms in the fourth residual model are merged into a matrix to obtain the first residual model.
3. The method according to claim 2, wherein, For each of the at least one feature point and each of the at least one camera states corresponding to the feature point, a third residual model is constructed based on the observation data of the feature point in the target observation data under the corresponding camera state, including: For each of the at least one feature point and each of the at least one camera states corresponding to the feature point, a residual function is constructed based on the observation data of the feature point in the target observation data under the camera state; the residual function is related to the target state vector and the three-dimensional coordinate vector of the feature point. The residual function is linearized to obtain the third residual model.
4. The method according to claim 2, wherein, The step of stacking the third residual models corresponding to multiple feature points to obtain a fourth residual model includes: The third residual models corresponding to the feature points in the at least one camera state are stacked to obtain a fifth residual model; The fourth residual models corresponding to the at least one feature point are stacked to obtain the fourth residual model.
5. The method according to claim 2, wherein, The acquisition of target observation data corresponding to the target state vector includes: Obtain the first observation data of at least one feature point in the current sliding window corresponding to the target state vector; Acquire second observation data for at least one of the feature points in at least one historical camera state outside the current sliding window; The first observation data and the second observation data are determined as the target observation data.
6. The method according to claim 5, wherein, The step of constructing the first residual model based on the target observation data further includes: For each of the at least one feature points and the corresponding feature point For each historical camera state in at least one historical camera state, a sixth residual model is constructed based on the observation data of the feature point in the target observation data under that historical camera state; the sixth residual model includes a product term of a third partial derivative matrix and the three-dimensional coordinate error vector of the feature point, and a third observation noise term; the third partial derivative matrix is the first-order partial derivative of the residual function constructed with the historical camera state as a constant with respect to the three-dimensional coordinate vector of the feature point; The step of stacking the third residual models corresponding to multiple feature points to obtain a fourth residual model includes: The third and sixth residual models corresponding to the multiple feature points are stacked to obtain the fourth residual model.
7. The method according to any one of claims 1 to 6, wherein, Before constructing the first residual model based on the target observation data, the method further includes: Acquire data from multiple IMUs between the previous time step and the current time step; Based on multiple IMU data, the system state vector of the previous moment is predicted to obtain the system state vector to be expanded; The system acquires the image frame at the current moment, calculates the camera state vector corresponding to the image frame, and expands the system state vector to be expanded based on the camera state vector to obtain the target state vector.
8. The method according to any one of claims 1 to 6, wherein, The step of eliminating the three-dimensional coordinate error vector of the second residual model to obtain the target residual model includes: Based on the Schur elimination method, the three-dimensional coordinate error vector of the second residual model is eliminated to obtain the target residual model.
9. A visual-inertial fusion positioning device, wherein, include: The acquisition module is used to acquire the target observation data corresponding to the target state vector; The target observation data includes observation data of at least one feature point; the observation data of each feature point includes observation data of the feature point in at least one camera state; the target state vector includes multiple camera states; The construction module is used to construct a first residual model based on the target observation data; the first residual model includes a product term of the first-order partial derivative matrix and the system state matrix, and a... An observation noise term; the first-order partial derivative matrix includes a first-order partial derivative matrix obtained by differentiating the target state vector and a second-order partial derivative matrix obtained by differentiating the three-dimensional coordinate vector of the at least one feature point; the system state matrix includes a target state error vector and a three-dimensional coordinate error vector of the feature point; the target state vector includes an inertial measurement unit (IMU) state vector and a camera state vector; The projection module is used to project the residual model based on the first-order partial derivative matrix to obtain a second residual model; the second residual model includes a product term of the second-order partial derivative matrix and the system state matrix, and a covariance term corresponding to the first observation noise term; The elimination module is used to eliminate the three-dimensional coordinate error vector of the second residual model to obtain the target residual model, and to update the target state vector according to the target residual model.
10. An electronic device, wherein, include: Processor and memory; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory, causing the processor to perform the visual-inertial fusion localization method as described in any one of claims 1 to 8.
11. A computer-readable storage medium, wherein, The computer-readable storage medium stores computer-executable instructions, which, when executed by the processor, implement the visual-inertial fusion positioning method as described in any one of claims 1 to 8.
12. A computer program product comprising a computer program, wherein, When the computer program is executed by a processor, it implements the visual-inertial fusion positioning method as described in any one of claims 1 to 8.