A visual inertial initialization method, a mobile device, and a storage medium
Patent Information
- Application Number
- CN202310193043.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-22
- Publication Date
- 2026-09-01
- Estimated Expiration
- 2043-02-22
AI Technical Summary
对于在二维空间运动的可移动设备来说,与三维空间运动的可移动设备相比,由于IMU缺少3个方向上的激励,所以采用上述初始化方式则需要进行长时间的联合优化才能完成初始化,同时视觉惯性初始化精度也较低
Smart Images

Figure CN116164743B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of visual initialization technology, and in particular to a visual inertial initialization method, a mobile device, and a storage medium. Background Technology
[0002] Robots and other mobile devices rely on visual-inertial systems for movement and positioning. VIO (Visual-Inertial Odometry) is the core component of a visual-inertial system, which includes a camera and an IMU (Inertial Measurement Unit). The camera pose is calculated from continuous image frames captured during movement, and then jointly optimized with the predicted pose obtained by pre-integrating the inertial data measured by the IMU to obtain the motion trajectory of the mobile device.
[0003] When a mobile device starts working, visual-inertial initialization is required. Current methods for visual-inertial initialization involve adding all IMU parameters to the objective function for joint optimization. However, for mobile devices moving in two-dimensional space, compared to those moving in three-dimensional space, the IMU lacks excitation in three directions. Therefore, the above initialization method requires a lengthy joint optimization process to complete the initialization, and the accuracy of visual-inertial initialization is also relatively low. Summary of the Invention
[0004] The purpose of this application is to provide a visual inertial initialization method, a movable device, and a storage medium to improve the efficiency and accuracy of visual inertial initialization. The specific technical solution is as follows:
[0005] In a first aspect, embodiments of this application provide a visual inertial initialization method, the method comprising:
[0006] Acquire wheel speed measurement information, inertial information, and multiple image frames captured by the camera from the mobile device;
[0007] If the image frame meets the visual initialization conditions, visual pose estimation is performed based on the image frame to obtain the initial pose.
[0008] If the keyframes included in the image frame satisfy the inertial initialization conditions, inertial pre-integration is performed based on the default parameters to obtain the inertial pre-integration result.
[0009] Based on the pre-integration result of the wheel speed meter information and the pose corresponding to the key frame, the scale parameter and gravity parameter are calculated.
[0010] Using the scale parameter and the gravity parameter as prior constraints, the inertial constraint bias is determined, and the inertial pre-integration result is updated based on the inertial constraint bias and the inertial information.
[0011] The camera pose is determined by joint optimization based on the updated inertial pre-integration results, the scale parameters, and the initial pose.
[0012] Optionally, the step of acquiring wheel speedometer information, inertial information, and multiple image frames captured by the camera of the mobile device includes:
[0013] During the curved motion of the mobile device, the wheel speed meter information, inertial information, and multiple image frames captured by the camera of the mobile device are acquired.
[0014] The steps of calculating scale parameters and gravity parameters based on the pre-integration result of the wheel speed meter information and the pose corresponding to the keyframe include:
[0015] The wheel speedometer information is pre-integrated to obtain the wheel speedometer pose;
[0016] Based on the wheel velocity meter pose and the pose corresponding to the key frame, the scale prior and scale parameter deviation are calculated.
[0017] Based on the first normal vector and the second normal vector, the gravity prior and gravity parameter deviation are calculated, wherein the first normal vector is the normal vector of the plane formed by the poses of the wheel velocity meter, and the second normal vector is the normal vector of the plane formed by the poses corresponding to the keyframes.
[0018] Optionally, the step of calculating the scale prior and scale parameter deviation based on the wheel velocity meter pose and the pose corresponding to the keyframe includes:
[0019] Based on the first position included in the first round of tachometric pose, the second position included in the second round of tachometric pose, the initial position included in the pose corresponding to the initial keyframe, the third position included in the pose corresponding to the first keyframe, and the fourth position included in the pose corresponding to the second keyframe, the scale prior s is calculated according to the following formula. pre and scale parameter deviation s rc :
[0020]
[0021]
[0022]
[0023]
[0024]
[0025]
[0026] Where (x1, y1) are the coordinates of the first position included in the first wheel velocity sensor pose, and (x2, y2) are the coordinates of the second position included in the second wheel velocity sensor pose. k0 y k0 (x) represents the coordinates of the initial position, including the pose corresponding to the initial keyframe. k1 y k1 (x) represents the coordinates of the third position included in the pose corresponding to the first keyframe. k2 y k2 () represents the coordinates of the fourth position included in the pose corresponding to the second keyframe. These are the compensation coefficients for the wheel velocity meter pose in the x and y directions, respectively, and T. rc This is the extrinsic parameter from the camera coordinate system to the mobile device coordinate system.
[0027] Optionally, the step of calculating the prior gravity and the deviation of gravity parameters based on the first normal vector and the second normal vector includes:
[0028] For the plane formed by the wheel speed meter pose, the normal vector of the plane formed by the wheel speed meter pose is calculated by singular value decomposition and used as the first normal vector;
[0029] For the plane formed by the poses corresponding to the keyframes, the normal vector of the plane formed by the poses corresponding to the keyframes is calculated by singular value decomposition and used as the second normal vector.
[0030] Based on the first normal vector and the second normal vector, the prior gravity is calculated according to the following formula. And the deviation of gravity parameter θ:
[0031]
[0032]
[0033] in, Let T be the second normal vector. rc This is the extrinsic parameter from the camera coordinate system to the mobile device coordinate system.
[0034] Optionally, the inertial constraint bias includes linear acceleration bias and angular acceleration bias;
[0035] The step of determining the inertial constraint bias using the scale parameter and the gravity parameter as prior constraints includes:
[0036] Using the scale parameter and the gravity parameter as prior constraints, and taking the minimization of the total error of the scale parameter, the gravity parameter, the linear acceleration bias, and the angular acceleration bias as the optimization direction, iterative optimization of the linear acceleration bias and the angular acceleration bias is performed.
[0037] The process continues until the total error converges, at which point the iteratively optimized linear acceleration bias and angular acceleration bias are obtained.
[0038] Optionally, the step of performing visual pose estimation based on the image frame to obtain an initial pose when the image frame satisfies the visual initialization conditions includes:
[0039] When the number of image frames reaches a first preset number, monocular visual pose estimation is performed based on the image frames to obtain the initial pose.
[0040] Optionally, the step of performing inertial pre-integration based on default parameters to obtain the inertial pre-integration result when the keyframes included in the image frame satisfy the inertial initialization conditions includes:
[0041] When the number of keyframes included in the image frame reaches a second preset number, inertial pre-integration is performed based on default parameters to obtain the inertial pre-integration result.
[0042] Secondly, embodiments of this application provide a mobile device, including a camera with an inertial measurement unit (IMU), a wheel speedometer, and a processor, wherein:
[0043] The camera is used to capture multiple image frames;
[0044] The IMU is used to collect the inertial information of the mobile device;
[0045] The wheel speed meter is used to collect wheel speed information of the mobile device;
[0046] The processor is configured to acquire wheel speed measurement information, inertial information, and multiple image frames captured by the camera of the mobile device; when the image frames meet the visual initialization conditions, perform visual pose estimation based on the image frames to obtain an initial pose; when the keyframes included in the image frames meet the inertial initialization conditions, perform inertial pre-integration based on default parameters to obtain an inertial pre-integration result; calculate scale parameters and gravity parameters based on the pre-integration result of the wheel speed measurement information and the pose corresponding to the keyframes; determine the inertial constraint bias using the scale parameters and the gravity parameters as prior constraints, and update the inertial pre-integration result based on the inertial constraint bias and the inertial information; and perform joint optimization based on the updated inertial pre-integration result, the scale parameters, and the initial pose to determine the pose of the camera.
[0047] Optionally, the processor is specifically configured to, during the curved motion of the mobile device, acquire wheel speed information, inertial information, and multiple image frames captured by the camera; pre-integrate the wheel speed information to obtain the wheel speed pose; calculate the scale prior and scale parameter deviation based on the wheel speed pose and the pose corresponding to the key frame; and calculate the gravity prior and gravity parameter deviation based on the first normal vector and the second normal vector, wherein the first normal vector is the normal vector of the plane formed by the wheel speed pose, and the second normal vector is the normal vector of the plane formed by the pose corresponding to the key frame.
[0048] Optionally, the processor is specifically configured to calculate the scale prior s according to the following formula, based on the first position included in the first round of tachometer pose, the second position included in the second round of tachometer pose, the initial position included in the pose corresponding to the initial keyframe, the third position included in the pose corresponding to the first keyframe, and the fourth position included in the pose corresponding to the second keyframe. pre and scale parameter deviation s rc :
[0049]
[0050]
[0051]
[0052]
[0053]
[0054]
[0055] Where (x1, y1) are the coordinates of the first position included in the first wheel velocity sensor pose, and (x2, y2) are the coordinates of the second position included in the second wheel velocity sensor pose. k0 y k0 (x) represents the coordinates of the initial position, including the pose corresponding to the initial keyframe. k1 y k1 (x) represents the coordinates of the third position included in the pose corresponding to the first keyframe. k2 y k2 () represents the coordinates of the fourth position included in the pose corresponding to the second keyframe. These are the compensation coefficients for the wheel velocity meter pose in the x and y directions, respectively, and T. rc This is the extrinsic parameter from the camera coordinate system to the mobile device coordinate system.
[0056] Optionally, the processor is specifically configured to: calculate the normal vector of the plane formed by the wheel speed meter poses through singular value decomposition, as a first normal vector; calculate the normal vector of the plane formed by the poses corresponding to the keyframes through singular value decomposition, as a second normal vector; and calculate the gravity prior according to the following formula based on the first normal vector and the second normal vector. And the deviation of gravity parameter θ:
[0057]
[0058]
[0059] in, Let T be the second normal vector. rc This is the extrinsic parameter from the camera coordinate system to the mobile device coordinate system.
[0060] Optionally, the inertial constraint bias includes linear acceleration bias and angular acceleration bias;
[0061] The processor is specifically configured to use the scale parameter and the gravity parameter as prior constraints, and take minimizing the total error of the scale parameter, the gravity parameter, the linear acceleration bias, and the angular acceleration bias as the optimization direction, to perform iterative optimization of the linear acceleration bias and the angular acceleration bias; until the total error converges, the iteratively optimized linear acceleration bias and angular acceleration bias are obtained.
[0062] Optionally, the processor is specifically configured to perform monocular visual pose estimation based on the image frames to obtain an initial pose when the number of image frames reaches a first preset number.
[0063] Optionally, the processor is specifically configured to perform inertial pre-integration based on default parameters to obtain an inertial pre-integration result when the number of keyframes included in the image frame reaches a second preset number.
[0064] Thirdly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements any of the methods described in the first aspect above.
[0065] Beneficial effects of the embodiments in this application:
[0066] The solution provided in this application embodiment allows the processor to acquire wheel speed measurement information, inertial information, and multiple image frames captured by the camera from a mobile device. If the image frames meet the visual initialization conditions, visual pose estimation is performed based on the image frames to obtain an initial pose. If the keyframes included in the image frames meet the inertial initialization conditions, inertial pre-integration is performed based on default parameters to obtain an inertial pre-integration result. Based on the pre-integration result of the wheel speed measurement information and the pose corresponding to the keyframes, scale parameters and gravity parameters are calculated. Using the scale parameters and gravity parameters as prior constraints, an inertial constraint bias is determined. Based on the inertial constraint bias and inertial information, the inertial pre-integration result is updated. Based on the updated inertial pre-integration result, scale parameters, and initial pose, joint optimization is performed to determine the camera pose. In visual inertial initialization on mobile devices, scale and gravity parameters can be calculated first. These parameters are then used as prior constraints to optimize the remaining IMU parameters, updating the inertial pre-integration results. Based on the updated pre-integration results, scale parameters, and initial pose, joint optimization is performed to determine the camera pose. This eliminates the need to include all IMU parameters in the objective function for joint optimization, reducing optimization time and allowing for camera pose correction, thus improving the efficiency and accuracy of visual inertial initialization. Of course, implementing any product or method of this application does not necessarily require achieving all of the above advantages simultaneously. Attached Figure Description
[0067] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other embodiments can be obtained based on these drawings.
[0068] Figure 1 A flowchart illustrating a visual inertial initialization method provided in an embodiment of this application;
[0069] Figure 2(a) shows the results based on Figure 1 A schematic diagram illustrating convergence of camera pose iterative optimization in the embodiment shown;
[0070] Figure 2(b) shows the results based on... Figure 1 Another schematic diagram illustrating convergence of camera pose iterative optimization in the embodiment shown;
[0071] Figure 3 Based on Figure 1 A schematic diagram of a mobile device performing curved motion according to the embodiment shown;
[0072] Figure 4 for Figure 1A specific flowchart of step S104 in the illustrated embodiment;
[0073] Figure 5 Based on Figure 1 The illustrated embodiment is a schematic diagram of a plane composed of three poses that are not on the same straight line;
[0074] Figure 6 for Figure 1 A specific flowchart of step S105 in the illustrated embodiment;
[0075] Figure 7 Based on Figure 1 A schematic diagram of a ground robot component module in the embodiment shown;
[0076] Figure 8 A specific flowchart of the visual inertial initialization method provided in the embodiments of this application;
[0077] Figure 9 This is a schematic diagram of the structure of a mobile device provided in an embodiment of this application. Detailed Implementation
[0078] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art based on this application are within the scope of protection of this application.
[0079] To improve the efficiency and accuracy of visual inertial initialization, embodiments of this application provide a visual inertial initialization method, a mobile device, a computer-readable storage medium, and a computer program product. The visual inertial initialization method provided in this application embodiment will be described first below.
[0080] The visual inertial initialization method provided in this application embodiment can be applied to mobile devices that require visual inertial initialization, wherein the mobile device includes a camera with an inertial measurement unit (IMU), a wheel speed meter, and a processor.
[0081] like Figure 1 As shown, a visual inertial initialization method includes:
[0082] S101, acquire wheel speed meter information, inertial information and multiple image frames captured by the camera from the mobile device;
[0083] S102, if the image frame satisfies the visual initialization conditions, visual pose estimation is performed based on the image frame to obtain the initial pose;
[0084] S103, if the keyframes included in the image frame satisfy the inertial initialization conditions, perform inertial pre-integration based on the default parameters to obtain the inertial pre-integration result;
[0085] S104, Calculate the scale parameters and gravity parameters based on the pre-integration result of the wheel speed meter information and the pose corresponding to the key frame;
[0086] S105, using the scale parameter and the gravity parameter as prior constraints, determine the inertial constraint bias, and update the inertial pre-integration result based on the inertial constraint bias and the inertial information;
[0087] S106, based on the updated inertial pre-integration results, the scale parameters, and the initial pose, perform joint optimization to determine the camera pose.
[0088] As can be seen, in the solution provided in this application embodiment, the processor can acquire wheel speedometer information, inertial information, and multiple image frames captured by the camera of the mobile device. When the image frames meet the visual initialization conditions, visual pose estimation is performed based on the image frames to obtain the initial pose. When the key frames included in the image frames meet the inertial initialization conditions, inertial pre-integration is performed based on default parameters to obtain the inertial pre-integration result. Based on the pre-integration result of the wheel speedometer information and the pose corresponding to the key frames, the scale parameters and gravity parameters are calculated. The scale parameters and gravity parameters are used as prior constraints to determine the inertial constraint bias. Based on the inertial constraint bias and inertial information, the inertial pre-integration result is updated. Based on the updated inertial pre-integration result, the scale parameters, and the initial pose, joint optimization is performed to determine the camera pose. Since visual inertial initialization can be performed on mobile devices, scale and gravity parameters can be calculated first. Then, the scale and gravity parameters can be used as prior constraints to optimize the remaining parameters of the IMU to update the inertial pre-integration results. Based on the updated pre-integration results, scale parameters, and initial pose, joint optimization can be performed to determine the camera pose. This eliminates the need to add all IMU parameters to the objective function for joint optimization, reducing the time required for joint optimization and allowing for correction of the camera pose, thereby improving the efficiency and accuracy of visual inertial initialization.
[0089] When a mobile device starts working, visual-inertial initialization is required. Initialization refers to quickly estimating the current pose of the camera in the initial stage.
[0090] Currently, visual inertial initialization (VII) involves jointly optimizing all IMU parameters within the objective function. If the mobile device is performing VIII in three-dimensional space, it typically requires the camera and IMU to undergo displacements in each direction or have sufficiently accurate estimates to achieve good initialization results. However, when performing VIII in two-dimensional space, the lack of excitation from the IMU in three directions necessitates lengthy joint optimization to complete the initialization process, resulting in lower VIII efficiency and accuracy.
[0091] To improve the efficiency and accuracy of visual inertial initialization, this application provides a visual inertial initialization method. In step S101, the processor can acquire wheel speedometer information, inertial information, and multiple image frames captured by the camera from the mobile device.
[0092] The wheel speed meter can be an encoder mounted on a motor, which calculates the distance and angle traveled by the mobile device based on the number of revolutions of the motor. Wheel speed meter information includes the position coordinates and angle data of the mobile device at any given location.
[0093] Inertial information can be acquired through an IMU, which is a sensor used to detect and measure acceleration and rotational motion. Inertial information includes angular velocity data collected by the gyroscope in the IMU and acceleration data collected by the accelerometer.
[0094] Image frames may include continuous frames captured by the camera of a mobile device during movement, or keyframes captured according to actual needs, without specific limitations.
[0095] Since the camera also begins capturing image frames when the mobile device starts operating, the processor can perform visual pose estimation based on the image frames, provided the image frames meet the visual initialization conditions, to obtain the initial pose, i.e., execute step S102. Visual pose estimation refers to the process of estimating the pose transformation using image data acquired by the visual sensor installed on the mobile device. For example, the initial pose can be estimated using the matching relationship between two adjacent image frames acquired by the mobile device's camera.
[0096] In one implementation, when the number of image frames captured by the camera reaches a preset number, the processor can perform visual pose estimation based on the image frames captured by the camera to obtain the initial pose of the camera.
[0097] For example, the visual initialization condition is that the number of image frames acquired by the camera is greater than or equal to a preset number of 2. When the camera acquires image frame 1 and image frame 2, that is, when the number of image frames reaches the preset number, the processor can perform visual pose estimation based on the acquired image frame 1 and image frame 2, and then obtain the initial pose of the camera.
[0098] Next, after obtaining the initial pose of the camera using image frames captured by the camera, the processor can pre-integrate the inertial data to obtain the predicted pose based on the inertial initialization conditions. Under the condition that the inertial initialization conditions are met, inertial pre-integration is performed using default parameters on one hand, and on the other hand, some parameters are calculated to optimize the remaining parameters in the inertial pre-integration. Then, the optimized parameters are used to replace the default parameters for inertial pre-integration.
[0099] In step S103, if the keyframes included in the image frame satisfy the inertial initialization conditions, the processor can perform inertial pre-integration based on default parameters to obtain the inertial pre-integration result. Here, a keyframe refers to an image frame spaced at a certain distance or possessing special representativeness, and is not specifically limited here. For example, a keyframe can be an image frame captured by the camera at its initial position, or it can be an image frame captured at positions spaced 1 meter apart.
[0100] In one implementation, when the number of keyframes included in the image frame reaches a preset number, the processor can perform inertial pre-integration based on default parameters to obtain the inertial pre-integration result.
[0101] For example, the mobile device is a ground robot, and the preset number of keyframes is 3. The keyframes are image frames captured by the camera at various positions with a distance of 1m. When the camera captures image frame 0 of the ground robot at the initial position, image frame 1 at the 1m position, and image frame 2 at the 2m position, that is, when the number of keyframes reaches the preset number of 3, the processor can perform inertial pre-integration based on the default parameters, and then obtain the inertial pre-integration result.
[0102] While the processor performs inertial pre-integration based on default parameters, in step S104, the processor can calculate scale parameters and gravity parameters based on the pre-integration results of the wheel velocity meter information and the pose corresponding to the key frame.
[0103] Since the wheel speed meter can collect wheel speed information of the mobile device at a preset frequency during its movement, the processor can acquire the wheel speed information when the number of keyframes in the image frame reaches a preset number. Then, based on this wheel speed information, pre-integration is performed to obtain the pre-integration result. For example, the pre-integration result can be the pose of the wheel speed meter at the location of the keyframe. Therefore, based on the wheel speed pre-integration result and the pose corresponding to the keyframe, scale parameters can be calculated.
[0104] The wheel speed meter pose is in two-dimensional space. Singular value decomposition can be performed on the plane formed by the wheel speed meter pose to obtain the first normal vector of the plane. The poses corresponding to the key frames captured by the camera can also form a plane. Singular value decomposition can be performed on the plane formed by the poses corresponding to the key frames to obtain the second normal vector of the plane. Then, gravity parameters can be calculated based on the first normal vector and the second normal vector.
[0105] For example, the camera captures keyframes 1, 2, and 3, which represent image frames captured by the camera when the ground robot is at its initial position, at a position of 1m, and at a position of 2m, respectively. When the number of keyframes reaches a preset number of 3, the processor can acquire the wheel velocity meter information of the ground robot. Based on the wheel velocity meter information, pre-integration can be performed to obtain the wheel velocity meter poses of the ground robot at the 1m and 2m positions. Then, based on the wheel velocity meter poses and the poses corresponding to the keyframes, the scale parameters can be calculated.
[0106] The wheel velocity sensor poses are in two-dimensional space. The poses of the wheel velocity sensor at the initial position, 1m position, and 2m position can form a plane. Singular value decomposition (SVD) can be performed on this plane to obtain the first normal vector. Similarly, the poses corresponding to keyframe 1, keyframe 2, and keyframe 3 can also form a plane. SVD can be performed on this plane to obtain the second normal vector. Then, gravity parameters can be calculated based on the first and second normal vectors.
[0107] Next, in step S105, the processor can determine the inertial constraint bias using the scale parameter and gravity parameter as prior constraints, and update the inertial pre-integration result based on the inertial constraint bias and inertial information. The inertial constraint bias may include linear acceleration bias and angular acceleration bias.
[0108] For example, the parameters in the inertial pre-integration include scale parameters, gravity parameters, linear acceleration bias, and angular acceleration bias. After calculating the scale parameters and gravity parameters, the linear acceleration bias and angular acceleration bias can be optimized using the scale parameters and gravity parameters as prior constraints. The optimized parameters replace the default parameters of the inertial pre-integration, thereby updating the inertial pre-integration results.
[0109] For example, as shown in Figure 2(a), from left to right are the scale parameter, gravity parameter, linear acceleration offset, and angular acceleration offset. Since IMU data degenerates from six-dimensional data to three-dimensional data in two-dimensional space, the inertial pre-integration results have significant errors. If the acquired parameters are used for inertial pre-integration to iteratively optimize the camera pose, with the optimization direction being the minimum total error, the above four parameters will diverge in all directions and fail to converge over a long period, making it difficult to obtain ideal initialization results.
[0110] In Figure 2(b), from left to right are the scale parameter, gravity parameter, linear acceleration bias, and angular acceleration bias. Using the obtained scale parameter and gravity parameter as prior conditions, the convergence of the linear acceleration bias and angular acceleration bias can be promoted, thus determining the inertial constraint bias. Furthermore, the processor can update the inertial pre-integration results based on the inertial constraint bias and inertial information to optimize the inertial pre-integration results.
[0111] After updating the inertial pre-integration results, in step S106, the processor can perform joint optimization based on the updated inertial pre-integration results, scale parameters, and initial pose to determine the camera pose. This joint optimization can be local bundle adjustment (BA) optimization, which is not specifically limited here. Local bundle adjustment refers to minimizing the error by incorporating the camera pose, map points, and the pose obtained from inertial pre-integration into the same objective function across several adjacent keyframes.
[0112] For example, the processor has obtained the initial pose of the camera through visual pose estimation, and calculated the scale parameters and gravity parameters, where the scale parameters include the scale prior s. pre and scale parameter deviation s rc Gravity parameters include prior gravity. And the gravity parameter deviation θ. Using the scale parameter and gravity parameter as prior constraints, the acceleration bias and angular acceleration bias can be determined. Based on the determined linear acceleration bias, angular acceleration bias, and inertial information, the inertial pre-integration results are updated. Then, based on the updated inertial pre-integration results, scale parameters, and initial pose, local BA optimization can be performed to determine the camera pose.
[0113] In this embodiment, when performing visual inertial initialization on a mobile device, scale and gravity parameters can be calculated first. These parameters are then used as prior constraints to optimize the remaining IMU parameters, updating the inertial pre-integration results. Based on the updated pre-integration results, scale parameters, and initial pose, joint optimization is performed to determine the camera pose. This eliminates the need to include all IMU parameters in the objective function for joint optimization, reducing optimization time and allowing for camera pose correction, thus improving the efficiency and accuracy of visual inertial initialization. Therefore, it accelerates camera pose convergence, achieving rapid initialization.
[0114] As one embodiment of this application, the steps of obtaining wheel speedometer information, inertial information, and multiple image frames captured by the camera of the mobile device may include:
[0115] During the curved motion of the mobile device, wheel speed meter information, inertial information, and multiple image frames captured by the camera are acquired.
[0116] To determine the gravity parameters, it is necessary to obtain the plane formed by the wheel speedometer pose and the plane formed by the poses corresponding to the keyframes, so as to perform singular value decomposition to calculate the normal vector of each plane. Since three non-collinear points determine a plane, in this embodiment, multiple poses that are not on the same straight line can determine a plane. Therefore, the mobile device can perform curved motion over a preset distance during initialization. The preset distance can be 2m, 3m, etc., and is not specifically limited here.
[0117] During the curved motion of the mobile device, the processor acquires wheel speed measurement information, inertial information, and multiple image frames captured by the camera.
[0118] For example, such as Figure 3 As shown, the mobile device is a ground robot. When the ground robot starts working, it can perform a 2m curved motion, that is, walk a curved trajectory 301. During the curved motion of the ground robot, the processor can acquire the wheel speed meter information, inertial information and multiple image frames captured by the camera.
[0119] like Figure 4 As shown, the steps for calculating scale parameters and gravity parameters based on the pre-integration results of the wheel velocity meter information and the pose corresponding to the keyframe may include:
[0120] S401, pre-integrate the wheel speed meter information to obtain the wheel speed meter pose;
[0121] The processor can begin acquiring wheel velocity meter information of the mobile device when the image frame meets the visual initialization conditions, and then perform pre-integration based on the wheel velocity meter information to obtain the wheel velocity meter pose when the key frames included in the image frame meet the inertial initialization conditions.
[0122] For example, during the 2m curved motion of a ground robot, the processor can acquire the wheel velocity meter information of the ground robot, integrate the wheel velocity meter information when the ground robot moves to the 1m position, and pre-integrate the wheel velocity meter information when it moves to the 2m position. Then, the wheel velocity meter pose of the ground robot when it moves to the 1m position and the wheel velocity meter pose when it moves to the 2m position can be obtained.
[0123] S402, Based on the wheel speed meter pose and the pose corresponding to the key frame, calculate the scale prior and scale parameter deviation;
[0124] After the processor obtains the wheel velocity meter pose, it can calculate the scale prior and scale parameter deviation based on the wheel velocity meter pose and the pose corresponding to the key frame, thus obtaining the scale parameters.
[0125] For example, such as Figure 3 As shown, the keyframes captured by the camera are the initial keyframe 0, keyframe 1 when the ground robot moves to a position of 1m, and keyframe 2 when the ground robot moves to a position of 2m. In the ground robot coordinate system, the wheel velocities of the ground robot at the 1m position and the 2m position are obtained by pre-integration based on the wheel velocities information, respectively, as p1 = (x1, y1, θ1) and p2 = (x2, y2, θ2). In the camera coordinate system, the keyframe poses can be obtained, namely the poses corresponding to the initial keyframe 0, keyframe 1, and keyframe 2, which can be represented as P ki =(x ki y ki θ ki ), i∈0,1,2, where P ko It is the pose corresponding to the initial keyframe 0, P k1 It is the pose corresponding to keyframe 1, P k2 This is the pose corresponding to keyframe 2. Therefore, the processor can calculate the scale prior and scale parameter deviation based on the wheel velocimeter pose and the poses corresponding to each keyframe.
[0126] S403 calculates the prior gravity and gravity parameter deviation based on the first and second normal vectors.
[0127] Wherein, the first normal vector is the normal vector of the plane formed by the poses of the wheel speedometer, and the second normal vector is the normal vector of the plane formed by the poses corresponding to the keyframes.
[0128] Since the mobile device moves along curved lines and the wheel velocity meter pose is in two-dimensional space, multiple wheel velocity meter poses that are not on the same straight line can form a plane. Similarly, the poses corresponding to multiple keyframes that are not on the same straight line can also form a plane. The normal vector of the plane formed by the wheel velocity meter poses is the first normal vector, and the normal vector of the plane formed by the poses corresponding to the keyframes is the second normal vector.
[0129] The processor can calculate the gravity prior and gravity parameter deviation based on the first normal vector and the second normal vector, thus obtaining the gravity parameters.
[0130] For example, such as Figure 5 As shown, the keyframes captured by the camera are keyframe 1, keyframe 2, and keyframe 3, and the processor can acquire the pose of each keyframe. Based on the acquired wheel velocity sensor information, the wheel velocity sensor poses of the ground robot at the initial position, at a position of 1m, and at a position of 2m can be obtained. The poses of each wheel velocity sensor can form a plane, and the normal vector of this plane is the first normal vector. The poses corresponding to keyframe 1, keyframe 2, and keyframe 3 can also form a plane, and the normal vector of this plane is the second normal vector. Therefore, the processor can calculate the gravity prior and gravity parameter deviation based on the first and second normal vectors.
[0131] As can be seen, in this embodiment, the processor can acquire wheel speed measurement information, inertial information, and multiple image frames captured by the camera during the curved motion of the mobile device. It pre-integrates the wheel speed measurement information to obtain the wheel speed measurement pose. Based on the wheel speed measurement pose and the poses corresponding to keyframes, it calculates the scale prior and scale parameter deviation. Based on the first normal vector and the second normal vector, it calculates the gravity prior and gravity parameter deviation, where the first normal vector is the normal vector of the plane formed by the wheel speed measurement poses, and the second normal vector is the normal vector of the plane formed by the poses corresponding to keyframes. Since the scale prior and gravity prior can be calculated simultaneously with the curved motion of the mobile device, the remaining IMU parameters can be better optimized. Determining the scale prior and gravity prior allows more weight to be assigned to the inertial constraint bias, thereby reducing iterative optimization time and optimizing the inertial pre-integration results. Therefore, it can accelerate the convergence of the camera pose, achieving rapid initialization and improving the efficiency and accuracy of visual inertial initialization.
[0132] As one embodiment of this application, the steps of calculating the scale prior and scale parameter deviation based on the wheel velocity meter pose and the pose corresponding to the keyframe may include:
[0133] Based on the first position included in the first round of tachometric pose, the second position included in the second round of tachometric pose, the initial position included in the pose corresponding to the initial keyframe, the third position included in the pose corresponding to the first keyframe, and the fourth position included in the pose corresponding to the second keyframe, the scale prior s is calculated according to the following formula. pre and scale parameter deviation s rc :
[0134]
[0135]
[0136]
[0137]
[0138]
[0139]
[0140] Where (x1, y1) are the coordinates of the first position included in the first wheel velocity sensor pose, and (x2, y2) are the coordinates of the second position included in the second wheel velocity sensor pose. k0 y k0 (x) represents the coordinates of the initial position, including the pose corresponding to the initial keyframe. k1 y k1 (x) represents the coordinates of the third position included in the pose corresponding to the first keyframe. k2 y k2 () represents the coordinates of the fourth position included in the pose corresponding to the second keyframe. These are the compensation coefficients for the wheel velocity meter pose in the x and y directions, respectively, and T. rc This is the extrinsic parameter from the camera coordinate system to the mobile device coordinate system.
[0141] For wheel velocity meters, relative pose estimation has higher accuracy than absolute pose estimation. Therefore, given the wheel velocity meter pose and the poses corresponding to keyframes, the processor can calculate the scale prior s based on the first position included in the first wheel velocity meter pose, the second position included in the second wheel velocity meter pose, the initial position included in the pose corresponding to the initial keyframe, the third position included in the pose corresponding to the first keyframe, and the fourth position included in the pose corresponding to the second keyframe, according to the scale prior formula and the scale parameter deviation formula. pre and scale parameter deviation s rc .
[0142] For example, following the example of step S402, in the ground robot coordinate system, based on the pre-integration of the wheel velocity sensor information, the wheel velocity sensor poses of the ground robot at the 1m position and the 2m position are obtained as p1 = (x1, y1, θ1) and p2 = (x2, y2, θ2), respectively. In the camera coordinate system, the poses corresponding to the initial keyframe 0, keyframe 1, and keyframe 2 can be obtained, i.e., P ki =(x ki y ki θ ki ), i∈0,1,2, where P k0 It is the pose corresponding to the initial keyframe 0, P k1 It is the pose corresponding to keyframe 1, P k2 This is the pose corresponding to keyframe 2. Therefore, the processor can calculate the scale prior s based on the wheel velocimeter pose and the poses corresponding to each keyframe using the following formula. pre and scale parameter deviation s rc :
[0143]
[0144]
[0145]
[0146]
[0147]
[0148]
[0149] Where (x1, y1) are the coordinates of the first position of the ground robot's wheel velocity sensor pose when it moves to a position of 1m, and (x2, y2) are the coordinates of the second position of the ground robot's wheel velocity sensor pose when it moves to a position of 2m. k0 y k0 (x) represents the coordinates of the initial position, including the pose corresponding to the initial keyframe 0. k1 y k1 (x) represents the coordinates of the third position included in the pose corresponding to keyframe 1. k2 y k2 () represents the coordinates of the fourth position in the pose corresponding to keyframe 2. These are the compensation coefficients for the wheel velocity meter pose in the x and y directions, respectively, and T. rc This is the extrinsic parameter from the camera coordinate system to the ground robot coordinate system. (D) r1 S r2This represents the Euclidean distances for the first 1m and the second 1m distances calculated based on the wheel velocities and pose of the ground robot. (D) k1 D k2 This represents the Euclidean distance for calculating the first 1m distance and the second 1m distance of the ground robot's motion based on the pose corresponding to the keyframe.
[0150] As can be seen, in this embodiment, the processor can calculate the scale prior s based on the first position included in the first round of tachometer pose, the second position included in the second round of tachometer pose, the initial position included in the pose corresponding to the initial keyframe, the third position included in the pose corresponding to the first keyframe, and the fourth position included in the pose corresponding to the second keyframe, according to the scale prior formula and the scale parameter deviation formula. pre and scale parameter deviation s rc Because of the computational scale prior and scale parameter bias, the remaining IMU parameters can be optimized along with the subsequently calculated gravity parameters. This allows for the allocation of more weight to the inertial constraint bias, thereby reducing iterative optimization time and improving the inertial pre-integration results. Therefore, it can accelerate camera pose convergence, achieving rapid initialization and improving the efficiency and accuracy of visual inertial initialization.
[0151] As one embodiment of this application, the step of calculating the prior gravity and gravity parameter deviation based on the first normal vector and the second normal vector may include:
[0152] For the plane formed by the wheel speed meter pose, the normal vector of the plane formed by the wheel speed meter pose is calculated by singular value decomposition and used as the first normal vector;
[0153] For the plane formed by the poses corresponding to the keyframes, the normal vector of the plane formed by the poses corresponding to the keyframes is calculated by singular value decomposition and used as the second normal vector.
[0154] Based on the first normal vector and the second normal vector, the prior gravity is calculated according to the following formula. And the deviation of gravity parameter θ:
[0155]
[0156]
[0157] in, Let T be the second normal vector. rc This is the extrinsic parameter from the camera coordinate system to the mobile device coordinate system.
[0158] The core of calculating gravity parameters is obtaining a set of data within a plane; that is, defining a plane using data that is not on a straight line. When the data meets the requirements for forming a plane, the plane's normal vector can be calculated using singular value decomposition. For example, ... Figure 5 As shown, a plane can be formed by the poses corresponding to multiple keyframes, and the normal vector of the plane can be calculated by singular value decomposition.
[0159] For a wheel speed meter, since its pose is in two-dimensional space, its plane is aligned with the horizontal plane, and the normal vector of the wheel speed meter is approximately the direction of gravity. For a camera, since its pose is six-dimensional, its plane is determined by the direction initialized by the pose corresponding to the first image frame, so the absolute direction of its normal vector cannot be determined.
[0160] To obtain prior gravity parameters, the normal vector of the plane formed by the wheel speedometer pose is calculated through singular value decomposition. As the first normal vector, the normal vector of the plane formed by the poses corresponding to the keyframes is calculated through singular value decomposition. As the second normal vector, the direction of gravity estimated by the camera is obtained according to the following formula:
[0161]
[0162] Among them, T rc This is the extrinsic parameter from the camera coordinate system to the mobile device coordinate system.
[0163] Under ideal conditions and They are parallel and both point vertically downwards. However, due to the lack of accuracy in motion estimation for mobile devices, the two vectors may have an angle, resulting in a gravity parameter deviation. This deviation θ can be calculated using the following formula:
[0164]
[0165] Therefore, the prior law of gravity is obtained as follows:
[0166] As can be seen, in this embodiment, the processor can calculate the normal vector of the plane formed by the wheel speed meter poses through singular value decomposition, which serves as the first normal vector. Similarly, it can calculate the normal vector of the plane formed by the poses corresponding to keyframes through singular value decomposition, which serves as the second normal vector. Based on the first and second normal vectors, the gravity prior is calculated according to the gravity direction estimated by the camera and the gravity parameter deviation formula. And the gravity parameter deviation θ. By calculating the gravity prior and gravity parameter deviation, the remaining parameters of the IMU can be optimized along with the calculated scale parameters, assigning more weight to the inertial constraint bias, thereby reducing the iteration optimization time and optimizing the inertial pre-integration results. Therefore, it can accelerate the convergence of camera pose, thereby achieving the goal of rapid initialization and improving the efficiency and accuracy of visual inertial initialization.
[0167] As one embodiment of this application, the above-mentioned inertial constraint bias includes linear acceleration bias and angular acceleration bias;
[0168] like Figure 6 As shown, the steps described above for determining the inertial constraint bias using the scale parameter and the gravity parameter as prior constraints may include:
[0169] S601, using the scale parameter and the gravity parameter as prior constraints, and taking the minimization of the total error of the scale parameter, the gravity parameter, the linear acceleration bias, and the angular acceleration bias as the optimization direction, perform iterative optimization of the linear acceleration bias and the angular acceleration bias;
[0170] S602, until the total error converges, the iteratively optimized linear acceleration bias and angular acceleration bias are obtained.
[0171] Since the scale prior and gravity prior are calculated, the scale parameter deviation and gravity parameter deviation are also calculated accordingly. Therefore, when estimating the camera pose by combining other parameters with inertial pre-integration, the pose convergence direction can be clearly defined.
[0172] In Figure 2(b), from left to right are the scale parameter, gravity parameter, linear acceleration bias, and angular acceleration bias. The processor can use the scale parameter and gravity parameter as prior constraints, and take minimizing the total error of the scale parameter, gravity parameter, linear acceleration bias, and angular acceleration bias as the optimization direction. Iterative optimization of the linear acceleration bias and angular acceleration bias is performed until the total error converges, thus obtaining the iteratively optimized linear acceleration bias and angular acceleration bias.
[0173] In other words, given that the scale parameter and gravity parameter are determined as prior constraints, the scale parameter and gravity parameter can promote the convergence of linear acceleration bias and angular acceleration bias, and thus obtain converged linear acceleration bias and angular acceleration bias when the total error tends to remain constant.
[0174] As can be seen, in this embodiment, scale parameters and gravity parameters are used as prior constraints. The optimization direction is to minimize the total error of scale parameters, gravity parameters, linear acceleration bias, and angular acceleration bias. Iterative optimization of linear acceleration bias and angular acceleration bias is performed until the total error converges, resulting in the iteratively optimized linear acceleration bias and angular acceleration bias. This not only reduces the iteration optimization time but also allows more weight to be assigned to linear acceleration bias and angular acceleration bias to optimize the inertial pre-integration results, thereby achieving the goal of rapid initialization and improving the efficiency and accuracy of visual inertial initialization.
[0175] As one embodiment of this application, the step of performing visual pose estimation based on the image frame to obtain the initial pose when the image frame satisfies the visual initialization conditions may include:
[0176] When the number of image frames reaches a first preset number, monocular visual pose estimation is performed based on the image frames to obtain the initial pose.
[0177] In one embodiment, the mobile device includes a monocular camera, and the mobile device is a ground robot, which refers to a mobile robot that moves in a two-dimensional plane. The ground robot includes a monocular camera with an IMU and is configured with logic operations. When the number of image frames acquired by the monocular camera reaches a first preset number, the processor can perform monocular visual pose estimation based on the acquired image frames to obtain an initial pose. The first preset number can be 2, 3, etc., and is not specifically limited here.
[0178] For example, such as Figure 7 As shown, the first preset number is 2, and the ground robot 701 includes a monocular camera 702 with an IMU sensor 703. The robot coordinate system is a three-dimensional coordinate system composed of x, y, and z. The monocular camera 702 can acquire image frames at times t1, t2, ..., tn during the movement of the ground robot 701. When the number of image frames reaches 2, the monocular camera can perform monocular visual pose estimation based on the acquired image frames, thereby obtaining the initial pose.
[0179] As can be seen, in this embodiment, when the number of image frames reaches a first preset number, the processor can perform monocular visual pose estimation based on the image frames to obtain the initial pose. In this way, the camera pose can be determined based on the updated pre-integration results, scale parameters, and the initial pose through joint optimization. This eliminates the need to include all IMU parameters in the objective function for joint optimization, reducing the time required for joint optimization and allowing for correction of the camera pose, thereby improving the efficiency and accuracy of visual inertial initialization.
[0180] As one embodiment of this application, the step of performing inertial pre-integration based on default parameters to obtain the inertial pre-integration result when the keyframes included in the image frame satisfy the inertial initialization conditions may include:
[0181] When the number of keyframes included in the image frame reaches a second preset number, inertial pre-integration is performed based on default parameters to obtain the inertial pre-integration result.
[0182] In one implementation, the processor can perform a curved motion over a certain distance on the mobile device. When the number of keyframes included in the image frame reaches a second preset number, inertial pre-integration is performed based on default parameters to obtain the inertial pre-integration result. The second preset number can be 3, 4, etc., and is not specifically limited here.
[0183] For example, if the mobile device is a ground robot and the preset number of keyframes is 3, when the camera captures keyframe 1, keyframe 2, and keyframe 3, that is, when the number of keyframes reaches the preset number of 3, the processor can perform inertial pre-integration based on the default parameters, and then obtain the inertial pre-integration result.
[0184] As can be seen, in this embodiment, when the number of keyframes included in the image frame reaches a second preset number, inertial pre-integration is performed based on default parameters to obtain the inertial pre-integration result. Thus, the inertial pre-integration result can be updated based on the inertial constraint bias and inertial information to optimize the inertial pre-integration result.
[0185] Figure 8 This is a specific flowchart of a visual inertial initialization method provided in an embodiment of this application. The following is in conjunction with... Figure 8 The visual inertial initialization method provided in the embodiments of this application will be described with examples. For instance... Figure 8 As shown, the visual inertial initialization method provided in this application embodiment may include the following steps:
[0186] S801, determine whether to acquire two frames of images;
[0187] When the mobile device starts working, the camera can capture image frames. The processor obtains the image frames captured by the camera and determines whether to acquire two image frames. If two image frames are acquired, pure visual pose estimation can be performed based on the image frames, i.e., step S802 is executed. Otherwise, image frames are acquired.
[0188] S802, pure vision pose estimation;
[0189] If the processor acquires two frames of images, pure visual pose estimation can be performed based on the acquired two frames of images.
[0190] S803, determine whether the IMU initialization conditions are met;
[0191] Determine whether the keyframes included in the image frames acquired by the processor meet the inertial initialization conditions. If yes, perform inertial pre-integration based on the default parameters and execute step S804; otherwise, execute step S802.
[0192] S804, IMU pre-integration;
[0193] If the keyframes in the image frames acquired by the processor meet the inertial initialization conditions, inertial pre-integration can be performed based on the default parameters to obtain the inertial pre-integration result.
[0194] S805, acquire wheel speed meter data;
[0195] With two frames of images acquired, the processor can obtain wheel speed meter information, i.e., wheel speed meter data.
[0196] S806, wheel speed meter pre-integral;
[0197] The processor can perform pre-integration of the wheel speed meter data to obtain the wheel speed meter integration result.
[0198] S807, estimating prior gravity parameters;
[0199] If the keyframes in the image frames acquired by the processor meet the inertial initialization conditions, the processor can perform multi-parameter step-by-step estimation while performing inertial pre-integration based on default parameters. Then, the processor can calculate gravity priors and gravity parameter deviations based on the normal vectors of the plane formed by the wheel velocimeter poses and the plane formed by the poses corresponding to the keyframes.
[0200] S808, estimating prior scale parameters;
[0201] After pre-integrating the wheel velocity meter information to obtain the wheel velocity meter pose, the processor can calculate the scale prior and scale parameter deviation based on the wheel velocity meter pose and the pose corresponding to the key frame.
[0202] S809, optimizes gravity, scale, and bias parameters;
[0203] After calculating the gravity and scale parameters, the processor can use these parameters as prior constraints to optimize the inertial constraint biases, thereby determining the inertial constraint biases, i.e., the linear acceleration bias and the angular acceleration bias. The optimized parameters are then used to replace the default parameters in the inertial pre-integration, updating the inertial pre-integration results.
[0204] S810, pose joint optimization.
[0205] The processor can perform joint optimization based on the updated inertial pre-integration results, scale parameters, and initial pose to determine the camera pose.
[0206] As can be seen, in the solution provided in this application embodiment, the processor can acquire wheel speedometer information, inertial information, and multiple image frames captured by the camera of the mobile device. When the image frames meet the visual initialization conditions, visual pose estimation is performed based on the image frames to obtain the initial pose. When the key frames included in the image frames meet the inertial initialization conditions, inertial pre-integration is performed based on default parameters to obtain the inertial pre-integration result. Based on the pre-integration result of the wheel speedometer information and the pose corresponding to the key frames, the scale parameters and gravity parameters are calculated. The scale parameters and gravity parameters are used as prior constraints to determine the inertial constraint bias. Based on the inertial constraint bias and inertial information, the inertial pre-integration result is updated. Based on the updated inertial pre-integration result, the scale parameters, and the initial pose, joint optimization is performed to determine the camera pose. Since visual inertial initialization can be performed on mobile devices, scale and gravity parameters can be calculated first. Then, the scale and gravity parameters can be used as prior constraints to optimize the remaining parameters of the IMU to update the inertial pre-integration results. Based on the updated pre-integration results, scale parameters, and initial pose, joint optimization can be performed to determine the camera pose. This eliminates the need to add all IMU parameters to the objective function for joint optimization, reducing the time required for joint optimization and allowing for correction of the camera pose, thereby improving the efficiency and accuracy of visual inertial initialization.
[0207] The multi-parameter step-by-step estimation method is not only fast and direct, but also simple in principle. It can be quickly applied to various ground mobile robot products and has the characteristics of strong adaptability, low cost, small code size, and fast implementation, making it suitable for a wide range of applications.
[0208] This application also provides a mobile device, such as... Figure 9 As shown, it includes a camera 901 with an inertial measurement unit (IMU) 902, a wheel speedometer 903, and a processor 904, wherein:
[0209] The camera 901 is used to acquire multiple image frames;
[0210] The IMU902 is used to collect the inertial information of the mobile device;
[0211] The wheel speed sensor 903 is used to collect wheel speed information of the mobile device;
[0212] The processor 904 is configured to acquire wheel speed measurement information, inertial information, and multiple image frames captured by the camera of the mobile device; when the image frames meet the visual initialization conditions, perform visual pose estimation based on the image frames to obtain an initial pose; when the keyframes included in the image frames meet the inertial initialization conditions, perform inertial pre-integration based on default parameters to obtain an inertial pre-integration result; calculate scale parameters and gravity parameters based on the pre-integration result of the wheel speed measurement information and the pose corresponding to the keyframes; determine the inertial constraint bias using the scale parameters and the gravity parameters as prior constraints, and update the inertial pre-integration result based on the inertial constraint bias and the inertial information; and perform joint optimization based on the updated inertial pre-integration result, the scale parameters, and the initial pose to determine the pose of the camera.
[0213] As can be seen, in the solution provided in this application embodiment, the processor can acquire wheel speedometer information, inertial information, and multiple image frames captured by the camera of the mobile device. When the image frames meet the visual initialization conditions, visual pose estimation is performed based on the image frames to obtain the initial pose. When the key frames included in the image frames meet the inertial initialization conditions, inertial pre-integration is performed based on default parameters to obtain the inertial pre-integration result. Based on the pre-integration result of the wheel speedometer information and the pose corresponding to the key frames, the scale parameters and gravity parameters are calculated. The scale parameters and gravity parameters are used as prior constraints to determine the inertial constraint bias. Based on the inertial constraint bias and inertial information, the inertial pre-integration result is updated. Based on the updated inertial pre-integration result, the scale parameters, and the initial pose, joint optimization is performed to determine the camera pose. Since visual inertial initialization can be performed on mobile devices, scale and gravity parameters can be calculated first. Then, the scale and gravity parameters can be used as prior constraints to optimize the remaining parameters of the IMU to update the inertial pre-integration results. Based on the updated pre-integration results, scale parameters, and initial pose, joint optimization can be performed to determine the camera pose. This eliminates the need to add all IMU parameters to the objective function for joint optimization, reducing the time required for joint optimization and allowing for correction of the camera pose, thereby improving the efficiency and accuracy of visual inertial initialization.
[0214] As one embodiment of this application, the processor 904 described above can be specifically used to acquire wheel speed information, inertial information, and multiple image frames captured by a camera of the mobile device during curved motion; pre-integrate the wheel speed information to obtain the wheel speed pose; calculate the scale prior and scale parameter deviation based on the wheel speed pose and the pose corresponding to the key frame; and calculate the gravity prior and gravity parameter deviation based on the first normal vector and the second normal vector, wherein the first normal vector is the normal vector of the plane formed by the wheel speed pose, and the second normal vector is the normal vector of the plane formed by the pose corresponding to the key frame.
[0215] As one embodiment of this application, the processor 904 described above can be specifically used to calculate the scale prior s according to the following formula based on the first position included in the first round of velocimetry pose, the second position included in the second round of velocimetry pose, the initial position included in the pose corresponding to the initial keyframe, the third position included in the pose corresponding to the first keyframe, and the fourth position included in the pose corresponding to the second keyframe. pre and scale parameter deviation s rc :
[0216]
[0217]
[0218]
[0219]
[0220]
[0221]
[0222] Where (x1, y1) are the coordinates of the first position included in the first wheel velocity sensor pose, and (x2, y2) are the coordinates of the second position included in the second wheel velocity sensor pose. k0 y k0 (x) represents the coordinates of the initial position, including the pose corresponding to the initial keyframe. k1 y k1 (x) represents the coordinates of the third position included in the pose corresponding to the first keyframe. k2 y k2 () represents the coordinates of the fourth position included in the pose corresponding to the second keyframe. These are the compensation coefficients for the wheel velocity meter pose in the x and y directions, respectively, and T. rc This is the extrinsic parameter from the camera coordinate system to the mobile device coordinate system.
[0223] As one embodiment of this application, the processor 904 can specifically be used to calculate the normal vector of the plane formed by the wheel speed meter poses through singular value decomposition, as a first normal vector; and to calculate the normal vector of the plane formed by the poses corresponding to the keyframes through singular value decomposition, as a second normal vector; and to calculate the gravity prior according to the following formula based on the first normal vector and the second normal vector. And the deviation of gravity parameter θ:
[0224]
[0225]
[0226] in, Let T be the second normal vector. rc This is the extrinsic parameter from the camera coordinate system to the mobile device coordinate system.
[0227] As one embodiment of this application, the above-mentioned inertial constraint bias includes linear acceleration bias and angular acceleration bias;
[0228] Specifically, the processor 904 described above can be used to perform iterative optimization of the linear acceleration bias and the angular acceleration bias, taking the scale parameter and the gravity parameter as prior constraints and minimizing the total error of the scale parameter, the gravity parameter, the linear acceleration bias, and the angular acceleration bias as the optimization direction, until the total error converges, and the iteratively optimized linear acceleration bias and angular acceleration bias are obtained.
[0229] As one embodiment of this application, the processor 904 described above can be used to perform monocular visual pose estimation based on the image frames when the number of image frames reaches a first preset number, so as to obtain an initial pose.
[0230] As one embodiment of this application, the processor 904 described above can be used to perform inertial pre-integration based on default parameters to obtain an inertial pre-integration result when the number of keyframes included in the image frame reaches a second preset number.
[0231] Furthermore, the aforementioned mobile device may also include a communication bus and / or a communication interface, with the processor 904, the communication interface, and the memory communicating with each other via the communication bus.
[0232] The communication bus mentioned in the aforementioned mobile devices can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not indicate that there is only one bus or one type of bus.
[0233] The communication interface is used for communication between the aforementioned mobile devices and other devices.
[0234] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.
[0235] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0236] In another embodiment provided in this application, a computer-readable storage medium is also provided, which stores a computer program that, when executed by a processor, implements the steps of any of the above-described visual inertial initialization methods.
[0237] In another embodiment provided in this application, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to execute any of the visual inertial initialization methods described above.
[0238] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a solid-state drive (SSD), etc.
[0239] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0240] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments for mobile devices, computer-readable storage media, and computer program products are basically similar to the method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0241] The above description is merely a preferred embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application are included within the scope of protection of this application.
Claims
1. A visual inertial initialization method, characterized in that, The method includes: Acquire wheel speed measurement information, inertial information, and multiple image frames captured by the camera from the mobile device; If the image frame meets the visual initialization conditions, visual pose estimation is performed based on the image frame to obtain the initial pose. If the keyframes included in the image frame satisfy the inertial initialization conditions, inertial pre-integration is performed based on the default parameters to obtain the inertial pre-integration result. The wheel speedometer information is pre-integrated to obtain the wheel speedometer pose; Based on the wheel velocity meter pose and the pose corresponding to the key frame, the scale prior and scale parameter deviation are calculated. For the plane formed by the wheel speed meter pose, the normal vector of the plane formed by the wheel speed meter pose is calculated by singular value decomposition and used as the first normal vector; For the plane formed by the poses corresponding to the keyframes, the normal vector of the plane formed by the poses corresponding to the keyframes is calculated by singular value decomposition and used as the second normal vector. Based on the first normal vector and the second normal vector, the prior gravity is calculated according to the following formula. and gravity parameter deviation : ; ; in, This is the second normal vector. The external parameter is the coordinate system from the camera coordinate system to the mobile device coordinate system; wherein, the first normal vector is the normal vector of the plane formed by the wheel speed meter pose, and the second normal vector is the normal vector of the plane formed by the poses corresponding to the key frames; Using the scale parameter and the gravity parameter as prior constraints, the inertial constraint bias is determined, and the inertial pre-integration result is updated based on the inertial constraint bias and the inertial information. The camera pose is determined by joint optimization based on the updated inertial pre-integration results, the scale parameters, and the initial pose.
2. The method according to claim 1, characterized in that, The steps of acquiring wheel speedometer information, inertial information, and multiple image frames captured by the camera from the mobile device include: During the curved motion of the mobile device, wheel speed meter information, inertial information, and multiple image frames captured by the camera are acquired.
3. The method according to claim 1, characterized in that, The step of calculating the scale prior and scale parameter deviation based on the wheel velocity meter pose and the pose corresponding to the keyframe includes: Based on the first position included in the first round of tachometric pose, the second position included in the second round of tachometric pose, the initial position included in the pose corresponding to the initial keyframe, the third position included in the pose corresponding to the first keyframe, and the fourth position included in the pose corresponding to the second keyframe, the scale prior is calculated according to the following formula. and scale parameter deviation : in, The coordinates of the first position included in the pose of the first wheel speedometer. The coordinates of the second position included in the pose of the second wheel speedometer. The initial keyframe corresponds to the coordinates of the initial position. The coordinates of the third position included in the pose corresponding to the first keyframe. The coordinates of the fourth position are included in the pose corresponding to the second keyframe. The wheel speed gauge position and pose are respectively , Compensation coefficients in both directions, This is the extrinsic parameter from the camera coordinate system to the mobile device coordinate system.
4. The method according to claim 1, characterized in that, The inertial constraint bias includes linear acceleration bias and angular acceleration bias; The step of determining the inertial constraint bias using the scale parameter and the gravity parameter as prior constraints includes: Using the scale parameter and the gravity parameter as prior constraints, and taking the minimization of the total error of the scale parameter, the gravity parameter, the linear acceleration bias, and the angular acceleration bias as the optimization direction, iterative optimization of the linear acceleration bias and the angular acceleration bias is performed. The process continues until the total error converges, at which point the iteratively optimized linear acceleration bias and angular acceleration bias are obtained.
5. The method according to any one of claims 1-4, characterized in that, The step of performing visual pose estimation based on the image frame to obtain the initial pose when the image frame meets the visual initialization conditions includes: When the number of image frames reaches a first preset number, monocular visual pose estimation is performed based on the image frames to obtain the initial pose.
6. The method according to any one of claims 1-4, characterized in that, The step of performing inertial pre-integration based on default parameters to obtain the inertial pre-integration result when the keyframes included in the image frame satisfy the inertial initialization conditions includes: When the number of keyframes included in the image frame reaches a second preset number, inertial pre-integration is performed based on default parameters to obtain the inertial pre-integration result.
7. A mobile device, characterized in that, It includes a camera with an inertial measurement unit (IMU), a wheel speedometer, and a processor, wherein: The camera is used to capture multiple image frames; The IMU is used to collect the inertial information of the mobile device; The wheel speed meter is used to collect wheel speed information of the mobile device; The processor is configured to acquire wheel speed sensor information, inertial information, and multiple image frames captured by the camera of the mobile device; when the image frames meet the visual initialization conditions, perform visual pose estimation based on the image frames to obtain an initial pose; when the keyframes included in the image frames meet the inertial initialization conditions, perform inertial pre-integration based on default parameters to obtain an inertial pre-integration result; perform pre-integration on the wheel speed sensor information to obtain the wheel speed sensor pose; calculate the scale prior and scale parameter deviation based on the wheel speed sensor pose and the poses corresponding to the keyframes; calculate the normal vector of the plane formed by the wheel speed sensor poses through singular value decomposition, as the first normal vector; calculate the normal vector of the plane formed by the poses corresponding to the keyframes through singular value decomposition, as the second normal vector; and calculate the gravity prior based on the first normal vector and the second normal vector according to the following formula. and gravity parameter deviation : ; ; in, This is the second normal vector. The external parameters are defined as follows: from the camera coordinate system to the mobile device coordinate system; wherein, the first normal vector is the normal vector of the plane formed by the wheel speedometer pose, and the second normal vector is the normal vector of the plane formed by the poses corresponding to the keyframes; using the scale parameter and the gravity parameter as prior constraints, the inertial constraint bias is determined, and the inertial pre-integration result is updated based on the inertial constraint bias and the inertial information; based on the updated inertial pre-integration result, the scale parameter, and the initial pose, joint optimization is performed to determine the camera pose.
8. The device according to claim 7, characterized in that, The processor is specifically used to acquire wheel speed information, inertial information, and multiple image frames captured by the camera of the mobile device during the curved motion of the mobile device.
9. The device according to claim 7, characterized in that, The processor is specifically configured to calculate the scale prior according to the following formula, based on the first position included in the first round of velocimetry pose, the second position included in the second round of velocimetry pose, the initial position included in the pose corresponding to the initial keyframe, the third position included in the pose corresponding to the first keyframe, and the fourth position included in the pose corresponding to the second keyframe. and scale parameter deviation : in, The coordinates of the first position included in the pose of the first wheel speedometer. The coordinates of the second position included in the pose of the second wheel speedometer. The initial keyframe corresponds to the coordinates of the initial position. The coordinates of the third position included in the pose corresponding to the first keyframe. The coordinates of the fourth position are included in the pose corresponding to the second keyframe. The wheel speed gauge position and pose are respectively , Compensation coefficients in both directions, This is the extrinsic parameter from the camera coordinate system to the mobile device coordinate system.
10. The device according to claim 7, characterized in that, The inertial constraint bias includes linear acceleration bias and angular acceleration bias; The processor is specifically configured to use the scale parameter and the gravity parameter as prior constraints, and take minimizing the total error of the scale parameter, the gravity parameter, the linear acceleration bias, and the angular acceleration bias as the optimization direction, to perform iterative optimization of the linear acceleration bias and the angular acceleration bias; until the total error converges, the iteratively optimized linear acceleration bias and angular acceleration bias are obtained.
11. The device according to any one of claims 7-10, characterized in that, The processor is specifically used to perform monocular visual pose estimation based on the image frames when the number of image frames reaches a first preset number, so as to obtain an initial pose.
12. The device according to any one of claims 7-10, characterized in that, The processor is specifically used to perform inertial pre-integration based on default parameters to obtain an inertial pre-integration result when the number of keyframes included in the image frame reaches a second preset number.
13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method described in any one of claims 1-6.
Citation Information
Patent Citations
Monocular vision inertial positioning method for automatic driving in closed park
CN113436261A
Visual inertial odometer initialization method and device, equipment and storage medium
CN113670327A