Positioning initialization method, apparatus, and computer-readable storage medium
By acquiring prior information and setting initial values for visual and inertial state variables when the positioning device is stationary, and combining inertial complementary filters and visual information, the problem of initialization complexity in visual-inertial tracking and positioning systems is solved, and a fast and stable initialization process is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ZHEJIANG SENSETIME TECH DEV CO LTD
- Filing Date
- 2021-11-16
- Publication Date
- 2026-05-19
AI Technical Summary
Existing visual-inertial tracking and positioning systems have complex initialization methods that require professional guidance from users. They cannot be initialized when the device is stationary or rotating, and the success rate of visual initialization and inertial sensor initialization is low, especially in environments with limited texture or large-scale outdoor environments.
The system acquires prior information about the state variables when the positioning device is stationary, sets initial values for the visual and inertial state variables, and determines convergence by updating them during operation. It also initializes the system using an inertial complementary filter and visual information.
It enables rapid and stable initialization of vision and inertial systems, simplifies user operations, and improves the initialization success rate, especially in stationary or near-stationary states.
Smart Images

Figure CN114022556B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of positioning technology, and in particular to a positioning initialization method, apparatus and computer-readable storage medium. Background Technology
[0002] Visual inertial tracking and positioning systems are an important underlying technology in fields such as computer vision, robotics, autonomous vehicles, 3D reconstruction, and augmented reality. Rapid initialization methods for visual inertial tracking and positioning systems have significant practical value in fields such as augmented reality and virtual reality.
[0003] However, the initialization methods for typical visual-inertial tracking and positioning systems are quite complex, requiring users to receive specialized technical guidance for proper operation. Generally, initialization schemes separate visual initialization from inertial sensor initialization. First, the vision system needs sufficient parallax to reconstruct the existing scene. After visual initialization, the results are aligned with the inertial sensor data to complete the inertial system initialization. This initialization process has several drawbacks: firstly, when the device is stationary or rotating, it cannot provide sufficient parallax, thus failing to complete visual initialization; secondly, due to the influence of visual initialization effectiveness and inertial sensor data, inertial system initialization still has a certain probability of failure; furthermore, the success rate of initialization is low in scenarios with limited texture or large-scale outdoor environments. Summary of the Invention
[0004] This application provides at least one positioning initialization method, apparatus, and computer-readable storage medium.
[0005] The first aspect of this application provides a positioning initialization method, the method comprising: acquiring prior information of state variables of a positioning device in a stationary state; determining an initial value of the state variables based on the prior information of the state variables; wherein the state variables are used for positioning and include visual state variables and inertial state variables; updating the state variables during positioning operation; and determining that positioning initialization is complete when the updated state variables satisfy convergence.
[0006] Therefore, by utilizing the prior information of the state variables obtained in a stationary state, initial values are set for the state variables, which are used for positioning and include visual state variables and inertial state variables. During the positioning process, if the updated state variables are found to be convergent, the positioning initialization is determined to be complete. This achieves the simultaneous initialization of the visual and inertial systems using prior information when the device is nearly stationary, making the positioning initialization process fast and stable, and easy for users to operate.
[0007] Wherein, the inertial state variables include at least one of gravity, position, angle, velocity, scale, and inertial system bias; and / or, the visual state variables include at least one of three-dimensional point inverse depth, position, and angle.
[0008] Therefore, since inertial state variables include at least one of gravity, position, angle, velocity, scale, and inertial system bias, and visual state variables include at least one of three-dimensional point inverse depth, position, and angle, the initial values of each state variable of the system can be set by utilizing the prior information of each state variable obtained in a stationary state, thereby achieving the initialization of the entire system.
[0009] The step of determining the initial value of the state variable based on the prior information of the state variable includes: using the prior information of the inertial state variable as the initial value of the inertial state variable; and / or, the step of obtaining the prior information of the state variable of the positioning device in a stationary state includes at least one of the following: when the state variable includes velocity, determining the velocity as a stationary velocity value and using the stationary velocity value as the prior information of the inertial state variable; when the state variable includes an inertial system bias, obtaining the bias or a preset calibration value during the positioning process before this positioning initialization, and using the bias or the preset calibration value during the positioning process before this positioning initialization as the prior information of the inertial system bias; using an inertial complementary filter to obtain the current value of the inertial state variable of the positioning device in the stationary state, and using the obtained current value of the inertial state variable as the prior information of the inertial state variable.
[0010] Therefore, by using the prior information of the inertial state variables of the positioning device in a stationary state as the initial values of the inertial state variables, the accuracy of the system initialization process is improved.
[0011] Before acquiring prior information about the state variables of the positioning device in a stationary state, the method further includes: acquiring two consecutive frames of first scene images collected by the positioning device; and determining that the positioning device is in a stationary state if the disparity value between all feature points in the two consecutive frames of the first scene images is less than a first preset threshold.
[0012] Therefore, by acquiring two consecutive frames of the first scene image collected by the positioning device, and determining that the disparity value between all feature points in the two consecutive frames of the first scene image is less than the first preset threshold, it is determined that the positioning device is in a stationary state. In other words, it is possible to accurately determine whether the positioning device is in a stationary state, thereby providing a suitable initialization time for the initialization process.
[0013] The step of updating the state variable during the positioning process includes: obtaining the current value of the state variable during the positioning process; and obtaining the updated value of the state variable based on the current value of the state variable.
[0014] Therefore, after setting initial values for the state variables, during the positioning process until the system completes initialization, the updated values of the state variables can be obtained to accurately determine whether the updated state variables meet the convergence requirement, thus determining whether the positioning initialization is complete.
[0015] Before updating the state variables, the method further includes decoupling the updates of several state variables.
[0016] Therefore, by decoupling the updates of several state variables, the mutual influence between them can be reduced. Thus, during the operation and positioning process after setting initial values for the state variables until the system completes initialization, the convergence speed of individual variables can be accelerated, thereby enabling the system to quickly and stably complete initialization.
[0017] The decoupling process for updating the plurality of state variables includes: determining the initial value confidence level of each state variable, wherein the initial confidence level is used to represent the degree of influence of the initial value on the updated state variable; and obtaining the updated value of the state variable based on the current value of the state variable includes: obtaining the updated value of the state variable based on the initial value, the initial value confidence level, and the current value of the state variable.
[0018] Therefore, by setting an initial confidence level for each state variable, where the initial confidence level represents the degree of influence of the initial value on the updated state variable, the mutual influence between the state variables is reduced, and each state variable has independent observability. This allows the system to converge toward the optimal value with smaller fluctuations during the convergence process, thereby improving the stability of static initialization and the system convergence speed.
[0019] The step of obtaining the current value of the state variable during the positioning process includes at least one of the following steps: when the state variable includes scale information, the positioning device is motion initialized during the positioning process, and the current value of the scale information is obtained using the scale information obtained after the motion initialization is completed; when the state variable includes three-dimensional point inverse depth, a second scene image collected by the positioning device is acquired during the positioning process, and a random inverse depth value of the corresponding three-dimensional point is determined along the depth direction of the point for the two-dimensional point in the second scene image, which is used as the current value of the three-dimensional point inverse depth.
[0020] Therefore, during the positioning process, by using the scale information obtained after the positioning device has completed motion initialization as the current value of the scale information, the scale information can be quickly converged after the static initialization. In addition, by observing the inverse depth value of the inverse depth 3D point during the positioning process, the stability of the system can be improved.
[0021] Before determining the random inverse depth value of the corresponding inverse depth 3D point for the 2D point in the second scene image along the depth direction of the point, the method further includes: obtaining device orientation information corresponding to at least two frames of the second scene image; and removing outgoing points in the second scene image based on the 2D points in the at least two frames of the second scene image and the device orientation information.
[0022] Therefore, based on the two-dimensional points in at least two frames of the second scene image and the device orientation information, the out-of-field points in the second scene image can be removed, thereby optimizing the scene image and improving the stability of system initialization.
[0023] The step of updating the state variables during the positioning process further includes: when the state variables include speed, determining whether the positioning device is stationary during the positioning process; when the positioning device is stationary, using the stationary speed value as the current speed value; and obtaining an updated speed value based on the current speed value.
[0024] Therefore, during the positioning process, by determining that the positioning device is stationary, the stationary speed value is used as the current speed value, directly constraining the system's speed information, which can effectively suppress large offsets in the system and improve the stability of system initialization.
[0025] In the context of updating the state variable during the positioning process, the method further includes at least one of the following steps: obtaining the uncertainty corresponding to the updated value of the state variable, and determining that the state variable satisfies convergence if the uncertainty of the state variable is lower than a second preset threshold; and determining the pose of the positioning device using the updated value of the state variable.
[0026] Therefore, since the lower the uncertainty of the state variable, the better the convergence of the state variable, by obtaining the uncertainty corresponding to the updated value of the state variable and judging whether the uncertainty of the state variable is lower than the second preset threshold, it can be determined whether the state variable meets the convergence requirement, and thus determine whether the positioning initialization is completed. In addition, based on the visual state variable and the inertial state variable, the pose information of the positioning device can be obtained to realize the positioning of the device.
[0027] To address the aforementioned problems, a second aspect of this application provides a positioning initialization device, comprising: an acquisition module for acquiring prior information of state variables of a positioning device in a stationary state; a setting module for determining an initial value of the state variables based on the prior information of the state variables; wherein the state variables are used for positioning and include visual state variables and inertial state variables; an update module for updating the state variables during positioning operation; and a determination module for determining that positioning initialization is complete when the updated state variables satisfy convergence.
[0028] To address the aforementioned problems, a third aspect of this application provides a positioning initialization apparatus, comprising a memory and a processor coupled to each other, wherein the processor is configured to execute program instructions stored in the memory to implement the positioning initialization method described in the first aspect.
[0029] To address the aforementioned problems, a fourth aspect of this application provides a computer-readable storage medium having program instructions stored thereon, which, when executed by a processor, implement the positioning initialization method described in the first aspect.
[0030] The above scheme uses prior information about the state variables acquired in a stationary state to set initial values for the state variables. These state variables are used for positioning and include both visual and inertial state variables. During the positioning process, the positioning initialization is considered complete when the updated state variables meet the convergence requirement. This achieves the simultaneous initialization of the visual and inertial systems using prior information when the device is nearly stationary, making the positioning initialization process fast and stable, and easy for users to operate. Attached Figure Description
[0031] Figure 1 This is a flowchart illustrating an embodiment of the positioning initialization method of this application;
[0032] Figure 2 yes Figure 1 A flowchart illustrating an embodiment of step S13;
[0033] Figure 3 This is a flowchart illustrating another embodiment of the positioning initialization method of this application;
[0034] Figure 4 This is a flowchart illustrating an application scenario of the positioning initialization method of this application;
[0035] Figure 5 This is a schematic diagram of the framework of an embodiment of the positioning initialization device of this application;
[0036] Figure 6 This is a schematic diagram of the framework of another embodiment of the positioning initialization device of this application;
[0037] Figure 7 This is a schematic diagram of a framework of an embodiment of the computer-readable storage medium of this application. Detailed Implementation
[0038] The embodiments of this application will now be described in detail with reference to the accompanying drawings.
[0039] In the following description, specific details such as particular system architectures, interfaces, and technologies are presented for illustrative purposes rather than for limiting purposes, in order to provide a thorough understanding of this application.
[0040] In this paper, the terms "system" and "network" are often used interchangeably. The term "and / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. Additionally, the character " / " generally indicates that the preceding and following related objects have an "or" relationship. Furthermore, "many" in this paper means two or more.
[0041] Please see Figure 1 , Figure 1 This is a flowchart illustrating an embodiment of the positioning initialization method of this application. Specifically, it may include the following steps:
[0042] Step S11: Obtain prior information about the state variables of the positioning device in a stationary state.
[0043] Step S12: Determine the initial value of the state variable based on the prior information of the state variable. The state variable is used for positioning and includes visual state variables and inertial state variables.
[0044] Visual-inertial tracking (VIS) localization systems are algorithms that fuse camera and IMU (Inertial Measurement Unit) data to achieve SLAM (Simultaneous Localization and Mapping). Due to the nonlinearity of VIS, the performance of the sensors, whether based on filtering or graph optimization, heavily depends on the accuracy of the initial values. Poor initialization can not only reduce convergence speed but also lead to incorrect estimations; therefore, robust initialization methods are crucial. Generally, in the initialization process of a VIS, visual-only initialization is performed first to calculate the relative pose of the camera; then, the initialization parameters are solved by aligning with IMU pre-integration to complete the joint visual-inertial initialization process. However, in practice, it is difficult for VIS to obtain an accurate initial state. On the one hand, the scale information of the camera cannot be directly observed; on the other hand, non-zero acceleration motion is required to initialize the scale information, but this initialization method fails when the localization device is stationary.
[0045] The execution entity of the positioning initialization method in this application can be a positioning initialization device. For example, the positioning initialization method can be executed by a positioning device, a server, or other processing devices. The positioning device can be a mobile device such as a robot, unmanned vehicle, or drone, or a user equipment (UE), user terminal, cordless phone, personal digital assistant (PDA), handheld device, computing device, in-vehicle device, wearable device, etc. In some possible implementations, the positioning initialization method can be implemented by a processor calling computer-readable instructions stored in memory. The positioning initialization method is performed when the positioning device is stationary. Before initialization, prior information about the state variables of the positioning device in its stationary state needs to be obtained. For example, IMU data can be used as prior information about inertial state variables, or camera data can be used as prior information about visual state variables. For example, IMU data can be used as prior information because the biases of the gyroscope and accelerometer can be estimated, and the gravity vector can be observed. In this case, the IMU has virtually no errors other than some small measurement errors and random walks. Furthermore, since the time between two frames is very short, the sensor deviation can be considered constant. Therefore, by pre-integrating the IMU measurement data, prior estimation of motion can be performed before the arrival of the new frame. As another example, camera data can be used as prior information. This application initializes the positioning device when it is stationary. Therefore, the visual information of the previous frame can be used as prior information for the visual information of the current frame, and image matching can be performed. This ensures that the system initialization is very stable even when the motion is very slow or when the device is stationary. Therefore, the stationary state in this application can be a state where the positioning device moves very slowly or a state where the positioning device is completely stationary.
[0046] Specifically, inertial state variables include at least one of gravity, position, angle, velocity, scale, and inertial system bias; visual state variables include at least one of three-dimensional point inverse depth, position, and angle.
[0047] In one embodiment, step S12 specifically includes: using the prior information of the inertial state variable as the initial value of the inertial state variable. It is understood that since this application is used for initialization when the positioning device is stationary, and the inertial state variable remains unchanged when the positioning device is stationary, the prior information of the inertial state variable can be used as the initial value of the inertial state variable.
[0048] In one implementation scenario, step S11 includes: when the state variable includes velocity, determining that the velocity is a stationary velocity value, and using the stationary velocity value as prior information for the inertial state variable. It is understood that the velocity in a stationary state is the stationary velocity value. Since this embodiment is used for initialization when the positioning device is stationary, the velocity of the positioning device is the stationary velocity value. Using the stationary velocity value as prior information for the inertial state variable gives the velocity variable a high degree of confidence, making the system initialization process more accurate.
[0049] In one implementation scenario, step S11 includes: when the state variable includes an inertial system bias, obtaining the bias or a preset calibration value from the positioning process prior to this positioning initialization, and using the bias or the preset calibration value as prior information for the inertial system bias. It is understood that IMUs all have a certain bias, which is generally related to temperature, factory accuracy, etc. When the positioning device is used for the first time, its inertial system bias can be a preset calibration value. Alternatively, when re-initializing the positioning system during use, since tracking and positioning have already been performed before this positioning initialization, the bias from the positioning process prior to this positioning initialization can also be used as the inertial system bias for this positioning initialization.
[0050] In one implementation scenario, step S11 includes: using an inertial complementary filter to obtain the current value of the inertial state variable of the positioning device when it is in the stationary state, and using the obtained current value of the inertial state variable as the prior information of the inertial state variable. It is understood that in an IMU, the accelerometer has better low-frequency characteristics because the angle of acceleration can be directly calculated without accumulated error, so it remains relatively accurate over a long period. However, the gyroscope, due to the accumulation of integration error over a long period, will result in a larger output error, or even become unusable. Therefore, using an inertial complementary filter, the gyroscope is used primarily for short periods, while the accelerometer is more accurate over long periods. At this time, the weight of the accelerometer is increased, different weights are given to the gyroscope and accelerometer, and a weighted sum is obtained to obtain the real-time value of the inertial state variable. When the real-time value of the inertial state variable obtained by the inertial complementary filter converges, and the positioning device is in a stationary state, the current value of the inertial state variable at this time can be obtained, and this current value of the inertial state variable can be used as the prior information of the inertial state variable.
[0051] Understandably, when the positioning device itself can calculate its gravity direction, and the calculated gravity direction is sufficiently reliable, the gravity direction calculated by the positioning device itself can be used to replace the gravity direction obtained by the aforementioned inertial complementary filter, as prior information about the gravity direction.
[0052] Therefore, by using the prior information of the inertial state variables of the positioning device in a stationary state, such as velocity and inertial system bias, as the initial values of the inertial state variables, the accuracy of the system initialization process can be improved.
[0053] Step S13: During the positioning process, update the state variables.
[0054] Step S14: If the updated state variables are found to be convergent, the positioning initialization is determined to be complete.
[0055] Understandably, from the time the state variables of the visual inertial tracking (VIS) positioning system are initially set until the system converges, the positioning device may be in a stationary or moving state; that is, during this period, the positioning device is in the process of positioning. After the initial values of the state variables of the VIS positioning system are set, each state variable updates its data in real time during the positioning process, and the convergence degree of the state variables will continuously change. Therefore, the state variables can be updated, and the positioning initialization is considered complete when the updated state variables meet the convergence requirement. It should be noted that the positioning initialization of the entire VIS positioning system is only considered complete when all state variables meet the convergence requirement.
[0056] The above scheme uses prior information about the state variables acquired in a stationary state to set initial values for the state variables. These state variables are used for positioning and include both visual and inertial state variables. During the positioning process, the positioning initialization is considered complete when the updated state variables meet the convergence requirement. This achieves the simultaneous initialization of the visual and inertial systems using prior information when the device is nearly stationary, making the positioning initialization process fast and stable, and easy for users to operate.
[0057] Please see Figure 2 , Figure 2 yes Figure 1 A flowchart illustrating an embodiment of step S13. In this embodiment, step S13 may specifically include the following steps:
[0058] Step S131: During the positioning process, obtain the current value of the state variable.
[0059] Step S132: Based on the current value of the state variable, obtain the updated value of the state variable.
[0060] Understandably, during the positioning process, for each state variable, there is a current value at every moment. The current value of the state variable is an observation. After obtaining the current value of the state variable, due to system cumulative error or observation error, the current value of the state variable may be inaccurate. Therefore, based on the obtained current value of the state variable and its historical values, a filter optimizer or nonlinear optimizer can be used to calculate the updated value of the state variable. The updated value of the state variable is more accurate than the current value of the state variable.
[0061] The above scheme, after setting initial values for the state variables, allows for accurate determination of whether the updated state variables meet convergence requirements during the positioning process until the system completes initialization.
[0062] In one implementation scenario, before updating the state variables in step S13 above, the location initialization method further includes: decoupling the updates of several state variables.
[0063] Understandably, before all state variables in a visual-inertial tracking and positioning system converge, these variables become coupled, such as scale, gravity direction, and inertial sensor bias. This coupling can cause the system to converge towards non-optimal values or experience significant fluctuations during convergence. Therefore, decoupling the updates of several state variables can reduce their mutual influence. This allows for faster convergence of individual variables during the positioning process after initializing them and before system initialization, leading to quicker and more stable system initialization.
[0064] In one implementation scenario, the above-described decoupling process for updating several state variables includes: determining the initial value confidence level for each state variable, wherein the initial confidence level is used to represent the degree of influence of the initial value on the updated state variable. In this case, step S132 includes:
[0065] Step S1321: Based on the initial value, the confidence level of the initial value, and the current value of the state variable, obtain the updated value of the state variable.
[0066] In practical applications, initial confidence levels can be set for each state variable based on experience. The initial confidence level represents the weight given to trusting the initial values or the current values during subsequent tracking observations after setting the initial values for the state variables of the visual-inertial tracking and positioning system. During the positioning process, a filtering optimizer or nonlinear optimizer can calculate the updated value at each moment based on the initial values, initial confidence levels, and observed values of the state variables. Therefore, for the current moment, the values and confidence levels of the state variables from the previous moment have been obtained—that is, the historical values and corresponding confidence levels of the state variables have been acquired. After obtaining the current value of the state variables, the filtering optimizer or nonlinear optimizer can calculate the updated value of the state variables at the current moment. Thus, after obtaining the updated value of the state variables, the confidence level of the updated value can be obtained, thereby reducing the impact of accumulated system errors or observation errors.
[0067] The above scheme reduces the mutual influence between state variables by setting an initial confidence level for each state variable. The initial confidence level is used to represent the degree of influence of the initial value on the updated state variable. This makes each state variable individually observable, thereby enabling the system to converge toward the optimal value with smaller fluctuations during the convergence process. This improves the stability of static initialization and the convergence speed of the system.
[0068] In one implementation scenario, step S13 may include: when the state variable includes scale information, during the positioning process, performing motion initialization on the positioning device, and using the scale information obtained after the motion initialization, obtaining the current value of the scale information. It is understood that since pure vision positioning systems cannot solve the scale information problem, and IMUs can estimate the scale information through their own inertial measurements, scale information is obtained by adding an IMU. Existing motion initialization methods can obtain the scale information of the visual inertial tracking positioning system relatively accurately. Therefore, in the positioning initialization process of this application, combined with existing motion initialization techniques, the motion initialization method can be run simultaneously in the background. After the motion initialization stabilizes, the system information from the motion initialization, such as relatively accurate scale information, is applied to the current visual inertial tracking positioning system, thereby facilitating rapid convergence of scale information during the positioning initialization process of this application.
[0069] In one implementation scenario, step S13 may include: when the state variable includes the inverse depth of a three-dimensional point, during the positioning process, acquiring a second scene image collected by the positioning device, and determining a random inverse depth value of the corresponding inverse depth three-dimensional point for the two-dimensional point in the second scene image along the depth direction of the point, so as to use it as the current value of the inverse depth of the three-dimensional point.
[0070] Understandably, when the positioning device is stationary, the 3D information of the scene cannot be correctly obtained. Without this 3D information, the visual-inertial tracking positioning system lacks constraints and is prone to deviations. Therefore, for feature points in the scene image, a random inverse depth value can be assigned along the depth direction of the point, forming an inverse depth 3D point. Thus, during the positioning process, the current value of the inverse depth of the 3D point can be obtained at any time. When the positioning device moves normally, the inverse depth value converges in the correct direction. When the positioning device does not move or moves only slightly, the inverse depth value can serve as prior information to constrain the system and prevent large deviations, thereby improving the stability of the visual-inertial tracking positioning system during initialization in scenes with unknown depth. Furthermore, when the positioning device jitters rapidly, the integration of discontinuous information in the inertial system will accumulate errors quickly. Inverse depth 3D points, as 3D information of the environment, can simultaneously constrain the current frame information and historical frame information, thereby reducing the accumulation error. For example, inverse depth 3D points can be converted into global Euclidean space 3D points, which can decouple the coordinates of the points from the pose information of the positioning device, thereby further reducing error accumulation and more effectively improving the stability of the positioning system in scenarios with small amplitude and rapid jitter.
[0071] In one implementation scenario, if the positioning device includes a depth sensor, the random inverse depth value of the aforementioned inverse depth 3D point can be replaced by the depth value obtained by the depth sensor.
[0072] Furthermore, before determining the random inverse depth value of the corresponding inverse depth 3D point for the 2D point in the second scene image along the depth direction of the point, the positioning initialization method further includes: obtaining device orientation information corresponding to at least two frames of the second scene image; and removing external points in the second scene image based on the 2D points in the at least two frames of the second scene image and the device orientation information.
[0073] Understandably, to obtain a more accurate scene image and ensure accurate scene image information during system positioning initialization, a two-point random sample consistency detection method can be used to remove obvious outliers in the scene image, given the device orientation information from at least two frames of scene images. The device orientation information can be calculated using an inertial complementary filter. Therefore, based on the two-dimensional points in at least two frames of scene images and the device orientation information, outliers in the scene image can be removed, thereby optimizing the scene image and improving the stability of positioning initialization in dynamic scenes.
[0074] In one implementation scenario, step S13 may further include: when the state variable includes speed, determining whether the positioning device is stationary during the positioning process; when the positioning device is stationary, using the stationary speed value as the current speed value; and obtaining an updated speed value based on the current speed value. Therefore, during the positioning process, by determining that the positioning device is stationary and using the stationary speed value as the current speed value, and then using a filter optimizer or nonlinear optimizer to calculate the updated speed value, the system's speed information can be directly constrained when the positioning device is stationary. This effectively suppresses large offsets in the system and improves the stability of system initialization.
[0075] In one implementation scenario, after step S13 above, the positioning initialization method may further include: obtaining the uncertainty corresponding to the updated value of the state variable, and determining that the state variable satisfies convergence if the uncertainty of the state variable is lower than a second preset threshold. It is understood that during the positioning process, while maintaining a state variable, the correlation between this state variable and other state variables is also maintained, i.e., the covariance matrix of the state variable. Information entropy is obtained based on the covariance matrix and the state variable information. The smaller the information entropy, the lower the uncertainty of the state variable, and the lower the uncertainty, the better the convergence of the state variable. Therefore, by obtaining the uncertainty corresponding to the updated value of the state variable and determining whether the uncertainty of the state variable is lower than the second preset threshold, it can be determined whether the state variable satisfies convergence, and thus whether the positioning initialization is complete.
[0076] In one implementation scenario, after step S13 above, the positioning initialization method may further include: determining the pose of the positioning device using the updated values of the state variables. It is understood that the pose information of the positioning device can be obtained based on the visual state variables and the inertial state variables, thereby enabling device positioning.
[0077] Please see Figure 3 , Figure 3 This is a flowchart illustrating another embodiment of the positioning initialization method of this application. Specifically, it may include the following steps:
[0078] Step S31: Obtain two consecutive frames of the first scene image captured by the positioning device.
[0079] Step S32: If the disparity value between all feature points in the first scene images of the two consecutive frames is less than the first preset threshold, then the positioning device is determined to be in a stationary state.
[0080] In this embodiment, visual information can be used to determine whether the positioning device is stationary. Specifically, by acquiring two consecutive frames of the first scene image collected by the positioning device, if the disparity value between all feature points in the two consecutive frames of the first scene image is less than a first preset threshold, it is determined that the positioning device is stationary. This allows for accurate determination of whether the positioning device is stationary, thereby providing a suitable initialization time for the initialization process.
[0081] Step S33: Obtain prior information about the state variables of the positioning device in a stationary state.
[0082] Step S34: Determine the initial value of the state variable based on the prior information of the state variable; wherein the state variable is used for positioning and includes visual state variables and inertial state variables.
[0083] Step S35: During the positioning process, update the state variables.
[0084] Step S36: If the updated state variables are determined to be convergent, the positioning initialization is determined to be complete.
[0085] In this embodiment, steps S33-S36 are basically similar to steps S11-S14 in the above embodiments of this application, and will not be described again here.
[0086] Please combine Figure 4 , Figure 4This is a flowchart illustrating an application scenario of the positioning initialization method of this application. In this application scenario, the positioning initialization method for the positioning device is divided into three stages: an initialization preparation stage, an initialization stage, and an initialization completion stage. The initialization preparation stage primarily aims to obtain prior information for positioning initialization and select an appropriate initialization time. For example, an inertial complementary filter can be used to estimate the direction of gravity as prior information. Then, after the inertial filter converges, visual information is used to determine whether the positioning device is stationary. If it is confirmed that the positioning device is stationary, the initialization stage can proceed. The initialization phase primarily involves system initialization, which utilizes prior information to set the initial values of each state variable. Simultaneously, the confidence levels of these state variables are set to decouple them. For example, setting the initial values includes: setting the gravity vector (using the gravity direction estimated during the initialization preparation phase); setting the velocity value (initializing the velocity at zero velocity); and setting the inertial system bias (using saved values from the previous stable system run or preset calibration values). Before system convergence, state variables such as scale, gravity direction, and inertial sensor biases are coupled together, potentially causing the system to converge towards non-optimal values or exhibiting significant fluctuations during convergence. Therefore, adding confidence levels to the initial values of each state variable reduces their mutual influence. A higher confidence level for a particular state variable accelerates its convergence.During the initialization phase, various strategies are employed to enhance system stability from initialization to system convergence. For example, out-of-frame removal is performed. Specifically, from the initialization of the system's state variables until system convergence, to obtain accurate scene images and ensure accurate scene image information during system positioning initialization, a two-point random sample consistency detection method can be used to remove obvious out-of-frame points, given the device orientation information of two scene images. Another strategy is adding and maintaining inverse depth 3D points. Specifically, for feature points in the scene image, a random inverse depth value can be assigned along the depth direction of the point, forming an inverse depth 3D point. When the positioning device moves normally, the values can guide the system to converge in the correct direction. When the positioning device does not move or moves only slightly, they can serve as prior information to constrain the system from producing large deviations. When the positioning device jitters rapidly, the inverse depth 3D points, as 3D information of the environment, can simultaneously constrain the current frame information and historical frame information, thereby reducing accumulated errors. For example, zero-velocity updates can be performed. Specifically, before the system converges, by determining that the positioning device is stationary, the stationary velocity value can be used as the velocity update value to directly constrain the system's velocity information, which can effectively suppress large offsets and improve the stability of system initialization. Finally, based on the various state variables, the pose information of the positioning device can be output, thereby achieving device positioning.
[0087] Please see Figure 5 , Figure 5 This is a schematic diagram of a framework of an embodiment of the positioning initialization device of this application. The positioning initialization device 50 includes: an acquisition module 500, used to acquire prior information of the state variables of the positioning device in a stationary state; a setting module 502, used to determine the initial value of the state variables based on the prior information of the state variables; wherein the state variables are used for positioning and include visual state variables and inertial state variables; an update module 504, used to update the state variables during the positioning process; and a determination module 506, used to determine that the positioning initialization is complete when the updated state variables satisfy convergence.
[0088] In the above scheme, the setting module 502 uses the prior information of the state variables acquired by the acquisition module 500 in a stationary state to set initial values for the state variables. The state variables are used for positioning and include visual state variables and inertial state variables. During the positioning process, the determining module 506 determines that the positioning initialization is complete when the updated state variables meet the convergence requirement. This realizes the simultaneous initialization of the visual and inertial systems using prior information when the device is nearly stationary, making the positioning initialization process fast and stable, and easy for users to operate.
[0089] In some embodiments, the setting module 502 is specifically used to use the prior information of the inertial state variable as the initial value of the inertial state variable. In this case, the acquisition module 500 can be used to: determine that the velocity is a stationary velocity value when the state variable includes velocity; and / or, when the state variable includes an inertial system bias, acquire the bias or a preset calibration value from the positioning process prior to this positioning initialization, and use the bias or the preset calibration value from the positioning process prior to this positioning initialization as prior information of the inertial system bias; and / or, use an inertial complementary filter to acquire the current value of the inertial state variable of the positioning device in the stationary state, and use the acquired current value of the inertial state variable as prior information of the inertial state variable.
[0090] In some embodiments, the positioning initialization device 50 further includes: a judgment module, configured to acquire two consecutive frames of first scene images collected by the positioning device; and determine that the positioning device is in a stationary state if the disparity value between all feature points in the two consecutive frames of the first scene images is less than a first preset threshold.
[0091] In some embodiments, the update module 504 is specifically used to obtain the current value of the state variable during the positioning process; and to obtain the updated value of the state variable based on the current value of the state variable.
[0092] In some embodiments, the setting module 502 is further configured to decouple the updates of the plurality of state variables; specifically, the setting module 502 is configured to determine the initial value confidence level of each state variable, wherein the initial confidence level is used to represent the degree of influence of the initial value on the updated state variable. At this time, the update module 504 is configured to obtain the updated value of the state variable based on the initial value, the initial value confidence level, and the current value of the state variable.
[0093] In some embodiments, the update module 504 performs the step of obtaining the current value of the state variable during the positioning process, which may specifically include at least the following steps: when the state variable includes scale information, during the positioning process, the positioning device is motion initialized, and the current value of the scale information is obtained using the scale information obtained after the motion initialization is completed; when the state variable includes three-dimensional point inverse depth, during the positioning process, a second scene image collected by the positioning device is acquired, and a random inverse depth value of the corresponding three-dimensional point is determined along the depth direction of the point for the two-dimensional point in the second scene image, as the current value of the three-dimensional point inverse depth.
[0094] In some embodiments, the update module 504 is further configured to obtain device orientation information corresponding to at least two frames of the second scene image; and to remove outgoing points in the second scene image based on the two-dimensional points in the at least two frames of the second scene image and the device orientation information.
[0095] In some embodiments, the update module 504 is further configured to determine whether the positioning device is stationary during the positioning process when the state variable includes speed; and to use the stationary speed value as the updated value of the speed when the positioning device is stationary.
[0096] In some embodiments, the determining module 506 may also be used to obtain the uncertainty corresponding to the updated value of the state variable, and determine that the state variable satisfies convergence if the uncertainty of the state variable is lower than a second preset threshold; and / or, determine the pose of the positioning device using the updated value of the state variable.
[0097] Please see Figure 6 , Figure 6 This is a schematic diagram of another embodiment of the positioning initialization device of this application. The positioning initialization device 60 includes a memory 61 and a processor 62 coupled to each other. The processor 62 is used to execute program instructions stored in the memory 61 to implement the steps of any of the above-described positioning initialization method embodiments. In a specific implementation scenario, the positioning initialization device 60 may include, but is not limited to, a microcomputer or a server.
[0098] Specifically, processor 62 controls itself and memory 61 to implement the steps in any of the above-described positioning initialization method embodiments. Processor 62 can also be referred to as a CPU (Central Processing Unit). Processor 62 may be an integrated circuit chip with signal processing capabilities. Processor 62 can also be a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor can be a microprocessor or any conventional processor. Furthermore, processor 62 can be implemented using integrated circuit chips.
[0099] The above scheme uses prior information about the state variables acquired in a stationary state to set initial values for the state variables. These state variables are used for positioning and include both visual and inertial state variables. During the positioning process, the positioning initialization is considered complete when the updated state variables meet the convergence requirement. This achieves the simultaneous initialization of the visual and inertial systems using prior information when the device is nearly stationary, making the positioning initialization process fast and stable, and easy for users to operate.
[0100] Please see Figure 7 , Figure 7 This is a schematic diagram of a framework of an embodiment of the computer-readable storage medium of this application. The computer-readable storage medium 70 stores program instructions 700 that can be executed by a processor. The program instructions 700 are used to implement the steps in any of the above-described positioning initialization method embodiments.
[0101] This disclosure relates to the field of augmented reality (AR). It involves acquiring image information of target objects in a real-world environment and then using various visual algorithms to detect or identify the relevant features, states, and attributes of these objects, thereby achieving an AR effect that combines virtual and real elements to suit specific applications. For example, target objects may include human features such as faces, limbs, gestures, and movements; objects such as signs and markers; or venues such as sand tables, display areas, or displayed items. Visual algorithms may include visual localization, SLAM, 3D reconstruction, image registration, background segmentation, keypoint extraction and tracking of objects, and pose or depth detection. Specific applications can include interactive scenarios related to real-world scenes or objects, such as guided tours, navigation, explanations, reconstruction, and virtual effect overlay displays, as well as human-related special effects processing, such as makeup enhancement, body enhancement, special effects displays, and virtual model displays.
[0102] Convolutional neural networks (CNNs) can be used to detect or identify the relevant features, states, and attributes of target objects. The aforementioned CNNs are network models obtained through training using deep learning frameworks.
[0103] In the embodiments provided in this application, it should be understood that the disclosed methods and apparatus can be implemented in other ways. For example, the apparatus implementations described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces, and the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms.
[0104] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.
[0105] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0106] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
Claims
1. A positioning initialization method, characterized in that, The method includes: Obtain prior information about the state variables of the positioning device when it is stationary; Based on the prior information of the state variables, the initial values of the state variables are determined; wherein, the state variables are used for positioning and include visual state variables and inertial state variables; During the positioning process, the state variables are updated; If the updated state variables are found to converge, the positioning initialization is considered complete. Wherein, determining the initial value of the state variable based on the prior information of the state variable includes: The prior information of the inertial state variable is used as the initial value of the inertial state variable.
2. The method according to claim 1, characterized in that, The inertial state variables include at least one of gravity, position, angle, velocity, scale, and inertial system bias. And / or, the visual state variables include at least one of three-dimensional point inverse depth, position, and angle.
3. The method according to claim 1 or 2, characterized in that, The acquisition of prior information on the state variables of the positioning device in a stationary state includes at least one of the following: When the state variable includes velocity, the velocity is determined to be a stationary velocity value, and the stationary velocity value is used as prior information of the inertial state variable. When the state variable includes the inertial system bias, the bias or preset calibration value obtained during the positioning process before this positioning initialization is obtained, and the bias or preset calibration value obtained during the positioning process before this positioning initialization is used as the prior information of the inertial system bias. The current value of the inertial state variable of the positioning device in the stationary state is obtained by using an inertial complementary filter, and the obtained current value of the inertial state variable is used as the prior information of the inertial state variable.
4. The method according to any one of claims 1 to 2, characterized in that, Before acquiring prior information about the state variables of the positioning device in a stationary state, the method further includes: Acquire two consecutive frames of the first scene image captured by the positioning device; If the disparity value between all feature points in the first scene images of the two consecutive frames is less than a first preset threshold, then the positioning device is determined to be in a stationary state.
5. The method according to any one of claims 1 to 2, characterized in that, The updating of the state variables during the positioning process includes: During the positioning process, the current value of the state variable is obtained; Based on the current value of the state variable, the updated value of the state variable is obtained.
6. The method according to claim 5, characterized in that, Before updating the state variable, the method further includes: The updates of several of the aforementioned state variables are decoupled.
7. The method according to claim 6, characterized in that, The decoupling process for updating the aforementioned state variables includes: Determine the initial value confidence level for each of the state variables, wherein the initial value confidence level is used to represent the degree of influence of the initial value on the updated state variable; The step of obtaining the updated value of the state variable based on its current value includes: The updated value of the state variable is obtained based on the initial value, the confidence level of the initial value, and the current value.
8. The method according to claim 5, characterized in that, Obtaining the current value of the state variable during the positioning process includes at least one of the following steps: When the state variable includes scale information, during the positioning process, the positioning device is motion initialized, and the current value of the scale information is obtained using the scale information obtained after the motion initialization is completed. When the state variable includes the inverse depth of a three-dimensional point, during the positioning process, a second scene image collected by the positioning device is acquired, and a random inverse depth value of the corresponding three-dimensional point is determined for the two-dimensional point in the second scene image along the depth direction of the point, so as to serve as the current value of the inverse depth of the three-dimensional point.
9. The method according to claim 8, characterized in that, Before determining the random inverse depth value of the corresponding inverse depth 3D point for the 2D point in the second scene image along the depth direction of the point, the method further includes: Obtain device orientation information corresponding to at least two frames of the second scene image; Based on the two-dimensional points in the at least two frames of the second scene image and the device orientation information, the outgoing points in the second scene image are removed.
10. The method according to claim 5, characterized in that, The step of updating the state variables during the positioning process also includes: When the state variable includes speed, during the positioning process, it is determined whether the positioning device is in a stationary state. When the positioning device is stationary, the stationary speed value is taken as the current speed value; Based on the current value of the speed, an updated value of the speed is obtained.
11. The method according to any one of claims 1 to 2, characterized in that, After updating the state variables during the positioning process, the method further includes at least one of the following steps: Obtain the uncertainty corresponding to the updated value of the state variable, and determine that the state variable satisfies convergence if the uncertainty of the state variable is lower than a second preset threshold; The pose of the positioning device is determined using the updated values of the state variables.
12. A positioning initialization device, characterized in that, include: The acquisition module is used to acquire prior information about the state variables of the positioning device when it is stationary. The setting module is used to determine the initial value of the state variable based on the prior information of the state variable; wherein the state variable is used for positioning and includes visual state variables and inertial state variables; An update module is used to update the state variables during the positioning process. The determination module is used to determine that the positioning initialization is complete when the updated state variables satisfy convergence. Wherein, determining the initial value of the state variable based on the prior information of the state variable includes: The prior information of the inertial state variable is used as the initial value of the inertial state variable.
13. A positioning initialization device, characterized in that, It includes a memory and a processor coupled to each other, the processor being used to execute program instructions stored in the memory to implement the positioning initialization method according to any one of claims 1 to 11.
14. A computer-readable storage medium having program instructions stored thereon, characterized in that, When the program instructions are executed by the processor, they implement the positioning initialization method according to any one of claims 1 to 11.