Terminal initialization method and apparatus, electronic device, and storage medium

By constructing an information matrix and performing Schur complement elimination and Cholesky decomposition, the covariance is quickly estimated, solving the problem of unstable initialization in visual inertial positioning systems and achieving more efficient system initialization and stability.

CN120950131BActive Publication Date: 2026-05-12QINGDAO PICO TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
QINGDAO PICO TECH CO LTD
Filing Date
2025-07-31
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

During initialization, visual inertial positioning systems may experience instability for a period of time due to discrepancies between the initial covariance value and the device or environment, affecting system stability and accuracy.

Method used

By acquiring the terminal's initialization parameters, inertial constraints, and visual constraints, an information matrix is ​​constructed. Schur complement elimination and Cholesky decomposition are then performed to quickly estimate the covariance, avoiding direct inversion. The covariance is calculated based on the actual constraints.

Benefits of technology

It improves the stability and accuracy of the system after initialization, shortens the initialization time, increases the calculation speed by more than 2 times, and reduces the fluctuation range of state variables.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120950131B_ABST
    Figure CN120950131B_ABST
Patent Text Reader

Abstract

The present disclosure provides a terminal initialization method, device, electronic equipment and storage medium. The terminal initialization method comprises: obtaining initialization parameters, inertia constraints and visual constraints of a terminal; constructing an information matrix according to the initialization parameters, inertia constraints and visual constraints; performing Schur complement elimination on the information matrix to eliminate the feature points from the information matrix to obtain an eliminated information matrix; performing Cholesky decomposition on the eliminated information matrix to obtain an upper triangular matrix and a lower triangular matrix; inverting the image frame state quantity corresponding to the nearest preset time in the upper triangular matrix to obtain the square root of the covariance, and then determining the covariance; and determining the initialization state of the terminal according to the covariance and the initialization parameters. The method of the present disclosure can quickly solve the covariance and reduce the state quantity fluctuation amplitude.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and in particular to a terminal initialization method, apparatus, electronic device, and storage medium. Background Technology

[0002] Visual-inertial positioning systems (VIS systems) are widely used in drones, extended reality, and autonomous driving. To improve the practicality and robustness of VIS systems, online calibration is often incorporated into state estimation. This allows the system to estimate multiple states, including those of the camera and the intrinsic and extrinsic parameters of the inertial measurement unit (IMU). Based on the state estimation method, these can be categorized into optimization-based and filtering-based frameworks. In the filtering-based framework, the covariance of the system state variables needs to be maintained, thus incorporating prior information from past time points during the solution process, resulting in smoother state variable changes.

[0003] The initial value for the covariance is usually given as an empirical value, and then the positioning function is executed normally. However, when using an empirical value for the initial covariance, there may be situations where the empirical value does not match the device or environment. This can cause instability in the visual inertial positioning system for a period of time after initialization, that is, the mean and covariance of the system state variables will fluctuate violently, which is detrimental to the stability and accuracy of the system. Summary of the Invention

[0004] This disclosure provides a terminal initialization method, apparatus, electronic device, and storage medium.

[0005] The following technical solution is adopted in this disclosure.

[0006] In some embodiments, this disclosure provides a terminal initialization method, including:

[0007] The initialization parameters, inertial constraints, and visual constraints of the terminal are obtained. The initialization parameters include the terminal's position, attitude, velocity, inertial offset, and the positions of multiple preset feature points in the environment during the initial time period. The inertial constraints include the relationship between the terminal's position, attitude, and velocity at different preset times and the inertial data. The visual constraints include the feature points in the image frames captured by the terminal at preset times. The initial time period includes multiple preset times.

[0008] An information matrix is ​​constructed based on the initialization parameters, inertial constraints, and visual constraints.

[0009] The information matrix is ​​subjected to Schur complement elimination to remove the feature points from the information matrix, resulting in an eliminated information matrix.

[0010] The eliminated information matrix is ​​decomposed by Cholliski to obtain an upper triangular matrix and a lower triangular matrix.

[0011] The square root of the covariance is obtained by inverting the state variables of the image frame corresponding to the nearest preset time in the upper triangular matrix, and then the covariance is determined.

[0012] The initialization state of the terminal is determined based on the covariance and the initialization parameters.

[0013] In some embodiments, this disclosure provides a terminal initialization apparatus, comprising:

[0014] The acquisition unit is used to acquire the initialization parameters, inertial constraints, and visual constraints of the terminal. The initialization parameters include: the position, attitude, velocity, inertial offset of the terminal in the initial time period, and the positions of multiple preset feature points in the environment in the initial time period. The inertial constraints include: the relationship between the position, attitude, and velocity of the terminal at different preset times and the inertial data. The visual constraints include: the feature points in the image frames captured by the terminal at preset times. The initial time period includes multiple preset times.

[0015] The control unit is used to construct an information matrix based on the initialization parameters, inertial constraints, and visual constraints.

[0016] The control unit is used to perform Schul complement elimination on the information matrix to remove the feature points from the information matrix and obtain the eliminated information matrix.

[0017] The control unit is used to perform Cholliski decomposition on the eliminated information matrix to obtain an upper triangular matrix and a lower triangular matrix.

[0018] The control unit is used to invert the state quantity of the image frame corresponding to the nearest preset time in the upper triangular matrix to obtain the square root of the covariance, and then determine the covariance;

[0019] The control unit is configured to determine the initialization state of the terminal based on the covariance and the initialization parameters.

[0020] In some embodiments, this disclosure provides an electronic device, including: at least one memory and at least one processor;

[0021] The memory is used to store program code, and the processor is used to call the program code stored in the memory to execute the above method.

[0022] In some embodiments, this disclosure provides a computer-readable storage medium for storing program code that, when run by a processor, causes the processor to perform the methods described above.

[0023] The terminal initialization method provided in this disclosure can quickly solve for the covariance. By using Schur complement elimination and Cholleysky decomposition, direct inversion is avoided, increasing the solution speed by more than 2 times and shortening the terminal initialization time. The covariance is calculated based on actual constraints rather than empirical values, resulting in a significant reduction in the fluctuation amplitude of state variables after initialization. Attached Figure Description

[0024] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.

[0025] Figure 1 This is a flowchart of a terminal initialization method according to an embodiment of the present disclosure.

[0026] Figure 2 This is a schematic diagram of a factor graph according to an embodiment of this disclosure.

[0027] Figure 3 This is a schematic diagram of the information matrix according to an embodiment of this disclosure.

[0028] Figure 4 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present disclosure. Detailed Implementation

[0029] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0030] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.

[0031] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0032] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0033] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0034] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0035] It should be understood that the various steps described in the method embodiments of this disclosure can be performed in sequence and / or in parallel. Furthermore, method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0036] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.

[0037] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0038] It should be noted that the use of the word "a" in this disclosure is illustrative rather than restrictive, and those skilled in the art should understand that it should be understood as "one or more" unless otherwise expressly indicated in the context.

[0039] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0040] The solutions provided by the embodiments of this disclosure will be described in detail below with reference to the accompanying drawings.

[0041] For terminals employing visual inertial positioning systems, the covariance of state variables needs to be determined during initialization. If the initial value of the covariance does not match the device or environment, it can lead to instability in the visual inertial positioning system for a period of time after initialization. Therefore, some embodiments of this disclosure propose a terminal initialization method that can quickly and accurately estimate the covariance, effectively improving the stability of the system after initialization. Compared with conventional methods for calculating covariance, this method is more efficient in terms of calculation.

[0042] like Figure 1 As shown, Figure 1 This is a flowchart of a terminal initialization method according to an embodiment of the present disclosure, which includes the following steps.

[0043] S11. Obtain the terminal's initialization parameters, inertial constraints, and visual constraints.

[0044] In some embodiments, the terminal in this disclosure is a terminal with a visual inertial positioning system, and the terminal has a camera and an inertial measurement unit (IMU). In some embodiments, the terminal is a virtual reality device, an augmented reality device, or a mixed reality device. In some embodiments, during initialization, the terminal continuously captures multiple image frames and detects data in the inertial measurement unit within an initial time period, including acceleration data and angular velocity data. The time point at which the terminal captures image frames is a preset time, for example, 24 image frames can be captured per second. The environment in which the terminal is located has multiple preset feature points, which may be easily identifiable points such as corners or protruding points in the environment. The initialization parameters of the terminal include: the terminal's position, attitude, velocity, inertial bias, and the positions of multiple preset feature points in the environment during the initial time period. The initial time period can be a period of time from the start of terminal initialization to the following period, including multiple preset moments. After the initialization begins, a preset moment is set every preset duration. The terminal captures image frames at the preset moments, and the initialization parameters are detected based on the image frames and the terminal's built-in inertial measurement unit. More specifically, the initialization parameters can be the terminal's position, attitude, velocity, inertial bias, and the positions of multiple preset feature points in the environment at each preset moment. In some embodiments, the inertial bias includes: accelerometer bias and angular velocity bias. In some embodiments, inertial constraints include the relationship between the terminal's position, attitude, and velocity at different preset times and inertial data. Specifically, the change in the terminal's position, attitude, or velocity at one preset time and the position, attitude, or velocity at another preset time can be represented by inertial data. Inertial data is data detected by the inertial measurement unit and may include acceleration data and angular velocity data. For example, the terminal's velocity at one preset time can be represented as the velocity at another previous time plus the product of acceleration and the time difference. Inertial constraints identify different preset times. In some embodiments, visual constraints include the feature points in the image frames captured by the terminal at each preset time. The terminal may be in motion during the initial period, so the feature points that the terminal can capture at different preset times may not be exactly the same. The feature points that the terminal can capture also reflect the relative positional relationship between the terminal and the feature points.

[0045] S12. Construct an information matrix based on initialization parameters, inertial constraints, and visual constraints.

[0046] In some embodiments, during the initialization of the visual-inertial positioning system, bundle adjustment (BA) is performed to jointly optimize the image frame state variables by combining inertial and visual constraints. For example... Figure 2 As shown, Figure 2The factor diagram is used for illustrative purposes, where squares and circles represent state variables to be optimized, and arrows represent constraints between state variables. The `body` represents an image frame captured by the terminal at a preset time, and `body1`, `body2`, and `body3` are image frames at different times within the initial time period. The state variables carried by the image frame are called image frame state variables. In some embodiments, the image frame state variables at any preset time include: the terminal's position P, attitude Q, velocity V, accelerometer bias Ba, and angular velocity meter bias Bg at that preset time. The IMU pre-integration constraint, also known as the inertial constraint, represents the relationship between the differences between the image frame state variables at different preset times and the data detected by the inertial measurement unit. Arrows from `body` to feature points represent visual constraints, i.e., seeing a certain feature point in the image frame. The factor diagram helps clarify the relationship between initialization parameters, inertial constraints, and visual constraints. The information matrix (i.e., the Hessian matrix) is formed based on the above relationships and combined with the Gauss-Newton optimization method. A schematic diagram of the information matrix is ​​shown below. Figure 3 As shown, the information matrix is ​​the second-order derivative matrix in the visual-inertial localization problem, quantifying the constraint strength between all state variables. Its function is to transform inertial constraints (motion continuity) and visual constraints (feature point projection) into mathematical matrices, achieving state optimization and covariance estimation through matrix operations (elimination, decomposition). In some embodiments, such as... Figure 3 As shown, the information matrix includes: constraint blocks H representing the image frame state variables. xx (Top left part) represents the constraint block H from the image frame to the feature point. xp (Upper right part) Represents the constraint block H from feature points to image frames. px (Lower left part) and constraint block H representing feature points pp (Lower right part). Upper left part, constraint block H. xx The dimension is 15N×15N, where N is equal to the number of image frames. Figure 3 (where N is 6), constraint block H xx Each square on the diagonal from the top left to the bottom right corresponds to an image frame at a preset time. The first square on the top left corresponds to an image frame (i.e., the first body). The square in an image frame contains sub-squares representing five quantities: position P, attitude Q, velocity V, accelerometer bias Ba, and angular velocity bias Bg. Each sub-square has a dimension of 3×3 (representing three dimensions in the XYZ directions). Therefore, the dimension of the square in an image frame is actually 15×15. Figure 3 The first and third bodies in the image appear to correspond to 5x5 squares, but their actual dimensions are 15x15. Because there are N image frames, H... xx The dimension is 15N×15N). In constraint block H xxOnly the green area has a value, the rest are 0. The two adjacent squares on the diagonal, plus the squares adjacent to these two squares, represent the inertial constraints of adjacent image frames. For example, in constraint block H xx The squares corresponding to the first body, the second body, and the two adjacent squares located in the first row and second column and the second row and first column, together represent the inertial constraint between the first and second image frames. In the information matrix, H... xp (the region where JptT*Jx is located) and H px The region containing JxT*Jpt is a transpose, with dimensions of 15N×3M and 3M×15N respectively (15N because the dimension of an image frame is 15, and 3M because the dimension of a feature point is 3). H xp and H px This represents the relationship between image frame state variables and visual constraints on feature points, where constraint block H... xp The horizontal and vertical axes of H represent the feature point and the image frame, respectively. If a square has a value, it indicates an observation, meaning the feature point it points to can be seen in the image frame it points to. For example, if feature point f can be observed in image frame k, then the corresponding 15×3 square is non-zero, representing the "visual constraint between image frame k and feature point f". Therefore, H... xp and H px Which blocks have values ​​depends on the observations in the image frame. The lower right constraint block H... pp The dimension is 3M×3M, and the number of M is equal to the number of feature points. Figure 3 In the context of M=15, a sub-block of a feature point has a dimension of 3×3, representing the coordinate position constraints between feature points. The constraint block H... pp The squares on the diagonal represent the three-dimensional position coordinates of the feature points, and the constraint block H pp The squares not on the diagonal are 0.

[0047] S13. Perform Schur complement elimination on the information matrix to remove feature points from the information matrix and obtain the eliminated information matrix.

[0048] In some embodiments, the desired state variables are a subset of the total state variables. Due to the correlation between different state variables, unnecessary state variables need to be marginalized. Since the information matrix has sparsity (specifically, the sparsity of Hxx: except for the diagonal, only the inertial constraints of adjacent image frames are non-zero, since the inertial constraints only constrain adjacent image frames, and non-adjacent frames have no direct correlation; the sparsity of Hpp: only the diagonal is non-zero; the sparsity of Hxp / Hpx: only the corresponding squares of "image frames observed feature points" are non-zero, which conforms to the sparsity of visual observation), Schur complement is first applied to it.

[0049] In some embodiments, the information matrix after elimination is obtained by performing Schur complement elimination using the following formula (1).

[0050]

[0051] The above formula (1)H xp H pp -1 H px The constraint strength passed on by feature points as intermediate variables is represented by eliminating them from the information matrix. This transforms the indirect constraints of feature points (position coordinates) on the image frame state variables into direct constraints. While eliminating feature points, their influence on the image frame state variables is preserved. For example, if feature point f is observed by image frames A and B, before elimination, image frames A and B are indirectly related through the position coordinates of f. After elimination, image frames A and B directly contain the constraint contribution of feature point f, equivalent to image frames A and B establishing direct constraints through feature point f. In this embodiment, the feature point position coordinates are deleted using Schur complement elimination, thereby reducing the matrix dimension. The dimension before elimination is 15N+3M, and the dimension after elimination is 15N (N is the number of image frames). This significantly reduces the required computational dimension and greatly improves the efficiency of subsequent Cholesky decomposition and inversion. In some embodiments, the information matrix after elimination is still a symmetric positive definite matrix that satisfies the conditions of the Choleski decomposition, and it still contains feature point constraints, making the calculation of covariance closer to actual observation and avoiding errors caused by empirical values.

[0052] S14. Perform Cholliski decomposition on the eliminated information matrix to obtain an upper triangular matrix and a lower triangular matrix.

[0053] In some embodiments, the information matrix after elimination contains information on multiple state variables, specifically including P, V, Q, Ba, and Bg of the terminal at all preset times. However, for visual inertial positioning, the required data is the P, V, Q, Ba, and Bg of the image frame at the terminal's most recent preset time, as well as the P and Q at other preset times. Therefore, it is necessary to perform edge-shifting operations on V, Ba, and Bg except for the most recent preset time, separating the state variables to be retained from those to be edge-shifted. This can be achieved by continuing to perform Schur complement elimination on the information matrix after elimination to eliminate the state variables to be edge-shifted, but this method has low computational efficiency. In some embodiments of this disclosure, a new method for fast covariance estimation is proposed to improve computational efficiency. In this embodiment, the information matrix after elimination is decomposed using Cholesky decomposition, making... L TLet L be the upper triangular matrix, and L be the lower triangular matrix, where the lower triangular matrix is ​​the transpose of the upper triangular matrix. After decomposition, the diagonal elements of L represent the independent constraint strength of each state variable, and the off-diagonal elements represent the correlation between state variables, serving as the basis for subsequent fast solution and covariance.

[0054] S15. Inverse the state variables of the image frame corresponding to the nearest preset time in the upper triangular matrix to obtain the square root of the covariance, and then determine the covariance.

[0055] In some embodiments, the lower right corner of the upper triangular matrix is ​​independent of other dimensions. The dimension of the image frame state quantity at any preset time is 15×15. The inverse of the lower right 15×15 part of the upper triangular matrix (which corresponds to the image frame at the most recent preset time) is obtained. The square root of the covariance is then obtained, and its square is used to obtain the covariance. The covariance characterizes the uncertainty of the state quantity of the image frame at the most recent preset time (e.g., position error range, attitude stability, etc.).

[0056] S16. Determine the initialization state of the terminal based on the covariance and initialization parameters.

[0057] In some embodiments, after obtaining the covariance, the initialization state of the terminal can be determined by combining it with the initialization parameters. This includes the terminal's position, attitude, velocity, and inertial bias at the most recent preset time (usually the current time), as well as the previous position, attitude, velocity, and inertial bias, and the uncertainty (covariance) of the terminal's position, attitude, velocity, and inertial bias at the most recent preset time.

[0058] In some embodiments of this disclosure, direct inversion is avoided by using Schur complement elimination and Cholleysky decomposition, improving the solution speed by more than 2 times and shortening the terminal initialization time. The covariance is calculated based on actual constraints rather than empirical values, significantly reducing the fluctuation amplitude of state variables after initialization. The method proposed in this disclosure is applicable to state estimation of terminals containing cameras and IMUs, and can be used in fields such as UAV autonomous navigation and visual-inertial positioning.

[0059] In some embodiments of this disclosure, after performing Cholleski decomposition on the eliminated information matrix to obtain an upper triangular matrix and a lower triangular matrix, and before determining the initialization state of the terminal based on the covariance and the initialization parameters, the method further includes:

[0060] Establish the equation shown in formula (2) below;

[0061]

[0062] Where, Δx x It is the increment of the image frame state quantity. It is the error vector after Schur complement elimination; the increment of the image frame state quantity is calculated using the properties of the upper triangular matrix and the lower triangular matrix; the initialization parameters are corrected based on the increment of the image frame state quantity.

[0063] In some embodiments, the optimization of visual-inertial positioning transforms the minimization of state variable errors into a system of linear equations, namely HΔx = -b, where H is the information matrix, Δx represents the increment of image frame state variables (such as corrections for velocity, position, etc.), and b is the Jacobian weighted sum of the error vectors, reflecting the bias between prediction and observation. In this embodiment, after Schur complement elimination, irrelevant state variables are removed, retaining only the eliminated information matrix of image frame state variables. Therefore, HΔx=-b becomes the form shown in formula (2). To solve Δx x First, solve the trigonometric system of equations Ly = -b x Let y = L T Δx x The original equation LL T Δx x =-b x It becomes: Ly = -b x Since L is a lower triangular matrix (values ​​below the diagonal and zeros above), we can quickly solve for y using forward substitution: starting from the first row, substitute the known values ​​row by row to calculate each element of y. Then solve the upper triangular system of equations L. T Δx x =y, substitute y into y=L T Δx x We get: L T Δx x =y, and L T It is an upper triangular matrix (values ​​above the diagonal, zeros below), and Δx can be solved using backward substitution. x Starting from the last row, substitute the values ​​of y in reverse order to calculate Δx. x Each element of Δx x This is the correction amount for the image frame state quantity, more specifically, the correction amount for the image frame state quantity at the most recent preset time. The step of correcting the initialization parameters based on the increment of the image frame state quantity includes: after calculation, adding it to the state quantity of the most recent image frame to obtain the corrected state quantity of the most recent image frame. Δx x It is a correction value for the state variables of the image frame. By calculating it, the state variables of the terminal's image frame can be made more accurate. Moreover, it directly uses the Choleski decomposition to obtain the upper triangular matrix and the lower triangular matrix, which reduces the amount of calculation and improves efficiency.

[0064] This disclosure also proposes a terminal initialization device, characterized in that it includes:

[0065] The acquisition unit is used to acquire the initialization parameters, inertial constraints, and visual constraints of the terminal. The initialization parameters include: the position, attitude, velocity, inertial offset of the terminal in the initial time period, and the positions of multiple preset feature points in the environment in the initial time period. The inertial constraints include: the relationship between the position, attitude, and velocity of the terminal at different preset times and the inertial data. The visual constraints include: the feature points in the image frames captured by the terminal at preset times. The initial time period includes multiple preset times.

[0066] The control unit is used to construct an information matrix based on the initialization parameters, inertial constraints, and visual constraints.

[0067] The control unit is used to perform Schul complement elimination on the information matrix to remove the feature points from the information matrix and obtain the eliminated information matrix.

[0068] The control unit is used to perform Cholliski decomposition on the eliminated information matrix to obtain an upper triangular matrix and a lower triangular matrix.

[0069] The control unit is used to invert the state quantity of the image frame corresponding to the nearest preset time in the upper triangular matrix to obtain the square root of the covariance, and then determine the covariance, and determine the initialization state of the terminal based on the covariance and the initialization parameters.

[0070] In some embodiments, the information matrix is ​​subjected to Schur complement elimination to remove the feature points from the information matrix, resulting in an eliminated information matrix, including:

[0071] The information matrix after elimination is obtained using the following formula (1).

[0072]

[0073] The information matrix includes: a constraint block H representing the image frame state variables. xx H represents the constraint block H from the image frame to the feature point. xp H represents the constraint block H from feature points to image frames. px and the constraint block H representing the feature points pp .

[0074] In some embodiments, after performing Cholleski decomposition on the eliminated information matrix to obtain upper and lower triangular matrices, and before determining the initialization state of the terminal based on the covariance and the initialization parameters, the method further includes:

[0075] Establish the equation shown in formula (2) below;

[0076]

[0077] Among them, L T Let L be the upper triangular matrix, and L be the lower triangular matrix, Δx x It is the increment of the image frame state quantity. It is the error vector after Schur complement elimination;

[0078] The increment of the image frame state quantity is calculated using the properties of the upper triangular matrix and the lower triangular matrix;

[0079] The initialization parameters are corrected based on the increment of the image frame state quantity.

[0080] In some embodiments, the lower triangular matrix is ​​the transpose of the upper triangular matrix.

[0081] In some embodiments, the inertial bias includes: accelerometer bias and angular velocity meter bias;

[0082] In some embodiments, the image frame state quantities at any preset time include: the position, attitude, velocity, accelerometer bias, and angular velocity bias of the terminal at that preset time;

[0083] The dimension of the image frame state quantity at any preset time is 15×15;

[0084] The square root of the covariance is obtained by inverting the state variables of the image frame corresponding to the nearest preset time in the upper triangular matrix, including: inverting the 15×15 part of the lower right corner of the upper triangular matrix.

[0085] In some embodiments, the terminal is a virtual reality device, an augmented reality device, or a mixed reality device.

[0086] For embodiments of the apparatus, since they basically correspond to the method embodiments, relevant details can be found in the descriptions of the method embodiments. The apparatus embodiments described above are merely illustrative, and the modules described as separate modules may or may not be separate. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0087] The methods and apparatus of this disclosure have been described above based on embodiments and application examples. Furthermore, this disclosure also provides an electronic device and a computer-readable storage medium, which are described below.

[0088] The following is for reference. Figure 4The figure illustrates a structural schematic of an electronic device (e.g., a terminal device or server) 800 suitable for implementing embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. The electronic device shown in the figure is merely an example and should not be construed as limiting the functionality and scope of the embodiments of the present disclosure.

[0089] Electronic device 800 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 801, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 802 or a program loaded from storage device 808 into random access memory (RAM) 803. RAM 803 also stores various programs and data required for the operation of electronic device 800. The processing device 801, ROM 802, and RAM 803 are interconnected via bus 804. Input / output (I / O) interface 805 is also connected to bus 804.

[0090] Typically, the following devices can be connected to I / O interface 805: input devices 806 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 807 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 808 including, for example, magnetic tapes, hard disks, etc.; and communication devices 809. Communication device 809 allows electronic device 800 to communicate wirelessly or wiredly with other devices to exchange data. Although an electronic device 800 with various devices is shown in the figure, it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0091] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 809, or installed from a storage device 808, or installed from a ROM 802. When the computer program is executed by a processing device 801, it performs the functions defined in the methods of embodiments of this disclosure.

[0092] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0093] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0094] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0095] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the methods of the present disclosure.

[0096] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0097] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0098] The units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units are not, in some cases, intended to limit the specific unit.

[0099] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0100] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0101] According to one or more embodiments of this disclosure, a terminal initialization method is provided, comprising:

[0102] The initialization parameters, inertial constraints, and visual constraints of the terminal are obtained. The initialization parameters include the terminal's position, attitude, velocity, inertial offset, and the positions of multiple preset feature points in the environment during the initial time period. The inertial constraints include the relationship between the terminal's position, attitude, and velocity at different preset times and the inertial data. The visual constraints include the feature points in the image frames captured by the terminal at preset times. The initial time period includes multiple preset times.

[0103] An information matrix is ​​constructed based on the initialization parameters, inertial constraints, and visual constraints.

[0104] The information matrix is ​​subjected to Schur complement elimination to remove the feature points from the information matrix, resulting in an eliminated information matrix.

[0105] The eliminated information matrix is ​​decomposed by Cholliski to obtain an upper triangular matrix and a lower triangular matrix.

[0106] The square root of the covariance is obtained by inverting the state variables of the image frame corresponding to the nearest preset time in the upper triangular matrix, and then the covariance is determined.

[0107] The initialization state of the terminal is determined based on the covariance and the initialization parameters.

[0108] According to one or more embodiments of this disclosure, a terminal initialization method is provided, which involves performing Schur complement elimination on the information matrix to remove the feature points from the information matrix and obtain an eliminated information matrix, including:

[0109] The information matrix after elimination is obtained using the following formula (1).

[0110]

[0111] The information matrix includes: a constraint block H representing the image frame state variables. xx H represents the constraint block H from the image frame to the feature point. xp H represents the constraint block H from feature points to image frames. px and the constraint block H representing the feature points pp .

[0112] According to one or more embodiments of this disclosure, a terminal initialization method is provided, which, after performing Cholleski decomposition on the eliminated information matrix to obtain an upper triangular matrix and a lower triangular matrix, and before determining the initialization state of the terminal based on the covariance and the initialization parameters, further includes:

[0113] Establish the equation shown in formula (2) below;

[0114]

[0115] Among them, L T Let L be the upper triangular matrix, and L be the lower triangular matrix, Δx x It is the increment of the image frame state quantity. It is the error vector after Schur complement elimination;

[0116] The increment of the image frame state quantity is calculated using the properties of the upper triangular matrix and the lower triangular matrix;

[0117] The initialization parameters are corrected based on the increment of the image frame state quantity.

[0118] According to one or more embodiments of this disclosure, a terminal initialization method is provided, wherein the lower triangular matrix is ​​the transpose of the upper triangular matrix.

[0119] According to one or more embodiments of this disclosure, a terminal initialization method is provided, wherein the inertial bias includes: accelerometer bias and angular velocity meter bias;

[0120] According to one or more embodiments of this disclosure, a terminal initialization method is provided, wherein the image frame state quantities at any preset time include: the terminal's position, attitude, velocity, accelerometer bias, and angular velocity bias at that preset time;

[0121] The dimension of the image frame state quantity at any preset time is 15×15;

[0122] The square root of the covariance is obtained by inverting the state variables of the image frame corresponding to the nearest preset time in the upper triangular matrix, including: inverting the 15×15 part of the lower right corner of the upper triangular matrix.

[0123] According to one or more embodiments of this disclosure, a method for initializing a terminal is provided, wherein the terminal is a virtual reality device, an augmented reality device, or a mixed reality device.

[0124] According to one or more embodiments of this disclosure, a terminal initialization apparatus is provided, comprising:

[0125] The acquisition unit is used to acquire the initialization parameters, inertial constraints, and visual constraints of the terminal. The initialization parameters include: the position, attitude, velocity, inertial offset of the terminal in the initial time period, and the positions of multiple preset feature points in the environment in the initial time period. The inertial constraints include: the relationship between the position, attitude, and velocity of the terminal at different preset times and the inertial data. The visual constraints include: the feature points in the image frames captured by the terminal at preset times. The initial time period includes multiple preset times.

[0126] The control unit is used to construct an information matrix based on the initialization parameters, inertial constraints, and visual constraints.

[0127] The control unit is used to perform Schul complement elimination on the information matrix to remove the feature points from the information matrix and obtain the eliminated information matrix.

[0128] The control unit is used to perform Cholliski decomposition on the eliminated information matrix to obtain an upper triangular matrix and a lower triangular matrix.

[0129] The control unit is used to invert the state quantity of the image frame corresponding to the nearest preset time in the upper triangular matrix to obtain the square root of the covariance, and then determine the covariance;

[0130] The control unit is used to determine the initialization state of the terminal based on the covariance and the initialization parameters.

[0131] According to one or more embodiments of the present disclosure, an electronic device is provided, including: at least one memory and at least one processor;

[0132] The at least one memory is used to store program code, and the at least one processor is used to call the program code stored in the at least one memory to execute the method described in any one of the above.

[0133] According to one or more embodiments of the present disclosure, a computer-readable storage medium is provided for storing program code that, when executed by a processor, causes the processor to perform the methods described above.

[0134] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0135] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0136] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.

Claims

1. A terminal initialization method, characterized in that, include: Obtain the terminal's initialization parameters, inertial constraints, and visual constraints; The initialization parameters include: the terminal's position, attitude, velocity, inertial bias, and the positions of multiple preset feature points in the environment during the initial time period; the inertial constraints include: the relationship between the terminal's position, attitude, and velocity at different preset times and the inertial data; the visual constraints include: the feature points in the image frames captured by the terminal at preset times, and the initial time period includes multiple preset times; An information matrix is ​​constructed based on the initialization parameters, inertial constraints, and visual constraints. The information matrix is ​​subjected to Schur complement elimination to remove the feature points from the information matrix, resulting in an eliminated information matrix. The eliminated information matrix is ​​decomposed by Cholliski to obtain an upper triangular matrix and a lower triangular matrix. The square root of the covariance is obtained by inverting the state variables of the image frame corresponding to the nearest preset time in the upper triangular matrix, and then the covariance is determined. The initialization state of the terminal is determined based on the covariance and the initialization parameters.

2. The method according to claim 1, characterized in that, The information matrix is ​​subjected to Schur complement elimination to remove the feature points from the information matrix, resulting in an eliminated information matrix, including: The information matrix after elimination is obtained using the following formula (1). ; The information matrix includes: a constraint block H representing the image frame state variables. xx H represents the constraint block H from the image frame to the feature point. xp H represents the constraint block H from feature points to image frames. px and the constraint block H representing the feature points pp .

3. The method according to claim 1, characterized in that, After performing Cholliski decomposition on the eliminated information matrix to obtain upper and lower triangular matrices, and before determining the initialization state of the terminal based on the covariance and the initialization parameters, the method further includes: Establish the equation shown in formula (2) below; Among them, L T Let L be the upper triangular matrix and L be the lower triangular matrix. It is the increment of the image frame state quantity. It is the error vector after Schur complement elimination; The increment of the image frame state quantity is calculated using the properties of the upper triangular matrix and the lower triangular matrix; The initialization parameters are corrected based on the increment of the image frame state quantity.

4. The method according to claim 1, characterized in that, The lower triangular matrix is ​​the transpose of the upper triangular matrix.

5. The method according to claim 1, characterized in that, The inertial bias includes: accelerometer bias and angular velocity meter bias.

6. The method according to claim 5, characterized in that, The image frame state quantities at any preset time include: the position, attitude, velocity, accelerometer bias, and angular velocity bias of the terminal at that preset time; The dimension of the image frame state quantity at any preset time is 15×15; The square root of the covariance is obtained by inverting the state variables of the image frame corresponding to the nearest preset time in the upper triangular matrix, including: inverting the 15×15 part of the lower right corner of the upper triangular matrix.

7. The method according to claim 1, characterized in that, The terminal is a virtual reality device, an augmented reality device, or a mixed reality device.

8. An initialization device for a terminal, characterized in that, include: The acquisition unit is used to acquire the terminal's initialization parameters, inertial constraints, and visual constraints. The initialization parameters include: the terminal's position, attitude, velocity, inertial bias, and the positions of multiple preset feature points in the environment during the initial time period; the inertial constraints include: the relationship between the terminal's position, attitude, and velocity at different preset times and the inertial data; the visual constraints include: the feature points in the image frames captured by the terminal at preset times, and the initial time period includes multiple preset times; The control unit is used to construct an information matrix based on the initialization parameters, inertial constraints, and visual constraints. The control unit is used to perform Schul complement elimination on the information matrix to remove the feature points from the information matrix and obtain the eliminated information matrix. The control unit is used to perform Cholliski decomposition on the eliminated information matrix to obtain an upper triangular matrix and a lower triangular matrix. The control unit is used to invert the state quantity of the image frame corresponding to the nearest preset time in the upper triangular matrix to obtain the square root of the covariance, and then determine the covariance; The control unit is used to determine the initialization state of the terminal based on the covariance and the initialization parameters.

9. An electronic device, comprising: At least one memory and at least one processor; The at least one memory is used to store program code, and the at least one processor is used to call the program code stored in the at least one memory to execute the method of any one of claims 1 to 7.

10. A computer-readable storage medium for storing program code, which, when executed by a processor, causes the processor to perform the method of any one of claims 1 to 7.