Positioning initialization method and related devices, equipment, and storage media
By acquiring a small number of initial images and combining them with inertial sensor information for visual initialization, the problems of long visual positioning initialization time and large computational load are solved, achieving faster and more accurate positioning initialization.
Patent Information
- Application Number
- CN202110331307.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-03-25
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2041-03-25
AI Technical Summary
Existing visual localization methods require a large number of initial images during initialization, resulting in long processing times and high computational costs. Furthermore, the inertial sensor bias can lead to errors in scale and gravity estimation.
By acquiring at least four initial images that are less than a preset frame number threshold, and combining the visual initialization results with inertial sensor information, alignment processing is performed using the preset bias of the inertial sensor to optimize the positioning state quantity and reduce the impact of erroneous bias.
It speeds up initialization, reduces computational load, improves initialization accuracy and stability, and reduces the impact of erroneous bias on positioning.
Smart Images

Figure CN113052897B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of positioning technology, and in particular to a positioning initialization method and related devices, equipment, and storage media. Background Technology
[0002] Currently, visual localization methods play an important role in computer vision, robotics, drones, 3D reconstruction, and augmented reality. Generally, the initial images required to initialize the state variables needed for visual localization are obtained after a long period of time, meaning that at least hundreds of initial images are needed. This results in long initialization times and excessive computational demands. Summary of the Invention
[0003] This application provides at least one positioning initialization method and related apparatus, equipment, and storage medium.
[0004] The first aspect of this application provides a positioning initialization method, comprising: acquiring a first preset number of initial image frames, wherein the first preset number is greater than or equal to 4 and less than a preset frame number threshold, the preset frame number threshold being determined by the computing power of the execution device; performing visual initialization using the first preset number of initial image frames to obtain a visual initialization result, wherein the visual initialization result includes the pose of the initial image; and initializing the positioning state quantity by combining the visual initialization result with inertial sensor information.
[0005] Therefore, by determining a preset frame number threshold based on the computing power of the execution device, initialization can be completed using at least four initial images that are less than the preset frame number threshold. Compared to the general approach that requires a large number of initial images to complete initialization, the former is faster and reduces the computational load of initialization.
[0006] The process of initializing the positioning state quantity by combining the visual initialization results and the inertial sensor information includes: aligning the inertial sensor information between each initial image with the visual initialization results to obtain the initialized positioning state quantity information, wherein the bias of the inertial sensor used in the alignment process is a preset value.
[0007] Therefore, by setting the bias of the inertial sensor to a preset value, the possibility of incorrect estimation of scale and / or gravity caused by excessive coupling between the bias of the inertial sensor and gravity and scale is reduced, thereby reducing the impact of incorrect bias on initialization stability.
[0008] The visual initialization process, which involves using a preset number of initial images to perform visual initialization operations and obtain visual initialization results, includes: using one of the initial images as a reference image and determining a first relative positional relationship between at least one first initial image and the reference image, wherein the first initial image is an initial image other than the reference image; determining the positional information of a first three-dimensional point corresponding to a common observation two-dimensional point in at least one first initial image and the reference image based on the first relative positional relationship; and determining a second relative positional relationship between a second initial image and the reference image using the positional information of the first three-dimensional point, wherein the second initial image is an initial image other than the first initial image and the reference image, and the total number of the first initial image, the second initial image, and the reference image is a first preset number.
[0009] Therefore, determining the positions of some initial images first, and then determining the positions of the remaining initial images, speeds up the visual initialization process.
[0010] The relative positional relationship includes distance and / or relative angle; determining the first relative positional relationship between at least one frame of the first initial image and the reference image includes: determining the distance between the first initial image and the reference image as a preset distance value; and / or, obtaining the relative angle between the first initial image and the reference image based on inertial sensor pre-integration between the first initial image and the reference image, or obtaining the relative angle between the first initial image and the reference image in a polar constraint manner.
[0011] Therefore, by using a preset distance value and by using an inertial sensor / epipole constraint, the distance and position between the reference image and the first initial image can be obtained, thus obtaining the relative positional relationship between the reference image and the first initial image.
[0012] The method further includes, after determining the second relative positional relationship between the second initial image and the reference image using the positional information of the first three-dimensional point, determining a number of two-dimensional points existing in the second preset number of frames of the initial image, wherein the number of two-dimensional points do not correspond to the first three-dimensional point; obtaining the positional information of the second three-dimensional point corresponding to the determined two-dimensional point; and optimizing at least one of the pose of the initial image, the positional information of the first three-dimensional point, and the positional information of the second three-dimensional point.
[0013] Therefore, by determining that there are at least two-dimensional points on the initial image of the second preset number of frames that do not correspond to the first three-dimensional point, and obtaining the corresponding three-dimensional point information, and then combining the position information of the first three-dimensional point and the second three-dimensional point with the pose of the initial image for joint optimization, the result of visual initialization is more accurate.
[0014] The initialization of the positioning state variables by combining the visual initialization results with the inertial sensor information is performed after the visual initialization process is determined to be successful. The method also includes at least one of the following steps to determine that the visual initialization process is successful: determining that the number of first three-dimensional points is greater than a first preset threshold; determining that the number of three-dimensional points with positive depth is greater than a second preset threshold, wherein the three-dimensional points include the first three-dimensional points and the second three-dimensional points; and determining that the average reprojection error of the three-dimensional points on the initial image is less than or equal to a third preset threshold.
[0015] Therefore, by determining that visual initialization is successful, and then combining the visual initialization results with the inertial sensor information, the positioning state variables are initialized, reducing the waste of computational resources caused by the incorrect combination between visual initialization failure results and inertial sensor information.
[0016] The method further includes at least one of the following steps to determine that the initialization of the positioning state quantities is successful: determining that at least a portion of the initialized positioning state quantities are within a first preset range; determining sensing data based on the initialized positioning state quantities, and determining that the sensing data are within a second preset range; wherein the sensing data includes one or more of the magnitude of gravity and the bias of the inertial sensor.
[0017] Therefore, by determining whether the positioning state variables obtained during initialization are within a reasonable range, we can determine whether the initialization was successful and reduce the occurrence of using incorrect initialization results for positioning.
[0018] After the positioning state variables are successfully initialized, the method further includes: determining the preset position of the third initial image in the world coordinate system, wherein the third initial image is an initial image that meets the preset requirements; and adjusting the positions of the other initial images in the world coordinate system based on the position determined by the third initial image and the relative positional relationship between adjacent initial images.
[0019] Therefore, by fixing the position of an initial image in world coordinates, the positions of the remaining initial images in the world coordinate system can be determined.
[0020] The preset requirement is the earliest shooting time; and / or, the method further includes at least one of the following steps to set the weights of the state variables: setting the weight of the position of the third initial image as a first weight; setting the weight of the heading angle of the third initial image as a second weight; setting the weight of the bias of the gyroscope in the inertial sensor as a third weight; setting the weights of the gravity direction, velocity, and bias of the accelerometer in the inertial sensor of the first preset number of initial images as fourth, fifth, and sixth weights, respectively; wherein, the larger the weight of the state variable, the lower the corresponding uncertainty.
[0021] Therefore, since the position of the first frame of the captured image and its heading angle in the world coordinate system are unknown, setting the position of the first frame of the captured image as a preset position in the world coordinate system and its heading angle and giving it a large weight makes it more reliable to obtain the positions of the remaining initial images by adjusting this position.
[0022] Among them, the first, second, and third weights are all greater than the fourth, fifth, and sixth weights.
[0023] Therefore, by determining the magnitude relationship between the first to sixth weights, we can rely more on the state variables with low uncertainty to update the state variables with high uncertainty when locating the target or updating the initialized state variables.
[0024] The method further includes at least one of the following steps: if visual initialization fails, then determine that positioning initialization has failed; if initialization of positioning state variables fails, then determine that positioning initialization has failed; before performing visual initialization using a first preset number of initial images, determine whether there is a first motion amplitude between adjacent initial images that is greater than a first preset amplitude; if there is, then perform visual initialization using a first preset number of initial images, otherwise determine that positioning initialization has failed.
[0025] Therefore, if the range of motion is too small, the number of corresponding 3D points may be too small, and the obtained visual initialization results may not be very accurate. In this case, the localization initialization is directly considered to have failed in order to reduce the amount of computation.
[0026] The method further includes: in the event that the positioning initialization fails, deleting the initial image with the earliest shooting time in the first preset number of frames and at least part of the processing results in the initialization process.
[0027] Therefore, by deleting the earliest captured initial image after an initialization failure, a new initial image can be used for initialization next time, thus improving the success rate of initialization.
[0028] The method of obtaining the first preset number of initial images includes at least one of the following steps: taking the images captured at preset time intervals as initial images; acquiring the captured images to be determined, and determining whether the second motion amplitude between the images to be determined and the fourth initial image is greater than the second preset amplitude; if the second motion amplitude is greater than the second preset amplitude, determining the images to be determined as initial images; wherein the fourth initial image is the initial image with the smallest time difference between it and the images to be determined.
[0029] Therefore, instead of fixing the interval for acquiring the initial image or acquiring the initial image through motion amplitude, using all captured image frames as the initial image can reduce the amount of computation during the initialization process.
[0030] The determination of whether the second motion amplitude between the image to be determined and the fourth initial image is greater than the second preset amplitude includes: determining whether the disparity between the image to be determined and the fourth initial image is greater than the preset disparity, and if the disparity is greater than the preset disparity, determining that the second motion amplitude is greater than the second preset amplitude; or, determining whether the pre-integration value of the inertial sensor between the image to be determined and the fourth initial image is greater than the preset integration value, and if the pre-integration value is greater than the preset integration value, determining that the second motion amplitude is greater than the second preset amplitude.
[0031] Therefore, judging the motion amplitude by using parallax and / or pre-integral values makes the judgment of motion amplitude more accurate.
[0032] A second aspect of this application provides a positioning initialization device, comprising: an acquisition module for acquiring a preset number of initial images, wherein the preset number is greater than or equal to 4 and less than a preset frame number threshold, the preset frame number threshold being determined by the computing power of the execution device; a visual initialization module for performing a visual initialization operation using the preset number of initial images to obtain a visual initialization result, the visual initialization result including the poses of all initial images and an initial map; and a joint initialization module for initializing positioning state variables by combining the visual initialization result with inertial sensor information.
[0033] A third aspect of this application provides an electronic device, including a memory and a processor, wherein the processor is configured to execute program instructions stored in the memory to implement the above-described positioning initialization method.
[0034] The fourth aspect of this application provides a computer-readable storage medium having program instructions stored thereon, which, when executed by a processor, implement the above-described positioning initialization method.
[0035] The above scheme determines the preset frame number threshold based on the computing power of the execution device and completes the initialization using at least four initial images that are less than the preset frame number threshold. Compared with the general scheme that requires a large number of initial images to complete the initialization, the former is faster and reduces the amount of computation during initialization.
[0036] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this application. Attached Figure Description
[0037] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the specification, serve to explain the technical solutions of this application.
[0038] Figure 1This is a first flowchart illustrating an embodiment of the positioning initialization method of this application;
[0039] Figure 2 This is a second flowchart of an embodiment of the positioning initialization method of this application;
[0040] Figure 3 This is a schematic diagram of the structure of an embodiment of the positioning initialization device of this application;
[0041] Figure 4 This is a schematic diagram of the structure of an embodiment of the electronic device of this application;
[0042] Figure 5 This is a schematic diagram of the structure of an embodiment of the computer-readable storage medium of this application. Detailed Implementation
[0043] The embodiments of this application will now be described in detail with reference to the accompanying drawings.
[0044] In the following description, specific details such as particular system architectures, interfaces, and technologies are presented for illustrative purposes rather than for limiting purposes, in order to provide a thorough understanding of this application.
[0045] In this document, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " generally indicates that the preceding and following related objects have an "or" relationship. Furthermore, "many" in this document means two or more. Moreover, the term "at least one" in this document means any combination of at least two of any one or more of a plurality of objects. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.
[0046] Please see Figure 1 , Figure 1 This is a first flowchart illustrating an embodiment of the positioning initialization method of this application. Specifically, it may include the following steps:
[0047] Step S11: Obtain the first preset number of initial images, wherein the first preset number is greater than or equal to 4 and less than the preset frame number threshold, and the preset frame number threshold is determined by the computing power of the execution device.
[0048] The initial image frames of the first preset number of frames can be captured by any device with a camera component, or they can be images acquired from other devices, or images that have undergone frame selection, brightness adjustment, resolution adjustment, etc. Here, "other devices" refers to devices that require different central processing units to operate. That is, the initial image can be captured by the device performing positioning initialization, or it can be captured by the device requiring positioning and then transmitted to the device performing positioning initialization via various communication methods. This embodiment of the disclosure takes the initial image being captured by the device performing positioning initialization as an example, where the executing device is referred to as the target (hereinafter, "target" refers to the executing device). That is, the initial image frames are acquired by the target. For example, during the operation of a drone, the drone is the target described in this embodiment of the disclosure. Or, during the operation of a robotic vacuum cleaner, the robotic vacuum cleaner is also the target described in this embodiment of the disclosure.
[0049] Optionally, the first preset number is greater than or equal to 4. For example, when the first preset number is equal to 4, the initialization of the state quantities required for positioning can be completed with four initial images. Of course, in other embodiments, the first preset number can also be 5, 6, or more. To control the computational load, the first preset number can be less than or equal to a preset frame count threshold. For example, a preset frame count threshold of 10 indicates that the first preset number is greater than or equal to 4 and less than 10. The preset frame count threshold can be determined based on the target's computational capabilities. A larger preset frame count threshold requires more initial images to be processed, resulting in greater computational power during positioning initialization. If the target's computational capabilities are strong enough, the preset frame count threshold can be set relatively larger; if the target's computational capabilities are weak, the preset frame count threshold can be set relatively smaller to prevent excessive initialization time or target lag due to insufficient computational capabilities. Optionally, the target's computational capabilities include the time per second the target can use for self-positioning, the number of positioning operations per second, and / or the time required to perform one positioning operation. For example, a target has 300ms per second to locate itself, and performs location 3 times per second. This means location initialization and target location operations need to be completed within 100ms. Although location initialization doesn't need to be performed before each target location operation, the initialization time needs to be controlled to ensure time for target location operations after initialization. Assuming one location operation takes 50ms, location initialization needs to be completed within another 50ms. Therefore, to ensure the target can complete location initialization within 50ms, a preset frame rate threshold is determined. The above numbers are just examples. The method for determining the first preset number is the same as the method for determining the preset frame rate threshold, and will not be repeated here.
[0050] Step S12: Perform visual initialization using the first preset number of initial images to obtain visual initialization results, wherein the visual initialization results include the pose of the initial images.
[0051] In this embodiment of the disclosure, after obtaining a first preset number of initial images, the position of one of the initial images can be fixed, and its relative positional relationship with adjacent initial images can be calculated. There are many ways to calculate the relative positional relationship between adjacent initial images. For example, the distance and angle between two initial images can be obtained using epipolar constraint technology. The main process of epipolar constraint technology includes extracting feature points from two adjacent initial images, matching them, calculating the fundamental matrix between the two initial images using the matched feature points, and finally decomposing the camera's rotation matrix and translation vector from the fundamental matrix. Here, the rotation matrix represents the rotation angle between the two initial image frames, and the translation vector represents the distance between the two initial images. Alternatively, the relative positional relationship between adjacent initial images can be calculated by assuming a relative distance between the two initial images, pre-integrating the gyroscope sensing data between the two initial images to obtain the relative angle between them, thereby obtaining the relative positional relationship between the two initial images. Using this method, the relative positional relationship between each initial image in the first preset number of frames can be obtained. Here, the relative positional relationship between the initial image frames is the pose of the initial image. Optionally, the visual initialization result may also include the coordinates of each three-dimensional point. Specifically, based on the pose of the initial image, feature points (i.e., two-dimensional points) in each initial image can be obtained. Triangulation of each two-dimensional point yields its corresponding three-dimensional coordinates. Generally, these three-dimensional points can form a simple initial map. Once the relative positions between the initial image frames and the coordinates of each three-dimensional point are calculated, visual initialization can be considered complete.
[0052] Step S13: Initialize the positioning state variables by combining the visual initialization results with the inertial sensor information.
[0053] The localization state variables are those required for target localization. Typically, these variables include the true-scale position, the angle aligned with the gravity direction, velocity, and sensor offset. However, in visual initialization, the relative distance error between the two initial images provided by the epipolar constraint may be significant, or the relative distance between two initial image frames may use an assumed distance. This means the calculated initial image position may not contain the true scale, and whether the calculated angle is aligned with gravity is unknown. Therefore, it is necessary to use the pre-integrated sensor information between two consecutive initial images to solve for the scale, the velocity corresponding to each initial image frame, and the gravity direction. Of course, if the target itself can provide an accurate gravity direction, it can be assumed that the gravity direction is known, and it is not necessary to solve for the gravity direction. The solution method can be linear or nonlinear.
[0054] If all state variables are solved successfully, the initialization of the positioning state variables is complete. After initialization, the target can be located using the obtained state variables.
[0055] The above scheme determines the preset frame number threshold based on the computing power of the execution device, and can complete the initialization by using at least four initial images that are less than the preset frame number threshold. Compared with the general scheme that requires a large number of initial images to complete the initialization, the former is faster and reduces the amount of computation for initialization.
[0056] In some disclosed embodiments, the method for obtaining the first preset number of initial images can be at least one of the following: First, images captured at preset time intervals are used as initial images. For example, the preset time is 100ms; in other embodiments, the preset time can also be 200ms. That is, one captured image is extracted as the initial image every 100ms. Because the frequency of camera image capture is fixed, many captured images within these 100ms will not participate in the initialization process. These initial images will not undergo image processing, i.e., feature extraction and matching will not be performed. Only the initial images are processed, which greatly reduces the computational load of image processing during initialization. Simultaneously, because the time difference between adjacent initial images increases, the sensing data between the two initial images increases, thereby improving the accuracy of scale calculation during initialization. Of course, if the computational load is not considered, each captured image can be used as the initial image, and each initial image can be processed. In this case, the preset time is related to the frequency of camera image capture.
[0057] Second, the process involves acquiring a captured image of the target but not yet determining whether it can be used as an initial image. The second motion amplitude between the image of the target and the fourth initial image is greater than a second preset amplitude. If the second motion amplitude is greater than the second preset amplitude, the image of the target is determined as the initial image. The image of the target is the one captured by the target but whose suitability as an initial image has not yet been determined. The fourth initial image is the initial image with the smallest time difference from the image of the target. In other words, the motion amplitude between the captured image of the target and the latest initial image is compared. If the motion amplitude is small, the image of the target will not be used as an initial image for initialization, nor will it undergo image processing. Optionally, the motion amplitude can be determined using disparity and / or sensor data. Specifically, it involves determining whether the disparity between the image of the target and the fourth initial image is greater than a preset disparity. If the disparity is greater than the preset disparity, the second motion amplitude is determined to be greater than the second preset amplitude. The disparity can be determined by extracting two-dimensional points from the image of the target and the fourth initial image, matching them to obtain several pairs of two-dimensional points, and using the disparity of the two-dimensional point pairs as the disparity between the image of the target and the fourth initial image. Optionally, the highest and lowest disparities among all two-dimensional point pairs can be discarded, and the mean of the remaining disparities can be calculated. This mean is then used as the disparity between the image to be determined and the fourth initial image. For example, if the moon is captured in both the image to be determined and the fourth initial image, but the distance between the moon and the target is too great, the disparity corresponding to the two-dimensional point extracted from the moon will be relatively small even if the target has undergone significant movement. This indicates that the two-dimensional point extracted from the moon is not suitable as a representation of the target's movement range. Therefore, the disparity with the smallest change can be discarded. Of course, the smallest disparity does not have to be discarded under all conditions. Instead, it can be discarded only when the smallest disparity is smaller than the average disparity including that disparity. The same principle applies to discarding the largest disparity.
[0058] Alternatively, it can be determined whether the pre-integrated value of the inertial sensor between the image to be determined and the fourth initial image is greater than a preset integration value. If the pre-integrated value is greater than the preset integration value, the second motion amplitude is determined to be greater than the second preset amplitude. Here, the pre-integrated value can be velocity, distance, or angle. Velocity is the pre-integration of the accelerometer reading, distance is the pre-integration of the accelerometer reading twice (i.e., the velocity pre-integration), and angle is the pre-integration of the gyroscope reading. Preset integration values can be set for velocity, distance, and angle respectively. If the velocity, distance, and angle between the image to be determined and the fourth initial image all satisfy their respective preset integration values, the second motion amplitude is considered greater than the second preset amplitude. Alternatively, the second motion amplitude can be considered greater than the second preset amplitude if any one of the velocity, distance, or angle between the image to be determined and the fourth initial image is greater than its corresponding preset integration value.
[0059] In some disclosed embodiments, the second motion amplitude is determined to be greater than the second preset amplitude only when the disparity between the pending image and the fourth initial image is greater than a preset disparity and the pre-integration value is greater than a preset integration value.
[0060] After acquiring a first preset number of initial images, visual initialization is performed using these initial images. To ensure the stability of subsequent initialization results, it is determined whether there exists a first motion amplitude greater than a first preset amplitude between adjacent initial images. If so, visual initialization using the first preset number of initial images is performed. If not, the localization initialization is considered a failure. Here, the motion amplitude can be determined using disparity and / or pre-integration values. This embodiment uses disparity as an example; if there is no disparity between adjacent initial images greater than a disparity threshold, the localization initialization is considered a failure. The reason for this failure determination is that if the disparities between initial images are very small, it is difficult to obtain the 3D coordinates of each matching 2D point through triangulation. Furthermore, if the number of 3D points is too small, the accuracy of subsequent initialization results is difficult to guarantee. Therefore, directly determining initialization failure reduces the problem of inaccurate localization caused by inaccurate initialization state quantities, while also saving computational resources and allowing the next initialization process to begin quickly.
[0061] The specific steps for visual initialization using a first preset number of initial frames may include: using one of the initial frames as a reference image, and determining a first relative positional relationship between at least one of the first initial frames and the reference image. Here, the first initial image is any initial image other than the reference image. That is, the position of one of the initial frames can be fixed, and then the relative positional relationship between the first initial image and the reference image can be calculated. The relative positional relationship includes distance and / or relative angle.
[0062] The selection of the reference image can be achieved by determining the motion amplitude between adjacent initial images in a preset first number of frames, identifying the two initial images corresponding to the maximum motion amplitude, and using the earliest captured initial image as the reference image. The motion amplitude can also be determined using parallax and / or sensor data. This embodiment selects parallax to determine the reference image. Specifically, the earlier captured frame among the two adjacent initial images with the largest parallax is selected as the reference image. This embodiment uses the first initial image as an adjacent frame of the reference image, specifically the other initial image frame corresponding to the maximum parallax. Of course, in other embodiments, the first initial image frame may also include another adjacent frame of the reference image frame. Optionally, the position of the reference image is temporarily fixed at the origin of the world coordinate system, and then the relative positional relationship between the initial images is determined, thus the obtained relative positional relationship is the positional relationship in the world coordinate system.
[0063] One way to determine the first relative positional relationship between the first initial image and the reference image is to define the distance between them as a preset distance value. That is, without knowing the true scale, the distance between the first initial image and the reference image is defined as a distance excluding the true scale; it's an assumed distance, which may not be truly accurate. There are several ways to determine the angle. For example, the relative angle between the first initial image and the reference image can be obtained based on inertial sensor pre-integration, or by using epipolar constraints. Here, inertial sensor pre-integration mainly refers to the pre-integration of gyroscope readings, because the gyroscope measures angular velocity, and the relative angle can be obtained through pre-integration. The method of obtaining the relative angle through epipolar constraints has been described above and will not be repeated here.
[0064] By using a preset distance value and by employing an inertial sensor / epipole constraint, the distance and position between the reference image and the first initial image can be obtained, thus revealing their relative positional relationship.
[0065] After obtaining the first relative positional relationship, the positional information of the first three-dimensional points corresponding to the commonly observed two-dimensional points in at least one frame of the first initial image and the reference image is determined based on the first relative positional relationship. Specifically, the commonly observed two-dimensional points in the first initial image and the reference image are triangulated to obtain the positional information of the first three-dimensional points corresponding to each two-dimensional point. After obtaining the positional information of the first three-dimensional points, the second relative positional relationship between the second initial image and the reference image is determined using the positional information of the first three-dimensional points, wherein the second initial image is an initial image other than the first initial image and the reference image. The total number of the first initial image, the second initial image, and the reference image is a first preset number. Here, the second relative positional relationship includes angle and position. Since the reference image is set as the origin in the world coordinate system, the angle and position here can be considered as the angle and position of the second initial image in the world coordinate system. Specifically, the two-dimensional points corresponding to the first three-dimensional points in each second initial image are found, and then the positions of each initial image can be solved using the commonly used method, which will not be elaborated here.
[0066] By first determining the positions of some initial images and then determining the positions of the remaining initial images, the speed of visual initialization is accelerated.
[0067] After determining the second relative positional relationship between the second initial image and the reference image using the positional information of the first three-dimensional point, several two-dimensional points existing in a second preset number of initial image frames are determined, wherein these two-dimensional points do not correspond to the first three-dimensional point. In this embodiment, the second preset number is greater than or equal to 3. That is, these two-dimensional points exist simultaneously in three initial image frames and have not undergone triangulation to obtain the corresponding first three-dimensional point. The positional information of the second three-dimensional point corresponding to the determined two-dimensional point is obtained. Specifically, the method for obtaining the positional information of the second three-dimensional point includes triangulation. Then, at least one of the pose of the initial image, the positional information of the first three-dimensional point, and the positional information of the second three-dimensional point is optimized. Here, the pose of the initial image includes the position and angle of the first and second initial images relative to the reference image. The optimization method includes combining the pose of each initial image frame and the positional information of the first and second three-dimensional points together for nonlinear optimization. Here, nonlinear optimization can simultaneously optimize the pose of each initial image frame and the positions of the first and second three-dimensional points to obtain the final visual initialization result.
[0068] By identifying at least two-dimensional points on the initial image that do not correspond to the first three-dimensional point and obtaining their corresponding three-dimensional point information, and then combining the position information of the first and second three-dimensional points with the pose of the initial image for joint optimization, the visual initialization results are made more accurate.
[0069] After obtaining the final visual initialization result, the success of the visual initialization process is determined. Specifically, initializing the positioning state variables by combining the visual initialization result with inertial sensor information is performed after confirming the visual initialization process's success. The steps for confirming successful visual initialization include at least one of the following: First, ensuring the number of first 3D points is greater than a first preset threshold. If the number of first 3D points is too small, the calculated second relative positional relationship between the second initial image and the reference image will be inaccurate. To reduce the possibility of using erroneous state variables for positioning, visual initialization is directly considered a failure in this case. Second, ensuring the number of 3D points with positive depth is greater than a second preset threshold. These 3D points include both first and second 3D points. Positive depth means the calculated 3D point is located in front of the target. According to imaging principles, 3D points behind the target (i.e., behind the camera) cannot be captured. If the triangulated 3D point is located behind the target, its depth is negative, indicating a problem. Therefore, only 3D points with positive depth are truly useful. If the number of 3D points with positive depth is too small, the pose error of the calculated initial images may be large, resulting in a large error in the final state variable initialization result. Third, the average reprojection error of the 3D points on the initial image is determined to be less than or equal to a third preset threshold. If the reprojection error is large, it indicates that there are significant problems with the pose calculation of each initial image, and the subsequent initialization process cannot continue.
[0070] After confirming successful visual initialization, the visual initialization results are combined with inertial sensor information to initialize the state variables required for positioning. This reduces the waste of computational resources caused by incorrect combinations between visual initialization failures and inertial sensor data.
[0071] After successful visual initialization, the positioning state variables are initialized by combining the visual initialization results with inertial sensor information. Specifically, the inertial sensor information between each initial image is aligned with the visual initialization results to obtain the initialized positioning state variables. The state variable information includes scale. The inertial sensor bias used in the alignment process is a preset value. Of course, if the target's camera is a binocular or multi-view camera, or if the target has a depth acquisition sensor, the depth of each 3D point can be directly obtained; that is, the scale can be directly determined from the initial image frames without needing to solve for the scale by aligning the visual initialization results with the inertial sensor information. The preset value of the inertial sensor bias can be a pre-calibrated value, the result from the previous positioning, or zero. The inertial sensor bias includes accelerometer bias and gyroscope bias. The method for determining the inertial sensor bias can be to check if there is a pre-calibrated value or the result from the previous positioning; if so, the pre-calibrated value or the result from the previous positioning is used; if not, the preset value is set to zero. Since the bias of a gyroscope is generally small and negligible relative to the gyroscope reading, the preset value can be 0. However, the bias of an accelerometer is highly correlated with information such as scale and gravity orientation. Giving the accelerometer a large degree of freedom in bias may lead to failure in solving for scale and gravity orientation. Therefore, setting the bias of the accelerometer to 0 can reduce the mutual influence between state variables.
[0072] The formula used in the alignment process can be a preset kinematic equation. Generally, when the target is detected to be in a non-stationary state, the positioning parameters include... Where X k X represents all positioning state variables at time k. C X represents the current location status variable being used. S This represents the localization state quantity corresponding to each of the first preset number of initial images, where T represents the transpose. Where G represents the world coordinate system and C represents the camera coordinate system. Ck R G This represents the rotation matrix from the world coordinate system to the camera coordinate system at time k. G p Ck This represents the position of the camera in the world coordinate system at time k. G v Ck Let b represent the velocity in the world coordinate system at time k. wk b represents the bias of the gyroscope at time k. ak This represents the accelerometer bias at time k. Assuming the first preset number is n frames, then... That is, the position, angle, and velocity of each initial image in the world coordinate system.
[0073] First, define a set of positioning parameters that need to be solved. Where s represents the scale and g represents the direction of gravity. This represents the velocity of the target in the world coordinate system at time k. Of course, if the target can provide an accurate direction of gravity, the direction of gravity can be taken as a known quantity, and then there is no need to solve for the direction of gravity.
[0074] The pre-defined kinematic equations are as follows:
[0075] in, This represents the error between the pre-integrated value from frame i to frame j and the prior value. express The corresponding pre-integrated covariance matrix. The kinematic equations consist of several of the aforementioned pre-defined kinematic equations.
[0076] In some disclosed embodiments, the alignment process can be achieved by linearly solving the pre-integrated information of the sensor data and the visually initialized localization state variables. The visually initialized localization state variables include 3D points and their reprojection errors. Nonlinear optimization using these two types of information is also possible. Therefore, the construction method of the preset kinematic equations is not limited to the aforementioned kinematic equations.
[0077] Specifically, the solution obtained for the scale, velocity, and / or gravity direction is the one that minimizes the error of the aforementioned set of pre-defined kinematic equations. In other words, it is the solution that minimizes the sum of the errors of each pre-defined kinematic equation. By using the solution that minimizes the error of the pre-defined kinematic equations as the obtained scale, velocity, and / or gravity direction, the obtained scale and velocity are more accurate compared to other solutions, leading to more accurate subsequent positioning results.
[0078] After obtaining the scale, velocity, and / or gravity direction, it is necessary to determine whether the initialization of the state variables was successful. The specific determination method includes at least one of the following steps: First, determine whether at least some of the initialized positioning state variables are within a first preset range. Specifically, determine whether the calculated scale is within the first preset range. The first preset range is determined by obtaining the maximum possible scale based on the target's fastest calibrated velocity and the pre-integration time. If the scale obtained through alignment processing is greater than this maximum scale, then the scale is considered to be outside the first preset range. Of course, some state variables may also include velocity; if the velocity exceeds the target's fastest calibrated velocity, then the velocity is considered to be outside the first preset range. If one or more of these variables are outside the first preset range, then the state variable initialization is considered to have failed.
[0079] Second, sensor data is determined based on the initialized positioning state variables, and the sensor data is determined to be within a second preset range. The sensor data includes one or more of the following: the magnitude of gravity and the bias of the inertial sensor. Specifically, the method for determining sensor data based on the initialization results required for positioning includes recalculating one or more of the following: gravity, gyroscope bias, and / or accelerometer bias. It is then determined whether the recalculated gravity, gyroscope bias, and / or accelerometer bias are within the corresponding second preset range. The second preset range for the magnitude of gravity is approximately 9.81 N plus or minus a certain number of Newtons. This number of Newtons can be set according to the positioning accuracy requirements. The second preset range for the gyroscope bias is a certain amount fluctuating around the bias of a typical gyroscope on the market. The second preset range for the accelerometer bias is also determined by a certain amount fluctuating around the bias of a typical accelerometer on the market, but can also be set according to the positioning accuracy requirements; generally, this amount is within one time factor.
[0080] By determining whether the state variables obtained during initialization are within a reasonable range, we can determine whether the initialization was successful, thus reducing the likelihood of using incorrect initialization results for location purposes.
[0081] The solved scale and target velocity are used as the initial scale and velocity results. Other positioning parameters, including angle, gyroscope, and accelerometer biases, can directly use the preset values from before initialization. Of course, preset values can be used, such as factory calibration values or values used in the previous positioning. That is, in this embodiment, the gyroscope and accelerometer biases are not initialized.
[0082] After successful initialization of the positioning state variables, the positions of each initial image in the world coordinate system are adjusted. Specifically, a preset position of the third initial image in the world coordinate system is determined, where the third initial image is the initial image that meets preset requirements. The preset requirement is the earliest capture time. That is, the position of the initial image with the earliest capture time among the first preset number is adjusted to the preset position in the world coordinate system. The preset position can be the origin. Based on the position determined by the third initial image and the relative positional relationships between adjacent initial images, the positions of the remaining initial images (excluding the third initial image) in the world coordinate system are adjusted. This process of adjusting the positions of the remaining initial images is essentially a translation of the entire initial image system.
[0083] By fixing the position of an initial image in world coordinates, the positions of the remaining initial images in the world coordinate system can be determined.
[0084] In some disclosed embodiments, after successful initialization of each state variable, the weights of the state variables are set. Specifically, this includes at least one of the following steps: setting the weight of the position of the third initial image as a first weight; setting the weight of the heading angle of the third initial image as a second weight; setting the weight of the bias of the gyroscope in the inertial sensor as a third weight; and setting the weights corresponding to the gravity direction, velocity, and accelerometer bias of the first preset number of initial images as fourth, fifth, and sixth weights, respectively. Wherein, the first weight is greater than or equal to the first preset weight, the second weight is greater than or equal to the second preset weight, the third weight is greater than or equal to the third preset weight, the sixth weight is greater than or equal to the sixth preset weight, the fourth weight is less than or equal to the fourth preset weight, and the fifth weight is less than or equal to the fifth preset weight. In this way, the pose of the third initial image, the heading angle of the third initial image, the bias of the gyroscope in the inertial sensor, and the bias of the accelerometer each have relatively large weights, while the gravity direction and velocity each have relatively small weights. In this embodiment, the weights corresponding to each state variable can be determined according to a first determination method. In the first determination method, the relationship between the weight of each state variable and the error between each state variable is: the reciprocal of the weight of the state variable * the unit of the state variable = the error of the state variable. For example, the third preset weight corresponding to the bias of the gyroscope is set to 10. 4 This indicates that the error between the set gyroscope bias and the actual gyroscope bias is 10. -4The value is expressed in degrees per second. This is just an example; in other embodiments, the bias weight of the gyroscope can be adjusted as needed. That is, the error between the set value and its corresponding true value is the reciprocal of the weight multiplied by the unit corresponding to that value. In other words, the larger the weight, the smaller the error between the value and its true value, meaning the more reliable the value is. The larger the weight of the state variable, the lower the corresponding uncertainty. Since a relatively small amount of information is used in the initialization calculation, there may still be some error between the initialized state variable and the actual value. If all state variables are weighted equally and set to a large uncertainty by default, the initialization error can easily affect other relatively correct state variables during positioning. Therefore, it is necessary to set a corresponding weight for each state variable. Furthermore, since the position of the initial image is unknown when the first frame is captured, a large weight is assigned to it, and this initial image is set as the origin of the world coordinate system. Since the heading angle of this initial image frame is also unknown, the heading angle of this frame is directly set to be known and given a large weight to avoid errors being incorrectly propagated to the heading angle. Furthermore, since the gyroscope bias is typically close to zero, setting it to zero and assigning it a large weight can prevent errors from being propagated incorrectly. Because the initial scale, gravity direction, and accelerometer bias are highly susceptible to mutual influence, to avoid these interactions and ensure each state variable converges as independently as possible, appropriate weights need to be assigned to the gravity direction, velocity, and accelerometer bias. For example, a larger weight can be assigned to the accelerometer bias, while smaller weights can be assigned to the gravity direction and velocity, thus reducing their mutual influence. Therefore, this approach makes the subsequent adjustments to the initial image positions more reliable. Moreover, during subsequent updates to the initialized positioning state variables, smaller-weighted state variables can be selectively adjusted, allowing for more precise target localization and improved positioning accuracy.
[0085] In some disclosed embodiments, the first, second, and third weights are all greater than the fourth, fifth, and sixth weights. The weights corresponding to each state variable can be determined according to a second determination method. This second determination method can be based on the error corresponding to each state variable. That is, it is determined according to the correspondence between error and weight. For example, if the error of a state variable is less than a first preset error, then the weight of that state variable is set to the first preset weight corresponding to the first preset error, for example, the first preset weight is 0.9. If the error of a state variable is greater than the first preset error but less than the second preset error, then the weight of that state variable is set to the second preset weight, for example, the second preset weight is 0.7, and so on. The first to sixth weights are determined according to this method. By setting the magnitude relationship between the first to sixth weights, in subsequent target localization or updating of initialized state variables, it is possible to rely more on state variables with low uncertainty to update state variables with high uncertainty.
[0086] In some disclosed embodiments, if the positioning initialization fails, the earliest captured initial image and at least a portion of the processing results during the initialization process are deleted from the first preset number of frames. The at least a portion of the processing results during the initialization process includes one or more of the following: the relative positions between the initial image frames, the position information of each 3D point, scale, and velocity. After deleting the earliest captured initial image, a new initial image is acquired, ensuring that the first preset number of initial images are still available for initialization in the next iteration.
[0087] After initialization fails, deleting the earliest captured initial image allows for the use of a new initial image for the next initialization, thus improving the success rate of initialization.
[0088] The method for determining that positioning initialization has failed includes at least the following steps: First, if visual initialization fails, positioning initialization is determined to have failed. Specifically, if visual initialization is determined to be unsuccessful, it is considered a failure. Second, if initialization of the positioning state variables fails, positioning initialization is determined to have failed. Specifically, if initialization of the positioning state variables is determined to be unsuccessful, it is considered a failure. Third, as mentioned above, before performing visual initialization using a first preset number of initial images, it is determined whether there is a first motion amplitude greater than a first preset amplitude between adjacent initial images; if so, visual initialization using the first preset number of initial images is performed; otherwise, positioning initialization is determined to have failed. If the motion amplitude is too small, the corresponding number of 3D points may be too small, and the obtained visual initialization result may be inaccurate; therefore, positioning initialization is directly determined to have failed to reduce the computational load.
[0089] The target can be located using the positioning state variables after initialization. This positioning can be achieved by determining the physical positional relationship between the target and the objects in the image based on the image captured by the target, or by using the target's previous positioning information and the currently captured image to determine the target's position, thereby realizing the positioning of the target and the tracking of the target's movement trajectory.
[0090] To better illustrate the technical solution proposed in this application, please refer to [reference needed]. Figure 2 , Figure 2 This is a schematic diagram of the second process in one embodiment of the positioning initialization method of this application. For example... Figure 2 As shown:
[0091] Step S11: Obtain the initial image of the first preset number of frames. The method for obtaining the initial image of the first preset number of frames is as described above, and will not be repeated here.
[0092] Step S14: Determine if there is a first motion amplitude greater than a first preset amplitude between adjacent initial images. If not, proceed to step S15: Determine if the positioning initialization has failed. If so, proceed to step S12: Perform visual initialization using a first preset number of initial images to obtain the visual initialization result. The specific method for performing visual initialization using the first preset number of initial images to obtain the visual initialization result is as described above and will not be repeated here.
[0093] Step S16: Determine if visual initialization was successful. If unsuccessful, proceed to step S15. If successful, proceed to step S13: Initialize the positioning state variables by combining the visual initialization results and inertial sensor information. The specific method for initializing the positioning state variables by combining the visual initialization results and inertial sensor information is as described above and will not be repeated here.
[0094] Execute step S17: Determine whether the initialization of the positioning status variables was successful. If unsuccessful, proceed to step S15. If successful, proceed to step S18: Determine that the positioning initialization was successful.
[0095] The above scheme can complete the initialization using at least four initial images that are less than a preset frame number threshold. Compared with the general scheme that requires a large number of initial images to complete the initialization, the former is faster and reduces the amount of computation during initialization.
[0096] The execution entity of the positioning initialization method can be a positioning initialization device. For example, the positioning initialization method can be executed by a terminal device, a server, or other processing devices. The terminal device can be: a virtual reality headset, augmented reality glasses, unmanned vehicle, mobile robot, robotic vacuum cleaner, flying equipment, user equipment (UE), mobile device, user terminal, terminal, cellular phone, cordless phone, personal digital assistant (PDA), handheld device, computing device, in-vehicle equipment, wearable device, or any other device that simultaneously possesses image sensors and inertial sensors. In some possible implementations, the positioning initialization method can be implemented by a processor calling computer-readable instructions stored in memory.
[0097] Please see Figure 3 , Figure 3 This is a schematic diagram of the structure of an embodiment of the positioning initialization device of this application. The positioning initialization device 30 includes: an acquisition module 31, a visual initialization module 32, and a joint initialization module 33. The acquisition module 31 is used to acquire a first preset number of initial image frames, wherein the first preset number is greater than or equal to 4 and less than a preset frame number threshold, and the preset frame number threshold is determined by the computing power of the execution device; the visual initialization module 32 is used to perform visual initialization using the first preset number of initial image frames to obtain a visual initialization result, wherein the visual initialization result includes the pose of the initial image; the joint initialization module 33 is used to combine the visual initialization result with inertial sensor information to initialize the positioning state quantity.
[0098] The above scheme can complete the initialization using at least four initial images that are less than a preset frame number threshold. Compared with the general scheme that requires a large number of initial images to complete the initialization, the former is faster and reduces the amount of computation during initialization.
[0099] In some disclosed embodiments, the joint initialization module 33 combines the visual initialization result with the inertial sensor information to initialize the positioning state quantity, including: aligning the inertial sensor information between each initial image with the visual initialization result to obtain the initialized positioning state quantity information, wherein the bias of the inertial sensor used in the alignment process is a preset value.
[0100] The above scheme reduces the impact of erroneous bias on initialization stability by setting the bias of the inertial sensor to a preset value, thereby reducing the possibility of incorrect estimation of scale and / or gravity caused by excessive coupling between the bias of the inertial sensor and gravity and scale.
[0101] In some disclosed embodiments, the visual initialization module 32 performs visual initialization operations using a preset number of initial images to obtain visual initialization results, including: using one of the initial images as a reference image and determining a first relative positional relationship between at least one first initial image and the reference image, wherein the first initial image is an initial image other than the reference image; determining the position information of a first three-dimensional point corresponding to a common observation two-dimensional point in at least one first initial image and the reference image based on the first relative positional relationship; and determining a second relative positional relationship between a second initial image and the reference image using the position information of the first three-dimensional point, wherein the second initial image is an initial image other than the first initial image and the reference image, and the total number of the first initial image, the second initial image, and the reference image is a first preset number.
[0102] The above scheme speeds up visual initialization by first determining the positions of some initial images and then determining the positions of the remaining initial images.
[0103] In some disclosed embodiments, the relative positional relationship includes distance and / or relative angle; the visual initialization module 32 determines the first relative positional relationship between at least one frame of the first initial image and the reference image, including: determining the distance between the first initial image and the reference image as a preset distance value; and / or, obtaining the relative angle between the first initial image and the reference image based on inertial sensor pre-integration between the first initial image and the reference image, or obtaining the relative angle between the first initial image and the reference image in a polar constraint manner.
[0104] The above scheme obtains the relative positional relationship between the reference image and the first initial image by acquiring the distance and position between the reference image and the first initial image through a preset distance value and by using an inertial sensor / epipole constraint.
[0105] In some disclosed embodiments, after the visual initialization module 32 determines the second relative positional relationship between the second initial image and the reference image using the position information of the first three-dimensional point, the visual initialization module 32 is further configured to: determine a plurality of two-dimensional points existing in the second preset number of frames of the initial image, wherein the plurality of two-dimensional points do not correspond to the first three-dimensional point; obtain the position information of the second three-dimensional point corresponding to the determined two-dimensional point; and optimize at least one of the pose of the initial image, the position information of the first three-dimensional point, and the position information of the second three-dimensional point.
[0106] The above scheme determines that there are at least two-dimensional points on the initial image of the second preset number of frames that do not correspond to the first three-dimensional point, obtains the corresponding three-dimensional point information, and then combines the position information of the first three-dimensional point and the second three-dimensional point with the pose of the initial image for joint optimization, so that the visual initialization result is more accurate.
[0107] In some disclosed embodiments, the initialization of the positioning state quantity, which combines the visual initialization result with the inertial sensor information, is performed after the visual initialization process is determined to be successful. The visual initialization module 32 is also used to perform at least one of the following steps to determine that the visual initialization process is successful: determining that the number of first three-dimensional points is greater than a first preset threshold; determining that the number of three-dimensional points with positive depth is greater than a second preset threshold, wherein the three-dimensional points include the first three-dimensional points and the second three-dimensional points; and determining that the average reprojection error of the three-dimensional points on the initial image is less than or equal to a third preset threshold.
[0108] The above scheme initializes the state variables required for positioning by combining the visual initialization result with the inertial sensor information after determining that the visual initialization was successful. This reduces the waste of computing resources caused by the incorrect combination between the visual initialization failure result and the inertial sensor.
[0109] In some disclosed embodiments, the joint initialization module 33 is further configured to perform at least one of the following steps to determine that the initialization of the state variables is successful: determining that at least a portion of the initialized state variables are within a first preset range; determining sensing data based on the initialized positioning state variables, and determining that the sensing data are within a second preset range; wherein the sensing data includes one or more of the magnitude of gravity and the bias of the inertial sensor.
[0110] The above scheme determines whether the initialization was successful by judging whether the positioning status quantity obtained during initialization is within a reasonable range, thereby reducing the occurrence of using incorrect initialization results for positioning.
[0111] In some disclosed embodiments, after the positioning state variables are successfully initialized, the joint initialization module 33 is further configured to: determine the preset position of the third initial image in the world coordinate system, wherein the third initial image is an initial image that meets the preset requirements; and adjust the positions of the remaining initial images other than the third initial image in the world coordinate system based on the position determined by the third initial image and the relative positional relationship between adjacent initial images.
[0112] The above scheme can determine the positions of the remaining initial images in the world coordinate system by fixing the position of an initial image in the world coordinate system.
[0113] In some disclosed embodiments, the preset requirement is the earliest shooting time; and / or, the joint initialization module 33 is further configured to perform at least one of the following steps to set the weights of the state variables: setting the weight of the position of the third initial image as a first weight; setting the weight of the heading angle of the third initial image as a second weight; setting the weight of the bias of the gyroscope in the inertial sensor as a third weight; setting the weights corresponding to the gravity direction, velocity, and bias of the accelerometer in the inertial sensor of the first preset number of initial images as a fourth weight, a fifth weight, and a sixth weight, respectively; wherein, the larger the weight of the state variable, the lower the corresponding uncertainty.
[0114] The above scheme, because the position of the first frame of the captured initial image and its heading angle in the world coordinate system are unknown, sets the position of the first frame of the initial image as a preset position in the world coordinate system and its heading angle and gives it a large weight, so that the position of the remaining initial images can be obtained more reliably by adjusting the position of the first frame of the initial image.
[0115] In some publicly available embodiments, the first weight, the second weight, and the third weight are all greater than the fourth weight, the fifth weight, and the sixth weight.
[0116] The above scheme, by determining the magnitude relationship between the first to the sixth weights, enables subsequent target localization or update of initialized state variables to rely more on state variables with low uncertainty to update state variables with high uncertainty.
[0117] In some disclosed embodiments, the joint initialization module 33 is further configured to perform at least one of the following steps: if visual initialization fails, then determine that positioning initialization has failed; if initialization of the positioning state quantity fails, then determine that positioning initialization has failed. Before performing visual initialization using a first preset number of initial images, it is determined whether there is a first motion amplitude between adjacent initial images that is greater than a first preset amplitude; if so, then perform visual initialization using the first preset number of initial images; otherwise, determine that positioning initialization has failed.
[0118] If the above scheme has too small a range of motion, the number of corresponding three-dimensional points may be too small, and the obtained visual initialization results may not be very accurate. Therefore, the positioning initialization is directly considered to have failed in order to reduce the amount of computation.
[0119] In some disclosed embodiments, the joint initialization module 33 is further configured to: delete the initial image with the earliest shooting time in the first preset number of frames and at least part of the processing results in the initialization process if it is determined that the positioning initialization has failed.
[0120] The above solution improves the success rate of initialization by deleting the earliest captured initial image after an initialization failure, allowing a new initial image to be used for initialization next time.
[0121] In some disclosed embodiments, the acquisition module 31 acquires a first preset number of initial images, including at least one of the following steps: taking an image captured at a preset time interval as an initial image; acquiring a captured pending image, and determining whether the second motion amplitude between the pending image and the fourth initial image is greater than a second preset amplitude; if the second motion amplitude is greater than the second preset amplitude, determining the pending image as an initial image; wherein, the fourth initial image is the initial image with the smallest time difference with the pending image.
[0122] The above scheme, instead of fixing the interval for acquiring the initial image or acquiring the initial image through the motion amplitude, uses all captured image frames as the initial image, which can reduce the amount of computation in the initialization process.
[0123] In some disclosed embodiments, the acquisition module 31 determines whether the second motion amplitude between the image to be determined and the fourth initial image is greater than the second preset amplitude, including: determining whether the parallax between the image to be determined and the fourth initial image is greater than the preset parallax, and if the parallax is greater than the preset parallax, determining that the second motion amplitude is greater than the second preset amplitude; or, determining whether the pre-integration value of the inertial sensor between the image to be determined and the fourth initial image is greater than the preset integration value, and if the pre-integration value is greater than the preset integration value, determining that the second motion amplitude is greater than the second preset amplitude.
[0124] The above scheme determines the motion amplitude by using parallax and / or pre-integration values, making the determination of motion amplitude more accurate.
[0125] The above scheme can complete the initialization using at least four initial images that are less than a preset frame number threshold. Compared with the general scheme that requires a large number of initial images to complete the initialization, the former is faster and reduces the amount of computation during initialization.
[0126] Please see Figure 4 , Figure 4 This is a schematic diagram of the structure of an embodiment of the electronic device of this application. The electronic device 40 includes a memory 41 and a processor 42. The processor 42 is used to execute program instructions stored in the memory 41 to implement the steps in any of the above-described positioning initialization method embodiments. In a specific implementation scenario, the electronic device 40 may include, but is not limited to: mobile robots, handheld mobile devices, robot vacuum cleaners, unmanned vehicles, virtual reality headsets, augmented reality glasses flying devices, microcomputers, desktop computers, servers, and other devices that simultaneously have image sensor and inertial sensor modules. In addition, the electronic device 40 may also include mobile devices such as laptops and tablets, which are not limited here.
[0127] Specifically, processor 42 controls itself and memory 41 to implement the steps in any of the above-described positioning initialization method embodiments. Processor 42 can also be referred to as a CPU (Central Processing Unit). Processor 42 may be an integrated circuit chip with signal processing capabilities. Processor 42 can also be a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor can be a microprocessor or any conventional processor. Furthermore, processor 42 can be implemented using integrated circuit chips.
[0128] The above scheme can complete the initialization using at least four initial images that are less than a preset frame number threshold. Compared with the general scheme that requires a large number of initial images to complete the initialization, the former is faster and reduces the amount of computation during initialization.
[0129] Please see Figure 5 , Figure 5 This is a schematic diagram of a computer-readable storage medium according to an embodiment of the present application. The computer-readable storage medium 50 stores program instructions 51 that can be executed by a processor. The program instructions 51 are used to implement the steps in any of the above-described positioning initialization method embodiments.
[0130] The above scheme determines the preset frame number threshold based on the computing power of the execution device and completes the initialization using at least four initial images that are less than the preset frame number threshold. Compared with the general scheme that requires a large number of initial images to complete the initialization, the former is faster and reduces the amount of computation during initialization.
[0131] In some embodiments, the functions or modules of the apparatus provided in this disclosure can be used to perform the methods described in the above method embodiments. The specific implementation can be referred to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.
[0132] The description of the various embodiments above tends to emphasize the differences between the various embodiments. The similarities or similarities between them can be referred to, and for the sake of brevity, they will not be repeated here.
[0133] In the several embodiments provided in this application, it should be understood that the disclosed methods and apparatus can be implemented in other ways. For example, the apparatus implementations described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms.
[0134] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
Claims
1. A positioning initialization method, characterized in that, include: Acquire a first preset number of initial image frames, wherein the first preset number is greater than or equal to 4 and less than a preset frame number threshold, the preset frame number threshold being determined by the computing power of the execution device; Visual initialization is performed using the first preset number of initial images to obtain visual initialization results, wherein the visual initialization results include the pose of the initial images; The positioning state quantity is initialized by combining the visual initialization result and the inertial sensor information; including: aligning the inertial sensor information between each initial image with the visual initialization result to obtain the initialized positioning state quantity, wherein the bias of the inertial sensor used in the alignment process is a preset value, and the bias of the accelerometer in the inertial sensor is zero, and the bias of the gyroscope in the inertial sensor is either pre-calibrated or the result of the previous positioning or zero; The method further includes setting weights for each of the positioning state quantities, wherein the larger the weight of the positioning state quantity, the lower the corresponding uncertainty.
2. The method according to claim 1, characterized in that, The step of performing visual initialization using the first preset number of initial images to obtain visual initialization results includes: Using one of the initial images as a reference image, a first relative positional relationship is determined between at least one first initial image and the reference image, wherein the first initial image is the initial image other than the reference image; Based on the first relative positional relationship, determine the positional information of the first three-dimensional point corresponding to the common observation two-dimensional point in the at least one frame of the first initial image and the reference image; Using the position information of the first three-dimensional point, a second relative positional relationship between the second initial image and the reference image is determined, wherein the second initial image is the initial image other than the first initial image and the reference image, and the total number of the first initial image, the second initial image and the reference image is the first preset number.
3. The method according to claim 2, characterized in that, The relative positional relationship includes distance and / or relative angle; Determining the first relative positional relationship between at least one frame of the first initial image and the reference image includes: The distance between the first initial image and the reference image is determined as a preset distance value; and / or, The relative angle between the first initial image and the reference image is obtained based on inertial sensor pre-integration between the first initial image and the reference image, or the relative angle between the first initial image and the reference image is obtained by epipolar constraint.
4. The method according to claim 2 or 3, characterized in that, After determining the second relative positional relationship between the second initial image and the reference image using the position information of the first three-dimensional point, the method further includes: Identify a number of two-dimensional points that exist in the initial image of a second preset number of frames, wherein the number of two-dimensional points do not correspond to the first three-dimensional point; Obtain the position information of the second three-dimensional point corresponding to the determined two-dimensional point; Optimize at least one of the pose of the initial image, the position information of the first three-dimensional point, and the position information of the second three-dimensional point.
5. The method according to claim 4, characterized in that, The initialization of the positioning state quantity by combining the visual initialization result with the inertial sensor information is performed after the visual initialization process is confirmed to be successful. The method further includes at least one of the following steps to determine if the visual initialization process was successful: Determine that the number of the first three-dimensional points is greater than a first preset threshold; The number of three-dimensional points with positive depth is determined to be greater than a second preset threshold, wherein the three-dimensional points include the first three-dimensional point and the second three-dimensional point; The average reprojection error of the three-dimensional points on the initial image is determined to be less than or equal to a third preset threshold.
6. The method according to any one of claims 1 to 5, characterized in that, The method further includes at least one of the following steps to determine that the initialization of the positioning state quantity was successful: Determine that at least some of the initialized positioning state quantities are within a first preset range; Sensing data is determined based on the initialized positioning state quantity, and the sensing data is determined to be within a second preset range; wherein the sensing data includes one or more of the magnitude of gravity and the bias of the inertial sensor.
7. The method according to any one of claims 1 to 6, characterized in that, After the location state variables are successfully initialized, the method further includes: Determine the preset position of the third initial image in the world coordinate system, wherein the third initial image is an initial image that meets preset requirements; Based on the position determined by the third initial image and the relative positional relationship between adjacent initial images, the positions of the remaining initial images other than the third initial image are adjusted in the world coordinate system.
8. The method according to claim 7, characterized in that, The preset requirement is the earliest shooting time; and / or, The method further includes at least one of the following steps to set the weights of the state variables: Set the weight of the position of the third initial image as the first weight; Set the weight of the heading angle of the third initial image as the second weight; Set the weight of the bias of the gyroscope in the inertial sensor as the third weight; The weights corresponding to the gravity direction, velocity, and accelerometer bias in the inertial sensor of the first preset number of initial images are respectively set as the fourth weight, the fifth weight, and the sixth weight.
9. The method according to claim 8, characterized in that, The first weight, the second weight, and the third weight are all greater than the fourth weight, the fifth weight, and the sixth weight.
10. The method according to any one of claims 1 to 9, characterized in that, The method further includes at least one of the following steps: If the visual initialization fails, then the positioning initialization is determined to have failed. If the initialization of the positioning status quantity fails, then the positioning initialization is determined to have failed. Before performing visual initialization using the first preset number of initial images, it is determined whether there is a first motion amplitude between adjacent initial images that is greater than a first preset amplitude; if so, the visual initialization using the first preset number of initial images is performed; otherwise, the positioning initialization is determined to have failed.
11. The method according to claim 10, characterized in that, The method further includes: If the positioning initialization fails, delete the initial image with the earliest shooting time in the first preset number of frames, as well as at least part of the processing results during the initialization process.
12. The method according to any one of claims 1 to 11, characterized in that, The process of obtaining the first preset number of initial image frames includes at least one of the following steps: The images captured at preset time intervals are used as the initial images; A pending image is acquired, and it is determined whether the second motion amplitude between the pending image and the fourth initial image is greater than a second preset amplitude. If the second motion amplitude is greater than the second preset amplitude, the pending image is determined as the initial image; wherein, the fourth initial image is the initial image with the smallest time difference between it and the pending image.
13. The method according to claim 12, characterized in that, The step of determining whether the second motion amplitude between the image to be determined and the fourth initial image is greater than the second preset amplitude includes: Determine whether the disparity between the image to be determined and the fourth initial image is greater than a preset disparity. If the disparity is greater than the preset disparity, determine that the second motion amplitude is greater than the second preset amplitude. Alternatively, determine whether the pre-integration value of the inertial sensor between the image to be determined and the fourth initial image is greater than a preset integration value. If the pre-integration value is greater than the preset integration value, determine that the second motion amplitude is greater than the second preset amplitude.
14. A positioning initialization device, characterized in that, include: The acquisition module is used to acquire a first preset number of initial image frames, wherein the first preset number is greater than or equal to 4 and less than a preset frame number threshold, and the preset frame number threshold is determined by the computing power of the execution device. A visual initialization module is used to perform visual initialization using the first preset number of initial images to obtain a visual initialization result, wherein the visual initialization result includes the pose of the initial images. The joint initialization module is used to initialize the positioning state quantity by combining the visual initialization result and the inertial sensor information; including: aligning the inertial sensor information between each initial image with the visual initialization result to obtain the initialized positioning state quantity, wherein the bias of the inertial sensor used in the alignment process is a preset value, and the bias of the accelerometer in the inertial sensor is zero, and the bias of the gyroscope in the inertial sensor is a pre-calibrated value or the result of the previous positioning or zero; The joint initialization module is also used to perform the following steps: setting the weight of each of the positioning state quantities, wherein the larger the weight of the positioning state quantity, the lower the corresponding uncertainty.
15. An electronic device, characterized in that, The method includes a memory and a processor, the processor being configured to execute program instructions stored in the memory to implement the method according to any one of claims 1 to 13.
16. A computer-readable storage medium having program instructions stored thereon, characterized in that, When the program instructions are executed by the processor, they implement the method described in any one of claims 1 to 13.
Citation Information
Patent Citations
Stable motion tracking method and stable motion tracking device based on integration of simple camera and IMU (inertial measurement unit) of smart cellphone
CN105953796A
Camera and inertial measurement unit online initialization and calibration method and system
CN111429524A