Instant positioning and map construction method and device, electronic equipment and readable storage medium
By fusing the initial pose information of the electronic device with interpolation variables, the problems of accuracy and latency in tracking pose in existing technologies are solved, achieving high-frequency and high-precision real-time positioning and map building.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- VIVO MOBILE COMM CO LTD
- Filing Date
- 2022-02-22
- Publication Date
- 2026-04-17
AI Technical Summary
In existing technologies, filtering-based methods result in poor accuracy in tracking the pose of electronic devices, while optimization-based methods involve large computational loads, leading to long delays and affecting the effectiveness of real-time localization and map building.
By fusing the initial pose information of the i-th frame image with the first interpolation variable, the final pose information is obtained, and real-time localization and map construction are performed based on this information. The interpolation variable is the variable before and after optimization of the keyframe image.
It improves the accuracy of pose tracking, shortens the pose tracking delay, ensures that electronic devices output high-frequency and high-precision pose information, and enhances the effect of real-time positioning and map building.
Smart Images

Figure CN115205419B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of communication technology, and specifically relates to a real-time positioning and map building method, device, electronic device and readable storage medium. Background Technology
[0002] With the continuous development of communication technology, electronic devices are becoming increasingly feature-rich. For example, electronic devices can perform real-time localization and map building by tracking the pose of each image frame of the current scene.
[0003] In related technologies, electronic devices can process the pose of image frames of the current scene acquired in real time by the electronic device based on filtering or optimization methods, and output the processed pose, thereby enabling the tracking of the pose of each image frame.
[0004] However, according to the above method, on the one hand, the filtering-based method cannot correct the pose of the image frame, resulting in poor accuracy of the electronic device tracking the pose; on the other hand, the optimization-based method has a large computational load, which makes it take a long time for the electronic device to process the pose of a single image frame, resulting in a long delay in the electronic device tracking the pose. Therefore, it may lead to poor performance of the electronic device in real-time localization and map building. Summary of the Invention
[0005] The purpose of this application is to provide a real-time positioning and mapping method, apparatus, electronic device, and readable storage medium that can solve the problem of poor performance of real-time positioning and mapping by electronic devices.
[0006] In a first aspect, embodiments of this application provide a real-time localization and map construction method, the method comprising: determining the initial pose information of the acquired i-th frame image based on the final pose information of the acquired (i-1)-th frame image, where i is an integer greater than 1; fusing the initial pose information of the i-th frame image with a first interpolation variable to obtain the final pose information of the i-th frame image, wherein the first interpolation variable is the last interpolation variable obtained before acquiring the i-th frame image, the first interpolation variable is the interpolation variable between the initial pose information of the first image and the target pose information, the first image is the image that is a keyframe in the images acquired before the i-th frame image, and the target pose information is the pose information optimized from the initial pose information of the first image; and performing real-time localization and map construction based on the final pose information of the i-th frame image and the i-th frame image.
[0007] Secondly, embodiments of this application provide an instant localization and map building apparatus, which includes an acquisition module, a determination module, a fusion module, and a processing module. The determination module is used to determine the initial pose information of the i-th frame image acquired by the acquisition module based on the final pose information of the (i-1)-th frame image acquired by the acquisition module, where i is an integer greater than 1. The fusion module is used to fuse the initial pose information of the i-th frame image with a first interpolation variable to obtain the final pose information of the i-th frame image. The first interpolation variable is the last interpolation variable obtained before the acquisition module acquires the i-th frame image. The first interpolation variable is the interpolation variable between the initial pose information of the first image and the target pose information. The first image is the keyframe image among the images acquired by the acquisition module before the i-th frame image, and the target pose information is the pose information optimized from the initial pose information of the first image. The processing module is used to perform instant localization and map building based on the final pose information of the i-th frame image and the i-th frame image.
[0008] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, wherein the memory stores programs or instructions executable on the processor, and the programs or instructions, when executed by the processor, implement the steps of the method described in the first aspect.
[0009] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.
[0010] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the method as described in the first aspect.
[0011] In a sixth aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the method described in the first aspect.
[0012] In this embodiment, the initial pose information of the acquired i-th frame image can be determined based on the final pose information of the acquired (i-1)-th frame image, where i is an integer greater than 1; and the initial pose information of the i-th frame image is fused with a first interpolation variable to obtain the final pose information of the i-th frame image. The first interpolation variable is the last interpolation variable obtained before acquiring the i-th frame image, and the first interpolation variable is the interpolation variable between the initial pose information of the first image and the target pose information. The first image is the keyframe image among the images acquired before the i-th frame image, and the target pose information is the pose information optimized from the initial pose information of the first image; and real-time localization and map construction are performed based on the final pose information of the i-th frame image and the i-th frame image. This scheme allows the electronic device to perform real-time localization and map construction based on the i-th frame image, the fused pose information obtained by fusing a first interpolation variable with the initial pose information of the i-th frame image determined from the final pose information of the (i-1)-th frame image. The first interpolation variable is the interpolation variable between the initial pose information of the keyframe images acquired by the electronic device before and after optimization. Therefore, on the one hand, the electronic device can correct the initial pose of the i-th frame image, thereby improving the accuracy of pose tracking; on the other hand, the electronic device only needs to calculate the interpolation variable for keyframe images, thus shortening the pose tracking delay. This ensures that the electronic device outputs high-frequency, high-precision pose information, thereby improving the effectiveness of real-time localization and map construction. Attached Figure Description
[0013] Figure 1 This is a flowchart of the real-time positioning and map building method provided in the embodiments of this application;
[0014] Figure 2 This is a schematic diagram of the real-time positioning and map building device provided in the embodiments of this application;
[0015] Figure 3 This is a schematic diagram of the electronic device provided in the embodiments of this application;
[0016] Figure 4 This is a hardware schematic diagram of the electronic device provided in the embodiments of this application. Detailed Implementation
[0017] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0018] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0019] The following description, in conjunction with the accompanying drawings, details the real-time positioning and map building method, apparatus, electronic device, and readable storage medium provided in this application through specific embodiments and application scenarios.
[0020] Simultaneous localization and mapping (SLAM) refers to the process by which a moving object calculates its own position and simultaneously constructs a map of its environment based on sensor information. Currently, SLAM is applied in robotics, virtual reality, and augmented reality, with uses including sensor localization and subsequent path planning and scene understanding. The mainstream SLAM methods are generally divided into two types: filtering-based and optimization-based. Filtering-based methods use the values of various states from the previous moment to estimate the next moment; while optimization-based methods treat all states as variables, the motion equations and observation equations as constraints between variables, construct an error function, and minimize the quadratic form of this error.
[0021] However, filtering-based methods, such as the Multi-State Constraint Kalman Filter (MSCKF), offer advantages like low power consumption, high output frequency, and good real-time positioning accuracy, making them suitable for mobile terminal applications. However, their theoretical framework is incomplete, lacking map-building functionality. Furthermore, the error in real-time spatial positioning increases with runtime. In practical applications, if the electronic device fails to locate, the entire MSCKF system ceases operation, potentially leading to poor robustness. Optimization-based methods, on the other hand, feature complete system modules (including both real-time positioning and map-building modules), high real-time positioning accuracy, and strong overall system robustness. However, these methods involve high computational cost, resulting in high power consumption and a low output frequency, making them unsuitable for the high output rate and low computational power consumption requirements of mobile terminals. Consequently, the real-time positioning and map-building performance of electronic devices is poor.
[0022] To address the aforementioned issues, in this embodiment, the electronic device can determine the initial pose information of the acquired i-th frame image based on the final pose information of the acquired (i-1)-th frame image, where i is an integer greater than 1. The initial pose information of the i-th frame image is then fused with a first interpolation variable to obtain the final pose information of the i-th frame image. The first interpolation variable is the last interpolation variable obtained before acquiring the i-th frame image, and it is the interpolation variable between the initial pose information of the first image and the target pose information. The first image is a keyframe image among the images acquired before the i-th frame image, and the target pose information is the pose information optimized from the initial pose information of the first image. Based on the final pose information of the i-th frame image and the i-th frame image, real-time localization and map construction are performed. This method improves the accuracy of pose tracking by correcting the initial pose of the i-th frame image and shortens the pose tracking delay by calculating the interpolation variable only for keyframe images. This ensures that electronic devices output high-frequency, high-precision pose information, thereby improving the effectiveness of real-time positioning and map building.
[0023] This application provides a method for real-time positioning and map construction. Figure 1 A flowchart illustrating the real-time positioning and map building method provided in an embodiment of this application is shown. Figure 1 As shown, the real-time positioning and map building method provided in this application embodiment may include the following steps 101 to 103.
[0024] Step 101: The electronic device determines the initial pose information of the i-th frame image based on the final pose information of the (i-1)-th frame image.
[0025] In this embodiment of the application, i is an integer greater than 1.
[0026] Optionally, in this embodiment of the application, during the real-time positioning and map building process, the electronic device can acquire images of the scene in which the electronic device is located in real time through the sensors in the electronic device.
[0027] In this embodiment, the pose information of the image can indicate the position of the image in three-dimensional space.
[0028] Optionally, in this embodiment, the pose information may include rotation coordinates and displacement coordinates.
[0029] For example, pose information T = {R, P}, where R is rotation coordinates, including rotation coordinates centered on the X-axis, Y-axis, and Z-axis in three-dimensional space; P is displacement coordinates, including coordinates on the X-axis, Y-axis, and Z-axis in three-dimensional space.
[0030] The following section details the specific method by which an electronic device determines the initial pose information of the acquired i-th frame image.
[0031] Optionally, in the embodiments of this application, the above step 101 can be implemented by the following step 101a, as well as step A or step B.
[0032] Step 101a: The electronic device uses a filtering algorithm to process the final pose information of the (i-1)th frame image to obtain the first pose information.
[0033] Optionally, in the embodiments of this application, the principle of the filtering algorithm can be:
[0034] x=f(x i-1 )+n
[0035] Where x is the first pose information, x i-1 Let f be the final pose information of the (i-1)th frame image, f be the transition matrix, and n be the noise term.
[0036] It can be seen that the electronic device can calculate the first pose information based on the final pose information of the (i-1)th frame image.
[0037] Optionally, in this embodiment of the application, after obtaining the first pose information, the electronic device can determine the matching status between the first pose information and the final pose information of the (i-1)th frame image, and determine whether to execute step A or step B based on the matching status.
[0038] Step A: When the matching degree between the first pose information and the final pose information of the (i-1)th frame image is less than or equal to the preset matching degree, the electronic device determines the final pose information of the (i-1)th frame image as the initial pose information of the i-th frame image.
[0039] Optionally, in this embodiment of the application, the preset matching degree can be the system default or set by the user according to actual usage needs.
[0040] It is understandable that the matching degree between the first pose information and the final pose information of the (i-1)th frame image is less than or equal to the preset matching degree, that is, the first pose information and the final pose information of the (i-1)th frame image differ too much.
[0041] For example, if the electronic device fails to locate during the process of acquiring images of the current scene, such as when the electronic device acquires images of a white wall for a long time, the calculated pose information will be abnormal due to the accumulation of errors when the electronic device uses the filtering algorithm. That is, the matching degree between the pose information (i.e., the first pose information) and the final pose information of the previous frame image (i.e., the i-1th frame image) is less than or equal to the preset matching degree.
[0042] Optionally, in this embodiment of the application, the electronic device can save the historical information of the entire positioning process. When the matching degree between the first pose information and the final pose information of the (i-1)th frame image is less than or equal to the preset matching degree, the final pose information of the (i-1)th frame image is determined as the initial pose information of the i-th frame image, thereby reducing errors and ensuring the accuracy of the initial pose information of the i-th frame image.
[0043] Step B: When the matching degree between the first pose information and the final pose information of the (i-1)th frame image is greater than the preset matching degree, the electronic device determines the first pose information as the initial pose information of the i-th frame image.
[0044] It can be understood that the matching degree between the first pose information and the final pose information of the (i-1)th frame image is greater than the preset matching degree, that is, the first pose information and the final pose information of the (i-1)th frame image are not much different. At this time, the electronic device can determine the first pose information as the initial pose information of the i-th frame image. That is, the initial pose information of the i-th frame image is the pose information obtained by the electronic device using a filtering algorithm to process the final pose information of the (i-1)th frame image.
[0045] In this embodiment of the application, since the electronic device can determine the initial pose information of the i-th frame image by judging the matching degree between the final pose information of the i-1th frame image and the pose information obtained by processing the final pose information of the i-1th frame image using a filtering algorithm (i.e., the first pose information), the accuracy of the initial pose information of the i-th frame image can be ensured.
[0046] Step 102: The electronic device fuses the initial pose information of the i-th frame image with the first interpolation variable to obtain the final pose information of the i-th frame image.
[0047] In this embodiment, the first interpolation variable is the last interpolation variable obtained before acquiring the i-th frame image.
[0048] It should be noted that the electronic device can obtain an interpolation variable by calculating the pose information of the keyframe image. During the calculation process, the electronic device will not calculate the pose information of the newly acquired keyframe image until the calculation is completed. If the electronic device acquires another keyframe image after the calculation is completed, the electronic device can obtain a new interpolation variable by calculating the pose information of that image.
[0049] In this embodiment of the application, the first interpolation variable is the interpolation variable between the initial pose information and the target pose information of the first image.
[0050] In this embodiment of the application, the target pose information can be the pose information optimized from the initial pose information of the first image.
[0051] In this embodiment of the application, the first image is the keyframe image captured before the i-th frame image.
[0052] Optionally, in the embodiments of this application, the keyframe image captured by the electronic device can be a single frame captured by the electronic device over a long period of time, or it can be an image that includes elements that have not been captured before.
[0053] Optionally, in this embodiment of the application, the electronic device can perform coordinate calculations to fuse the initial pose information of the i-th frame image with the first interpolation variable to obtain the final pose information of the i-th frame image.
[0054] For example, if the initial pose information of the i-th frame image is The first interpolation variable is Electronic devices can then perform coordinate calculations. The initial pose information of the i-th frame image is fused with the first interpolation variable to obtain the final pose information T of the i-th frame image. out ={R out ,P out}
[0055] Optionally, in this embodiment of the application, the electronic device can fuse the initial pose information of the i-th frame image with the first interpolation variable using the Kalman filtering method to obtain the final pose information of the i-th frame image.
[0056] The real-time positioning and map construction method provided in the embodiments of this application will be described by way of example below.
[0057] For example, assuming the a-th frame is a keyframe among images captured before the i-th frame, the electronic device can calculate an interpolation variable between the initial pose information of the a-th frame and the optimized pose information (i.e., the target pose information). If the i-th frame is an image captured after the electronic device calculates the interpolation variable, the electronic device can fuse the initial pose information of the i-th frame with the interpolation variable (i.e., the first interpolation variable) to obtain the final pose information of the i-th frame. If the i-th frame is an image captured before the electronic device calculates the interpolation variable, the electronic device can fuse the initial pose information of the i-th frame with the previously calculated interpolation variable (i.e., the first interpolation variable) to obtain the final pose information of the i-th frame. This ensures the accuracy of the final pose information of the i-th frame.
[0058] The specific methods for determining the first interpolation variable and optimizing the initial pose information of the first image by the electronic device will be described in detail in the following embodiments, and will not be repeated here to avoid repetition.
[0059] Step 103: The electronic device performs real-time localization and map construction based on the final pose information of the i-th frame image and the i-th frame image.
[0060] In this embodiment of the application, the electronic device can perform real-time positioning and map construction based on the final pose information of the i-th frame image and the i-th frame image, so as to construct a three-dimensional map corresponding to the i-th frame image.
[0061] Optionally, in this embodiment of the application, the electronic device can overlay the three-dimensional map corresponding to the i-th frame image with the three-dimensional map corresponding to each frame image constructed before the i-th frame image, thereby enabling real-time positioning and map construction of the current scene.
[0062] It should be noted that if the electronic device does not obtain any interpolation variables before acquiring the i-th frame image, the electronic device can directly perform real-time localization and map construction based on the initial pose information of the i-th frame image and the i-th frame image itself; or, the electronic device can optimize the initial pose information of the i-th frame image, determine a difference variable, and fuse the initial pose information of the i-th frame image with the interpolation variable to obtain the final pose information of the i-th frame image. Thus, the electronic device can perform real-time localization and map construction based on the final pose information of the i-th frame image and the i-th frame image itself.
[0063] In the real-time localization and mapping method provided in this application embodiment, the electronic device can perform real-time localization and mapping based on the i-th frame image, the fused pose information obtained by fusing a first interpolation variable with the initial pose information of the i-th frame image determined based on the final pose information of the (i-1)-th frame image, and the first interpolation variable being the interpolation variable between the initial pose information before and after optimization of the keyframe images acquired by the electronic device before the i-th frame image. Therefore, on the one hand, the electronic device can correct the initial pose of the i-th frame image, thereby improving the accuracy of pose tracking; on the other hand, the electronic device only needs to calculate the interpolation variable for the keyframe images, thereby shortening the pose tracking delay. Thus, it can ensure that the electronic device outputs high-frequency, high-precision pose information, thereby improving the effect of real-time localization and mapping.
[0064] Optionally, in the embodiments of this application, before step 101 above, the real-time positioning and map building method provided in the embodiments of this application may further include the following steps 104 and 105.
[0065] Step 104: The electronic device optimizes the initial pose information of the first image based on M pose information, M sets of offset information and the initial pose information of the first image to obtain the target pose information.
[0066] In this embodiment, the aforementioned M pose information are the pose information of the most recently optimized M-frame images, where the M-frame images are keyframes among the images acquired before the first image.
[0067] In this embodiment of the application, each set of offset information in the above M sets is the offset of the feature points of the first image relative to the feature points of one frame of the above M frames.
[0068] Optionally, in the embodiments of this application, the feature points of the image can be any possible points such as vertices, corners, or center points in the image.
[0069] Optionally, in the embodiments of this application, the number of feature points in the image can be one or more, which can be determined by the electronic device based on the acquired image.
[0070] Optionally, in the embodiments of this application, the number of feature points in different images can be the same or different.
[0071] Optionally, in this embodiment of the application, the number of feature points corresponding to the first image and one of the M frames can be N, where N is an integer greater than or equal to 0; it can be understood that the offset information in this case includes N offsets.
[0072] The specific method for determining the aforementioned M sets of offset information by electronic devices will be described in detail in the following embodiments, and will not be repeated here to avoid repetition.
[0073] The following section details the specific methods for optimizing the initial pose information of the first image using electronic devices.
[0074] Optionally, in the embodiments of this application, step 104 can be implemented by steps 104a and 104b as described below.
[0075] Step 104a: The electronic device determines the three-dimensional position information of the M groups based on the above M groups of offset information.
[0076] In this embodiment of the application, the above M sets of offset information correspond one-to-one with the above M sets of three-dimensional position information. Each set of three-dimensional position information can be used to indicate feature points in a three-dimensional map constructed based on one of the above M frames of images.
[0077] Optionally, in this embodiment of the application, when one set of offset information in the above M sets of offset information includes N offset values, the electronic device can indicate N feature points in the constructed three-dimensional map based on a set of three-dimensional position information determined by the set of offset information.
[0078] Optionally, in this embodiment of the application, for each set of offset information in the above M sets of offset information, the electronic device can determine a set of three-dimensional position information for indicating the feature points in the three-dimensional map constructed based on the frame image, according to the offset of the feature points of the first image relative to the feature points of one of the M frames of images.
[0079] Step 104b: The electronic device uses a preset bundle adjustment algorithm to process the above M pose information, the initial pose information of the first image, and the above M sets of three-dimensional position information to obtain the target pose information.
[0080] Optionally, in the embodiments of this application, the principle of the preset bundle adjustment algorithm can be:
[0081]
[0082] Where T is the initial pose information of the image, P is the three-dimensional position information, Z is the two-dimensional observation, π is the projection equation, M is the number of images that are keyframes before the first image, and N is the number of feature points in the three-dimensional map indicated by the M sets of three-dimensional position information.
[0083] It can be seen that the electronic device can process the above M pose information, the initial pose information of the first image, and the above M sets of three-dimensional position information to obtain the target pose information.
[0084] It should be noted that in the embodiments of this application, the electronic device adopts a preset bundle adjustment algorithm. While obtaining the target pose information, it can optimize the above M pose information so that when the electronic device obtains the target pose information of a new first image using the preset bundle adjustment algorithm, it can use the optimized M pose information to improve the real-time positioning accuracy of the electronic device.
[0085] In this embodiment, since the electronic device can determine the M sets of three-dimensional position information corresponding one-to-one with the M sets of offset information and can indicate the feature points in the constructed three-dimensional map according to the M sets of offset information, and calculate the optimized pose information of the initial pose information of the first image using a preset bundle adjustment algorithm, the accuracy of the electronic device in optimizing the initial pose information of the first image can be improved. Thus, the electronic device can improve the accuracy of map construction when performing real-time positioning and map construction.
[0086] Step 105: The electronic device determines the first interpolation variable based on the first rotation coordinate and the first displacement coordinate in the initial pose information of the first image, and the second rotation coordinate and the second displacement coordinate in the target pose information.
[0087] For a detailed description of rotational and displacement coordinates, please refer to the relevant descriptions in the above embodiments. To avoid repetition, they will not be repeated here.
[0088] The following section provides a detailed explanation of the specific method for determining the first interpolation variable in electronic devices.
[0089] Optionally, in the embodiments of this application, step 105 can be implemented by steps 105a to 105c as described below.
[0090] Step 105a: The electronic device determines the target rotation coordinates based on the first rotation coordinates and the second rotation coordinates.
[0091] Optionally, in this embodiment of the application, the electronic device can perform a multiplication operation on the transpose of the first rotation coordinate and the second rotation coordinate to obtain the target rotation coordinate.
[0092] Step 105b: The electronic device determines the target displacement coordinates based on the target rotation coordinates, the first displacement coordinates, and the second displacement coordinates.
[0093] Optionally, in this embodiment of the application, the electronic device can multiply the target rotation coordinate and the first displacement coordinate to obtain the intermediate displacement coordinate; and subtract the second displacement coordinate from the intermediate displacement coordinate to obtain the target displacement coordinate.
[0094] Step 105c: The electronic device determines the target rotation coordinates and target displacement coordinates as the first interpolation variables.
[0095] The real-time positioning and map construction method provided in the embodiments of this application will be described by way of example below.
[0096] For example, suppose the first rotation coordinate in the initial pose information of the first image is... The first displacement coordinate is The second rotation coordinate in the target pose information of the first image is: The second displacement coordinate is Therefore, the electronic device can determine the target rotation coordinates based on the first rotation coordinate and the second rotation coordinate. The electronic device can determine the target displacement coordinates based on the target rotation coordinates, the first displacement coordinates, and the second displacement coordinates. In this way, electronic devices can rotate the target coordinates. and target displacement coordinates It is determined to be the first interpolation variable.
[0097] In this embodiment, the electronic device can determine the target rotation coordinates based on the first rotation coordinates in the initial pose information of the first image and the second rotation coordinates in the target pose information of the first image, and determine the target displacement coordinates based on the target rotation coordinates, the first displacement coordinates in the initial pose information of the first image, and the second displacement coordinates in the target pose information of the first image. Thus, the electronic device can determine the target rotation coordinates and the target displacement coordinates as first interpolation variables to facilitate further fusion processing.
[0098] In this embodiment, the electronic device can optimize the initial pose information of the first image based on the most recently optimized pose information of the keyframes in the M frames acquired before the first image, the M sets of offset information, and the initial pose information of the first image. It can also determine the first interpolation variable based on the rotation and displacement coordinates in the optimized pose information and the rotation and displacement coordinates in the initial pose information of the first image. That is, the electronic device can determine the first interpolation variable based on the coordinates in the pose information of the first image before and after optimization, thus improving the accuracy of the electronic device in determining the first interpolation variable.
[0099] The following is a detailed explanation of the specific method by which the electronic device determines the aforementioned M sets of offset information.
[0100] Optionally, in this embodiment of the application, before step 104 above, the real-time positioning and map building method provided in this embodiment of the application may further include step 106 below.
[0101] Step 106: The electronic device determines the above-mentioned M sets of offset information based on the two-dimensional position information of the feature points of the first image and the two-dimensional position information of the feature points of the above-mentioned M frame images.
[0102] Optionally, in this embodiment of the application, the two-dimensional position information of the feature points of the first image can be determined by an electronic device through a filtering method.
[0103] Optionally, in this embodiment of the application, the two-dimensional location information can be used to indicate the position of feature points of the first image in the first image.
[0104] Optionally, in this embodiment of the application, the principle by which the electronic device determines the two-dimensional position information of the feature points of the first image is as follows:
[0105] z k =h(x k )+r k
[0106] Among them, z k For the two-dimensional location information (i.e., two-dimensional observation) of the feature points in the first image, x k For the initial pose information of the first image, r k denoted as the noise term, and h as the observation matrix.
[0107] It can be seen that the electronic device can determine the two-dimensional position information of the feature points of the first image by using a filtering method based on the initial pose information of the first image.
[0108] Optionally, in this embodiment, each set of offset information in the M sets of offset information can be the difference between the image coordinates of the feature points of the first image and the feature points of one frame in the M frames of images.
[0109] For example, the offset of feature point A(x1, y1) of the first image relative to feature point A'(x2, y2) of one of the M frames is x = |x1 - x2|, y = |y1 - y2|.
[0110] It should be noted that, in the embodiments of this application, each set of offset information includes the offset of all feature points in the first image corresponding to one of the M frames of images.
[0111] For example, assuming the first image includes feature point 1, feature point 2 and feature point 3, and one frame a in the above M-frame images includes feature point 1' corresponding to feature point 1, feature point 3' corresponding to feature point 3 and feature point 4, then this set of offset information can include the difference between the image coordinates of feature point 1 and feature point 1', and the difference between the image coordinates of feature point 3 and feature point 3'.
[0112] In this embodiment, since the electronic device can determine M sets of offset information based on the two-dimensional position information of the feature points of the first image and the two-dimensional position information of the feature points of the M frame images, the electronic device can obtain the offset information between the first image and each image that was acquired as a key frame before the acquisition of the first image, so that the electronic device can obtain accurate target pose information based on the offset information, thereby improving the accuracy of real-time positioning and map construction of the electronic device.
[0113] The real-time positioning and mapping method provided in this application can be executed by a real-time positioning and mapping device. This application uses an example of a real-time positioning and mapping device executing the method to illustrate the real-time positioning and mapping device provided in this application.
[0114] Combination Figure 2 This application provides a real-time localization and mapping (RTL) device 20, which may include: an acquisition module 21, a determination module 22, a fusion module 23, and a processing module 24. The determination module 22 is used to determine the initial pose information of the i-th frame image acquired by the acquisition module 21 based on the final pose information of the (i-1)-th frame image acquired by the acquisition module 21, where i is an integer greater than 1. The fusion module 23 is used to fuse the initial pose information of the i-th frame image with a first interpolation variable to obtain the final pose information of the i-th frame image. The first interpolation variable is the last interpolation variable obtained before the acquisition module acquires the i-th frame image, and is an interpolation variable between the initial pose information of the first image and the target pose information. The first image is a keyframe image among the images acquired by the acquisition module before the i-th frame image, and the target pose information is the pose information optimized from the initial pose information of the first image. The processing module 24 is used to perform RTL and map construction based on the final pose information of the i-th frame image and the i-th frame image.
[0115] In one possible implementation, the aforementioned real-time localization and mapping device 20 may further include an optimization module. The optimization module may be used to optimize the initial pose information of the first image based on M pose information, M sets of offset information, and the initial pose information of the first image before the determining module 22 determines the initial pose information of the i-th frame image acquired by the acquisition module 21 based on the final pose information of the (i-1)-th frame image acquired by the acquisition module 21, to obtain target pose information. The M pose information refers to the pose information of the most recently optimized M-frame images, and the M-frame images are keyframe images among the images acquired by the acquisition module 21 before the first image. Each set of offset information is the offset of a feature point in the first image relative to a feature point in one of the M-frame images. The determining module 22 may also be used to determine a first interpolation variable based on the first rotation coordinates and the first displacement coordinates in the initial pose information of the first image, and the second rotation coordinates and the second displacement coordinates in the target pose information.
[0116] In one possible implementation, the determining module 22 can be used to determine M sets of 3D position information based on M sets of offset information. The M sets of offset information correspond one-to-one with the M sets of 3D position information, and each set of 3D position information is used to indicate feature points in a 3D map constructed based on one frame of the M frames. The optimization module can be used to process the M pose information, the initial pose information of the first image, and the M sets of 3D position information using a preset bundle adjustment algorithm to obtain the target pose information.
[0117] In one possible implementation, the determining module 22 can also be used to determine M sets of offset information based on the two-dimensional position information of the feature points of the first image and the two-dimensional position information of the feature points of the M frames before the optimization module optimizes the initial pose information of the first image based on M pose information, M sets of offset information and the initial pose information of the first image to obtain the target pose information.
[0118] In one possible implementation, the determining module 22 can be specifically used to determine the target rotation coordinates based on the first rotation coordinates and the second rotation coordinates. The determining module 22 can also be used to determine the target displacement coordinates based on the target rotation coordinates, the first displacement coordinates, and the second displacement coordinates. Furthermore, the determining module 22 can be used to determine the target rotation coordinates and the target displacement coordinates as first interpolation variables.
[0119] In one possible implementation, processing module 24 can be used to process the final pose information of the (i-1)th frame image using a filtering algorithm to obtain the first pose information. Determining module 22 can be used to determine the final pose information of the (i-1)th frame image as the initial pose information of the i-th frame image when the matching degree between the first pose information and the final pose information of the (i-1)th frame image is less than or equal to a preset matching degree. Determining module 22 can also be used to determine the first pose information as the initial pose information of the i-th frame image when the matching degree between the first pose information and the final pose information of the (i-1)th frame image is greater than a preset matching degree.
[0120] In the real-time localization and mapping (RTL) device provided in this application embodiment, since the RTL device can perform RTL and mapping based on the i-th frame image, and the fused pose information obtained by fusing a first interpolation variable with the initial pose information of the i-th frame image determined according to the final pose information of the (i-1)-th frame image, and the first interpolation variable is the interpolation variable between the initial pose information before and after optimization of the keyframe images acquired by the RTL device before the i-th frame image, on the one hand, the RTL device can correct the initial pose of the i-th frame image, thereby improving the accuracy of pose tracking; on the other hand, the RTL device only needs to calculate the interpolation variable for the keyframe images, thereby shortening the pose tracking delay. Thus, it can ensure that the RTL device outputs high-frequency, high-precision pose information, thereby improving the effect of RTL and mapping in RTL.
[0121] For details on the beneficial effects of the various implementation methods in this embodiment, please refer to the beneficial effects of the corresponding implementation methods in the above method embodiments. To avoid repetition, these will not be repeated here.
[0122] The real-time positioning and mapping device in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the device.
[0123] The real-time positioning and mapping device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit it.
[0124] The real-time positioning and mapping device provided in this application embodiment can achieve... Figure 1 The various processes implemented in the method implementation examples will not be described again here to avoid repetition.
[0125] Optionally, such as Figure 3 As shown, this application embodiment also provides an electronic device 300, including a processor 301 and a memory 302. The memory 302 stores a program or instructions that can run on the processor 301. When the program or instructions are executed by the processor 301, they implement the various steps of the above-described real-time positioning and map building method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.
[0126] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.
[0127] Figure 4 A schematic diagram of the hardware structure of an electronic device to implement an embodiment of this application.
[0128] The electronic device 1000 includes, but is not limited to, components such as: radio frequency unit 1001, network module 1002, audio output unit 1003, input unit 1004, sensor 1005, display unit 1006, user input unit 1007, interface unit 1008, memory 1009, and processor 1010.
[0129] Those skilled in the art will understand that the electronic device 1000 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 1010 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 4 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.
[0130] The processor 1010 can be used to determine the initial pose information of the i-th frame image acquired by the sensor 1005 based on the final pose information of the (i-1)-th frame image acquired by the sensor 1005, where i is an integer greater than 1. The processor 1010 can also be used to fuse the initial pose information of the i-th frame image with a first interpolation variable to obtain the final pose information of the i-th frame image. The first interpolation variable is the last interpolation variable obtained before the acquisition module acquires the i-th frame image; it is the interpolation variable between the initial pose information of the first image and the target pose information. The first image is the keyframe image among the images acquired by the acquisition module before the i-th frame image, and the target pose information is the pose information optimized from the initial pose information of the first image. The processor 1010 can also be used to perform real-time localization and map construction based on the final pose information of the i-th frame image and the i-th frame image itself.
[0131] In one possible implementation, the processor 1010 can further be used to optimize the initial pose information of the first image based on M pose information, M sets of offset information, and the initial pose information of the first image before determining the initial pose information of the i-th frame image acquired by the sensor 1005 according to the final pose information of the (i-1)-th frame image acquired by the sensor 1005, to obtain target pose information. The M pose information refers to the pose information of the M frames after the most recent optimization, and the M frames are keyframes among the images acquired by the sensor 1005 before the first image. Each set of offset information is the offset of a feature point of the first image relative to a feature point of one of the M frames. The processor 1010 can also be used to determine a first interpolation variable based on the first rotation coordinates and the first displacement coordinates in the initial pose information of the first image, and the second rotation coordinates and the second displacement coordinates in the target pose information.
[0132] In one possible implementation, the processor 1010 can be used to determine M sets of 3D position information based on M sets of offset information. The M sets of offset information correspond one-to-one with the M sets of 3D position information, and each set of 3D position information is used to indicate feature points in a 3D map constructed based on one frame of the M frames. Specifically, the processor 1010 can be used to process the M pose information, the initial pose information of the first image, and the M sets of 3D position information using a preset bundle adjustment algorithm to obtain the target pose information.
[0133] In one possible implementation, the processor 1010 can also be used to determine the M sets of offset information based on the two-dimensional position information of the feature points of the first image and the two-dimensional position information of the feature points of the M frames before optimizing the initial pose information of the first image based on the M pose information, the M sets of offset information and the initial pose information of the first image to obtain the target pose information.
[0134] In one possible implementation, processor 1010 can be specifically used to determine the target rotation coordinates based on the first rotation coordinates and the second rotation coordinates. Processor 1010 can also be used to determine the target displacement coordinates based on the target rotation coordinates, the first displacement coordinates, and the second displacement coordinates. Furthermore, processor 1010 can be used to determine the target rotation coordinates and the target displacement coordinates as first interpolation variables.
[0135] In one possible implementation, processor 1010 can be used to process the final pose information of the (i-1)th frame image using a filtering algorithm to obtain the first pose information. Specifically, processor 1010 can be used to determine the final pose information of the (i-1)th frame image as the initial pose information of the i-th frame image when the matching degree between the first pose information and the final pose information of the (i-1)th frame image is less than or equal to a preset matching degree. Alternatively, processor 1010 can be used to determine the first pose information as the initial pose information of the i-th frame image when the matching degree between the first pose information and the final pose information of the (i-1)th frame image is greater than a preset matching degree.
[0136] In the electronic device provided in this application embodiment, since the electronic device can perform real-time localization and map construction based on the i-th frame image, and the fused pose information obtained by fusing the first interpolation variable with the initial pose information of the i-th frame image determined according to the final pose information of the (i-1)-th frame image, and since the first interpolation variable is the interpolation variable between the initial pose information before and after optimization of the keyframe images captured by the electronic device before the i-th frame image, the electronic device can correct the initial pose of the i-th frame image, thereby improving the accuracy of pose tracking. Furthermore, the electronic device only needs to calculate the interpolation variable for the keyframe images, thereby shortening the pose tracking delay. Thus, it can ensure that the electronic device outputs high-frequency, high-precision pose information, thereby improving the effect of real-time localization and map construction.
[0137] For details on the beneficial effects of the various implementation methods in this embodiment, please refer to the beneficial effects of the corresponding implementation methods in the above method embodiments. To avoid repetition, these will not be repeated here.
[0138] It should be understood that, in this embodiment, the input unit 1004 may include a graphics processing unit (GPU) 10041 and a microphone 10042. The GPU 10041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 1006 may include a display panel 10061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 1007 includes a touch panel 10071 and at least one of other input devices 10072. The touch panel 10071 is also called a touch screen. The touch panel 10071 may include a touch detection device and a touch controller. Other input devices 10072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.
[0139] The memory 1009 can be used to store software programs and various data. The memory 1009 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 1009 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 1009 in this embodiment includes, but is not limited to, these and any other suitable types of memory.
[0140] The processor 1010 may include one or more processing units; optionally, the processor 1010 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into the processor 1010.
[0141] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described real-time positioning and map building method embodiments and achieve the same technical effects. To avoid repetition, they will not be described again here.
[0142] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0143] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described real-time positioning and map building method embodiments, and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0144] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.
[0145] This application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the above-described real-time positioning and map building method embodiments, and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0146] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0147] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0148] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. A method for real-time positioning and map construction, characterized in that, The method includes: Based on the final pose information of the (i-1)th frame image, determine the initial pose information of the i-th frame image, where i is an integer greater than 1. The initial pose information of the i-th frame image is fused with the first interpolation variable to obtain the final pose information of the i-th frame image. The first interpolation variable is the last interpolation variable obtained before acquiring the i-th frame image. The first interpolation variable is the interpolation variable between the initial pose information of the first image and the target pose information. The first image is the keyframe image among the images acquired before the i-th frame image. The target pose information is the pose information after optimizing the initial pose information of the first image. Based on the final pose information of the i-th frame image and the i-th frame image, real-time localization and map construction are performed; Before determining the initial pose information of the acquired i-th frame image based on the final pose information of the acquired (i-1)-th frame image, the method further includes: Based on M pose information, M sets of offset information and the initial pose information of the first image, the initial pose information of the first image is optimized to obtain the target pose information. The M pose information are the pose information of the most recently optimized M frame images. The M frame images are the key frames in the images acquired before the first image. Each set of offset information is the offset of the feature points of the first image relative to the feature points of one frame in the M frame images. The first interpolation variable is determined based on the first rotation coordinate and the first displacement coordinate in the initial pose information of the first image, and the second rotation coordinate and the second displacement coordinate in the target pose information.
2. The method according to claim 1, characterized in that, The optimization of the initial pose information of the first image based on M pose information, M sets of offset information, and the initial pose information of the first image to obtain the target pose information includes: Based on the M sets of offset information, M sets of three-dimensional position information are determined. The M sets of offset information correspond one-to-one with the M sets of three-dimensional position information. Each set of three-dimensional position information is used to indicate feature points in a three-dimensional map constructed based on one of the M frames of images. The target pose information is obtained by processing the M pose information, the initial pose information of the first image, and the M sets of three-dimensional position information using a preset bundle adjustment algorithm.
3. The method according to claim 2, characterized in that, Before optimizing the initial pose information of the first image based on M pose information, M sets of offset information, and the initial pose information of the first image to obtain the target pose information, the method further includes: The M sets of offset information are determined based on the two-dimensional position information of the feature points of the first image and the two-dimensional position information of the feature points of the M frames.
4. The method according to any one of claims 1 to 3, characterized in that, The step of determining the first interpolation variable based on the first rotation coordinates and first displacement coordinates in the initial pose information of the first image, and the second rotation coordinates and second displacement coordinates in the target pose information, includes: Determine the target rotation coordinates based on the first rotation coordinates and the second rotation coordinates; The target displacement coordinates are determined based on the target rotation coordinates, the first displacement coordinates, and the second displacement coordinates. The target rotation coordinates and the target displacement coordinates are determined as the first interpolation variables.
5. The method according to claim 1, characterized in that, The step of determining the initial pose information of the acquired i-th frame image based on the final pose information of the acquired (i-1)-th frame image includes: A filtering algorithm is used to process the final pose information of the (i-1)th frame image to obtain the first pose information; If the matching degree between the first pose information and the final pose information of the (i-1)th frame image is less than or equal to a preset matching degree, the final pose information of the (i-1)th frame image is determined as the initial pose information of the i-th frame image. If the matching degree between the first pose information and the final pose information of the (i-1)th frame image is greater than a preset matching degree, the first pose information is determined as the initial pose information of the i-th frame image.
6. A real-time positioning and map-building device, characterized in that, The device includes an acquisition module, a determination module, a fusion module, and a processing module; The determining module is used to determine the initial pose information of the i-th frame image acquired by the acquisition module based on the final pose information of the (i-1)-th frame image acquired by the acquisition module, where i is an integer greater than 1. The fusion module is used to fuse the initial pose information of the i-th frame image with a first interpolation variable to obtain the final pose information of the i-th frame image. The first interpolation variable is the last interpolation variable obtained before the acquisition module acquires the i-th frame image. The first interpolation variable is the interpolation variable between the initial pose information of the first image and the target pose information. The first image is the keyframe image among the images acquired by the acquisition module before the i-th frame image. The target pose information is the pose information after optimizing the initial pose information of the first image. The processing module is used to perform real-time localization and map construction based on the final pose information of the i-th frame image and the i-th frame image; The device also includes an optimization module; The optimization module is used to optimize the initial pose information of the first image based on M pose information, M sets of offset information, and the initial pose information of the first image before the determining module determines the initial pose information of the i-th frame image acquired by the acquisition module based on the final pose information of the (i-1)-th frame image acquired by the acquisition module, to obtain the target pose information. The M pose information are the pose information of the most recently optimized M-frame images. The M-frame images are the images that are keyframes in the images acquired by the acquisition module before the first image. Each set of offset information is the offset of the feature points of the first image relative to the feature points of one frame in the M-frame images. The determining module is further configured to determine the first interpolation variable based on the first rotation coordinates and the first displacement coordinates in the initial pose information of the first image, and the second rotation coordinates and the second displacement coordinates in the target pose information.
7. An electronic device, characterized in that, It includes a processor and a memory, the memory storing programs or instructions that can run on the processor, the programs or instructions being executed by the processor to implement the steps of the instantaneous positioning and mapping method as described in any one of claims 1-5.
8. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the real-time positioning and mapping method as described in any one of claims 1-5.
Citation Information
Patent Citations
Global map positioning method and device based on vision, storage medium and device
CN110246182A
Simultaneous localization and map construction method and device, electronic equipment and storage medium
CN112967340A