Pose estimation methods, devices, terminals, and storage media
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-01
- Publication Date
- 2026-08-14
AI Technical Summary
[0004]但是,上述三种位姿估计方法均有不同程度的精度损失
[0057]本公开的实施例提供的技术方案可以包括以下有益效果:该方法中,不再仅仅根据深度摄像模组确定位姿估计数据,而是引入了第一位姿估计数据,然后根据由深度摄像模组确定的第一位姿数据和第二位姿数据,以及引入的第一位姿估计数据,确定终端的目标位姿估计数据,以准确地反映终端由第一时刻至第二时刻的位姿变化情况,提升位姿估计的准确性,提升用户使用体验。
Smart Images

Figure CN116091591B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of terminal technology, and in particular to a pose estimation method, apparatus, terminal and storage medium. Background Technology
[0002] Pose estimation has wide applications in visual navigation and visual positioning.
[0003] Currently, there are three main methods for pose estimation. First, pose estimation is based on color cameras (RGB cameras). Second, pose estimation is based on depth cameras. Third, pose estimation is based on motion sensors (such as accelerometers and gyroscopes).
[0004] However, all three pose estimation methods mentioned above suffer from varying degrees of accuracy loss. Summary of the Invention
[0005] To overcome the problems existing in related technologies, this disclosure provides a pose estimation method, apparatus, terminal and storage medium.
[0006] According to a first aspect of the present disclosure, a pose estimation method is provided, applied to a terminal, the method comprising:
[0007] The first pose estimation data, the first pose data at the first moment, and the second pose data at the second moment are determined. The first pose data and the second pose data represent the pose determined by the depth camera module of the terminal. The first pose estimation data represents the pose change of the terminal from the first moment to the second moment.
[0008] Based on the first pose data, the second pose data, and the first pose estimation data, the target pose estimation data is determined.
[0009] Optionally, determining the target pose estimation data based on the first pose data, the second pose data, and the first pose estimation data includes:
[0010] The third pose data is determined based on the first pose data and the first pose estimation data;
[0011] If it is determined that the third pose data and the second pose data satisfy the first similarity condition, then the first pose estimation data is determined as the target pose estimation data.
[0012] Optionally, the first pose estimation data represents the pose change determined based on the color camera module of the terminal, and the step of determining the target pose estimation data based on the first pose data, the second pose data, and the first pose estimation data includes:
[0013] The third pose data is determined based on the first pose data and the first pose estimation data;
[0014] Determine second pose estimation data from the first time point to the second time point. The second pose estimation data represents the pose change determined based on the non-visual sensors of the terminal.
[0015] If it is determined that the third pose data and the second pose data do not meet the first similarity condition, then the target pose estimation data is determined based on the first pose data, the second pose data, and the second pose estimation data.
[0016] Optionally, determining the target pose estimation data based on the first pose data, the second pose data, and the second pose estimation data includes:
[0017] Based on the first pose data and the second pose estimation data, the fourth pose data is determined;
[0018] If it is determined that the fourth pose data and the second pose data satisfy the second similarity condition, then the second pose estimation data is determined as the target pose estimation data.
[0019] Optionally, determining the target pose estimation data based on the first pose data, the second pose data, and the second pose estimation data includes:
[0020] Based on the first pose data and the second pose estimation data, the fourth pose data is determined;
[0021] If it is determined that the fourth pose data and the second pose data do not meet the second similarity condition, then it is determined whether the number of iterations meets the set threshold.
[0022] If it is determined that the number of iterations meets the set threshold, then the target pose estimation data is determined based on the first pose data and the second pose data.
[0023] Optionally, determining the first pose estimation data, the first pose data at the first moment, and the second pose data at the second moment includes:
[0024] If the set conditions are met, then the first pose estimation data, the first pose data, and the second pose data are determined based on the feature point set and the set calibration information.
[0025] The determination that the set conditions are met includes at least one of the following methods:
[0026] Method 1: Determine the feature point set based on the first color image at the first time point and the second color image at the second time point;
[0027] Method 2: Determine the feature point set based on the first depth image at the first time point and the second depth image at the second time point;
[0028] Method 3: If it is determined that the number of iterations does not meet the set threshold, then the feature point set and the number of iterations are adjusted.
[0029] According to a second aspect of the present disclosure, a pose estimation apparatus is provided for use in a terminal, the apparatus comprising:
[0030] The determination module is used to determine the first pose estimation data, the first pose data at the first moment, and the second pose data at the second moment. The first pose data and the second pose data represent the pose determined based on the depth camera module of the terminal. The first pose estimation data represents the pose change of the terminal from the first moment to the second moment.
[0031] Used to determine target pose estimation data based on the first pose data, the second pose data, and the first pose estimation data.
[0032] Optionally, the determining module is configured to:
[0033] The third pose data is determined based on the first pose data and the first pose estimation data;
[0034] If it is determined that the third pose data and the second pose data satisfy the first similarity condition, then the first pose estimation data is determined as the target pose estimation data.
[0035] Optionally, the first pose estimation data characterizes the pose changes determined based on the color camera module of the terminal, and the determining module is used to:
[0036] The third pose data is determined based on the first pose data and the first pose estimation data;
[0037] Determine second pose estimation data from the first time point to the second time point. The second pose estimation data represents the pose change determined based on the non-visual sensors of the terminal.
[0038] If it is determined that the third pose data and the second pose data do not meet the first similarity condition, then the target pose estimation data is determined based on the first pose data, the second pose data, and the second pose estimation data.
[0039] Optionally, the determining module is configured to:
[0040] Based on the first pose data and the second pose estimation data, the fourth pose data is determined;
[0041] If it is determined that the fourth pose data and the second pose data satisfy the second similarity condition, then the second pose estimation data is determined as the target pose estimation data.
[0042] Optionally, the determining module is configured to:
[0043] Based on the first pose data and the second pose estimation data, the fourth pose data is determined;
[0044] If it is determined that the fourth pose data and the second pose data do not meet the second similarity condition, then it is determined whether the number of iterations meets the set threshold.
[0045] If it is determined that the number of iterations meets the set threshold, then the target pose estimation data is determined based on the first pose data and the second pose data.
[0046] Optionally, the determining module is configured to:
[0047] If the set conditions are met, then the first pose estimation data, the first pose data, and the second pose data are determined based on the feature point set and the set calibration information.
[0048] The determination that the set conditions are met includes at least one of the following methods:
[0049] Method 1: Determine the feature point set based on the first color image at the first time point and the second color image at the second time point;
[0050] Method 2: Determine the feature point set based on the first depth image at the first time point and the second depth image at the second time point;
[0051] Method 3: If it is determined that the number of iterations does not meet the set threshold, then the feature point set and the number of iterations are adjusted.
[0052] According to a third aspect of the present disclosure, a terminal is provided, the terminal comprising:
[0053] processor;
[0054] Memory used to store the processor's executable instructions;
[0055] The processor is configured to perform the method as described in any one of the first aspects.
[0056] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium is provided, wherein when instructions in the storage medium are executed by a processor of a terminal, the terminal is enabled to perform the method as described in any of the first aspects.
[0057] The technical solutions provided by the embodiments of this disclosure may include the following beneficial effects: In this method, the pose estimation data is no longer determined solely based on the depth camera module, but a first pose estimation data is introduced. Then, based on the first pose data and the second pose data determined by the depth camera module, as well as the introduced first pose estimation data, the target pose estimation data of the terminal is determined to accurately reflect the pose change of the terminal from the first moment to the second moment, thereby improving the accuracy of pose estimation and enhancing the user experience.
[0058] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0059] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.
[0060] Figure 1 This is a flowchart illustrating a pose estimation method according to an exemplary embodiment.
[0061] Figure 2 This is a flowchart illustrating a pose estimation method according to an exemplary embodiment.
[0062] Figure 3 This is a flowchart illustrating a pose estimation method according to an exemplary embodiment.
[0063] Figure 4 This is a flowchart illustrating a pose estimation method according to an exemplary embodiment.
[0064] Figure 5 This is a flowchart illustrating a pose estimation method according to an exemplary embodiment.
[0065] Figure 6 This is a flowchart illustrating a pose estimation method according to an exemplary embodiment.
[0066] Figure 7 This is a flowchart illustrating a pose estimation method according to an exemplary embodiment.
[0067] Figure 8 This is a block diagram illustrating a pose estimation apparatus according to an exemplary embodiment.
[0068] Figure 9 This is a block diagram of a terminal according to an exemplary embodiment. Detailed Implementation
[0069] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0070] This disclosure provides a pose estimation method applied to a terminal. Instead of solely relying on a depth camera module to determine pose estimation data, this method introduces a first pose estimation data set. Then, based on the first and second pose data determined by the depth camera module, along with the introduced first pose estimation data, the target pose estimation data of the terminal is determined. This accurately reflects the pose change of the terminal from a first moment to a second moment, improving the accuracy of pose estimation and enhancing the user experience.
[0071] When this method is applied to positioning scenarios such as visual navigation and visual positioning, the determined target pose estimation data can more accurately reflect the pose change of the terminal from the first moment to the second moment. Therefore, it can improve the positioning accuracy in positioning scenarios such as visual navigation and visual positioning, and enhance the user experience.
[0072] In one exemplary embodiment, a pose estimation method is provided, applied to a terminal. (Reference) Figure 1 As shown, the method includes:
[0073] S110, determine the first pose estimation data, the first pose data at the first moment, and the second pose data at the second moment;
[0074] S120, determine the target pose estimation data based on the first pose data, the second pose data, and the first pose estimation data.
[0075] In step S110, the first pose data and the second pose data represent the pose determined by the terminal's depth camera module (e.g., a depth camera). Specifically, the pose data of the terminal at the first moment determined by the depth camera module is denoted as the first pose data, representing the pose of the terminal at the first moment as determined by the depth camera module. The pose data of the terminal at the second moment determined by the depth camera module is denoted as the second pose data, representing the pose of the terminal at the second moment as determined by the depth camera module.
[0076] Example 1,
[0077] At the first moment, the terminal's depth camera captures the first depth image, which can be simply referred to as Depth_1.
[0078] At the second moment, the terminal's depth camera captures a second depth image, which can be simply referred to as Depth_2.
[0079] After the terminal's processor obtains Depth_1 and Depth_2, it can perform feature point detection and matching to determine the matching feature points in Depth_1 and Depth_2, and determine all matching feature points as a feature point set.
[0080] Feature points are representative points in an image that remain unchanged when the image changes, such as through rotation or scaling. A feature point consists of a keypoint and a discriptor. A keypoint indicates the feature point's location in the image, while the discriptor is typically a vector describing the pixel information surrounding the keypoint.
[0081] Then, based on Depth_1 and the set of feature points, the first pose data of the terminal at the first moment can be determined. The first pose data includes rotation data and translation data. The rotation data can be simply referred to as R_depth_1, and the translation data can be simply referred to as T_depth_1.
[0082] Based on Depth_2 and the set of feature points, the second pose data of the terminal at the second time point can be determined. The second pose data includes rotation data and translation data. The rotation data can be simply referred to as R_depth_2 and the translation data can be simply referred to as T_depth_2.
[0083] The first pose estimation data represents the pose change of the terminal from the first time step to the second time step.
[0084] The first pose estimation data can be determined based on the terminal's color camera module (e.g., an RGB camera, i.e., a color camera). Alternatively, the first pose estimation data can be determined based on the terminal's non-visual sensors (e.g., motion sensors such as accelerometers and gyroscopes).
[0085] It should be noted that in this step, the first pose estimation data is not determined based on the depth camera module. Apart from that, as long as the first pose estimation data can characterize the pose change of the terminal from the first moment to the second moment, there are no restrictions on the method for determining the first pose estimation data.
[0086] Example 2,
[0087] The first pose estimation data is determined based on the terminal's color camera.
[0088] At the first moment, the terminal's color camera captures the first color image, which can be simply referred to as Image_1; the terminal's depth camera captures the first depth image, which can be simply referred to as Depth_1.
[0089] At the second moment, the terminal's depth camera captures a second depth image, which can be simply referred to as Image_2; the terminal's depth camera captures a second depth image, which can be simply referred to as Depth_2.
[0090] After the terminal's processor acquires Image_1 and Image_2, it can perform feature point detection and matching to determine the matching feature points in Image_1 and Image_2, and define all matching feature points as the first feature point set. Then, pose estimation is performed to obtain the homography matrix H.
[0091] Homography transformation can be simply understood as describing the positional mapping between an object in the world coordinate system and the pixel coordinate system. The corresponding transformation matrix is called the homography matrix.
[0092] Then, based on Image_1, Image_2, and the homography matrix H, the first pose estimation data of the terminal from the first time to the second time can be determined. The first pose estimation data may include rotation estimation data and translation estimation data. The rotation estimation data can be simply referred to as R_rgb, and the translation estimation data can be simply referred to as T_rgb.
[0093] After determining the first set of feature points, the second set of feature points, Depth_1 and Depth_2, are determined based on the offline calibration information from the color camera and depth camera. The offline calibration information can be determined using Zhang Zhengyou's stereo calibration method. This offline calibration information can be set before or after the terminal leaves the factory, and it can be modified afterward to better meet user needs.
[0094] Then, based on Depth_1 and the second set of feature points, the first pose data of the terminal at the first moment can be determined. The first pose data includes rotation data and translation data. The rotation data can be simply referred to as R_depth_1 and the translation data can be simply referred to as T_depth_1.
[0095] Based on Depth_2 and the second set of feature points, the second pose data of the terminal at the second time point can be determined. The second pose data includes rotation data and translation data. The rotation data can be simply referred to as R_depth_2 and the translation data can be simply referred to as T_depth_2.
[0096] Example 3,
[0097] The first pose estimation data is determined based on the terminal's color camera.
[0098] At the first moment, the terminal's color camera captures the first color image, which can be simply referred to as Image_1; the terminal's depth camera captures the first depth image, which can be simply referred to as Depth_1.
[0099] At the second moment, the terminal's depth camera captures a second depth image, which can be simply referred to as Image_2; the terminal's depth camera captures a second depth image, which can be simply referred to as Depth_2.
[0100] After the terminal's processor obtains Depth_1 and Depth_2, it can perform feature point detection and matching to determine the matching feature points in Depth_1 and Depth_2, and determine all matching feature points as the second feature point set.
[0101] After determining the second set of feature points, the first set of feature points for Image_1 and Image_2 is determined based on the offline calibration information of the color camera and the depth camera.
[0102] Then, pose estimation is performed based on Image_1 and Image_2 to obtain the homography matrix H.
[0103] Homography transformation can be simply understood as describing the positional mapping between an object in the world coordinate system and the pixel coordinate system. The corresponding transformation matrix is called the homography matrix.
[0104] Then, based on Image_1, Image_2, and the homography matrix H, the first pose estimation data of the terminal from the first time to the second time can be determined. The first pose estimation data may include rotation estimation data and translation estimation data. The rotation estimation data can be simply referred to as R_rgb, and the translation estimation data can be simply referred to as T_rgb.
[0105] After determining the second set of feature points, the first pose data of the terminal at the first moment can be determined based on Depth_1 and the second set of feature points. The first pose data includes rotation data and translation data. The rotation data can be simply referred to as R_depth_1 and the translation data can be simply referred to as T_depth_1.
[0106] Based on Depth_2 and the second set of feature points, the second pose data of the terminal at the second time point can be determined. The second pose data includes rotation data and translation data. The rotation data can be simply referred to as R_depth_2 and the translation data can be simply referred to as T_depth_2.
[0107] Example 4,
[0108] The first pose estimation data is determined based on the terminal's accelerometer and gyroscope, which are non-visual sensors.
[0109] It should be noted that, generally speaking, the sampling frequency of accelerometers and gyroscopes is much higher than the frame rate of cameras (such as depth cameras and color cameras).
[0110] In this example, the first pose estimation data can be determined based on the non-visual detection data detected by the non-visual sensor.
[0111] The accelerometer can collect multiple translation data from the first moment to the second moment. Then, the terminal processes all the above translation data to determine the translation estimation data of the terminal from the first moment to the second moment. This translation estimation data can be simply referred to as T'.
[0112] The gyroscope can collect multiple rotation data from the first moment to the second moment. Then the terminal processes all the above rotation data to determine the rotation estimate data of the terminal from the first moment to the second moment. This rotation estimate data can be simply referred to as α'.
[0113] The above T' and α' together constitute the first pose estimation data.
[0114] In step S120, the first pose data and the second pose data are determined based on the depth camera module, while the first pose estimation data is not determined based on the depth camera module. Therefore, the accuracy of the first pose estimation data can be determined based on the first pose data and the second pose data, thereby determining the final target pose estimation data. After determining the target pose estimation data, the pose estimation method process can be terminated.
[0115] Example 5,
[0116] The first pose estimation data is determined based on the terminal's color camera. The first pose estimation data includes rotation estimation data and translation estimation data, where rotation estimation data can be abbreviated as R_rgb and translation estimation data can be abbreviated as T_rgb.
[0117] The first pose data includes rotation data and translation data. The rotation data can be simply referred to as R_depth_1, and the translation data can be simply referred to as T_depth_1.
[0118] The second pose data includes rotation data and translation data. The rotation data can be simply referred to as R_depth_2, and the translation data can be simply referred to as T_depth_2.
[0119] In this example, the third pose estimation data can be determined based on the first pose data and the second pose data. The third pose estimation data may include rotation estimation data and translation estimation data, where the rotation estimation data can be simply referred to as R_depth and the translation estimation data can be simply referred to as T_depth.
[0120] R_depth and T_depth can be obtained based on the formulas for determining the rotation estimation data and the translation estimation data. The formula for determining the rotation estimation data is: The formula for determining the translation estimation data is: When determining R_depth and T_depth, R_depth_1, R_depth_2, T_depth_1, T_depth_2, R_depth and T_depth can be substituted into the two formulas above, and used as r1, r2, t1, t2, R and T in the formulas in turn.
[0121] Then, it is determined whether the first pose estimation data and the third pose estimation data meet the similarity condition. This similarity condition can be denoted as the third similarity condition. If the similarity condition is met, the first pose estimation data can be determined as the target pose estimation data; if the similarity condition is not met, the third pose estimation data can be determined as the target pose estimation data to improve the accuracy of pose estimation.
[0122] Only if R_rgb and R_depth both satisfy the similarity condition, and T_rgb and T_depth also satisfy the similarity condition, are the first pose estimation data and the third pose estimation data considered to satisfy the similarity condition. Otherwise, the first pose estimation data and the third pose estimation data are considered not to satisfy the similarity condition.
[0123] To determine whether R_rgb and R_depth meet the similarity criteria, the difference between R_rgb and R_depth is first determined. This difference is then divided by R_depth to obtain a ratio. If this ratio is less than a difference threshold (e.g., 10%), R_rgb and R_depth are considered to meet the similarity criteria. Otherwise, R_rgb and R_depth are considered not to meet the similarity criteria.
[0124] The difference threshold can be set before or after the terminal leaves the factory. Furthermore, the difference threshold can be modified after it is set to better meet the needs of users.
[0125] The method for determining whether T_rgb and T_depth meet the similarity criteria is similar to the method for determining whether R_rgb and R_depth meet the similarity criteria, and will not be repeated here.
[0126] It should be noted that, in addition to the methods for determining similarity conditions mentioned above, other methods can also be used to determine similarity conditions, and there are no restrictions on these methods. For example, a neural network-based similarity determination model can be used to determine similarity.
[0127] Example 6,
[0128] The difference between Example 6 and Example 5 above is that in Example 5, the first pose estimation data is determined based on the terminal's color camera. The first pose estimation data includes rotation estimation data and translation estimation data, where the rotation estimation data can be simply referred to as α' and the translation estimation data can be simply referred to as T'.
[0129] The method is similar to that in Example 5, and will not be repeated here.
[0130] It should be noted that, in addition to the above-mentioned method of determining the target pose estimation data based on the first pose data, the second pose data, and the first pose estimation data, other methods can also be used to determine the target pose estimation data based on the first pose data, the second pose data, and the first pose estimation data, and no restrictions are imposed here.
[0131] In this method, the pose estimation data is no longer determined solely by the depth camera module. Instead, a first pose estimation data is introduced. Then, based on the first and second pose data determined by the depth camera module, as well as the introduced first pose estimation data, the target pose estimation data of the terminal is determined to accurately reflect the pose change of the terminal from the first moment to the second moment, thereby improving the accuracy of pose estimation and enhancing the user experience.
[0132] When this method is applied to positioning scenarios such as visual navigation and visual positioning, the determined target pose estimation data can more accurately reflect the pose change of the terminal from the first moment to the second moment. Therefore, it can improve the positioning accuracy in positioning scenarios such as visual navigation and visual positioning, and enhance the user experience.
[0133] In one exemplary embodiment, a pose estimation method is provided, applied to a terminal. (Reference) Figure 2 As shown, in this method, determining the target pose estimation data based on the first pose data, the second pose data, and the first pose estimation data may include:
[0134] S210, determine the third pose data based on the first pose data and the first pose estimation data;
[0135] S220, determine whether the third pose data and the second pose data satisfy the first similarity condition; if the determination result is yes, then proceed to step S230; otherwise, proceed to step S240.
[0136] S230, the first pose estimation data is determined as the target pose estimation data;
[0137] S240, determine the target pose estimation data based on the first pose data and the second pose data.
[0138] In step S210, the third pose data may include rotation data and translation data. The third pose data may be determined based on the formula for determining the rotation estimation data and the formula for determining the translation estimation data.
[0139] Example 1,
[0140] The first pose data includes rotation data and translation data. The rotation data can be simply referred to as R_depth_1, and the translation data can be simply referred to as T_depth_1.
[0141] The first pose estimation data can be determined based on the terminal's color camera. The first pose estimation data includes rotation estimation data and translation estimation data, where rotation estimation data can be abbreviated as R_rgb and translation estimation data can be abbreviated as T_rgb.
[0142] The third pose data may include rotation data and translation data, where rotation data can be abbreviated as R2' and translation data can be abbreviated as T2'.
[0143] The third pose data can be obtained based on the formulas for determining the rotation estimation data and the translation estimation data. The formula for determining the rotation estimation data is: The formula for determining the translation estimation data is: Based on the two formulas above, we can obtain: r2 = R·r1,
[0144] By substituting R_depth_1, R2', T_depth_1, T2', R_rgb, and T_rgb into the newly determined formula, and using them as r1, r2, t1, t2, R, and T in the formula, the final R2' and T2' can be determined, thus determining the third pose data.
[0145] Example 2,
[0146] The difference between Example 2 and Example 1 above is that in Example 2, the first pose estimation data includes rotation estimation data and translation estimation data, where the rotation estimation data can be simply referred to as α' and the translation estimation data can be simply referred to as T'; the third pose data can include rotation data and translation data, where the rotation data can be simply referred to as R2" and the translation data can be simply referred to as T2".
[0147] The remaining methods in Example 2 are similar to those in Example 1 above. When determining the third pose data, R_depth_1, α', T_depth_1, T', R2”, and T2” can be substituted into the newly determined formula, which will then be used as r1, r2, t1, t2, R, and T in the formula, respectively. This will determine the final R2” and T2” to define the third pose data.
[0148] In steps S220 to S240, if the third pose data and the second pose data meet the similarity condition, the first pose estimation data can be determined as the target pose estimation data; if the third pose data and the second pose data do not meet the similarity condition, the target pose estimation data can be determined based on the first pose data and the second pose data, that is, the third pose estimation data can be determined based on the first pose data and the second pose data, and the third pose estimation data is determined as the target pose estimation data to improve the accuracy of pose estimation.
[0149] The aforementioned similarity condition can be referred to as the first similarity condition. The third pose estimation data may include rotation estimation data and translation estimation data, where rotation estimation data can be abbreviated as R_depth and translation estimation data can be abbreviated as T_depth.
[0150] Example 3,
[0151] Example 3 is the same as Example 1 above. In Example 3, the third pose estimation data may include rotation estimation data and translation estimation data, where the rotation estimation data may be simply referred to as R_depth and the translation estimation data may be simply referred to as T_depth.
[0152] Only if R2' and R_depth_2 both satisfy the first similarity condition, and T2' and T_depth_2 also satisfy the first similarity condition, are the third pose data and the second pose data considered to satisfy the first similarity condition. Otherwise, the third pose data and the second pose data are considered not to satisfy the first similarity condition.
[0153] In determining whether R2' and R_depth_2 meet the similarity criteria, the difference between R2' and R_depth_2 is first determined. Then, this difference is divided by R_depth_2 to obtain a ratio. If this ratio is less than a first difference threshold (e.g., 10%), then R2' and R_depth_2 are considered to meet the first similarity criteria. Otherwise, R2' and R_depth_2 are considered not to meet the first similarity criteria.
[0154] The first difference threshold can be set before the terminal leaves the factory or after the terminal leaves the factory. Furthermore, after the first difference threshold is set, it can be modified later to better meet the needs of users.
[0155] The method for determining whether T2' and T_depth_2 satisfy the first similarity condition is similar to the method for determining whether R2' and R_depth_2 satisfy the similarity condition, and will not be elaborated here.
[0156] In Example 3, if it is determined that R2' and R_depth_2 satisfy the first similarity condition, and T2' and T_depth_2 also satisfy the first similarity condition, then the first pose estimation data can be determined as the target pose estimation data, that is, R_rgb and T_rgb can be determined as the target pose estimation data.
[0157] If it is determined that R2' and R_depth_2 do not meet the first similarity condition, or T2' and T_depth_2 do not meet the first similarity condition, or R2' and R_depth_2 do not meet the first similarity condition and T2' and T_depth_2 do not meet the first similarity condition, then the third pose estimation data can be determined as the target pose estimation data, that is, R_depth and T_depth can be determined as the target pose estimation data.
[0158] It should be noted that, in addition to the method described above for determining the first similarity condition, other methods can also be used, and there are no restrictions on this. For example, a neural network-based similarity judgment model can be used to determine similarity.
[0159] Example 4,
[0160] The difference between Example 4 and Example 3 above is that in Example 4, the first pose estimation data includes rotation estimation data and translation estimation data, where the rotation estimation data can be simply referred to as α' and the translation estimation data can be simply referred to as T'; the third pose data can include rotation data and translation data, where the rotation data can be simply referred to as R2" and the translation data can be simply referred to as T2".
[0161] The remaining methods in Example 4 are the same as those in Examples 2 and 3 above, and will not be repeated here.
[0162] It should be noted that in this method, the entire pose estimation process can be terminated after the target pose estimation data is determined.
[0163] This method no longer relies solely on the depth camera module to determine pose estimation data. Instead, it introduces a first pose estimation data set. Then, based on the first pose data determined by the depth camera module and the introduced first pose estimation data, the terminal's pose data at the second moment, i.e., the third pose data, is estimated. Finally, based on whether the third pose data and the second pose data satisfy a first similarity condition, the target pose estimation data for the terminal is selected. This accurately reflects the pose change of the terminal from the first moment to the second moment, improving the accuracy of pose estimation and enhancing the user experience.
[0164] In one exemplary embodiment, a pose estimation method is provided, applied to a terminal. In this method, the first pose estimation data can characterize the pose change determined based on the color camera module of the terminal. That is, the first sub-estimation data is determined based on the color camera module. The first pose estimation data includes rotation estimation data and translation estimation data, wherein the rotation estimation data can be abbreviated as R_rgb, and the translation estimation data can be abbreviated as T_rgb.
[0165] refer to Figure 3 As shown, in this method, determining the target pose estimation data based on the first pose data, the second pose data, and the first pose estimation data may include:
[0166] S310, determine the third pose data based on the first pose data and the first pose estimation data;
[0167] S320, determine whether the third pose data and the second pose data satisfy the first similarity condition; if the determination result is yes, then proceed to step S330; otherwise, proceed to step S340.
[0168] S330, the first pose estimation data is determined as the target pose estimation data;
[0169] S340, determine the second pose estimation data from the first time point to the second time point;
[0170] S350, determine the target pose estimation data based on the first pose data, the second pose data, and the second pose estimation data.
[0171] In this embodiment, steps S310 to S330 can be referred to sequentially as steps S210 to S230 in the above embodiment, and will not be repeated here. Furthermore, in this method, the entire pose estimation process can end after the target pose estimation data is determined. That is, the process can end after step S330 or step S350 determines the target pose estimation process.
[0172] In step S340, the second pose estimation data represents the pose change determined based on the terminal's non-visual sensors. That is, the second pose estimation data is determined based on non-visual sensors. The second pose estimation data may include rotation estimation data and translation estimation data, wherein the rotation estimation data may be simply referred to as α', and the translation estimation data may be simply referred to as T'.
[0173] In step S350, the first pose data and the second pose data are determined based on the depth camera module, while the second pose estimation data is determined based on a non-visual sensor. Therefore, if the third pose data and the second pose data do not meet the first similarity condition, it indicates that the accuracy of the first pose estimation data is poor and cannot be determined as the target pose estimation data. Then, based on the first pose data and the second pose data, the accuracy of the second pose estimation data can be determined, and the final target pose estimation data can be determined to further improve the accuracy of the target pose estimation data.
[0174] The method for determining the accuracy of the second pose estimation data can refer to the method for determining the accuracy of the first pose estimation data in the above embodiments.
[0175] Example 1,
[0176] The second pose estimation data may include rotation estimation data and translation estimation data, where rotation estimation data can be simply referred to as α' and translation estimation data can be simply referred to as T'.
[0177] The first pose data includes rotation data and translation data. The rotation data can be simply referred to as R_depth_1, and the translation data can be simply referred to as T_depth_1.
[0178] The second pose data includes rotation data and translation data. The rotation data can be simply referred to as R_depth_2, and the translation data can be simply referred to as T_depth_2.
[0179] In this example, the third pose estimation data can be determined based on the first pose data and the second pose data. The third pose estimation data may include rotation estimation data and translation estimation data, where the rotation estimation data can be simply referred to as R_depth and the translation estimation data can be simply referred to as T_depth.
[0180] Then, it is determined whether the second pose estimation data and the third pose estimation data meet the similarity condition. This similarity condition can be denoted as the fourth similarity condition. If the similarity condition is met, the second pose estimation data can be determined as the target pose estimation data; if the similarity condition is not met, the third pose estimation data can be determined as the target pose estimation data to improve the accuracy of pose estimation.
[0181] Only when α' and R_depth satisfy the similarity condition, and T' and T_depth also satisfy the similarity condition, are the second pose estimation data and the third pose estimation data considered to satisfy the similarity condition. Otherwise, the second pose estimation data and the third pose estimation data are considered not to satisfy the similarity condition.
[0182] In determining whether α' and R_depth meet the similarity criteria, the difference between α' and R_depth' is first determined. Then, this difference is divided by R_depth to obtain a ratio. If this ratio is less than the difference threshold (e.g., 10%), then α' and R_depth are considered to meet the similarity criteria. Otherwise, α' and R_depth are considered not to meet the similarity criteria.
[0183] The difference threshold can be set before or after the terminal leaves the factory. Furthermore, the difference threshold can be modified after it is set to better meet the needs of users.
[0184] The method for determining whether T' and T_depth satisfy the similarity condition is similar to the method for determining whether α' and R_depth satisfy the similarity condition, and will not be repeated here.
[0185] It should be noted that, in addition to the methods for determining similarity conditions mentioned above, other methods can also be used to determine similarity conditions, and there are no restrictions on these methods. For example, a neural network-based similarity determination model can be used to determine similarity.
[0186] This method incorporates both first pose estimation data and second pose estimation data. The first pose estimation data is determined based on a color camera module, while the second pose estimation data is determined based on a non-visual sensor. In cases where the accuracy of the first pose estimation data is poor, the accuracy of the second pose estimation data can be determined based on both the first and second pose data. This, in turn, determines the final target pose estimation data, further improving the accuracy of the target pose estimation and enhancing the user experience.
[0187] In one exemplary embodiment, a pose estimation method is provided, applied to a terminal. (Reference) Figure 4 As shown, in this method, determining the target pose estimation data based on the first pose data, the second pose data, and the second pose estimation data may include:
[0188] S410, determine the fourth pose data based on the first pose data and the second pose estimation data;
[0189] S420, determine whether the fourth pose data and the second pose data satisfy the second similarity condition; if the determination result is yes, then proceed to step S430; otherwise, proceed to step S440.
[0190] S430, the second pose estimation data is determined as the target pose estimation data;
[0191] S440, determine the target pose estimation data based on the first pose data and the second pose data.
[0192] In the other embodiments described above, steps S210 to S240 are similar to steps S410 to S440 when the third pose estimation data is determined by a non-visual sensor. Therefore, steps S410 to S440 will not be described in detail here.
[0193] In addition, in this method, the entire pose estimation process can be terminated once the target pose estimation data is determined.
[0194] Example 1,
[0195] The first pose data includes rotation data and translation data. The rotation data can be simply referred to as R_depth_1, and the translation data can be simply referred to as T_depth_1.
[0196] The second pose estimation data may include rotation estimation data and translation estimation data, where rotation estimation data can be simply referred to as α' and translation estimation data can be simply referred to as T'.
[0197] The fourth pose data may include rotation data and translation data, where rotation data can be abbreviated as R2 and translation data can be abbreviated as T2.
[0198] The third pose estimation data may include rotation estimation data and translation estimation data, where rotation estimation data can be simply referred to as R_depth and translation estimation data can be simply referred to as T_depth.
[0199] The fourth pose data can be obtained based on the formulas for determining the rotation estimation data and the translation estimation data. The formula for determining the rotation estimation data is: The formula for determining the translation estimation data is: Based on the two formulas above, we can obtain: r2 = R·r1,
[0200] By substituting R_depth_1, α', T_depth_1, T', R2” and T2” into the newly determined formula, and using them as r1, r2, t1, t2, R and T in the formula, the final R2” and T2” can be determined to determine the fourth pose data.
[0201] In this example, the fourth pose data and the second pose data are considered to meet the second similarity condition only if R2” and R_depth_2 satisfy the first similarity condition, and T2” and T_depth_2 also satisfy the second similarity condition. Otherwise, the fourth pose data and the second pose data are considered not to meet the second similarity condition.
[0202] In determining whether R2” and R_depth_2 satisfy the second similarity condition, the difference between R2” and R_depth_2 can be determined first. Then, this difference is divided by R_depth_2 to obtain the ratio. If the ratio is less than the second difference threshold (e.g., 10%), then R2” and R_depth_2 are considered to satisfy the second similarity condition. Otherwise, R2” and R_depth_2 are considered not to satisfy the second similarity condition.
[0203] The second difference threshold can be set before the terminal leaves the factory or after the terminal leaves the factory. Furthermore, the second difference threshold can be modified after it is set to better meet the needs of users.
[0204] The method for determining whether T2” and T_depth_2 satisfy the second similarity condition is similar to the method for determining whether R2” and R_depth_2 satisfy the second similarity condition, and will not be elaborated here.
[0205] In this example, if it is determined that R2” and R_depth_2 satisfy the second similarity condition, and T2” and T_depth_2 also satisfy the second similarity condition, then the second pose estimation data can be determined as the target pose estimation data, and α' and T' can be determined as the target pose estimation data.
[0206] If it is determined that R2” and R_depth_2 do not meet the second similarity condition, or T2” and T_depth_2 do not meet the second similarity condition, or R2” and R_depth_2 do not meet the second similarity condition and T2” and T_depth_2 do not meet the second similarity condition, then the third pose estimation data can be determined as the target pose estimation data, that is, R_depth and T_depth can be determined as the target pose estimation data.
[0207] It should be noted that, in addition to the methods described above for determining the second similarity condition, other methods can also be used, and no restrictions are placed here. For example, a neural network-based similarity judgment model can be used to determine similarity.
[0208] In this method, when the accuracy of the first pose estimation data is poor, the target pose estimation data of the terminal can be selected based on whether the fourth pose data and the second pose data meet the second similarity condition. This can accurately reflect the pose change of the terminal from the first time to the second time, improve the accuracy of pose estimation, and enhance the user experience.
[0209] In one exemplary embodiment, a pose estimation method is provided, applied to a terminal. (Reference) Figure 5 As shown, in this method, determining the target pose estimation data based on the first pose data, the second pose data, and the second pose estimation data may include:
[0210] S510, determine the fourth pose data based on the first pose data and the second pose estimation data;
[0211] S520, determine whether the fourth pose data and the second pose data satisfy the second similarity condition; if the determination result is yes, then proceed to step S530; otherwise, proceed to step S540.
[0212] S530, the second pose estimation data is determined as the target pose estimation data;
[0213] S540, determine that the number of iterations meets the set threshold, and determine the target pose estimation data based on the first pose data and the second pose data.
[0214] In this embodiment, steps S510 to S530 can be referred to sequentially as steps S410 to S430 in the above embodiment, and will not be repeated here. Furthermore, in this method, the entire pose estimation process can be terminated after the target pose estimation data is determined.
[0215] In step S540, if the fourth pose data and the second pose data do not meet the second similarity condition, it indicates that the accuracy of the second pose estimation data is poor and the second pose estimation data cannot be determined as the target pose estimation data. Then, the number of iterations can be used to determine whether the target pose estimation data can be determined by the first pose data and the second pose data, so as to further improve the accuracy of the target pose estimation data.
[0216] Among them, the number of iterations meeting the set threshold can mean that the number of iterations is greater than the set threshold.
[0217] The threshold can be set before or after the terminal leaves the factory. Furthermore, the threshold can be modified after it is set to better meet the user's needs.
[0218] The threshold can be determined based on the number of feature points in the first feature point set or the second feature point set. For example, the threshold can be set to 10% of the number of feature points. For instance, if the number of matching feature point pairs in the first color image and the second color image is 100, then the number of feature points in the first feature point set is 100, and the threshold can be set to 10.
[0219] Example 1,
[0220] A threshold of 10 is set. If the number of iterations is greater than 10, it means the number of iterations meets the set threshold, and the target pose estimation data can be determined based on the first pose data and the second pose data. That is, the third pose estimation data can be determined based on the first pose data and the second pose data. The third pose estimation data may include rotation estimation data and translation estimation data, where rotation estimation data can be simply referred to as R_depth, and translation estimation data can be simply referred to as T_depth. In other words, the finally determined target pose estimation data may include R_depth and T_depth.
[0221] In this method, when the accuracy of both the first pose estimation data and the second pose estimation data is poor, the method can determine whether to determine the target pose estimation data based on the first pose data and the second pose data based on whether the number of iterations meets the set threshold. This ensures that the final target pose estimation data can accurately reflect the pose change of the terminal from the first time to the second time, thereby improving the accuracy of pose estimation and enhancing the user experience.
[0222] In one exemplary embodiment, a pose estimation method is provided, applied to a terminal. (Reference) Figure 6 As shown, in this method, determining the first pose estimation data, the first pose data at the first time step, and the second pose data at the second time step may include:
[0223] S610, if the set conditions are met, then determine the first pose estimation data, the first pose data, and the second pose data based on the feature point set and the set calibration information.
[0224] Among them, the setting calibration information refers to the offline calibration information of the color camera module and the depth camera module. It can be set before the terminal leaves the factory or after the terminal leaves the factory. Furthermore, after the setting calibration information is completed, it can be modified to better meet user needs.
[0225] The determination that the set conditions are met may include at least one of method 1, method 2 and method 3.
[0226] Method 1: Determine the feature point set based on the first color image at the first time step and the second color image at the second time step.
[0227] Example 1,
[0228] At the first moment, the terminal's color camera captures the first color image, which can be simply referred to as Image_1.
[0229] At the second moment, the terminal's depth camera captures a second depth image, which can be simply referred to as Image_2.
[0230] After the terminal's processor obtains Image_1 and Image_2, it can perform feature point detection and matching to determine the matching feature points in Image_1 and Image_2, and then determine all the matching feature points as a feature point set.
[0231] Method 2: Determine the feature point set based on the first depth image at the first time step and the second depth image at the second time step.
[0232] Example 2,
[0233] At the first moment, the terminal's depth camera captures the first depth image, which can be simply referred to as Depth_1.
[0234] At the second moment, the terminal's depth camera captures a second depth image, which can be simply referred to as Depth_2.
[0235] After the terminal's processor obtains Depth_1 and Depth_2, it can perform feature point detection and matching to determine the matching feature points in Depth_1 and Depth_2, and determine all matching feature points as a feature point set.
[0236] Method 3: If the number of iterations does not meet the set threshold, adjust the feature point set and the number of iterations.
[0237] Example 3,
[0238] At the first moment, the terminal's color camera captures the first color image, which can be simply referred to as Image_1.
[0239] At the second moment, the terminal's depth camera captures a second depth image, which can be simply referred to as Image_2.
[0240] After the terminal's processor obtains Image_1 and Image_2, it can perform feature point detection and matching to determine the matching feature points in Image_1 and Image_2, and then determine all the matching feature points as a feature point set.
[0241] After performing pose estimation based on the feature point set, if the number of iterations is less than or equal to the set threshold, the feature point set is readjusted and the number of iterations is incremented by 1. Then, based on the adjusted feature point setting calibration information, the first pose estimation data, the first pose data, and the second pose data are re-determined, and pose estimation is performed again until the target pose estimation data is determined.
[0242] It should be noted that in this method 3, there are no restrictions on how the feature point set is adjusted. The number of feature points can be reduced from the outside to the inside according to the FOV (field of view), or the feature points can be divided into multiple parts. The feature point set can be adjusted by determining the feature points of different parts as the feature point set.
[0243] Example 4,
[0244] The threshold is set to 3. The initial number of iterations is set to 0. The initial set of feature points determined through feature point detection and matching is as follows: Figure 1 As shown, it is simply referred to as P0. The initial feature point set can be divided into 2*2=4 equal parts, which are respectively referred to as P1, P2, P3 and P4.
[0245] In this example, when the number of iterations is 0, the corresponding feature point set is P0. At this time, the target pose estimation data has not been determined. Since 0 is less than 3, the feature point set can be adjusted to P1, and the number of iterations can be adjusted to 1 to determine the target pose estimation data again.
[0246] If the target pose estimation data is still not determined, since 1 is less than 3, the feature point set and the number of iterations are adjusted again. The feature point set is adjusted to P2, and the number of iterations is adjusted to 2. The target pose estimation data is determined again.
[0247] This process continues until the target pose estimation data is determined. When the iteration count is 4, if the target pose estimation data is still not determined after judging the first and second similarity conditions, then since 4 is greater than 3, the target pose estimation data can be determined based on the first and second pose data, thus ending the target pose estimation data determination process.
[0248] In this method, by adjusting the feature point set, the target pose estimation data is selected using the first similarity condition and the second similarity condition under different feature point sets, so as to further determine more accurate target pose estimation data. This ensures that the final target pose estimation data can accurately reflect the pose change of the terminal from the first time to the second time, thereby improving the accuracy of pose estimation and enhancing the user experience.
[0249] In one exemplary embodiment, a pose estimation method is provided, applied to a terminal. (Reference) Figure 7As shown, the method may include:
[0250] S710, determine the first color image and the first depth image at the first time, determine the second color image and the second depth image at the second time, and determine the non-visual detection data of the non-visual sensor from the first time to the second time;
[0251] S720, determine the feature point set based on the first color image and the second color image;
[0252] S730 determines the first pose estimation data, the first pose data, and the second pose data based on the feature point set and the set calibration information;
[0253] S740 determines the third pose estimation data based on the first pose data and the second pose data, and determines the second pose estimation data based on the non-visual detection data.
[0254] S750 determines the third pose data based on the first pose data and the first pose estimation data;
[0255] S760, determine whether the third pose data and the second pose data satisfy the first similarity condition; if the determination result is yes, then proceed to step S770; otherwise, proceed to step S780.
[0256] S770, the first pose estimation data is determined as the target pose estimation data;
[0257] S780, based on the first pose data and the second pose estimation data, determines the fourth pose data;
[0258] S790, determine whether the fourth pose data and the second pose data satisfy the second similarity condition; if the determination result is yes, then execute step S7100; otherwise, execute step S7110.
[0259] S7100, the second pose estimation data is determined as the target pose estimation data;
[0260] S7110, determine whether the number of iterations meets the set threshold; if the result is yes, proceed to step S7120; otherwise, proceed to step S7130.
[0261] S7120, the third pose estimation data is determined as the target pose estimation data.
[0262] S7130, adjust the feature point set and record one iteration, then update the iteration count and return to step S730.
[0263] In this method, in step S730, when the iteration count is the initial value, the first pose estimation data, the first pose data, and the second pose data can be determined directly based on the feature point set determined in step S720 and the set calibration information, and then the subsequent steps are executed. When the iteration count is updated in step S7130, it indicates that the terminal has adjusted the feature point set. At this point, the first pose estimation data, the first pose data, and the second pose data can be determined based on the adjusted feature point set and the set calibration information, and then the subsequent steps are executed. This process continues until the target pose estimation data is determined, ending the entire pose estimation process.
[0264] In addition, in this method, the entire pose estimation process can be terminated once the target pose estimation data is determined.
[0265] It should be noted that if pose estimation is based solely on a color camera module, although the accuracy of feature point detection and matching is high, the pose estimation error at a single moment is large, meaning the reliability of the pose data at a single moment is insufficient. If pose estimation is based solely on a depth camera module, although the accuracy of pose estimation at a single moment is high, the error is large when detecting and matching feature points in depth images at different moments. If pose estimation is based solely on non-visual sensors, although the pose estimation accuracy between different moments is high, the non-visual detection data itself has accuracy drift between different moments, resulting in a large error.
[0266] This method, by setting a first similarity condition, a second similarity condition, and a threshold, can effectively avoid the shortcomings of various pose estimation methods mentioned above. It can also combine the advantages of various pose estimation methods, making the final target pose estimation data more accurate, improving the accuracy of pose estimation, and enhancing the user experience.
[0267] When this method is applied to positioning scenarios such as visual navigation and visual positioning, the determined target pose estimation data can more accurately reflect the pose change of the terminal from the first moment to the second moment. Therefore, it can improve the positioning accuracy in positioning scenarios such as visual navigation and visual positioning, and enhance the user experience.
[0268] In one exemplary embodiment, a pose estimation apparatus is provided, applied to a terminal. This apparatus is used to implement the pose estimation method described above. For example, refer to... Figure 8 As shown, the device may include a determining module 101, wherein the determining module 101 is configured to:
[0269] Determine the first pose estimation data, the first pose data at the first moment, and the second pose data at the second moment. The first pose data and the second pose data are represented by the pose determined by the depth camera module of the terminal. The first pose estimation data is used to represent the pose change of the terminal from the first moment to the second moment.
[0270] Based on the first pose data, the second pose data, and the first pose estimation data, the target pose estimation data is determined.
[0271] In one exemplary embodiment, a pose estimation apparatus is provided, applied to a terminal. (Reference) Figure 8 As shown, in this device, the determining module 101 is used for:
[0272] The third pose data is determined based on the first pose data and the first pose estimation data;
[0273] If the third pose data and the second pose data are determined to satisfy the first similarity condition, then the first pose estimation data is determined as the target pose estimation data.
[0274] In one exemplary embodiment, a pose estimation apparatus is provided, applied to a terminal. (Reference) Figure 8 As shown, in this device, the first pose estimation data characterizes the pose change determined by the terminal's color camera module. The determination module 101 is used for:
[0275] The third pose data is determined based on the first pose data and the first pose estimation data;
[0276] Determine the second pose estimation data from the first time point to the second time point. The second pose estimation data represents the pose change determined by the terminal's non-visual sensors.
[0277] If it is determined that the third pose data and the second pose data do not meet the first similarity condition, then the target pose estimation data is determined based on the first pose data, the second pose data, and the second pose estimation data.
[0278] In one exemplary embodiment, a pose estimation apparatus is provided, applied to a terminal. (Reference) Figure 8 As shown, in this device, the determining module 101 is used for:
[0279] The fourth pose data is determined based on the first pose data and the second pose estimation data.
[0280] If the fourth pose data and the second pose data are determined to satisfy the second similarity condition, then the second pose estimation data is determined as the target pose estimation data.
[0281] In one exemplary embodiment, a pose estimation apparatus is provided, applied to a terminal. (Reference) Figure 8As shown, in this device, the determining module 101 is used for:
[0282] The fourth pose data is determined based on the first pose data and the second pose estimation data.
[0283] If it is determined that the fourth pose data and the second pose data do not meet the second similarity condition, then it is determined whether the number of iterations meets the set threshold.
[0284] If the number of iterations meets the set threshold, then the target pose estimation data is determined based on the first pose data and the second pose data.
[0285] In one exemplary embodiment, a pose estimation apparatus is provided, applied to a terminal. (Reference) Figure 8 As shown, in this device, the determining module 101 is used for:
[0286] If the set conditions are met, then the first pose estimation data, the first pose data, and the second pose data are determined based on the feature point set and the set calibration information.
[0287] The determination that the set conditions are met includes at least one of the following methods:
[0288] Method 1: Determine the feature point set based on the first color image at the first time step and the second color image at the second time step;
[0289] Method 2: Determine the feature point set based on the first depth image at the first time step and the second depth image at the second time step;
[0290] Method 3: If the number of iterations does not meet the set threshold, adjust the feature point set and the number of iterations.
[0291] In one exemplary embodiment, a terminal is provided, such as a mobile phone, a laptop computer, a tablet computer, and a wearable device.
[0292] refer to Figure 9 As shown, terminal 400 may include one or more of the following components: processing component 402, memory 404, power supply component 406, multimedia component 408, audio component 410, input / output (I / O) interface 412, sensor component 414, and communication component 416.
[0293] Processing component 402 typically controls the overall operation of terminal 400, such as operations associated with display, telephone calls, data communication, camera operation, and recording. Processing component 402 may include one or more processors 420 to execute instructions to complete all or part of the steps of the methods described above. Furthermore, processing component 402 may include one or more modules to facilitate interaction between processing component 402 and other components. For example, processing component 402 may include a multimedia module to facilitate interaction between multimedia component 408 and processing component 402.
[0294] Memory 404 is configured to store various types of data to support operation on terminal 400. Examples of this data include instructions for any application or method operating on terminal 400, contact data, phonebook data, messages, pictures, videos, etc. Memory 404 can be implemented by any type of volatile or non-volatile storage terminal or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0295] Power supply component 406 provides power to various components of terminal 400. Power supply component 406 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to terminal 400.
[0296] Multimedia component 408 includes a screen that provides an output interface between terminal 400 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of touch or swipe actions but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 408 includes a front-facing camera module and / or a rear-facing camera module. When terminal 400 is in an operating mode, such as shooting mode or video mode, the front-facing camera module and / or rear-facing camera module may receive external multimedia data. Each front-facing camera module and rear-facing camera module may be a fixed optical lens system or have focal length and optical zoom capabilities.
[0297] Audio component 410 is configured to output and / or input audio signals. For example, audio component 410 includes a microphone (MIC) configured to receive external audio signals when terminal 400 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 404 or transmitted via communication component 416. In some embodiments, audio component 410 also includes a speaker for outputting audio signals.
[0298] I / O interface 412 provides an interface between processing component 402 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.
[0299] Sensor assembly 414 includes one or more sensors for providing state assessments of various aspects of terminal 400. For example, sensor assembly 414 may detect the on / off state of terminal 400, the relative positioning of components such as the display and keypad of terminal 400, changes in the position of terminal 400 or a component of terminal 400, the presence or absence of user contact with terminal 400, the orientation or acceleration / deceleration of terminal 400, and temperature changes of terminal 400. Sensor assembly 414 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 414 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 414 may also include an accelerometer, a gyroscope, a magnetometer, a pressure sensor, or a temperature sensor.
[0300] Communication component 416 is configured to facilitate wired or wireless communication between terminal 400 and other terminals. Terminal 700 can access wireless networks based on communication standards, such as WiFi, 2G, 3G, 4G, 5G, or combinations thereof. In one exemplary embodiment, communication component 416 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 416 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0301] In an exemplary embodiment, terminal 400 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing terminals (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.
[0302] In one exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 404 including instructions, which can be executed by a processor 420 of a terminal 400 to perform the described method. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage terminal, etc. When the instructions in the storage medium are executed by the terminal's processor, the terminal is able to perform the method shown in the above embodiments.
[0303] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the claims.
[0304] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.
[0305] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. A pose estimation method applied to a terminal, characterized in that, The method includes: The first pose estimation data, the first pose data at the first moment, and the second pose data at the second moment are determined. The first pose data represents the pose of the terminal at the first moment as determined by the depth camera module of the terminal. The second pose data represents the pose of the terminal at the second moment as determined by the depth camera module. The first pose estimation data represents the pose change of the terminal from the first moment to the second moment. The first pose estimation data is determined based on a color camera module or a non-visual sensor. Based on the first pose data, the second pose data, and the first pose estimation data, target pose estimation data characterizing the pose change of the terminal from the first time moment to the second time moment is determined.
2. The method according to claim 1, characterized in that, The step of determining the target pose estimation data based on the first pose data, the second pose data, and the first pose estimation data includes: The third pose data is determined based on the first pose data and the first pose estimation data. If it is determined that the third pose data and the second pose data satisfy the first similarity condition, then the first pose estimation data is determined as the target pose estimation data.
3. The method according to claim 1, characterized in that, The first pose estimation data represents the pose change determined based on the color camera module of the terminal. The step of determining the target pose estimation data based on the first pose data, the second pose data, and the first pose estimation data includes: The third pose data is determined based on the first pose data and the first pose estimation data; Determine second pose estimation data from the first time point to the second time point. The second pose estimation data represents the pose change determined based on the non-visual sensors of the terminal. If it is determined that the third pose data and the second pose data do not meet the first similarity condition, then the target pose estimation data is determined based on the first pose data, the second pose data, and the second pose estimation data.
4. The method according to claim 3, characterized in that, The step of determining the target pose estimation data based on the first pose data, the second pose data, and the second pose estimation data includes: Based on the first pose data and the second pose estimation data, the fourth pose data is determined; If it is determined that the fourth pose data and the second pose data satisfy the second similarity condition, then the second pose estimation data is determined as the target pose estimation data.
5. The method according to claim 3, characterized in that, The step of determining the target pose estimation data based on the first pose data, the second pose data, and the second pose estimation data includes: Based on the first pose data and the second pose estimation data, the fourth pose data is determined; If it is determined that the fourth pose data and the second pose data do not meet the second similarity condition, then it is determined whether the number of iterations meets the set threshold. If it is determined that the number of iterations meets the set threshold, then the target pose estimation data is determined based on the first pose data and the second pose data.
6. The method according to claim 5, characterized in that, The determination of the first pose estimation data, the first pose data at the first moment, and the second pose data at the second moment includes: If the set conditions are met, then the first pose estimation data, the first pose data, and the second pose data are determined based on the feature point set and the set calibration information. The determination that the set conditions are met includes at least one of the following methods: Method 1: Determine the feature point set based on the first color image at the first time point and the second color image at the second time point; Method 2: Determine the feature point set based on the first depth image at the first time point and the second depth image at the second time point; Method 3: If it is determined that the number of iterations does not meet the set threshold, then the feature point set and the number of iterations are adjusted.
7. A pose estimation device, applied to a terminal, characterized in that, The device includes: The determination module is used to determine first pose estimation data, first pose data at a first moment, and second pose data at a second moment. The first pose data represents the pose of the terminal at the first moment as determined by the depth camera module of the terminal. The second pose data represents the pose of the terminal at the second moment as determined by the depth camera module. The first pose estimation data represents the pose change of the terminal from the first moment to the second moment. The first pose estimation data is determined based on a color camera module or a non-visual sensor. The target pose estimation data is used to determine the pose change of the terminal from the first time moment to the second time moment based on the first pose data, the second pose data, and the first pose estimation data.
8. The apparatus according to claim 7, characterized in that, The determining module is used for: The third pose data is determined based on the first pose data and the first pose estimation data; If it is determined that the third pose data and the second pose data satisfy the first similarity condition, then the first pose estimation data is determined as the target pose estimation data.
9. The apparatus according to claim 7, characterized in that, The first pose estimation data characterizes the pose changes determined by the color camera module of the terminal. The determining module is used to: The third pose data is determined based on the first pose data and the first pose estimation data; Determine second pose estimation data from the first time point to the second time point. The second pose estimation data represents the pose change determined based on the non-visual sensors of the terminal. If it is determined that the third pose data and the second pose data do not meet the first similarity condition, then the target pose estimation data is determined based on the first pose data, the second pose data, and the second pose estimation data.
10. The apparatus according to claim 9, characterized in that, The determining module is used for: Based on the first pose data and the second pose estimation data, the fourth pose data is determined; If it is determined that the fourth pose data and the second pose data satisfy the second similarity condition, then the second pose estimation data is determined as the target pose estimation data.
11. The apparatus according to claim 9, characterized in that, The determining module is used for: Based on the first pose data and the second pose estimation data, the fourth pose data is determined; If it is determined that the fourth pose data and the second pose data do not meet the second similarity condition, then it is determined whether the number of iterations meets the set threshold. If it is determined that the number of iterations meets the set threshold, then the target pose estimation data is determined based on the first pose data and the second pose data.
12. The apparatus according to claim 11, characterized in that, The determining module is used for: If the set conditions are met, then the first pose estimation data, the first pose data, and the second pose data are determined based on the feature point set and the set calibration information. The determination that the set conditions are met includes at least one of the following methods: Method 1: Determine the feature point set based on the first color image at the first time point and the second color image at the second time point; Method 2: Determine the feature point set based on the first depth image at the first time point and the second depth image at the second time point; Method 3: If it is determined that the number of iterations does not meet the set threshold, then the feature point set and the number of iterations are adjusted.
13. A terminal, characterized in that, The terminal includes: processor; Memory used to store the processor's executable instructions; The processor is configured to perform the method as described in any one of claims 1 to 6.
14. A non-transitory computer-readable storage medium, characterized in that, When the instructions in the storage medium are executed by the processor of the terminal, the terminal is able to perform the method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Pose estimation method and apparatus
CN107123142A
Relative pose estimation method and device, electronic device and medium
CN112184810A