Positioning method and apparatus, electronic device, and storage medium

By using a binocular imaging system and feature point enhancement technology, the problem of inaccurate positioning caused by inaccurate image feature extraction was solved, achieving higher positioning accuracy.

CN115187634BActive Publication Date: 2025-11-11GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210852112.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-19
Publication Date
2025-11-11
Estimated Expiration
2042-07-19

AI Technical Summary

Technical Problem

In existing technologies, the accuracy of image feature extraction is low, resulting in poor localization accuracy.

Method used

A binocular imaging system is used to acquire target images and depth images through a first imaging module and a second imaging module. By using feature point tracking and feature point enhancement technology, a preset number of feature points are extracted in the feature extraction area to avoid mismatches and improve the accuracy of feature extraction and matching.

Benefits of technology

This improves the accuracy and stability of feature points, thereby enhancing the accuracy of localization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115187634B_ABST
    Figure CN115187634B_ABST
Patent Text Reader

Abstract

The positioning method, the positioning device, the electronic equipment and the nonvolatile computer readable storage medium of the application, the method comprises: in the case that the number of feature points of the previous frame is greater than 0, obtaining the second feature point and the fourth feature point corresponding to the first feature point and the third feature point in the current frame respectively; in the case that the total number of the second feature point and the fourth feature point of the current frame is less than the preset number, extracting feature points in the feature extraction area so that the total number reaches the preset number, the feature extraction area is an image area in the current frame, which is a first area containing the obtained second feature point or fourth feature point, and a second area other than the second area where the first target image and the second target image overlap; calculating the pose information according to the second feature point and the fourth feature point extracted to the preset number. The accuracy of feature extraction can be improved, and the accuracy of positioning can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of consumer electronics technology, and in particular to a positioning method, positioning device, electronic device, and non-volatile computer-readable storage medium. Background Technology

[0002] Currently, real-time positioning typically requires extracting feature points from captured images. Based on these extracted feature points and a pre-defined positioning algorithm, positioning is performed. However, due to the low accuracy of feature extraction, the positioning accuracy is also poor. Summary of the Invention

[0003] This application provides a positioning method, a positioning device, an electronic device, and a non-volatile computer-readable storage medium.

[0004] This application provides a positioning method. The positioning method is applied to an electronic device, which includes a first imaging module and a second imaging module. The first imaging module is used to acquire a first target image and a first depth image, and the second imaging module is used to acquire a second target image and a second depth image. The positioning method includes: if the number of first feature points in the first target image of the previous frame is greater than 0, acquiring a second feature point corresponding to the first feature point in the first target image of the current frame; and if the number of third feature points in the second target image of the previous frame is greater than 0, acquiring a fourth feature point corresponding to the third feature point in the second target image of the current frame; if the total number of the second feature points in the first target image of the current frame and the fourth feature points in the second target image of the current frame is less than a preset number, extracting the second feature points and / or the fourth feature points in a feature extraction region, such that the total number reaches the preset number. The feature extraction region is a first region in the first target image of the current frame and the second target image of the current frame, including the acquired second feature points or the fourth feature points, and an image region other than a second region overlapping the first target image and the second target image; and calculating pose information based on the extracted preset number of second feature points and the fourth feature points.

[0005] This application provides a positioning device. The positioning device is applied to an electronic device, which includes a first imaging module and a second imaging module. The first imaging module is used to acquire a first target image and a first depth image, and the second imaging module is used to acquire a second target image and a second depth image. The positioning device includes a first acquisition module, an extraction module, and a positioning module. The first acquisition module is used to acquire a second feature point corresponding to the first feature point in the current frame of the first target image when the number of first feature points in the previous frame is greater than 0, and to acquire a fourth feature point corresponding to the third feature point in the current frame of the second target image when the number of third feature points in the previous frame is greater than 0; the extraction module is used to extract the second feature point and / or the fourth feature point in the feature extraction region when the total number of the second feature points and the fourth feature points in the current frame of the first target image is less than a preset number, so that the total number reaches the preset number, wherein the feature extraction region is a first region in the current frame of the first target image and the current frame of the second target image that includes the acquired second feature point or the fourth feature point, and an image region other than a second region in which the first target image and the second target image overlap; the positioning module is used to calculate pose information based on the extracted preset number of second feature points and the fourth feature points.

[0006] This application provides an electronic device including a first imaging module, a second imaging module, and a processor. The first imaging module is used to acquire a first target image and a first depth image. The second imaging module is used to acquire a second target image and a second depth image. The processor is used to: acquire a second feature point corresponding to the first feature point in the first target image of the current frame when the number of first feature points in the first target image of the previous frame is greater than 0; and acquire a fourth feature point corresponding to the third feature point in the second target image of the current frame when the number of third feature points in the second target image of the previous frame is greater than 0; extract the second feature point and / or the fourth feature point in a feature extraction region when the total number of the second feature points and the fourth feature points in the first target image of the current frame is less than a preset number, so that the total number reaches the preset number. The feature extraction region is a first region in the first target image and the second target image of the current frame that includes the acquired second feature point or the fourth feature point, and an image region other than a second region in which the first target image and the second target image overlap; and calculate pose information based on the extracted second feature points and the fourth feature points.

[0007] This application provides a non-volatile computer-readable storage medium containing a computer program that implements a location method when the computer program is executed by one or more processors. The positioning method is applied to an electronic device, which includes a first imaging module and a second imaging module. The first imaging module is used to acquire a first target image and a first depth image, and the second imaging module is used to acquire a second target image and a second depth image. The positioning method includes: if the number of first feature points in the first target image of the previous frame is greater than 0, acquiring a second feature point corresponding to the first feature point in the first target image of the current frame; and if the number of third feature points in the second target image of the previous frame is greater than 0, acquiring a fourth feature point corresponding to the third feature point in the second target image of the current frame; if the total number of the second feature points in the first target image of the current frame and the fourth feature points in the second target image of the current frame is less than a preset number, extracting the second feature points and / or the fourth feature points in a feature extraction region, such that the total number reaches the preset number. The feature extraction region is a first region in the first target image of the current frame and the second target image of the current frame, including the acquired second feature points or the fourth feature points, and an image region other than a second region overlapping the first target image and the second target image; and calculating pose information based on the extracted preset number of second feature points and the fourth feature points.

[0008] In the positioning method, positioning device, electronic device, and non-volatile computer-readable storage medium of this application, it is first determined whether there are feature points in the previous frame image. If feature points exist, it is only necessary to track the feature points (such as optical flow tracking) to obtain the feature points corresponding to the feature points in the previous frame in the current frame. It is understood that the feature points in different frames may change, resulting in a decrease in the number of feature points obtained in the current frame. Therefore, it is necessary to extract more feature points in the current frame so that the number of feature points meets the requirements for map construction. By performing feature extraction in the current frame outside the region of the tracked feature points and the overlapping region of the left and right eye images, it is possible to avoid feature point extraction that is too close and mismatch, resulting in the extraction of duplicate feature points. This can improve the accuracy of feature extraction and the stability of feature matching, thereby improving the accuracy of subsequent positioning based on feature points.

[0009] Additional aspects and advantages of the embodiments of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0010] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, wherein:

[0011] Figure 1 This is a flowchart illustrating the positioning method of some embodiments of this application;

[0012] Figure 2 This is a schematic diagram of the structure of an electronic device according to certain embodiments of this application;

[0013] Figure 3 This is a schematic diagram of a positioning method according to certain embodiments of this application;

[0014] Figure 4 This is a schematic diagram showing the placement angle of the first imaging module and the second imaging module in certain embodiments of this application;

[0015] Figure 5 This is a schematic diagram of a positioning method according to certain embodiments of this application;

[0016] Figure 6 This is a flowchart illustrating the positioning method of some embodiments of this application;

[0017] Figure 7 This is a schematic diagram of a positioning method according to certain embodiments of this application;

[0018] Figure 8 This is a flowchart illustrating the positioning method of some embodiments of this application;

[0019] Figure 9 This is a flowchart illustrating the positioning method of some embodiments of this application;

[0020] Figure 10 This is a flowchart illustrating the positioning method of some embodiments of this application;

[0021] Figure 11 This is a flowchart illustrating the positioning method of some embodiments of this application;

[0022] Figure 12 This is a flowchart illustrating the positioning method of some embodiments of this application;

[0023] Figure 13 This is a flowchart illustrating the positioning method of some embodiments of this application;

[0024] Figure 14 This is a schematic diagram of the positioning device according to certain embodiments of this application; and

[0025] Figure 15 This is a schematic diagram illustrating the interaction between a non-volatile computer-readable storage medium and a processor in certain embodiments of this application. Detailed Implementation

[0026] The embodiments of this application are described in detail below. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the embodiments of this application, and should not be construed as limiting the embodiments of this application.

[0027] Please see Figure 1 and Figure 2 The positioning method of this application is applied to an electronic device 100, which includes a first imaging module 20 and a second imaging module 30. The first imaging module 20 is used to acquire a first target image and a first depth image, and the second imaging module 30 is used to acquire a second target image and a second depth image. The positioning method includes:

[0028] Step 011: If the number of first feature points in the first target image of the previous frame is greater than 0, obtain the second feature point corresponding to the first feature point in the first target image of the current frame; and if the number of third feature points in the second target image of the previous frame is greater than 0, obtain the fourth feature point corresponding to the third feature point in the second target image of the current frame.

[0029] Specifically, the electronic device 100 has a first imaging module 20 and a second imaging module 30 for capturing scene images. The first imaging module 20 is used to acquire a first target image and a first depth image, and the second imaging module 30 is used to acquire a second target image and a second depth image. For example, the first imaging module 20 includes a first target camera 21 and a first depth camera 22, and the second imaging module 30 includes a second target camera 31 and a second depth camera 32.

[0030] The first target camera 21 and the second target camera 31 can be visible light cameras, infrared cameras, etc., without limitation, as long as they can perform feature extraction. The first depth camera 22 and the second depth camera 32 generally include a light emitter and a light receiver. The light emitter emits light of a specific frequency, which is received by the light receiver, thereby obtaining the depth of different objects within the field of view. The first depth camera 22 and the second depth camera 32 can be binocular cameras based on the binocular imaging principle, TOF cameras based on the time-of-flight (TOF) principle, or structured light cameras based on the structured light principle.

[0031] Taking the first depth camera 22 and the second depth camera 32 as examples of binocular cameras based on the binocular imaging principle, the first depth camera 22 includes a first light receiver 221, a second light receiver 222, and a first light emitter 223. The second depth camera 32 includes a third light receiver 321, a fourth light receiver 322, and a second light emitter 323. Scene images are acquired through the first light receiver 221 and the second light receiver 222, respectively. Then, based on the binocular imaging principle, a first depth image is obtained. Light can also be emitted through the first light emitter 223 to provide more additional features, ensuring feature extraction effectiveness even for textureless scenes where feature points are difficult to extract. Similarly, the third light receiver 321, the fourth light receiver 322, and the second light emitter 323 can work together to obtain a second depth image. Optionally, the first light emitter 223 and the second light emitter 323 can emit structured light, using the binocular structured light principle to obtain the first depth image and the second depth image respectively.

[0032] Optionally, the first depth image and the second depth image acquired by the first depth camera 22 and the second depth camera 32 respectively can be two-dimensional depth images or three-dimensional point cloud images.

[0033] Please see Figure 3 The scene image RGB represents the office scene, and the scene image is the scene image of the composite field of view of the first target camera and the second target camera. Different colors in the two-dimensional depth map corresponding to the scene image represent different depths (i.e., different pixel values ​​represent different depths). The three-dimensional point cloud image corresponding to the scene image can intuitively display the depth of different objects in the scene.

[0034] The first imaging module 20 and the second imaging module 30 form a binocular system, capable of acquiring scene images with a wider field of view. The first imaging module 20 and the second imaging module 30 may have overlapping portions. Generally, due to the close installation positions of the first imaging module 20 and the second imaging module 30, the first horizontal field of view of the first imaging module 20 and the second horizontal field of view of the second imaging module 30 overlap. By identifying the overlapping portions, the images acquired by the first imaging module 20 and the second imaging module 30 are fused to form a scene image with a wider field of view, thereby constructing a map with a larger field of view. Alternatively, the first imaging module 20 and the second imaging module 30 may not have overlapping portions, thus maximizing the combined field of view.

[0035] Please see Figure 4Optionally, the angle γ between the imaging plane S1 of the first imaging module 20 and the imaging plane S2 of the second imaging module 30 is 180 degrees - 1 / 2 (first horizontal field of view α + second horizontal field of view β). If the left boundary of the first horizontal field of view α of the first imaging module 20 is parallel to the right boundary of the second horizontal field of view β of the second imaging module 30, then the first imaging module 20 and the second imaging module 30 can see the maximum field of view. Furthermore, the first imaging module 20 and the second imaging module 30 form an overlapping portion.

[0036] like Figure 4 As shown, the horizontal field of view of the first imaging module 20 and the second imaging module 30 is 65 degrees, and the vertical field of view is 90 degrees. Therefore, the combined horizontal field of view α of the first imaging module 20 and the second imaging module 30 is 130 degrees, and the vertical field of view is 90 degrees. This is beneficial for obtaining maps with a larger field of view.

[0037] For example, the first target image captured by the first imaging module 20 at a preset shooting frequency (e.g., 60 frames per second) and the second target image captured by the second imaging module 30 at a preset shooting frequency, and the first target camera 21 and the second target camera 31 can shoot simultaneously, so that the first target image and the second target image correspond one-to-one, and the frame alignment of the first target image and the second target image is achieved; similarly, the first depth camera 22 and the second depth camera 32 can also shoot simultaneously at the same shooting frequency, so as to achieve the frame alignment of the first depth image and the second depth image.

[0038] Optionally, the first imaging module 20 and the second imaging module 30 can be connected via a circuit, and the shooting time of the first target camera 21 and the second target camera 31 can be controlled by the same pulse signal, so that the first target camera 21 and the second target camera 31 can shoot simultaneously; and the shooting time of the first depth camera 22 and the second depth camera 32 can be controlled by the same pulse signal, so that the first depth camera 22 and the second depth camera 32 can shoot simultaneously. In this way, frame synchronization through circuit connection and the same pulse signal results in high frame synchronization accuracy.

[0039] Of course, frame synchronization can also be achieved through image acquisition time. Specifically, the first target image and the second target image can be frame-aligned based on their acquisition times. For example, the first target image and the second target image with the smallest acquisition time interval can be matched together for frame alignment.

[0040] After frame alignment, the electronic device 100 performs feature extraction on each frame during positioning. For extracted feature points, since the differences between adjacent frames are generally not too large, feature point tracking (such as optical flow tracking) can be performed on the feature points of the previous frame to find feature points matching those of the previous frame in the current frame, thus establishing a connection between different frames. Therefore, when extracting feature points, it can first determine whether feature points have been extracted in the previous frame. If feature points have been extracted in the previous frame, feature points matching those of the previous frame can be found in the current frame. However, if there is no previous frame (i.e., the current frame is the first frame of a multi-frame image) or if feature point extraction was not performed in the previous frame (i.e., the number of feature points in the previous frame is equal to 0), feature extraction is performed directly in the current frame to extract a preset number of feature points.

[0041] For example, in this application, if the number of first feature points in the first target image of the previous frame is greater than 0, the second feature point corresponding to the first feature point in the first target image of the current frame is obtained; and if the number of third feature points in the second target image of the previous frame is greater than 0, the fourth feature point corresponding to the third feature point in the second target image of the current frame is obtained. Thus, feature points in the current frame can be quickly obtained by tracking feature points in the previous frame.

[0042] Step 012: If the total number of the second feature points of the first target image in the current frame and the fourth feature points of the second target image in the current frame is less than a preset number, extract the second feature points and / or the fourth feature points in the feature extraction region so that the total number reaches the preset number. The feature extraction region is the first region in the first target image and the second target image in the current frame that contains the acquired second feature points or fourth feature points, and the image region other than the second region that overlaps with the first target image and the second target image.

[0043] Specifically, after quickly acquiring the second feature points and the fourth feature points of the first target image in the current frame based on feature point tracking, the shooting scene of the current frame and the previous frame changes due to the movement of the electronic device 100. The total number of the second and fourth feature points in the current frame is generally less than the total number of the first and third feature points in the previous frame, resulting in a total number of the second and fourth feature points being less than a preset number, such as 150 or 200. In this case, it is necessary to add more feature points based on the already tracked second and fourth feature points to ensure that the total number of the second and fourth feature points reaches the preset number.

[0044] It is understandable that when adding feature points, if the extraction is performed near existing second and fourth feature points, the extracted features will cluster together. This will make it very easy for mismatches to occur when feature tracking is performed in the current frame and the next frame, thus affecting the accuracy of feature tracking and localization.

[0045] Similarly, in the overlapping areas of the first and second target images, the corresponding actual scenes are the same. If feature extraction is performed in the overlapping areas, it will result in the extraction of duplicate features (i.e., two feature points correspond to the same local area of ​​the actual scene), which will affect the accuracy of localization.

[0046] Therefore, please refer to Figure 5 Before performing feature extraction, it is necessary to first determine the feature extraction regions that can be extracted. First, a first region A1 containing the acquired second or fourth feature points can be determined. For example, in the first target image P1, a predetermined range around each second feature point (e.g., a circular area within a predetermined radius (e.g., 10 pixels, 15 pixels, etc.) centered on the second feature point is defined as the first region A1; in the second target image P2, a predetermined range around each fourth feature point (e.g., a circular area within a predetermined radius (e.g., 10 pixels, 15 pixels, etc.) centered on the second feature point is defined as the first region A1. Then, the image region overlapping between the first target image P1 and the second target image P2 is defined as the second region A2. Finally, in both the first and second target images P1 and P2, all image regions outside of the first and second regions A1 and A2 are defined as the feature extraction regions A3. Thus, by performing feature extraction in the feature extraction region A3, the total number of extracted second and fourth feature points reaches a preset number. The second feature points extracted in the first target image P1 are far away from the tracked second feature points, and the fourth feature points extracted in the second target image P2 are far away from the tracked fourth feature points. This ensures the accuracy of the extracted feature points, which is beneficial to improving the accuracy and stability of feature matching between the current frame and the next frame, and also to improving the accuracy of localization.

[0047] Please see Figure 6 and Figure 7 Optionally, step 012 includes:

[0048] Step 0121: Stitch together the first target image P1 of the current frame and the second target image P2 of the current frame to generate a third target image P3. The third target image P3 includes the image regions of the first target image P1 and the second target image P2 except for the second region A2, or the third target image P3 includes the image regions of the second target image P2 and the first target image P1 except for the second region A2.

[0049] Step 0122: In the third target image P3, extract feature points within the feature extraction region A3 to serve as the second or fourth feature points. The feature extraction region A3 is the image region in the third target image P3 outside of the first region A1 and the second region A2.

[0050] Specifically, during feature extraction, the first target image P1 and the second target image P2 of the current frame can be stitched together to obtain the third target image P3. During stitching, only the second region A2 of one of the first target image P1 and the second target image P2 of the current frame is retained. That is, only the overlapping region of one of the first target image P1 and the second target image P2 of the current frame is retained.

[0051] After image stitching, feature extraction can be performed on the third target image P3. Similarly, a feature extraction region A3 can be first determined in the third target image P3. For example, the image area outside the first region A1 and the second region A2 in the third target image P3 can be defined as the feature extraction region A3. For instance, in the third target image P3, the second feature point is extracted corresponding to the feature extraction region A3 of the first target image P1, and the fourth feature point is extracted corresponding to the feature extraction region A3 of the second target image P2. In this way, a preset number of feature points can be extracted quickly.

[0052] Please see Figure 8 Optionally, feature points can be added based on the response values ​​of pixels in the feature extraction region. Step 012 includes:

[0053] Step 0123: Calculate the response value of each pixel based on the pixel value of each pixel within the feature extraction region and the pixel values ​​of the pixels surrounding that pixel;

[0054] Step 0124: Extract the second feature point and / or the fourth feature point based on the response value of each pixel.

[0055] Specifically, during feature extraction, the response value of each pixel in the feature extraction region can be calculated first. This can be done by calculating the response value of each pixel based on the pixel values ​​of each pixel and its surrounding pixels, such as using the mean or median of the pixel values ​​of each pixel and its surrounding pixels. For example, pixels within a predetermined radius (e.g., 5 pixels, 8 pixels, etc.) centered on each pixel can be considered as the pixels surrounding that pixel, and thus the response value of each pixel can be calculated based on the pixel values ​​of each pixel and its surrounding pixels.

[0056] It's understandable that the larger the response value, the more prominent the pixel's features, and the better it performs as a feature point. After calculating the response value for each pixel, they can be sorted according to their magnitude, allowing for the extraction of second and / or fourth feature points based on these values. For example, if there are 150 existing feature points and the preset quantity is 200, then 50 additional second and / or fourth feature points need to be extracted from the feature extraction area. In this case, the top 50 pixels by response value can be extracted as second and / or fourth feature points. The extracted feature points located in the first target image are designated as second feature points, and those located in the second target image are designated as fourth feature points.

[0057] Step 013: Calculate pose information based on the extracted second and fourth feature points up to a preset number.

[0058] Specifically, after the feature points are extracted, a preset number of second and fourth feature points are obtained. The pose information can then be calculated based on these preset number of second and fourth feature points. For example, the pose information of the electronic device 100 can be obtained based on the second and fourth feature points from multiple consecutive frames. Thus, positioning using more accurate feature points improves the accuracy of the localization.

[0059] Please see Figure 2 and Figure 9 Optionally, the electronic device 100 also includes two attitude sensors 40, which are respectively disposed in the first imaging module 20 and the second imaging module 30. The attitude sensors 40 can collect attitude data of the first imaging module 20 and the second imaging module 30. During positioning, either the first imaging module 20 or the second imaging module 30 needs to be used as a reference. When using the first imaging module 20 as a reference, the attitude data of the first imaging module 20 needs to be collected; when using the second imaging module 30 as a reference, the attitude data of the second imaging module 30 needs to be collected. This application uses the first imaging module 20 as a reference for positioning as an example. The principle of positioning using the second imaging module 30 as a reference is basically similar and will not be repeated here. Step 013 includes:

[0060] Step 0131: Calculate pose information based on the extracted second and fourth feature points (up to a preset number) and pose data.

[0061] Specifically, the reprojection error between frames can be calculated using the second and fourth feature points of multiple consecutive frames, and combined with the error of the inertial measurement unit (IMU), the pose information of the current frame can be solved by tightly coupled optimization. The pose information can include three-dimensional coordinates and attitude information. The three-dimensional coordinates represent the current position of the electronic device 100, and the attitude information represents the current attitude of the electronic device.

[0062] Please see Figure 10 Step 0131 includes:

[0063] Step 01311: Based on the preset first extrinsic parameter between the first imaging module 20 and the second imaging module 30, the fourth feature point located in the feature extraction area is mapped to the coordinate system of the first imaging module 20 to obtain multiple fifth feature points.

[0064] Step 01312: Construct the visual reprojection error based on the second and fifth feature points;

[0065] Step 01313: Output pose information based on visual reprojection error and pose data.

[0066] Specifically, when positioning based on the first imaging module 20, feature point conversion can be performed first. Based on the preset first extrinsic parameter between the first imaging module 20 and the second imaging module 30, the fourth feature point located in the feature extraction area is mapped to the coordinate system of the first imaging module 20 to obtain multiple fifth feature points. For example, each fourth feature point can be mapped to a fifth feature point according to the following formula (1), P L =T LR1 *d r *P r_norm (1).

[0067] Among them, P L T represents the three-dimensional coordinates corresponding to the fifth feature point. LR1 As the first preset extrinsic parameter, d r P represents the depth value corresponding to the fourth feature point in the second depth image. r_norm The three-dimensional coordinates corresponding to the fourth feature point can be determined based on the image coordinates of the fourth feature point and the preset intrinsic and extrinsic parameters of the first target camera 21. Optionally, for ease of calculation, P can be... r_norm Normalization is performed to obtain the normalized P. r_norm .

[0068] Then, based on the 3D coordinates of the second feature point and the fifth feature point in multiple consecutive frames, the visual reprojection error can be constructed. Combining the visual reprojection error with pose data, the pose information can be calculated. For example, the reprojection error between frames can be calculated, and combined with the IMU error, to tightly couple and optimize the pose information of the current frame.

[0069] Thus, when the fourth feature point in the second target image is mapped to the coordinate system of the first imaging module 20, even if the first target image does not have feature points, the second target image can provide sufficient and effective feature points, which can also establish correct visual constraints and couple the IMU to achieve robust and accurate positioning.

[0070] Please see Figure 11 Optionally, since the shooting times of the first target camera 21 and the second target camera 31 may not be consistent, the posture of the first target camera 21 when acquiring the first target image and the posture of the second target camera 31 when acquiring the second target image may be different, causing the preset first extrinsic parameter to become inaccurate. Therefore, it is necessary to compensate for the first extrinsic parameter. The positioning method also includes:

[0071] Step 014: Calculate the extrinsic parameter compensation parameters based on the attitude data between the first shooting time of the first target image in the current frame and the second shooting time of the second target image in the current frame;

[0072] Step 015: Calculate the third extrinsic parameter based on the first extrinsic parameter, the extrinsic parameter compensation parameter, and the second extrinsic parameter preset between the attitude sensor 40 and the first imaging module 20;

[0073] Step 01311 includes:

[0074] Step 01314: Based on the third extrinsic parameter, map the fourth feature point located in the feature extraction region to the coordinate system of the first imaging module 20 to obtain multiple fifth feature points.

[0075] Specifically, the attitude data between the first shooting time of the first target image in the current frame and the second shooting time of the second target image in the current frame can be obtained first. Then, the attitude data can be integrated to obtain the extrinsic parameter compensation parameters. The attitude data may include acceleration and angular velocity. By integrating the acceleration between the first and second shooting times, the translation parameter can be obtained; by integrating the angular velocity between the first and second shooting times, the rotation parameter can be obtained. Finally, the extrinsic parameter compensation parameters can be obtained based on the translation and rotation parameters. Specifically, the extrinsic parameter compensation parameters can be obtained according to the following formulas (2), (3), and (4). T LR_comp ={r comp |t comp} (4).

[0076] Among them, T LR_comp For external parameter compensation parameters, t comp r is the translation parameter. comp Here are the rotational parameters, a is the acceleration, w is the angular velocity, and V is the rotational velocity. L dt represents the velocity corresponding to the first target image, and dt represents the time interval between the two IMU data points.

[0077] After calculating the extrinsic compensation parameters, the extrinsic compensation parameters can be transformed into the coordinate system of the first imaging module 20 based on the second preset extrinsic parameter between the attitude sensor 40 and the first imaging module 20. Specifically, the transformation can be performed according to the following formula (5). Among them, T comp T is the converted extrinsic compensation parameter. ci As the second external parameter, It is the inverse matrix of the second extrinsic parameter.

[0078] Then, the third external parameter can be calculated based on the converted external parameter compensation parameter and the first external parameter. Specifically, it can be calculated according to the following formula (5), T LR2 =T comp *T LR1 (5), where T LR2 T is the third external parameter. LR1 It is the first external parameter.

[0079] After obtaining the third extrinsic parameter after attitude data compensation, the fourth feature point located in the feature extraction region can be mapped to the coordinate system of the first imaging module 20 according to the third extrinsic parameter to obtain multiple fifth feature points. Specifically, the mapping from each fourth feature point to the fifth feature point can be performed according to the following formula (6), P L =T LR2 *d r *P r_norm (6). The specific mapping scheme is basically the same as the aforementioned scheme of mapping the fourth feature point to the fifth feature point based on the first extrinsic parameter, and will not be repeated here.

[0080] Please see Figure 12 The positioning methods also include:

[0081] Step 016: Obtain the depth value corresponding to the pixel based on the first depth image and the second depth image;

[0082] Step 017: If the depth value corresponding to a pixel is within the preset depth range, determine that the pixel is a valid pixel;

[0083] Step 0124 includes:

[0084] Step 01241: Extract the second feature point and / or the fourth feature point from all valid pixels based on the response value of each valid pixel.

[0085] Specifically, when acquiring pose information for localization, it is necessary to obtain the depth value of each feature point. Therefore, if the depth value is not accurate enough, it means that the accuracy of the feature point is also low, which affects the accuracy of the pose information calculation.

[0086] Therefore, when performing feature extraction, in addition to calculating the response value of each pixel, it is also necessary to obtain the depth value corresponding to each pixel from the first depth image and the second depth image. It can be understood that the first target image and the first depth image are corresponding, and the second target image and the second depth image are corresponding. The depth value of each pixel in the first target image can be obtained from the corresponding depth value in the first depth image, and the depth value of each pixel in the second target image can be obtained from the corresponding depth value in the second depth image.

[0087] It is understood that both the first depth camera 22 and the second depth camera 32 have a detection range, that is, the depth value has a corresponding preset depth range. If the depth value is outside the preset depth range, it means that the depth value exceeds the range and the accuracy of the depth value is low. Therefore, pixels with depth values ​​within the preset depth range can be determined as valid pixels.

[0088] During feature extraction, features can be extracted from valid pixels, and second and / or fourth feature points can be extracted based on the response values ​​of the valid pixels. This ensures that the depth value of each feature point not only has a high response value but also a valid depth value, thereby further guaranteeing the accuracy of feature extraction.

[0089] Please see Figure 13 The positioning methods also include:

[0090] 018: Construct a map based on pose information, the first depth image of the current frame, and the second depth image of the current frame.

[0091] Specifically, both the first and second depth images are 3D point cloud images, each containing one or more point clouds. To improve the field of view of the map, the first and second depth images can be fused. For example, the second point cloud can be projected onto the coordinate system of the first imaging module 20 according to a first or third extrinsic parameter to obtain a third point cloud, thus obtaining a 3D point cloud image that combines the first and third point clouds. The field of view of the fused 3D point cloud image is wider. The third point cloud can be specifically defined according to the formula PointCloud. LR =T LR *PointCloud R (6) We obtain, where PointCloudLR Third point cloud, T LR First or third external parameter, PointCloud R This is the second point cloud.

[0092] The blended 3D point cloud image is a 3D image in the coordinate system of the first imaging module 20. After obtaining the pose information, the blended 3D point cloud image (i.e., the first point cloud and the third point cloud) can be converted to the world coordinate system according to the pose information, thereby realizing map construction. The constructed map is a 3D point cloud map in the world coordinate system. Specifically, each first point cloud and third point cloud can be converted into map points in the map according to the following formula (7), thereby realizing map construction. PointCloud world =T WL *PointCloud cur (7), where PointCloud world For map points, T WL PointCloud provides the pose information for the current frame. cur It is either the first point cloud or the third point cloud.

[0093] Optionally, map construction is not performed for every frame, as this would result in a significant waste of computational power. Therefore, a map can be constructed once every preset time interval based on the pose information, the first depth image of the current frame, and the second depth image of the current frame; or, a map can be constructed once if the difference between the pose information of the current frame and the pose information of the previous frame is greater than a preset difference.

[0094] Specifically, within a certain timeframe, the pose information of the electronic device 100 generally does not change significantly, and the map during this period is consistently applicable. Therefore, it is not necessary to build a map for every single frame. Instead, every preset time interval (e.g., the time required to acquire 3, 5, or 10 frames of images, or a preset time interval of 2 or 3 seconds), a map is constructed based on the second and fourth feature points, the first depth image of the current frame, and the second depth image of the current frame. This eliminates the need for map construction for every single frame, thereby reducing the computational load required for map construction while maintaining map accuracy.

[0095] Similarly, if the pose information of electronic device 100 changes significantly, it indicates that the scene within the field of view of electronic device 100 has changed, and the map is no longer accurate. If the pose information includes three-dimensional coordinates and attitude information, it can be determined whether the difference between the pose information of the current frame and the pose information of the previous frame is greater than a preset difference. If the difference is greater than the preset difference, it can be determined that the pose information has changed significantly.

[0096] Specifically, the difference between the pose information of the current frame and the pose information of the previous frame can be determined by judging whether the 3D coordinates of the current frame and the 3D coordinates of the previous frame are greater than a preset coordinate difference; or, the difference between the pose information of the current frame and the pose information of the previous frame can be determined by judging whether the attitude information (such as pitch angle, roll angle and yaw angle) of the current frame is greater than a preset attitude difference.

[0097] When the pose information changes significantly, the map needs to be reconstructed based on the second and fourth feature points, the first depth image of the current frame, and the second depth image of the current frame. This ensures the accuracy of the map construction and reduces the number of map constructions required, as the map construction is only performed when the pose information changes significantly.

[0098] To facilitate better implementation of the positioning method of this application embodiment, this application embodiment also provides a positioning device 10. Please refer to the figure. Figure 2 and Figure 14 , Figure 14 This is a schematic diagram of the structure of the positioning device 10 provided in an embodiment of this application. The positioning device 10 may include:

[0099] The first acquisition module 11 is used to acquire a second feature point corresponding to the first feature point in the current frame of the first target image when the number of first feature points in the first target image of the previous frame is greater than 0, and to acquire a fourth feature point corresponding to the third feature point in the current frame of the second target image when the number of third feature points in the second target image of the previous frame is greater than 0.

[0100] Extraction module 12 is used to extract second feature points and / or fourth feature points in a feature extraction region when the total number of second feature points in the first target image of the current frame and fourth feature points in the second target image of the current frame is less than a preset number, so that the total number reaches the preset number. The feature extraction region is a first region in the first target image of the current frame and the second target image of the current frame that contains the acquired second feature points or fourth feature points, and an image region other than the second region that overlaps with the first target image and the second target image.

[0101] The extraction module 12 is specifically used to stitch together the first target image of the current frame and the second target image of the current frame to generate a third target image. The third target image includes the image regions in the first and second target images except for the second region, or the third target image includes the second target image and the image regions in the first target image except for the second region. In the third target image, feature points within the feature extraction region are extracted as the second feature points or the fourth feature points. The feature extraction region is the image region in the third target image that is outside the first and second regions.

[0102] The extraction module 12 is further configured to calculate the response value of each pixel based on the pixel value of each pixel in the feature extraction region and the pixel values ​​of the pixels surrounding the pixel; and to extract the second feature point and / or the fourth feature point based on the response value of each pixel.

[0103] The positioning module 13 is used to calculate pose information based on the extracted second and fourth feature points to a preset number.

[0104] The positioning device 10 also includes:

[0105] The second acquisition module 14 acquires the depth value corresponding to the pixel based on the first depth image and the second depth image;

[0106] The determination module 15 determines a pixel as a valid pixel if the depth value corresponding to the pixel is within a preset depth range.

[0107] The extraction module 12 is further configured to extract a second feature point and / or a fourth feature point from all valid pixels based on the response value of each valid pixel.

[0108] The positioning device 10 also includes a frame alignment module 16, which is used to perform frame alignment between the first target image and the second target image according to the acquisition time of the first target image and the acquisition time of the second target image; or to synchronize the first imaging module 20 and the second imaging module 30 through the same pulse signal so that the first imaging module 20 and the second imaging module 30 can capture images simultaneously.

[0109] The positioning module 13 is also used to calculate pose information based on the extracted preset number of second and fourth feature points and pose data.

[0110] The positioning module 13 is further configured to map the fourth feature point located in the feature extraction area to the coordinate system of the first imaging module according to the preset first extrinsic parameter between the first imaging module 20 and the second imaging module 30, so as to obtain multiple fifth feature points; construct the visual reprojection error according to the second feature point and the fifth feature point; and output the pose information according to the visual reprojection error and the pose data.

[0111] The positioning device 10 also includes:

[0112] The compensation module 17 is used to calculate the extrinsic parameter compensation parameters based on the attitude data between the first shooting time of the first target image in the current frame and the second shooting time of the second target image in the current frame; and to calculate the third extrinsic parameter based on the first extrinsic parameter, the extrinsic parameter compensation parameters, and the second extrinsic parameter preset between the attitude sensor 40 and the first imaging module 20.

[0113] The compensation module 17 is specifically used to integrate the acceleration between the first shooting moment and the second shooting moment to obtain the translation parameters; to integrate the angular velocity between the first shooting moment and the second shooting moment to obtain the rotation parameters; and to calculate the external parameter compensation parameters based on the translation parameters and the rotation parameters.

[0114] The positioning module 13 is further used to map the fourth feature point located in the feature extraction area to the coordinate system of the first imaging module 20 according to the third extrinsic parameter, so as to obtain multiple fifth feature points.

[0115] The positioning device 10 also includes:

[0116] Module 18 is used to construct a map based on pose information, the first depth image of the current frame, and the second depth image of the current frame.

[0117] The construction module 18 is also specifically used to project the second point cloud onto the coordinate system of the first imaging module 20 according to the third extrinsic parameter to obtain the third point cloud; and to transform the first point cloud and the third point cloud into the world coordinate system according to the pose information to construct a map.

[0118] The extraction module 12 is further configured to extract a second feature point in the first target image of the current frame and a fourth feature point in the second target image of the current frame when the number of the first feature points in the first target image of the previous frame is equal to 0, so that the total number reaches a preset number.

[0119] The construction module 18 is further configured to construct a map once every preset time interval based on the pose information, the first depth image of the current frame, and the second depth image of the current frame; or, if the difference between the pose information of the current frame and the pose information of the previous frame is greater than a preset difference, construct the map once based on the pose information, the first depth image of the current frame, and the second depth image of the current frame.

[0120] Each module in the positioning device 10 described above can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0121] Please refer to it again. Figure 2 The electronic device 100 of this application includes a first imaging module 20, a second imaging module 30, and a processor 50. The first imaging module 20 is used to acquire a first target image and a first depth image, and the second imaging module 30 is used to acquire a second target image and a second depth image. The processor 50 can also implement the steps of the positioning method of any of the above embodiments, which will not be described in detail here for the sake of brevity.

[0122] The electronic device 100 may be a robot, mobile phone, smartphone, personal digital assistant (PDA), tablet computer and video game device, portable terminal (e.g., laptop computer), or larger device (e.g., desktop computer and television), or any other type of device that can be configured with an imaging module.

[0123] Please see Figure 15 This application also provides a computer-readable storage medium 300 storing a computer program 310. When the computer program 310 is executed by the processor 50, it implements the steps of the positioning method of any of the above embodiments. For the sake of brevity, these steps will not be repeated here.

[0124] It is understood that a computer program 310 includes computer program code. Computer program code can be in the form of source code, object code, executable files, or certain intermediate forms. Computer-readable storage media can include: any entity or device capable of carrying computer program code, recording media, USB flash drives, external hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), and software distribution media, etc.

[0125] In the description of this specification, the terms "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with an embodiment or example that are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0126] Any process or method described in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing a particular logical function or process, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the function involved, as will be understood by those skilled in the art to which embodiments of this application pertain.

[0127] Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of this application.

Claims

1. A positioning method, characterized in that, Applied to an electronic device, the electronic device includes a first imaging module and a second imaging module, the first imaging module being used to acquire a first target image and a first depth image, and the second imaging module being used to acquire a second target image and a second depth image, the positioning method comprising: If the number of first feature points in the first target image of the previous frame is greater than 0, obtain the second feature point corresponding to the first feature point in the first target image of the current frame; and if the number of third feature points in the second target image of the previous frame is greater than 0, obtain the fourth feature point corresponding to the third feature point in the second target image of the current frame. If the total number of the second feature points and the fourth feature points of the first target image in the current frame is less than a preset number, the second feature points and / or the fourth feature points are extracted within the feature extraction region so that the total number reaches the preset number. The feature extraction region is a first region in the first target image and the second target image in the current frame that includes the acquired second feature points or the fourth feature points, and an image region other than the second region where the first target image and the second target image overlap. Based on the extracted preset number of second and fourth feature points, calculate pose information; The electronic device further includes an attitude sensor for acquiring attitude data of the first imaging module; the step of calculating pose information based on the extracted preset number of second feature points and the fourth feature points includes: Based on the extracted preset number of second feature points and fourth feature points, and the pose data, the pose information is calculated; The step of calculating pose information based on the extracted preset number of second and fourth feature points and the pose data includes: Based on a preset first extrinsic parameter between the first imaging module and the second imaging module, the fourth feature point located in the feature extraction region is mapped to the coordinate system of the first imaging module to obtain multiple fifth feature points. Based on the second feature point and the fifth feature point, a visual reprojection error is constructed; The pose information is output based on the visual reprojection error and the pose data.

2. The positioning method according to claim 1, characterized in that, The step of extracting the second feature point and / or the fourth feature point within the feature extraction region, so that the total number reaches the preset number, includes: The first target image of the current frame and the second target image of the current frame are stitched together to generate a third target image. The third target image includes image regions in the first target image and the second target image other than the second region, or the third target image includes image regions in the second target image and the first target image other than the second region. In the third target image, feature points within the feature extraction region are extracted as the second feature point or the fourth feature point. The feature extraction region is the image region in the third target image other than the first region and the second region.

3. The positioning method according to claim 1 or 2, characterized in that, The step of extracting the second feature point and / or the fourth feature point within the feature extraction region, so that the total number reaches the preset number, includes: The response value of each pixel is calculated based on the pixel value of each pixel within the feature extraction region and the pixel values ​​of the pixels surrounding the pixel. Based on the response value of each pixel, the second feature point and / or the fourth feature point are extracted.

4. The positioning method according to claim 3, characterized in that, Also includes: The depth value corresponding to the pixel is obtained based on the first depth image and the second depth image; If the depth value corresponding to the pixel is within a preset depth range, the pixel is determined to be a valid pixel. The step of extracting the second feature point and / or the fourth feature point based on the response value of each pixel includes: Based on the response value of each of the valid pixels, the second feature point and / or the fourth feature point are extracted from all the valid pixels.

5. The positioning method according to claim 1, characterized in that, Also includes: Based on the acquisition time of the first target image and the acquisition time of the second target image, frame alignment is performed on the first target image and the second target image; or The first imaging module and the second imaging module are synchronized by the same pulse signal so that the first imaging module and the second imaging module can take pictures simultaneously.

6. The positioning method according to claim 1, characterized in that, Also includes: The extrinsic parameter compensation parameters are calculated based on the attitude data between the first shooting time of the first target image in the current frame and the second shooting time of the second target image in the current frame. The third extrinsic parameter is calculated based on the first extrinsic parameter, the extrinsic parameter compensation parameter, and the second extrinsic parameter preset between the attitude sensor and the first imaging module; The fourth feature point located in the feature extraction region is mapped to the coordinate system of the first imaging module based on a preset first extrinsic parameter between the first imaging module and the second imaging module, to obtain multiple fifth feature points, including: The fourth feature point located in the feature extraction region is mapped to the coordinate system of the first imaging module according to the third extrinsic parameter to obtain multiple fifth feature points.

7. The positioning method according to claim 6, characterized in that, The attitude data includes acceleration and angular velocity. The calculation of extrinsic parameter compensation based on the attitude data between the first capture time of the first target image in the current frame and the second capture time of the second target image in the current frame includes: The acceleration between the first and second shooting times is integrated to obtain the translation parameters; Integrate the angular velocity between the first and second shooting moments to obtain the rotation parameters; The extrinsic compensation parameters are calculated based on the translation parameters and the rotation parameters.

8. The positioning method according to claim 6, characterized in that, Also includes: A map is constructed based on the pose information, the first depth image of the current frame, and the second depth image of the current frame.

9. The positioning method according to claim 8, characterized in that, The first depth image includes a first point cloud, and the second depth image includes a second point cloud. The step of constructing a map based on the pose information, the first depth image of the current frame, and the second depth image of the current frame includes: The second point cloud is projected onto the coordinate system of the first imaging module according to the third extrinsic parameter to obtain the third point cloud; The first point cloud and the third point cloud are transformed into the world coordinate system based on the pose information to construct the map.

10. The positioning method according to claim 8, characterized in that, The step of constructing a map based on the pose information, the first depth image of the current frame, and the second depth image of the current frame includes: Every preset time interval, the map is constructed once based on the pose information, the first depth image of the current frame, and the second depth image of the current frame; or, If the difference between the pose information of the current frame and the pose information of the previous frame is greater than a preset difference, the map is constructed once based on the pose information, the first depth image of the current frame, and the second depth image of the current frame.

11. The positioning method according to claim 1, characterized in that, Also includes: If the number of first feature points in the first target image of the previous frame is equal to 0, then in the first target image of the current frame, the second feature point is extracted, and in the second target image of the current frame, the fourth feature point is extracted, so that the total number reaches the preset number.

12. A positioning device, characterized in that, Applied to an electronic device, the electronic device includes a first imaging module and a second imaging module, the first imaging module being used to acquire a first target image and a first depth image, and the second imaging module being used to acquire a second target image and a second depth image, the positioning device comprising: The first acquisition module is used to acquire a second feature point corresponding to the first feature point in the current frame of the first target image when the number of first feature points in the previous frame is greater than 0, and to acquire a fourth feature point corresponding to the third feature point in the current frame of the second target image when the number of third feature points in the previous frame is greater than 0. An extraction module is configured to extract the second feature point and / or the fourth feature point within a feature extraction region when the total number of the second feature point in the first target image of the current frame and the fourth feature point in the second target image of the current frame is less than a preset number, so that the total number reaches the preset number. The feature extraction region is a first region in the first target image of the current frame and the second target image of the current frame that includes the acquired second feature point or the fourth feature point, and an image region other than the second region that overlaps between the first target image and the second target image. The positioning module is used to calculate pose information based on the extracted preset number of second feature points and fourth feature points; The electronic device further includes an attitude sensor, which is used to collect attitude data of the first imaging module; the positioning module is specifically used to: calculate pose information based on the extracted preset number of second feature points and fourth feature points, and the attitude data; The positioning module is specifically configured to: map the fourth feature point located in the feature extraction area to the coordinate system of the first imaging module according to a preset first extrinsic parameter between the first imaging module and the second imaging module, so as to obtain multiple fifth feature points; construct a visual reprojection error based on the second feature point and the fifth feature point; and output pose information based on the visual reprojection error and the pose data.

13. An electronic device, characterized in that, The system includes a first imaging module, a second imaging module, and a processor. The first imaging module is used to acquire a first target image and a first depth image. The second imaging module is used to acquire a second target image and a second depth image. The processor is used to: If the number of first feature points in the first target image of the previous frame is greater than 0, obtain the second feature point corresponding to the first feature point in the first target image of the current frame; and if the number of third feature points in the second target image of the previous frame is greater than 0, obtain the fourth feature point corresponding to the third feature point in the second target image of the current frame. If the total number of the second feature points in the first target image of the current frame and the fourth feature points in the second target image of the current frame is less than a preset number, the second feature points and / or the fourth feature points are extracted in the feature extraction region so that the total number reaches the preset number. The feature extraction region is the first region in the first target image of the current frame and the second target image of the current frame that includes the acquired second feature points or the fourth feature points, and the image region other than the second region that overlaps between the first target image and the second target image. Based on the extracted preset number of second and fourth feature points, calculate pose information; The electronic device further includes an attitude sensor for acquiring attitude data of the first imaging module; the step of calculating pose information based on the extracted preset number of second feature points and the fourth feature points includes: Based on the extracted preset number of second feature points and fourth feature points, and the pose data, the pose information is calculated; The step of calculating pose information based on the extracted preset number of second and fourth feature points and the pose data includes: Based on a preset first extrinsic parameter between the first imaging module and the second imaging module, the fourth feature point located in the feature extraction region is mapped to the coordinate system of the first imaging module to obtain multiple fifth feature points. Based on the second feature point and the fifth feature point, a visual reprojection error is constructed; The pose information is output based on the visual reprojection error and the pose data.

14. The electronic device according to claim 13, characterized in that, The angle between the imaging surface of the first imaging module and the imaging surface of the second imaging module is 180 degrees - 1 / 2 (the first horizontal field of view of the first imaging module + the second horizontal field of view of the second imaging module).

15. A non-volatile computer-readable storage medium containing a computer program, characterized in that, When the computer program is executed by one or more processors, it implements the positioning method according to any one of claims 1-11.

Citation Information

Patent Citations

  • Robot and pose estimation method and device thereof

    CN111160298A

  • Feature extraction method, visual positioning method and device, medium and electronic equipment

    CN113902932A