Pose determination method and apparatus, computer-readable storage medium, and electronic device

CN117994333BActive Publication Date: 2026-09-04GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211337302.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-28
Publication Date
2026-09-04
Estimated Expiration
2042-10-28

AI Technical Summary

Technical Problem

[0004]本公开提供一种位姿确定方法、位姿确定装置、计算机可读存储介质和电子设备,进而至少在一定程度上克服视觉定位精度差的问题

Benefits of technology

[0009]In some embodiments of this disclosure, a first two-dimensional feature point matching the previous frame of a color image captured by the first camera is determined. A second two-dimensional feature point matching the previous frame of a color image captured by the second camera is also determined. The second two-dimensional feature points are then transformed into the coordinate system of the first camera to obtain a third two-dimensional feature point. Based on the first two-dimensional feature point, the third two-dimensional feature point, the three-dimensional feature point of the previous frame of a color image captured by the first camera in the world coordinate system, and the three-dimensional feature point of the previous frame of a color image captured by the second camera in the world coordinate system, the pose of the first camera when capturing the current frame of the color image is determined. This disclosure transforms the feature points captured by the second camera into the coordinate system of the first camera for pose calculation together with the feature points captured by the first camera. Since the feature points come from at least two cameras and the coordinate systems are unified, more feature points are collected, resulting in more comprehensive feature points participating in the unified processing. This leads to a more accurate pose determination and improved positioning precision. In addition, the pose determination process disclosed herein takes into account the correlation between frames, combines the feature information of the previous frame image, and uses the data of the previous frame for constraints, which further improves the accuracy of positioning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117994333B_ABST
    Figure CN117994333B_ABST
Patent Text Reader

Abstract

The present disclosure provides a pose determination method, a pose determination device, a computer readable storage medium and an electronic device, and relates to the technical field of computer vision. The pose determination method comprises the following steps: determining first two-dimensional feature points on a current frame color image collected by a first camera and matched with a previous frame color image collected by the first camera, determining second two-dimensional feature points on a current frame color image collected by a second camera and matched with a previous frame color image collected by the second camera, converting the second two-dimensional feature points to a first camera coordinate system to obtain third two-dimensional feature points, and determining a pose of the first camera when collecting the current frame color image according to the first two-dimensional feature points, the third two-dimensional feature points, three-dimensional feature points of the previous frame color image collected by the first camera in a world coordinate system, and three-dimensional feature points of the previous frame color image collected by the second camera in the world coordinate system. The present disclosure can improve the positioning accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer vision technology, and more specifically, to a pose determination method, a pose determination device, a computer-readable storage medium, and an electronic device. Background Technology

[0002] In the field of computer vision technology, visual positioning is a technique that uses images captured by a camera to determine the camera's pose in the real world. It has important application value in fields such as augmented reality, virtual reality, robotics, and intelligent transportation.

[0003] In scenarios where multiple cameras perform visual positioning, poor positioning accuracy may occur. Summary of the Invention

[0004] This disclosure provides a pose determination method, a pose determination device, a computer-readable storage medium, and an electronic device, thereby overcoming, at least to some extent, the problem of poor visual positioning accuracy.

[0005] According to a first aspect of this disclosure, a pose determination method is provided, applied to a terminal device, the terminal device being configured with a first camera and at least one second camera. The pose determination method includes: acquiring a current frame color image captured by the first camera, and determining a first two-dimensional feature point on the current frame color image captured by the first camera that matches a previous frame color image captured by the first camera; acquiring a current frame color image captured by the second camera, and determining a second two-dimensional feature point on the current frame color image captured by the second camera that matches a previous frame color image captured by the second camera; converting the second two-dimensional feature point into a third two-dimensional feature point in the first camera coordinate system using a transformation matrix between the first camera coordinate system and the second camera coordinate system; and determining the pose of the first camera when acquiring the current frame color image based on the first two-dimensional feature point, the third two-dimensional feature point, a three-dimensional feature point of the previous frame color image captured by the first camera in the world coordinate system, and a three-dimensional feature point of the previous frame color image captured by the second camera in the world coordinate system.

[0006] According to a second aspect of this disclosure, a pose determination device is provided, configured in a terminal device, the terminal device further comprising a first camera and at least one second camera. The pose determination device includes: a first feature point determination module, configured to acquire a current frame color image captured by the first camera and determine a first two-dimensional feature point on the current frame color image captured by the first camera that matches a previous frame color image captured by the first camera; a second feature point determination module, configured to acquire a current frame color image captured by the second camera and determine a second two-dimensional feature point on the current frame color image captured by the second camera that matches a previous frame color image captured by the second camera; a feature point transformation module, configured to convert the second two-dimensional feature point into a third two-dimensional feature point in the first camera coordinate system using a transformation matrix between the first camera coordinate system and the second camera coordinate system; and a pose determination module, configured to determine the pose of the first camera when acquiring the current frame color image based on the first two-dimensional feature point, the third two-dimensional feature point, a three-dimensional feature point of the previous frame color image captured by the first camera in the world coordinate system, and a three-dimensional feature point of the previous frame color image captured by the second camera in the world coordinate system.

[0007] According to a third aspect of this disclosure, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the pose determination method described above.

[0008] According to a fourth aspect of this disclosure, an electronic device is provided, including a processor; and a memory for storing one or more programs, which, when executed by the processor, cause the processor to implement the pose determination method described above.

[0009] In some embodiments of this disclosure, a first two-dimensional feature point matching the previous frame of a color image captured by the first camera is determined. A second two-dimensional feature point matching the previous frame of a color image captured by the second camera is also determined. The second two-dimensional feature points are then transformed into the coordinate system of the first camera to obtain a third two-dimensional feature point. Based on the first two-dimensional feature point, the third two-dimensional feature point, the three-dimensional feature point of the previous frame of a color image captured by the first camera in the world coordinate system, and the three-dimensional feature point of the previous frame of a color image captured by the second camera in the world coordinate system, the pose of the first camera when capturing the current frame of the color image is determined. This disclosure transforms the feature points captured by the second camera into the coordinate system of the first camera for pose calculation together with the feature points captured by the first camera. Since the feature points come from at least two cameras and the coordinate systems are unified, more feature points are collected, resulting in more comprehensive feature points participating in the unified processing. This leads to a more accurate pose determination and improved positioning precision. In addition, the pose determination process disclosed herein takes into account the correlation between frames, combines the feature information of the previous frame image, and uses the data of the previous frame for constraints, which further improves the accuracy of positioning.

[0010] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0011] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure. It is obvious that the drawings described below are merely some embodiments of this disclosure, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort. In the drawings:

[0012] Figure 1 A schematic diagram of the system architecture of the pose determination system according to an embodiment of the present disclosure is shown;

[0013] Figure 2 A schematic diagram showing the arrangement of dual cameras on a terminal device according to an embodiment of the present disclosure is provided.

[0014] Figure 3 A schematic diagram showing the placement angle of the dual cameras according to an embodiment of the present disclosure is provided.

[0015] Figure 4 A schematic diagram of the various processing stages involved in the pose determination scheme of this disclosure embodiment is shown;

[0016] Figure 5 A flowchart illustrating an exemplary embodiment of the pose determination method of this disclosure is shown schematically.

[0017] Figure 6 A schematic diagram of dual-camera point-to-point matching according to an embodiment of the present disclosure is shown;

[0018] Figure 7 A flowchart illustrating the positioning initialization process according to an embodiment of this disclosure is shown;

[0019] Figure 8 A schematic diagram illustrating the determination of two planes according to an embodiment of the present disclosure is shown;

[0020] Figure 9 A schematic diagram illustrating the determination of the ground plane according to an embodiment of the present disclosure is shown;

[0021] Figure 10 A flowchart illustrating the process of determining the transformation matrix between the first camera coordinate system and the world coordinate system according to an embodiment of this disclosure is shown.

[0022] Figure 11 A block diagram of a pose determination apparatus according to a first exemplary embodiment of the present disclosure is shown schematically.

[0023] Figure 12 A block diagram of a pose determination apparatus according to a second exemplary embodiment of the present disclosure is shown schematically.

[0024] Figure 13 A block diagram of a pose determination apparatus according to a third exemplary embodiment of the present disclosure is shown schematically.

[0025] Figure 14 A block diagram of a pose determination apparatus according to a fourth exemplary embodiment of the present disclosure is shown schematically.

[0026] Figure 15 A block diagram of an electronic device according to an exemplary embodiment of the present disclosure is shown schematically. Detailed Implementation

[0027] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided to make this disclosure more comprehensive and complete, and to fully convey the concept of the example embodiments to those skilled in the art. The described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a full understanding of embodiments of this disclosure. However, those skilled in the art will recognize that the technical solutions of this disclosure can be practiced with one or more of the specific details omitted, or other methods, components, apparatus, steps, etc., can be employed. In other instances, well-known technical solutions are not shown or described in detail to avoid obscuring various aspects of this disclosure.

[0028] Furthermore, the accompanying drawings are merely illustrative of this disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0029] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all steps. For example, some steps may be broken down, while others may be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances. Furthermore, all terms such as "first," "second," and "third" used below are for distinction purposes only and should not be construed as limiting the scope of this disclosure.

[0030] Visual positioning technology enables computer devices to autonomously perceive their own position and orientation within their environment, allowing them to perform any user-initiated task, such as tracking, monitoring, interaction, displaying images, or playing audio. The accuracy of this positioning greatly impacts the functionality of the computer device.

[0031] To improve the accuracy of device visual positioning, this disclosure provides a new positioning scheme.

[0032] Figure 1 A schematic diagram of the system architecture of a pose determination system according to an embodiment of the present disclosure is shown. (Reference) Figure 1 The terminal device 1 may include a processor 100, a first camera 110, and at least one second camera 120.

[0033] Terminal device 1 may include, for example, robots, intelligent monitoring equipment, intelligent tracking equipment, etc. It can be a single device or a device system composed of multiple physical units.

[0034] For example, terminal device 1 could be a robot dog. A robot dog is a type of robot with advantages such as flexibility and strong mobility, and can perform tasks such as security patrol, transporting goods, and providing emotional companionship.

[0035] The first camera 110 and at least one second camera 120 serve as input sensors for the pose determination scheme of this disclosure embodiment, and can transmit the sensed color image and depth image to the processor 100.

[0036] For example, the first camera 110 and the second camera 120 can be Realsense D455 cameras. A Realsense D455 camera consists of one RGB camera, two IR (infrared) cameras, and one IR transmitter. The RGB camera outputs a color image, and the two IR cameras can output a dense depth map aligned with the color image. The Realsense D455 camera has a field of view (FOV) of 90° horizontally and 65° vertically.

[0037] In the case where terminal device 1 includes a first camera 110 and a second camera 120, the first camera 110 may be a left-eye camera and the second camera 120 may be a right-eye camera. In the following embodiments, the left-eye camera may be understood as the first camera 110 and the right-eye camera may be understood as the second camera 120. However, it should be understood that "left," "right," "first," and "second" are merely exemplary descriptions for distinction. In other embodiments of this disclosure, the first camera 110 may be a right-eye camera and the second camera 120 may be a left-eye camera, and this disclosure does not impose any limitations on this.

[0038] Taking two cameras, a first camera 110 and a second camera 120, as an example, Figure 2 A schematic diagram illustrating the placement of the dual cameras according to an embodiment of this disclosure on a terminal device is shown. It should be understood that... Figure 2 The arrangement shown is merely an example. Depending on the type of terminal device and the available space for the camera, there may be various other arrangements, and this disclosure does not limit them.

[0039] Figure 3 A schematic diagram showing the placement angle of the dual cameras according to an embodiment of this disclosure is provided. For the first camera 110 and the second camera 120, both vertically positioned, their viewing angles are both 65°, corresponding to... Figure 3 Angles A and B are shown in the diagram. When positioning the camera, the leftmost line of sight of the first camera 110 can be parallel to the rightmost line of sight of the second camera 120. At this point, both cameras can achieve their maximum field of view, which is 130°. Figure 3 Angle C in the diagram. There is a small shared viewing area between the first camera 110 and the second camera 120. Based on the above angle design, the angle between the placement of the first camera 110 and the second camera 120 can be determined to be 115°, corresponding to... Figure 3 Angle D in the middle.

[0040] Therefore, the first camera 110 and the second camera 120 are placed vertically side by side at an angle of 115°, with the field of view of the two cameras being 130° horizontally and 90° vertically. This maximizes the superposition of the fields of view of the two cameras, effectively increasing the field of view of the terminal device 1 and providing more sufficient accuracy for subsequent positioning algorithms.

[0041] Furthermore, the first camera 110 and the second camera 120 support multi-camera hardware synchronization. They can be connected via wires, and the same pulse signal can trigger simultaneous exposure, achieving hardware synchronization across multiple cameras. After hardware synchronization, the images input into the subsequent positioning algorithm are those captured at the same time. This avoids additional errors caused by inconsistent shooting times from multiple cameras.

[0042] After positioning the first and second cameras as described above, the intrinsic and extrinsic parameters of each camera can be calibrated for use in subsequent algorithms. This disclosure does not restrict the calibration process.

[0043] In the pose determination scheme of this embodiment, the processor 100 can acquire the current frame color image captured by the first camera 110 and determine the first two-dimensional feature point on the current frame color image captured by the first camera 110 that matches the previous frame color image captured by the first camera 110.

[0044] The processor 100 can acquire the current frame color image captured by the second camera 120 and determine a second two-dimensional feature point on the current frame color image captured by the second camera 120 that matches the previous frame color image captured by the second camera 120. The processor 100 uses a transformation matrix between the first camera coordinate system of the first camera 110 and the second camera coordinate system of the second camera to convert the second two-dimensional feature point into a third two-dimensional feature point in the first camera coordinate system.

[0045] Next, the processor 100 can determine the pose of the first camera 110 when acquiring the current frame of color image based on the first two-dimensional feature points, the third two-dimensional feature points, the three-dimensional feature points of the previous frame of color image acquired by the first camera 110 in the world coordinate system, and the three-dimensional feature points of the previous frame of color image acquired by the second camera 120 in the world coordinate system.

[0046] When the terminal device 1 is equipped with multiple second cameras 120, the feature point data of each second camera 120 can be mapped to the first camera coordinate system of the first camera 110 for processing.

[0047] It is understandable that the positions of the first camera 110 and the second camera 120 on the terminal device 1 are fixed. Once the current pose of the first camera 110 is determined, the current pose of the second camera 120 and the current pose of the terminal device 1 can be obtained.

[0048] Furthermore, when the terminal device 1 is equipped with two or more cameras, any one camera can be designated as the first camera 110 in the algorithm implementation, and the remaining cameras can be designated as the second camera 120.

[0049] Based on the pose determination scheme of this disclosure, feature points acquired by the second camera 120 are transformed into the coordinate system of the first camera, and then used together with feature points acquired by the first camera 110 for pose calculation. Since the feature points come from at least two cameras and the coordinate systems are unified, more feature points are acquired, meaning more comprehensive feature points participate in the unified processing, resulting in a more accurate pose determination and improved positioning precision. Furthermore, the pose determination process of this disclosure considers the inter-frame correlation, combining feature information from the previous frame and using data from the previous frame for constraints, further improving positioning precision.

[0050] The pose determination process for implementing the embodiments of this disclosure involves multiple processing stages. (See reference...) Figure 4 The processing stages involved include, but are not limited to, coordinate system alignment, positioning initialization, and real-time positioning.

[0051] During the coordinate system alignment phase, the terminal device determines the transformation matrix between the first camera coordinate system and the world coordinate system.

[0052] First, the terminal device can construct a point cloud using the depth images output by the first camera and the second camera. Specifically, the corresponding 3D spatial points from the two depth images can be merged to obtain a point cloud of 3D feature points.

[0053] Next, the terminal device uses a plane detection algorithm to extract plane information from the point cloud and filters out a specified plane (such as the ground plane) based on the extracted plane information.

[0054] Then, the terminal device can calculate the transformation matrix based on the normal vector and gravity vector of the specified plane to align the first camera coordinate system with the world coordinate system.

[0055] Furthermore, it is understandable that, based on the pre-calibrated intrinsic and extrinsic parameters, the transformation matrix between the first camera coordinate system and the second camera coordinate system can be obtained. In this case, the transformation matrix between the second camera coordinate system and the world coordinate system can also be obtained, achieving alignment between the first camera coordinate system, the second camera coordinate system, and the world coordinate system.

[0056] During the positioning initialization phase, the terminal device can determine the pose of the first camera when it initially captures a color image. It should be understood that the pose of the camera when capturing an image, as described in this disclosure, refers to the pose in the world coordinate system.

[0057] On the one hand, the terminal device can determine the three-dimensional feature points corresponding to the initial frame color image captured by the first camera. These three-dimensional feature points are feature points in the coordinate system of the first camera.

[0058] On the other hand, an initial rotation matrix and an initial translation vector can be set. For example, the initial rotation matrix can be the identity matrix, and the initial translation vector can be [0,0,0].

[0059] After determining the three-dimensional feature points corresponding to the initial frame color image, as well as the initial rotation matrix and initial translation vector, the positioning initialization in the first camera coordinate system is completed.

[0060] Next, by combining the transformation matrix between the first camera coordinate system and the world coordinate system determined in the coordinate system alignment stage, the positioning initialization result in the first camera coordinate system can be converted into the positioning initialization result in the world coordinate system, that is, the pose of the first camera when acquiring the initial frame color image can be determined.

[0061] For the real-time positioning phase, the terminal device can combine the initial pose determined during the positioning initialization phase to obtain the pose of the current frame in real time. During this process, the features of the second camera can be transferred to the coordinate system of the first camera and combined with the features of the first camera to solve for the pose, thus completing the pose prediction of the current frame.

[0062] The pose determination method of this disclosure will be described by way of example below.

[0063] Figure 5 A flowchart illustrating an exemplary embodiment of the pose determination method of this disclosure is shown schematically. (Reference) Figure 5 The pose determination method may include the following steps:

[0064] S52. Obtain the current frame color image captured by the first camera, and determine the first two-dimensional feature point on the current frame color image captured by the first camera that matches the previous frame color image captured by the first camera.

[0065] In an exemplary embodiment of this disclosure, the current frame color image is the color image captured by the camera at the current moment, and the previous frame color image is the color image captured by the camera in the previous frame. This disclosure does not limit the image size, shooting scene, etc.

[0066] After acquiring the current frame color image captured by the first camera, the terminal device can extract the feature points of the current color image captured by the first camera.

[0067] The feature extraction algorithms used in the exemplary embodiments of this disclosure may include, but are not limited to, the FAST feature point detection algorithm, the DOG feature point detection algorithm, the Harris feature point detection algorithm, the SIFT feature point detection algorithm, and the SURF feature point detection algorithm. Feature descriptors may include, but are not limited to, the BRIEF feature point descriptor, the BRISK feature point descriptor, and the FREAK feature point descriptor.

[0068] According to one embodiment of this disclosure, the combination of feature extraction algorithm and feature descriptor can be the FAST feature point detection algorithm and the BRIEF feature point descriptor. According to other embodiments of this disclosure, the combination of feature extraction algorithm and feature descriptor can be the DOG feature point detection algorithm and the FREAK feature point descriptor.

[0069] It should be understood that different combinations can be used for different texture scenes. For example, for strong texture scenes, the FAST feature point detection algorithm and BRIEF feature point descriptor can be used for feature extraction; for weak texture scenes, the DOG feature point detection algorithm and FREAK feature point descriptor can be used for feature extraction.

[0070] In the processing of the previous color image corresponding to the current color image, there is also a process of extracting feature points. Therefore, the terminal device can use the feature points of the current color image captured by the first camera and the feature points of the previous color image captured by the first camera to determine the matching two-dimensional feature points between the two images, that is, the first two-dimensional feature points as described in this disclosure.

[0071] Specifically, optical flow can be used to determine the matching relationship of feature points. That is, optical flow tracking is performed using feature points of the current frame color image captured by the first camera and feature points of the previous frame color image captured by the first camera to determine the first two-dimensional feature points. In addition, other image matching methods can also be used to determine 2D-2D feature point pairs, and this disclosure does not limit this.

[0072] S54. Acquire the current frame color image captured by the second camera, and determine the second two-dimensional feature points on the current frame color image captured by the second camera that match the previous frame color image captured by the second camera.

[0073] It should be understood that, compared with step S52, although both steps include descriptions of the current frame color image and the previous frame color image, the current frame color image and the previous frame color image in step S52 are acquired by the first camera, while the current frame color image and the previous frame color image in step S54 are acquired by the second camera.

[0074] After acquiring the current frame color image captured by the second camera, the terminal device can extract feature points from the current color image captured by the second camera. The method for extracting feature points can be the same as that for extracting feature points in step S52, and will not be described again.

[0075] The terminal device can use the feature points of the current frame color image captured by the second camera and the feature points of the previous frame color image captured by the second camera to perform optical flow tracking in order to determine the second two-dimensional feature points.

[0076] S56. Use the transformation matrix between the first camera coordinate system of the first camera and the second camera coordinate system of the second camera to convert the second two-dimensional feature points into third two-dimensional feature points in the first camera coordinate system.

[0077] In an exemplary embodiment of this disclosure, for the purpose of distinction, the camera coordinate system of the first camera is referred to as the first camera coordinate system, and the camera coordinate system of the second camera is referred to as the second camera coordinate system.

[0078] With the placement of the first and second cameras on the terminal device fixed, the intrinsic and extrinsic parameters of the first and second cameras are calibrated in advance. From the calibration results, the transformation matrix between the first camera coordinate system of the first camera and the second camera coordinate system of the second camera can be determined.

[0079] The terminal device can acquire the transformation matrix between the first camera coordinate system and the second camera coordinate system, as well as the depth information of the second two-dimensional feature points, and determine the third two-dimensional feature points based on the transformation matrix, the depth information of the second two-dimensional feature points, and the second feature points. The third two-dimensional feature point is a two-dimensional feature point transformed from the second two-dimensional feature point into the first camera coordinate system.

[0080] Specifically, the transformation matrix, the depth information of the second two-dimensional feature points, and the second two-dimensional feature points themselves can be multiplied, and the result of the multiplication can be normalized to determine the third two-dimensional feature points. Here, the second two-dimensional feature points in the multiplication operation refer to the position coordinate information of these feature points. The third two-dimensional feature points can be determined using Formula 1.

[0081]

[0082] Among them, T lr Let d be the transformation matrix between the first camera coordinate system and the second camera coordinate system. j This represents the depth value of the second two-dimensional feature point. These are the second two-dimensional feature points.

[0083] S58. Based on the first two-dimensional feature points, the third two-dimensional feature points, the three-dimensional feature points of the previous color image captured by the first camera in the world coordinate system, and the three-dimensional feature points of the previous color image captured by the second camera in the world coordinate system, determine the pose of the first camera when capturing the current color image.

[0084] In an exemplary embodiment of this disclosure, the first two-dimensional feature points and the third two-dimensional feature points constitute two-dimensional coordinate information, and the three-dimensional feature points of the previous frame color image captured by the first camera in the world coordinate system and the three-dimensional feature points of the previous frame color image captured by the second camera in the world coordinate system constitute three-dimensional coordinate information.

[0085] The terminal device can associate two-dimensional coordinate system information with three-dimensional coordinate information to obtain point-pair information, and use this point-pair information to solve the perspective n-point (PnP) problem, and combine the solution results to determine the pose of the first camera when acquiring the current frame color image.

[0086] PnP is a method in the field of machine vision that determines the relative pose of a camera based on n feature points in a scene. Specifically, it determines the camera's rotation matrix and translation vector based on n feature points in the scene.

[0087] It should be noted that the process of determining the three-dimensional feature points of the previous frame color image in the world coordinate system can be performed during the processing of the current frame or during the processing of the previous frame, and this disclosure does not limit it in this way.

[0088] The process of determining the three-dimensional feature points of the previous color image captured by the first camera in the world coordinate system is explained below.

[0089] First, the terminal device can acquire the previous frame of color image captured by the first camera and extract the feature points of the previous frame of color image captured by the first camera. The process of extracting feature points is the same as that in step S52, and will not be described again here.

[0090] Next, the terminal device can use the previous frame depth image, which is aligned with the previous frame color image captured by the first camera, to spatially project the feature points of the previous frame color image captured by the first camera, so as to obtain the three-dimensional feature points of the previous frame color image captured by the first camera in the first camera coordinate system. This previous frame depth image can be output by the first camera, or it can be obtained by other depth cameras equipped on the terminal device; this disclosure does not limit this.

[0091] Furthermore, to further improve the positioning accuracy of this disclosure, constraints can be imposed on the spatial projection process. Specifically, the terminal device can use a previous frame depth image aligned with the previous frame color image captured by the first camera to spatially project feature points within a predetermined depth range from the feature points of the previous frame color image captured by the first camera, thereby obtaining the three-dimensional feature points of the previous frame color image captured by the first camera in the first camera coordinate system.

[0092] The predetermined depth range is determined based on the depth measurement range. The value of the predetermined depth range may vary depending on the type and model of the depth camera. This disclosure does not limit the specific value of the predetermined depth range. For example, feature points with depth values ​​greater than 0.5m and less than 6m are spatially projected.

[0093] Then, the terminal device can transform the 3D feature points in the first camera coordinate system based on the pose of the first camera when it captured the previous color image, to obtain the 3D feature points of the previous color image captured by the first camera in the world coordinate system. (Refer to Formula 2:)

[0094]

[0095] in, These are the three-dimensional feature points of the previous color image captured by the first camera in the world coordinate system. T represents the three-dimensional feature points of the previous color image captured by the first camera in the first camera coordinate system. w_last The pose of the first camera when it captured the previous frame of the color image.

[0096] It should be noted that the pose of the first camera when acquiring the previous color image can be determined during the processing of the previous image. In other words, the pose corresponding to the previous frame is known during the processing of the current frame. The initial pose is explained in the localization initialization process of this disclosure.

[0097] The process of determining the three-dimensional feature points of the previous color image captured by the second camera in the world coordinate system is explained below.

[0098] First, the terminal device can acquire the previous frame of color image captured by the second camera and extract the feature points of the previous frame of color image captured by the second camera. The process of extracting feature points is the same as that in step S52, and will not be described again here.

[0099] Next, the terminal device can use the previous frame depth image, which is aligned with the previous frame color image captured by the second camera, to spatially project the feature points of the previous frame color image captured by the second camera, so as to obtain the three-dimensional feature points of the previous frame color image captured by the second camera in the second camera coordinate system. This previous frame depth image can be output by the second camera, or it can be obtained by other depth cameras equipped on the terminal device; this disclosure does not limit this.

[0100] Similarly, to further improve the positioning accuracy of this disclosure, constraints can be imposed on the spatial projection process. Specifically, the terminal device can use a previous frame depth image aligned with the previous frame color image captured by the second camera to spatially project feature points within a predetermined depth range from the feature points of the previous frame color image captured by the second camera, thereby obtaining the three-dimensional feature points of the previous frame color image captured by the second camera in the second camera coordinate system.

[0101] The predetermined depth range is determined based on the depth measurement range. The value of the predetermined depth range may vary depending on the type and model of the depth camera. This disclosure does not limit the specific value of the predetermined depth range. For example, feature points with depth values ​​greater than 0.5m and less than 6m are spatially projected.

[0102] Subsequently, the terminal device can use the transformation matrix between the first camera coordinate system and the second camera coordinate system to convert the three-dimensional feature points of the previous frame color image captured by the second camera in the second camera coordinate system into three-dimensional feature points in the first camera coordinate system.

[0103] Then, the terminal device can transform the three-dimensional feature points in the first camera coordinate system based on the pose of the first camera when it captured the previous color image, so as to obtain the three-dimensional feature points of the previous color image captured by the second camera in the world coordinate system.

[0104] The above process will be explained below with reference to Formula 3:

[0105]

[0106] in, The three-dimensional feature points of the previous color image captured by the second camera in the world coordinate system. T represents the three-dimensional feature points of the previous color image captured by the second camera in the second camera coordinate system. w_last T represents the pose of the first camera when it captured the previous color image. lr This is the transformation matrix between the first camera coordinate system and the second camera coordinate system.

[0107] Based on the above point-to-point matching relationships, Figure 6 A schematic diagram is given to achieve PnP pose solving by matching point pairs between the first camera and the second camera, which involves the matching relationship of 2D-2D feature points in the current frame and the matching relationship of 3D-2D feature points.

[0108] In the process of determining the three-dimensional feature points of the previous color image in the world coordinate system, the pose of the first camera when acquiring the previous color image was utilized. The process of determining the initial pose of the first camera is described below.

[0109] According to some embodiments of this disclosure, firstly, the terminal device can acquire an initial frame color image captured by the first camera and extract feature points from the initial frame color image captured by the first camera. The process of extracting feature points is the same as that in step S52, and will not be described again here.

[0110] Next, the terminal device can use the initial frame depth image, which is aligned with the initial frame color image captured by the first camera, to spatially project the feature points of the initial frame color image captured by the first camera, so as to obtain the three-dimensional feature points of the initial frame color image captured by the first camera in the first camera coordinate system.

[0111] Similarly, to further improve the positioning accuracy of this disclosure, constraints can be imposed on the spatial projection process. Specifically, the terminal device can use feature points within a predetermined depth range from the feature points of the initial frame color image acquired by the first camera to perform spatial projection, thereby obtaining the three-dimensional feature points of the initial frame color image acquired by the first camera in the first camera coordinate system.

[0112] The predetermined depth range is determined based on the depth measurement range. The value of the predetermined depth range may vary depending on the type and model of the depth camera. This disclosure does not limit the specific value of the predetermined depth range. For example, feature points with depth values ​​greater than 0.5m and less than 6m are spatially projected.

[0113] Subsequently, the terminal device can determine the initial positioning result of the first camera in the first camera coordinate system based on the three-dimensional feature points, initial rotation matrix, and initial translation vector of the initial frame color image captured by the first camera in the first camera coordinate system.

[0114] In one embodiment of this disclosure, the initial rotation matrix can be set to the identity matrix, and the translation vector can be set to [0,0,0].

[0115] It should be noted that, given the 3D feature points, initial rotation matrix, and initial translation vector of the initial frame color image acquired by the first camera in the first camera coordinate system, only the pose of the first camera in the first camera coordinate system is determined at this point. To obtain the pose applicable to subsequent current frame processing, this pose needs to be transformed to obtain the pose of the first camera in the world coordinate system.

[0116] Specifically, the terminal device can use the transformation matrix between the first camera coordinate system and the world coordinate system to transform the initial positioning result of the first camera in the first camera coordinate system, so as to determine the pose of the first camera when acquiring the initial frame color image.

[0117] According to some other embodiments of this disclosure, the process of determining the initial pose of the first camera can also be combined with the feature data of the second camera, and this process will be described below.

[0118] On the one hand, the terminal device can determine the three-dimensional feature points of the initial frame color image captured by the first camera in the first camera coordinate system.

[0119] On the other hand, the terminal device can acquire the initial frame color image captured by the second camera and extract the feature points of the initial frame color image captured by the second camera. The process of extracting feature points is the same as that in step S52, and will not be described again here.

[0120] The terminal device can use the initial frame depth image, which is aligned with the initial frame color image captured by the second camera, to spatially project the feature points of the initial frame color image captured by the second camera, so as to obtain the three-dimensional feature points of the initial frame color image captured by the second camera in the second camera coordinate system.

[0121] Similarly, constraints can be imposed on the spatial projection process. Specifically, the terminal device can use feature points within a predetermined depth range from the feature points of the initial frame color image captured by the second camera to perform spatial projection, thereby obtaining the three-dimensional feature points of the initial frame color image captured by the second camera in the second camera coordinate system.

[0122] The predetermined depth range is determined based on the depth measurement range. The value of the predetermined depth range may vary depending on the type and model of the depth camera. This disclosure does not limit the specific value of the predetermined depth range. For example, feature points with depth values ​​greater than 0.5m and less than 6m are spatially projected.

[0123] Next, the terminal device can use the transformation matrix between the first camera coordinate system of the first camera and the second camera coordinate system of the second camera to transform the three-dimensional feature points of the initial frame color image captured by the second camera in the second camera coordinate system to the three-dimensional feature points in the first camera coordinate system.

[0124] The transformed 3D feature points and the 3D feature points of the initial frame color image acquired by the first camera in the first camera coordinate system can be merged to obtain merged 3D feature points. It can be understood that the merged 3D feature points are 3D feature points in the first camera coordinate system.

[0125] Subsequently, the terminal device can determine the initial positioning result of the first camera in the first camera coordinate system based on the merged 3D feature points, the initial rotation matrix, and the initial translation vector. For example, the initial rotation matrix can be set to the identity matrix, and the translation vector can be set to [0, 0, 0].

[0126] Then, the terminal device can use the transformation matrix between the first camera coordinate system and the world coordinate system to transform the initial positioning result of the first camera in the first camera coordinate system, so as to determine the pose of the first camera when acquiring the initial frame color image.

[0127] The following will refer to Figure 7 The positioning initialization process of the embodiments of this disclosure will be described.

[0128] In step S702, the terminal device can acquire the initial frame color image captured by the first camera and extract the feature points of the initial frame color image captured by the first camera.

[0129] In step S704, the terminal device can perform spatial projection by combining a depth image aligned with the initial frame color image captured by the first camera, to obtain the three-dimensional feature points of the initial frame color image captured by the first camera in the first camera coordinate system. As described in the above embodiment, the three-dimensional feature points determined in step S704 may also include the three-dimensional feature points corresponding to the initial frame color image captured by the second camera.

[0130] In step S706, the terminal device can determine the initial positioning result of the first camera in the first camera coordinate system based on the three-dimensional feature points, initial rotation matrix and initial translation vector determined in step S704.

[0131] In step S708, the terminal device can use the transformation matrix between the first camera coordinate system and the world coordinate system to transform the initial positioning result, so as to determine the pose of the first camera when acquiring the initial frame color and complete the positioning initialization.

[0132] In the above processing, a transformation matrix between the first camera coordinate system and the world coordinate system is used. For this predetermined transformation matrix, the embodiments of this disclosure provide a coordinate system alignment scheme. Specifically, coordinate system alignment is achieved by combining depth information. For distinction, in the following embodiments, the terminology of a reference depth image is used to describe the coordinate system alignment process.

[0133] First, the terminal device can acquire the reference depth image output by the first camera.

[0134] Next, after determining the existence of a specified plane in the scene by combining the reference depth image output by the first camera, the terminal device can determine the transformation matrix between the first camera coordinate system and the world coordinate system based on the normal vector and gravity vector of the specified plane.

[0135] Wherein, the gravity vector can be N g (0, 0, 1), in this case, the specified plane is usually the ground plane to match the scenario where the terminal device is, for example, a robot dog. However, it is understood that the specified plane can also be a plane manually specified in a specific scenario, such as a wall, a desktop, etc., and this disclosure does not limit it.

[0136] If the normal vector of the specified plane is denoted as n c , will n c Rotation R wc After that, it can be used with N gBy aligning the coordinates, the first camera coordinate system can be aligned with the world coordinate system. Where R... wc R is the transformation matrix between the first camera coordinate system and the world coordinate system. wc The axis of rotation ω can be determined by N g With n c The cross product yields the result, as shown in Formula 4:

[0137] ω=N g ×n c (Formula 4)

[0138] R wc The rotation angle θ can be determined by N g With n c The dot product is shown in Formula 5:

[0139]

[0140] The rotation axis ω and the rotation angle θ constitute the rotation vector between the first camera coordinate system and the world coordinate system. According to the Rodriguez formula, the terminal device can calculate the transformation matrix R between the first camera coordinate system and the world coordinate system. wc Therefore, the coordinate system alignment process ends.

[0141] During the above processing, if the specified plane does not exist in the scene, the terminal device can return to the step of acquiring the reference depth image, reacquire the reference depth image, and perform the process of determining whether the specified plane exists.

[0142] The process of determining the specified plane is explained below.

[0143] First, the terminal device can combine the reference depth image output by the first camera to determine the point cloud corresponding to the first camera, which is denoted as the reference point cloud.

[0144] According to some embodiments of this disclosure, the terminal device determines the three-dimensional spatial point of each pixel in the reference depth image output by the first camera, based on the pixel, the pixel's depth value, and the camera intrinsic parameters of the first camera. Formula 6 illustrates the method for determining the three-dimensional spatial point here:

[0145] P = z * K -1 *p (Formula 6)

[0146] Where P represents a three-dimensional spatial point projected into space, z represents the depth value of that pixel, and K... -1 represents the inverse of the camera intrinsic parameter matrix, and p represents the coordinate position of the pixel.

[0147] In these embodiments, a reference point cloud corresponding to the first camera can be constructed from the three-dimensional spatial points obtained through this process.

[0148] According to some other embodiments of this disclosure, on one hand, the terminal device determines the three-dimensional spatial point of each pixel in the reference depth image output by the first camera based on the pixel, the depth value of the pixel, and the camera intrinsic parameters of the first camera.

[0149] On the other hand, the terminal device can acquire the reference depth image output by the second camera and determine the three-dimensional spatial point of each pixel in the reference depth image output by the second camera by combining the above formula 6.

[0150] The terminal device can transform the three-dimensional spatial points of each pixel in the reference depth image output by the second camera according to the transformation matrix between the first camera coordinate system and the second camera coordinate system to obtain the transformed three-dimensional spatial points.

[0151] Therefore, the 3D spatial points of each pixel in the reference depth image output by the first camera are merged with the transformed 3D spatial points to construct the reference point cloud corresponding to the first camera. (Refer to Formula 7:)

[0152] PC_mixture = PC_left + T lr *PC_right (Formula 7)

[0153] Where PC_mixture is the determined reference point cloud, PC_right is the 3D spatial point of each pixel in the reference depth image output by the second camera, PC_left is the 3D spatial point of each pixel in the reference depth image output by the first camera, and T... lr This is the transformation matrix between the first camera coordinate system and the second camera coordinate system.

[0154] In these embodiments, the construction of the reference point cloud incorporates information from the depth image output by the second camera, thereby providing a more comprehensive spatial feature point and improving the accuracy of the algorithm.

[0155] After determining the reference point cloud corresponding to the first camera, the terminal device can extract the planar information of the reference point cloud. This disclosure does not limit the planar extraction method; it can employ methods such as Ransac fitting, normal vector region growing, hierarchical clustering, etc., as long as the planar information in the scene can be extracted. Some embodiments of this disclosure use the hierarchical clustering-based planar extraction algorithm PEAC, see reference... Figure 8 This algorithm can extract two planes. Figure 8 This is just an example; the algorithm described above can be used to extract all planes in a scene.

[0156] It is understandable that the extracted planar information includes, but is not limited to, the planar ID, the planar normal vector, and the distance between the planar and the camera.

[0157] After extracting the plane from the reference point cloud, the terminal device can filter for a specific plane based on the plane information in the reference point cloud. Specifically, the terminal device can filter for a specific plane based on the distance information between the plane and the first camera contained in the plane information of the reference point cloud.

[0158] If the distance information includes a distance within a predetermined distance range, the terminal device can determine a candidate plane corresponding to that distance, and the number of candidate planes determined can be one or more.

[0159] When there is only one candidate plane, the terminal device can identify that candidate plane as the designated plane.

[0160] When there are multiple candidate planes, the terminal device can determine the candidate plane whose distance from the first camera is closest to a distance threshold as the designated plane. This distance threshold is within the aforementioned predetermined distance range.

[0161] Figure 9 The diagram shows the filtering of ground planes. Compared with the results of plane detection, the above distance-based filtering process eliminates planes such as ceilings.

[0162] Taking a robot dog as an example, the terminal device is equipped with a first camera and a second camera. The positions of the two cameras are fixed. In the implementation, the robot dog is controlled to move for a short period of time, only moving on the ground plane. Based on this prior condition, the position of the ground plane in the coordinate system of the first camera is basically fixed. The height of the ground plane from the camera is approximately the same as the height of the robot dog, about 0.3m. Therefore, the predetermined distance range can be set to 0.25m to 0.35m as the ground plane. If multiple candidate planes are selected, the plane closest to 0.3m is taken as the ground plane.

[0163] It should be understood that if the ground plane is not detected during this process, the control terminal device will continuously repeat the above process of determining the plane from the depth image and filtering the plane until the terminal device detects the ground plane.

[0164] The following is for reference. Figure 10 The process of aligning the coordinate system according to the embodiments of this disclosure will be described.

[0165] In step S1002, the terminal device acquires the reference depth image output by the first camera and backprojects the reference depth image to obtain three-dimensional spatial points in space.

[0166] In step S1004, the terminal device acquires the reference depth image output by the second camera and backprojects the reference depth image to obtain three-dimensional spatial points in space.

[0167] In step S1006, the terminal device transforms the three-dimensional spatial points obtained in step S1004 into three-dimensional spatial points in the first camera coordinate system.

[0168] In step S1008, the terminal device merges the three-dimensional spatial points obtained in step S1002 with the three-dimensional spatial points obtained in step S1006 to obtain the reference point cloud corresponding to the first camera.

[0169] In step S1010, the terminal device can extract planar information based on the reference point cloud.

[0170] In step S1012, the terminal device can filter the extracted planes to determine the ground plane;

[0171] In step S1014, the terminal device can use the normal vector of the ground plane and the gravity vector to determine the transformation matrix between the first camera coordinate system and the world coordinate system, so as to complete the alignment between the first camera coordinate system and the world coordinate system.

[0172] Furthermore, since the relationship between the first camera coordinate system and the second camera coordinate system has been determined through calibration, the transformation matrix between the second camera coordinate system and the world coordinate system can also be obtained to achieve alignment of the first camera coordinate system, the second camera coordinate system, and the world coordinate system. Therefore, the coordinate system alignment result can be applied to the pose determination process described above in this disclosure.

[0173] It should be noted that although the steps of the method in this disclosure are described in a specific order in the accompanying drawings, this does not require or imply that the steps must be performed in that specific order, or that all the steps shown must be performed to achieve the desired result. Additional or alternative steps may be omitted, multiple steps may be combined into one step, and / or a step may be broken down into multiple steps.

[0174] Furthermore, this example embodiment also provides a pose determination device. This pose determination device is configured in a terminal device, which is further configured with a first camera and at least one second camera.

[0175] Figure 11 A block diagram of a pose determination apparatus according to an exemplary embodiment of the present disclosure is shown schematically. (Reference) Figure 11 The pose determination device 11 according to an exemplary embodiment of the present disclosure may include a first feature point determination module 111, a second feature point determination module 113, a feature point conversion module 115, and a pose determination module 117.

[0176] Specifically, the first feature point determination module 111 can be used to acquire the current frame color image captured by the first camera and determine the first two-dimensional feature points on the current frame color image captured by the first camera that match the previous frame color image captured by the first camera; the second feature point determination module 113 can be used to acquire the current frame color image captured by the second camera and determine the second two-dimensional feature points on the current frame color image captured by the second camera that match the previous frame color image captured by the second camera; the feature point conversion module 115 can be used to convert the second two-dimensional feature points into third two-dimensional feature points in the first camera coordinate system using the conversion matrix between the first camera coordinate system of the first camera and the second camera coordinate system of the second camera; the pose determination module 117 can be used to determine the pose of the first camera when acquiring the current frame color image based on the first two-dimensional feature points, the third two-dimensional feature points, the three-dimensional feature points of the previous frame color image captured by the first camera in the world coordinate system, and the three-dimensional feature points of the previous frame color image captured by the second camera in the world coordinate system.

[0177] According to an exemplary embodiment of the present disclosure, the first feature point determination module 111 may be configured to perform: extracting feature points of the current frame color image captured by the first camera; and performing optical flow tracking using the feature points of the current frame color image captured by the first camera and the feature points of the previous frame color image captured by the first camera to determine the first two-dimensional feature points.

[0178] According to an exemplary embodiment of the present disclosure, the feature point conversion module 115 can be configured to perform: obtaining the transformation matrix between the first camera coordinate system of the first camera and the second camera coordinate system of the second camera and the depth information of the second two-dimensional feature points; and determining the third two-dimensional feature points based on the transformation matrix, the depth information of the second two-dimensional feature points and the second two-dimensional feature points.

[0179] According to an exemplary embodiment of the present disclosure, the feature point conversion module 115 can be configured to perform: multiplying the conversion matrix, the depth information of the second two-dimensional feature point and the second two-dimensional feature point, and normalizing the result of the multiplication to determine the third two-dimensional feature point.

[0180] According to an exemplary embodiment of this disclosure, a first two-dimensional feature point and a third two-dimensional feature point constitute two-dimensional coordinate information, and three-dimensional feature points of the previous frame color image captured by the first camera in the world coordinate system and three-dimensional feature points of the previous frame color image captured by the second camera in the world coordinate system constitute three-dimensional coordinate information. In this case, the pose determination module 117 can be configured to perform: associating the two-dimensional coordinate information with the three-dimensional coordinate information to obtain point pair information; using the point pair information to solve the perspective n-point problem, and combining the solution results to determine the pose of the first camera when capturing the current frame color image.

[0181] According to exemplary embodiments of this disclosure, reference is made to Figure 12 Compared to the pose determination device 11, the pose determination device 12 may also include a third feature point determination module 121.

[0182] Specifically, the third feature point determination module 121 can be configured to perform the following: acquire the previous frame color image captured by the first camera, extract the feature points of the previous frame color image captured by the first camera; use the previous frame depth image aligned with the previous frame color image captured by the first camera to spatially project the feature points of the previous frame color image captured by the first camera to obtain the three-dimensional feature points of the previous frame color image captured by the first camera in the first camera coordinate system; and transform the three-dimensional feature points in the first camera coordinate system according to the pose of the first camera when acquiring the previous frame color image to obtain the three-dimensional feature points of the previous frame color image captured by the first camera in the world coordinate system.

[0183] According to an exemplary embodiment of the present disclosure, the third feature point determination module 121 may be configured to perform: spatial projection of feature points in the previous frame color image acquired by the first camera that are within a predetermined depth range using a previous frame depth image aligned with the previous frame color image acquired by the first camera, so as to obtain three-dimensional feature points of the previous frame color image acquired by the first camera in the first camera coordinate system; wherein the predetermined depth range is determined based on the range of depth measurement.

[0184] According to an exemplary embodiment of this disclosure, the third feature point determination module 121 may further be configured to perform: acquiring a previous frame color image captured by the second camera, and extracting feature points from the previous frame color image captured by the second camera; spatially projecting the feature points of the previous frame color image captured by the second camera using a previous frame depth image aligned with the previous frame color image captured by the second camera, to obtain three-dimensional feature points of the previous frame color image captured by the second camera in the second camera coordinate system; converting the three-dimensional feature points of the previous frame color image captured by the second camera in the second camera coordinate system into three-dimensional feature points in the first camera coordinate system using a transformation matrix between the first camera coordinate system and the second camera coordinate system; and transforming the three-dimensional feature points in the first camera coordinate system according to the pose of the first camera when capturing the previous frame color image, to obtain the three-dimensional feature points of the previous frame color image captured by the second camera in the world coordinate system.

[0185] According to an exemplary embodiment of the present disclosure, the third feature point determination module 121 may also be configured to perform: spatial projection of feature points in the previous frame color image acquired by the second camera that are within a predetermined depth range using a previous frame depth image aligned with the previous frame color image acquired by the second camera, so as to obtain three-dimensional feature points of the previous frame color image acquired by the second camera in the second camera coordinate system; wherein the predetermined depth range is determined based on the range of depth measurement.

[0186] According to exemplary embodiments of this disclosure, reference is made to Figure 13 Compared to the pose determination device 11, the pose determination device 13 may also include a positioning initialization module 131.

[0187] Specifically, the positioning initialization module 131 can be configured to perform the following: acquire an initial frame color image captured by the first camera, extract feature points from the initial frame color image captured by the first camera; spatially project the feature points of the initial frame color image captured by the first camera using an initial frame depth image aligned with the initial frame color image captured by the first camera, to obtain three-dimensional feature points of the initial frame color image captured by the first camera in the first camera coordinate system; determine the initial positioning result of the first camera in the first camera coordinate system based on the three-dimensional feature points of the initial frame color image captured by the first camera in the first camera coordinate system, the initial rotation matrix, and the initial translation vector; and transform the initial positioning result of the first camera in the first camera coordinate system using a transformation matrix between the first camera coordinate system and the world coordinate system, to determine the pose of the first camera when acquiring the initial frame color image.

[0188] According to exemplary embodiments of this disclosure, reference is made to Figure 14 Compared to the pose determination device 13, the pose determination device 14 may also include a transformation matrix determination module 141.

[0189] Specifically, the transformation matrix determination module 141 can be configured to perform: acquiring a reference depth image output by the first camera; and, if a specified plane is determined by combining the reference depth image output by the first camera, determining the transformation matrix between the first camera coordinate system and the world coordinate system based on the normal vector and gravity vector of the specified plane.

[0190] According to an exemplary embodiment of the present disclosure, the transformation matrix determination module 141 can be configured to perform: determining a reference point cloud corresponding to the first camera by combining the reference depth image output by the first camera; extracting planar information of the reference point cloud; and filtering a specified plane based on the planar information of the reference point cloud.

[0191] According to an exemplary embodiment of the present disclosure, the process of the transformation matrix determination module 141 determining the reference point cloud can be configured to perform: for each pixel in the reference depth image output by the first camera, determine the three-dimensional spatial point of the pixel based on the pixel, the depth value of the pixel and the camera intrinsic parameters of the first camera; and construct the reference point cloud corresponding to the first camera by combining the three-dimensional spatial points of each pixel in the reference depth image output by the first camera.

[0192] According to an exemplary embodiment of this disclosure, the process of the transformation matrix determination module 141 determining the reference point cloud can also be configured to perform: acquiring a reference depth image output by a second camera; determining the three-dimensional spatial point of each pixel in the reference depth image output by the second camera; transforming the three-dimensional spatial point of each pixel in the reference depth image output by the second camera according to the transformation matrix between the first camera coordinate system and the second camera coordinate system to obtain the transformed three-dimensional spatial point; merging the three-dimensional spatial point of each pixel in the reference depth image output by the first camera with the transformed three-dimensional spatial point to construct the reference point cloud corresponding to the first camera.

[0193] According to an exemplary embodiment of the present disclosure, the process of the transformation matrix determination module 141 filtering a specified plane can be configured to perform: filtering the specified plane based on the distance information of the plane from the first camera contained in the plane information of the reference point cloud.

[0194] According to an exemplary embodiment of the present disclosure, the process of the transformation matrix determination module 141 filtering a specified plane can be configured to perform: if the distance information contains a distance within a predetermined distance range, determine a candidate plane corresponding to a distance within the predetermined distance range in the distance information; if the number of candidate planes is one, determine the candidate plane as the specified plane; if the number of candidate planes is multiple, determine the candidate plane whose distance from the first camera is closest to a distance threshold as the specified plane; wherein the distance threshold is within the predetermined distance range.

[0195] According to an exemplary embodiment of this disclosure, the designated plane is the ground plane.

[0196] Since the various functional modules of the pose determination device in this embodiment are the same as those in the above-described method embodiment, they will not be described again here.

[0197] Figure 15 A schematic diagram is shown that is suitable for implementing exemplary embodiments of the present disclosure. The terminal device of the exemplary embodiments of the present disclosure can be configured as follows: Figure 15 In the form of. It should be noted that, Figure 15 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0198] The electronic device disclosed herein includes at least a processor and a memory, the memory being used to store one or more programs, which, when executed by the processor, enable the processor to implement the pose determination method of the exemplary embodiments of this disclosure.

[0199] Specifically, such as Figure 15As shown, the electronic device 150 includes at least: a processor 1510, internal memory 1521, external memory interface 1522, Universal Serial Bus (USB) interface 1530, charging management module 1540, power management module 1541, battery 1542, antenna, wireless communication module 1550, audio module 1560, display screen 1570, sensor module 1580, and camera module 1590. The sensor module 1580 may include depth sensors, pressure sensors, gyroscope sensors, barometric pressure sensors, magnetic sensors, accelerometers, distance sensors, proximity sensors, fingerprint sensors, temperature sensors, touch sensors, ambient light sensors, and bone conduction sensors.

[0200] It is understood that the structures illustrated in the embodiments of this disclosure do not constitute a specific limitation on the electronic device 150. In other embodiments of this disclosure, the electronic device 150 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0201] Processor 1510 may include one or more processing units, such as an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural network processing unit (NPU). Different processing units may be independent devices or integrated into one or more processors. Additionally, processor 1510 may include memory for storing instructions and data.

[0202] The electronic device 150 can implement shooting functions through an ISP, camera module 1590, video codec, GPU, display screen 1570, and application processor. In some embodiments, the electronic device 150 may include at least two camera modules 1590. In implementing the present disclosure, one camera module is designated as a reference camera, and the feature data acquired by the other camera modules is transferred to the coordinate system of the reference camera for processing. For example, the electronic device 150 is equipped with two RealSense D455 cameras.

[0203] Internal memory 1521 can be used to store computer executable program code, which includes instructions. Internal memory 1521 may include a program storage area and a data storage area. External memory interface 1522 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of electronic device 150.

[0204] This disclosure also provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiments; or it may exist independently and not assembled into the electronic device.

[0205] Computer-readable storage media can be, for example—but not limited to—electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0206] A computer-readable storage medium can be sent, propagated, or transmitted for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable storage medium can be transmitted using any suitable medium, including but not limited to: wireless, wireline, optical fiber, RF, etc., or any suitable combination thereof.

[0207] A computer-readable storage medium carries one or more programs that, when executed by an electronic device, cause the electronic device to perform the methods described in the embodiments of this disclosure.

[0208] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0209] The units described in the embodiments of this disclosure can be implemented in software or hardware, and the described units can also be located in a processor. The names of these units do not necessarily limit the unit itself.

[0210] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, terminal device, or network device, etc.) to execute the methods according to the embodiments of this disclosure.

[0211] Furthermore, the above figures are merely illustrative of the processes included in the method according to exemplary embodiments of this disclosure and are not intended to be limiting. It is readily understood that the processes shown in the above figures do not indicate or limit the temporal order of these processes. Additionally, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.

[0212] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to embodiments of this disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0213] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the disclosure herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the claims.

[0214] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. A pose determination method, characterized in that, Applied to a terminal device, the terminal device being configured with a first camera and at least one second camera, the pose determination method includes: Acquire the current frame color image captured by the first camera, and determine the first two-dimensional feature point on the current frame color image captured by the first camera that matches the previous frame color image captured by the first camera; Acquire the current frame color image captured by the second camera, and determine the second two-dimensional feature points on the current frame color image captured by the second camera that match the previous frame color image captured by the second camera; The transformation matrix between the first camera coordinate system of the first camera and the second camera coordinate system of the second camera is used to convert the second two-dimensional feature point into a third two-dimensional feature point in the first camera coordinate system; wherein, the first two-dimensional feature point and the third two-dimensional feature point constitute two-dimensional coordinate information, and the three-dimensional feature points of the previous frame color image captured by the first camera in the world coordinate system and the three-dimensional feature points of the previous frame color image captured by the second camera in the world coordinate system constitute three-dimensional coordinate information. The two-dimensional coordinate information is associated with the three-dimensional coordinate information to obtain point-to-point information; The point-pair information is used to solve the perspective n-point problem, and the pose of the first camera when acquiring the current frame color image is determined by combining the solution results.

2. The pose determination method according to claim 1, characterized in that, The first two-dimensional feature point that matches the previous frame color image captured by the first camera in the current frame color image includes: Extract feature points from the current frame color image captured by the first camera; Optical flow tracking is performed using feature points from the current frame color image captured by the first camera and feature points from the previous frame color image captured by the first camera to determine the first two-dimensional feature points.

3. The pose determination method according to claim 1, characterized in that, Converting the second two-dimensional feature point into a third two-dimensional feature point in the first camera coordinate system using the transformation matrix between the first camera's first camera coordinate system and the second camera's second camera coordinate system includes: Obtain the transformation matrix between the first camera coordinate system of the first camera and the second camera coordinate system of the second camera, as well as the depth information of the second two-dimensional feature point; The third two-dimensional feature point is determined based on the transformation matrix, the depth information of the second two-dimensional feature point, and the second two-dimensional feature point.

4. The pose determination method according to claim 3, characterized in that, Determining the third two-dimensional feature point based on the transformation matrix, the depth information of the second two-dimensional feature point, and the second two-dimensional feature point includes: The transformation matrix, the depth information of the second two-dimensional feature point, and the second two-dimensional feature point are multiplied together, and the result of the multiplication is normalized to determine the third two-dimensional feature point.

5. The pose determination method according to claim 1, characterized in that, The pose determination method further includes: Acquire the previous frame of color image captured by the first camera, and extract the feature points of the previous frame of color image captured by the first camera. Using the previous frame depth image aligned with the previous frame color image captured by the first camera, the feature points of the previous frame color image captured by the first camera are spatially projected to obtain the three-dimensional feature points of the previous frame color image captured by the first camera in the first camera coordinate system. Based on the pose of the first camera when it captured the previous color image, the three-dimensional feature points in the first camera coordinate system are transformed to obtain the three-dimensional feature points of the previous color image captured by the first camera in the world coordinate system.

6. The pose determination method according to claim 5, characterized in that, Using a previous frame depth image aligned with the previous frame color image captured by the first camera, spatial projection is performed on the feature points of the previous frame color image captured by the first camera to obtain the three-dimensional feature points of the previous frame color image captured by the first camera in the first camera coordinate system, including: Using the previous frame depth image aligned with the previous frame color image captured by the first camera, the feature points in the previous frame color image captured by the first camera that are within a predetermined depth range are spatially projected to obtain the three-dimensional feature points of the previous frame color image captured by the first camera in the first camera coordinate system. The predetermined depth range is determined based on the depth measurement range.

7. The pose determination method according to claim 1, characterized in that, The pose determination method further includes: Acquire the previous frame of color image captured by the second camera, and extract the feature points of the previous frame of color image captured by the second camera; Using the previous frame depth image aligned with the previous frame color image captured by the second camera, the feature points of the previous frame color image captured by the second camera are spatially projected to obtain the three-dimensional feature points of the previous frame color image captured by the second camera in the second camera coordinate system. The transformation matrix between the first camera coordinate system and the second camera coordinate system is used to convert the three-dimensional feature points of the previous frame color image captured by the second camera in the second camera coordinate system into three-dimensional feature points in the first camera coordinate system. Based on the pose of the first camera when it captured the previous color image, the three-dimensional feature points in the first camera coordinate system are transformed to obtain the three-dimensional feature points of the previous color image captured by the second camera in the world coordinate system.

8. The pose determination method according to claim 7, characterized in that, Using a previous frame depth image aligned with the previous frame color image captured by the second camera, spatial projection is performed on the feature points of the previous frame color image captured by the second camera to obtain the three-dimensional feature points of the previous frame color image captured by the second camera in the second camera coordinate system, including: Using the previous frame depth image aligned with the previous frame color image captured by the second camera, spatial projection is performed on the feature points of the previous frame color image captured by the second camera that are within a predetermined depth range, so as to obtain the three-dimensional feature points of the previous frame color image captured by the second camera in the second camera coordinate system. The predetermined depth range is determined based on the depth measurement range.

9. The pose determination method according to any one of claims 1 to 8, characterized in that, The pose determination method further includes: Acquire the initial frame color image captured by the first camera, and extract the feature points of the initial frame color image captured by the first camera; Using an initial frame depth image aligned with the initial frame color image captured by the first camera, the feature points of the initial frame color image captured by the first camera are spatially projected to obtain the three-dimensional feature points of the initial frame color image captured by the first camera in the first camera coordinate system. Based on the three-dimensional feature points, initial rotation matrix, and initial translation vector of the initial frame color image acquired by the first camera in the first camera coordinate system, the initial positioning result of the first camera in the first camera coordinate system is determined. The initial positioning result of the first camera in the first camera coordinate system is transformed using the transformation matrix between the first camera coordinate system and the world coordinate system to determine the pose of the first camera when acquiring the initial frame color image.

10. The pose determination method according to claim 9, characterized in that, The pose determination method further includes: Obtain the reference depth image output by the first camera; When a specified plane is determined by combining the reference depth image output by the first camera, the transformation matrix between the first camera coordinate system and the world coordinate system is determined based on the normal vector and gravity vector of the specified plane.

11. The pose determination method according to claim 10, characterized in that, The pose determination method further includes: By combining the reference depth image output by the first camera, the reference point cloud corresponding to the first camera is determined. Extract the planar information of the reference point cloud; The specified plane is selected based on the planar information of the reference point cloud.

12. The pose determination method according to claim 11, characterized in that, Based on the reference depth image output by the first camera, a reference point cloud corresponding to the first camera is determined, including: For each pixel in the reference depth image output by the first camera, the three-dimensional spatial point of the pixel is determined based on the pixel, the depth value of the pixel, and the camera intrinsic parameters of the first camera. By combining the three-dimensional spatial points of each pixel on the reference depth image output by the first camera, a reference point cloud corresponding to the first camera is constructed.

13. The pose determination method according to claim 12, characterized in that, By combining the three-dimensional spatial points of each pixel in the reference depth image output by the first camera, a reference point cloud corresponding to the first camera is constructed, including: Obtain the reference depth image output by the second camera; Determine the three-dimensional spatial point of each pixel in the reference depth image output by the second camera; The three-dimensional spatial points of each pixel in the reference depth image output by the second camera are transformed according to the transformation matrix between the first camera coordinate system and the second camera coordinate system to obtain the transformed three-dimensional spatial points. The three-dimensional spatial points of each pixel on the reference depth image output by the first camera are merged with the transformed three-dimensional spatial points to construct the reference point cloud corresponding to the first camera.

14. The pose determination method according to claim 11, characterized in that, Filtering the specified plane based on the planar information of the reference point cloud includes: The specified plane is selected based on the distance information between the plane and the first camera contained in the planar information of the reference point cloud.

15. The pose determination method according to claim 14, characterized in that, Filtering the specified plane based on the distance information of the plane from the first camera contained in the planar information of the reference point cloud includes: If the distance information includes distances within a predetermined distance range, a candidate plane corresponding to a distance within the predetermined distance range in the distance information is determined; If there is only one candidate plane, then the candidate plane is determined as the designated plane. When there are multiple candidate planes, the candidate plane whose distance from the first camera is closest to the distance threshold is determined as the designated plane; The distance threshold is within the predetermined distance range.

16. The pose determination method according to claim 10, characterized in that, The designated plane is the ground plane.

17. A pose determination device, characterized in that, Configured in a terminal device, the terminal device further comprising a first camera and at least one second camera, the pose determination device comprising: The first feature point determination module is used to acquire the current frame color image captured by the first camera and determine the first two-dimensional feature point on the current frame color image captured by the first camera that matches the previous frame color image captured by the first camera. The second feature point determination module is used to acquire the current frame color image captured by the second camera and determine the second two-dimensional feature points on the current frame color image captured by the second camera that match the previous frame color image captured by the second camera. The feature point conversion module is used to convert the second two-dimensional feature point into a third two-dimensional feature point in the first camera coordinate system using the transformation matrix between the first camera coordinate system of the first camera and the second camera coordinate system of the second camera; wherein, the first two-dimensional feature point and the third two-dimensional feature point constitute two-dimensional coordinate information, and the three-dimensional feature points of the previous frame color image captured by the first camera in the world coordinate system and the three-dimensional feature points of the previous frame color image captured by the second camera in the world coordinate system constitute three-dimensional coordinate information. The pose determination module is used to associate the two-dimensional coordinate information with the three-dimensional coordinate information to obtain point pair information; use the point pair information to solve the perspective n-point problem, and combine the solution results to determine the pose of the first camera when acquiring the current frame color image.

18. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the pose determination method as described in any one of claims 1 to 16.

19. An electronic device, characterized in that, include: processor; A memory for storing one or more programs, which, when executed by the processor, cause the processor to implement the pose determination method as described in any one of claims 1 to 16.

Citation Information

Patent Citations

  • Camera pose determination method and device, electronic equipment and storage medium

    CN111415387A

  • Visual positioning method and terminal

    CN111415388A