Image-based pose estimation method and device, electronic equipment and storage medium
By using key point pairs and matching point pairs in scenarios such as head-mounted display devices, combining deep learning and non-deep learning, the effectiveness of pose estimation is solved, efficient and reliable pose estimation is achieved, and the display effect of display devices is improved.
Patent Information
- Application Number
- CN202410031108.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-08
- Publication Date
- 2025-07-08
AI Technical Summary
In the prior art, there are still challenges in how to effectively implement pose estimation, especially in scenarios such as head-mounted display devices, robot grasping and robot navigation.
By acquiring the original image and pending image of the tracking object, key point pairs and matching point pairs are determined, combined with deep learning and non-deep learning methods, the target position of the tracking object is estimated using the conversion relationship between the reference coordinate system and the camera coordinate system.
实现了在不同场景下的高效、可靠的位姿估计,提升了头戴显示设备的显示效果和用户体验,并适用于单目和双目图像传感器。
Smart Images

Figure CN120279088A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of machine vision technology, and in particular, to an image-based pose estimation method, apparatus, electronic device, and storage medium. Background Art
[0002] In some scenarios, there is a need for pose estimation. How to effectively implement pose estimation is a problem worthy of attention for those skilled in the art. Summary of the Invention
[0003] According to one aspect of the embodiments of the present disclosure, an image-based pose estimation method is provided, including: obtaining a to-be-processed image including a tracking object; determining key point pairs of the tracking object based on a pre-obtained original image of the tracking object and the to-be-processed image, where the key point pairs include corner points of the tracking object on the original image and two-dimensional key points of the tracking object in the to-be-processed image corresponding to the corner points, and the corner points have three-dimensional coordinates in a reference coordinate system where the original image is located; determining matching point pairs of the tracking object based on the original image and the to-be-processed image, where the matching point pairs include feature points of the tracking object on the original image and two-dimensional feature points of the tracking object in the to-be-processed image matching the feature points on the original image, and the feature points on the original image have three-dimensional coordinates in the reference coordinate system; obtaining the target pose of the tracking object based on the key point pairs and the matching point pairs.
[0004] According to another aspect of the embodiments of the present disclosure, an image-based pose estimation apparatus is provided, including: a first obtaining module, configured to obtain a to-be-processed image including a tracking object; a first determining module, configured to determine key point pairs of the tracking object based on a pre-obtained original image of the tracking object and the to-be-processed image, where the key point pairs include corner points of the tracking object on the original image and two-dimensional key points of the tracking object in the to-be-processed image corresponding to the corner points, and the corner points have three-dimensional coordinates in a reference coordinate system where the original image is located; a second determining module, configured to determine matching point pairs of the tracking object based on the original image and the to-be-processed image, where the matching point pairs include feature points of the tracking object on the original image and two-dimensional feature points of the tracking object in the to-be-processed image matching the feature points on the original image, and the feature points on the original image have three-dimensional coordinates in the reference coordinate system; a second obtaining module, configured to obtain the target pose of the tracking object based on the key point pairs and the matching point pairs.
[0005] According to still another aspect of the present disclosure, a computer-readable storage medium is provided, and the storage medium stores a computer program for executing the above-mentioned image-based pose estimation method.
[0006] According to another aspect of the present disclosure, there is provided an electronic device, including: a processor; a memory for storing processor-executable instructions; and the processor for reading the executable instructions from the memory and executing the instructions to implement the above-mentioned image-based pose estimation method.
[0007] The technical solutions of the present disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Figure 1 is a flowchart of an image-based pose estimation method provided by some exemplary embodiments of the present disclosure.
[0009] Figure 2-1 is a schematic diagram of a left-eye image in some exemplary embodiments of the present disclosure.
[0010] Figure 2-2 is a schematic diagram of a right-eye image in some exemplary embodiments of the present disclosure.
[0011] Figure 2-3 is a schematic diagram of an original image in some exemplary embodiments of the present disclosure.
[0012] Figure 2-4 is a schematic diagram of a left-eye image in some other exemplary embodiments of the present disclosure.
[0013] Figure 2-5 is a schematic diagram of a left-eye image in some further exemplary embodiments of the present disclosure.
[0014] Figure 2-6 is a schematic diagram of a right-eye image in some other exemplary embodiments of the present disclosure.
[0015] Figure 2-7 is a schematic diagram of a right-eye image in some further exemplary embodiments of the present disclosure.
[0016] Figure 3 is a flowchart of an image-based pose estimation method provided by some other exemplary embodiments of the present disclosure.
[0017] Figure 4 is a flowchart of a method for determining matching point pairs of a tracking object provided by some exemplary embodiments of the present disclosure.
[0018] Figure 5 is a flowchart of a method for determining pose estimation values provided by some exemplary embodiments of the present disclosure.
[0019] Figure 6 is a flowchart of a method for determining key point pairs of a tracking object provided by some exemplary embodiments of the present disclosure.
[0020] Figure 7It is a schematic flowchart of a method for obtaining a target pose provided by some exemplary embodiments of the present disclosure.
[0021] Figure 8 It is a schematic flowchart of a method for rendering a virtual image provided by some exemplary embodiments of the present disclosure.
[0022] Figure 9 It is a schematic flowchart of a method for determining feature points on an original image provided by some exemplary embodiments of the present disclosure.
[0023] Figure 10 It is a schematic flowchart of a method for image-based pose estimation provided by some further exemplary embodiments of the present disclosure.
[0024] Figure 11 It is a schematic structural diagram of an image-based pose estimation device provided by some exemplary embodiments of the present disclosure.
[0025] Figure 12 It is a schematic diagram of a module involved in determining matching point pairs in some exemplary embodiments of the present disclosure.
[0026] Figure 13 It is a schematic structural diagram of a first determination module in some exemplary embodiments of the present disclosure.
[0027] Figure 14 It is a schematic structural diagram of an estimation module in some exemplary embodiments of the present disclosure.
[0028] Figure 15 It is a schematic structural diagram of a second acquisition module in some exemplary embodiments of the present disclosure.
[0029] Figure 16 It is a schematic structural diagram of an image-based pose estimation device provided by some other exemplary embodiments of the present disclosure.
[0030] Figure 17 It is a schematic diagram of a module involved in determining feature points on an original image in some exemplary embodiments of the present disclosure.
[0031] Figure 18 It is a schematic structural diagram of an electronic device provided by some exemplary embodiments of the present disclosure. Detailed Description of the Embodiments
[0032] To explain the present disclosure, exemplary embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. It should be understood that the present disclosure is not limited by the exemplary embodiments.
[0033] It should be noted that: Unless otherwise specifically stated, the relative arrangements, numerical expressions, and numerical values of the components and steps set forth in these embodiments do not limit the scope of the present disclosure.
[0034] Exemplary Overview
[0035] In some scenarios, there is a need for pose estimation.
[0036] For example, in the usage scenario of a head-mounted display device, it is necessary to estimate the pose of a predetermined object in the camera coordinate system corresponding to the image sensor provided in the head-mounted display device. The head-mounted display device can also be referred to as a Head-Mounted Display (HMD) or a head-mounted display. The head-mounted display device can be presented in the form of glasses, helmets, etc. The head-mounted display device can include, but is not limited to, Augmented Reality (AR) glasses, Virtual Reality (VR) glasses, etc. The specific object can be either a movable object in the real physical world or an immovable object in the real physical world. According to the pose of the predetermined object in the camera coordinate system corresponding to the image sensor, the virtual image displayed through the head-mounted display device can be rendered, adjusted, controlled, etc., to achieve a specific display effect.
[0037] The situation introduced in the above paragraph is that the image sensor is provided in the head-mounted display device, and it can be considered that the image sensor and the head-mounted display device are integrally provided. In specific implementation, the image sensor and the head-mounted display device can also be separately provided. For the sake of easy understanding, the following description will take the case where the image sensor and the head-mounted display device are integrally provided as an example.
[0038] It should be pointed out that the scenarios with the need for pose estimation are not limited to the usage scenario of the head-mounted display device. For example, they can also include scenarios such as robot grasping and robot navigation, which will not be listed one by one here.
[0039] Therefore, how to effectively implement pose estimation is a problem worthy of attention for those skilled in the art.
[0040] Exemplary Method
[0041] Figure 1 It is a schematic flowchart of an image-based pose estimation method provided by some exemplary embodiments of the present disclosure. Figure 1 The method shown can include step 110, step 120, and step 130.
[0042] Step 110, obtain a to-be-processed image including a tracking object.
[0043] In some alternative embodiments of the present disclosure, the object to be tracked may be an object for which pose estimation is required. Optionally, the predetermined object described above may be used as the object to be tracked. To facilitate the implementation of pose estimation for the object to be tracked, the object to be tracked may be in the shape of a thin sheet and may have a specific pattern. The specific pattern may include both patterns with regular shapes such as rectangles, rhombuses, and regular hexagons, and patterns with irregular shapes. The number of objects to be tracked may be one, or two or more. Since the processing method for each object to be tracked is similar, for ease of understanding, the following mainly describes the processing method for a single object to be tracked.
[0044] In some alternative embodiments of the present disclosure, image acquisition may be performed by an image sensor at regular or irregular intervals to obtain a to-be-processed image including the object to be tracked. The image sensor may be either a monocular image sensor or a binocular image sensor. Correspondingly, the to-be-processed image may be either a monocular image or a binocular image. A binocular image may include two images, namely a left image and a right image. In an alternative example, the left image may be as Figure 2-1 shown, and the right image may be as Figure 2-2 shown.
[0045] Step 120: Determine the target point pairs of the object to be tracked between the original image of the object to be tracked obtained in advance and the to-be-processed image.
[0046] In some alternative embodiments of the present disclosure, before performing the image-based pose estimation method provided by the embodiments of the present disclosure, the original image of the object to be tracked may be obtained. The original image of the object to be tracked may be understood as a close-up image of the object to be tracked, which is used to clearly and completely reflect various characteristics of the object to be tracked, such as its shape and texture. The background part in the original image may be a white background, and the foreground part in the original image may only include the object to be tracked. In an alternative example, the original image may be as Figure 2-3 shown.
[0047] In some alternative embodiments of the present disclosure, pixel points that are relatively special in terms of color, texture, etc. on the original image can be extracted to obtain multiple pixel points of the tracking object on the original image. Similarly, pixel points that are relatively special in terms of color, texture, etc. in the image to be processed can be extracted to obtain multiple pixel points of the tracking object in the image to be processed. By combining the multiple pixel points extracted from the original image and the multiple pixel points extracted from the image to be processed, multiple target point pairs of the tracking object can be obtained. Each target point pair can include a pixel point extracted from the original image and a pixel point corresponding to this pixel point extracted from the image to be processed. Optionally, the pixel points of the tracking object extracted from the original image include but are not limited to key points, feature points, and corner points. Optionally, the pixel points of the tracking object extracted from the image to be processed include but are not limited to key points, feature points, and corner points.
[0048] Step 130: Obtain the target pose of the tracking object based on the target point pairs.
[0049] In step 130, based on the multiple target point pairs obtained in step 120, the target pose of the tracking object can be determined.
[0050] In some alternative embodiments of the present disclosure, the image sensor can be a monocular image sensor. The monocular image sensor can only correspond to one camera coordinate system, and the target pose can be the pose of the tracking object in this camera coordinate system. Of course, the target pose can also be the pose of the tracking object in other coordinate systems other than this camera coordinate system, as long as the conversion relationship between the other coordinate systems and this camera coordinate system is known or can be calculated.
[0051] In some other alternative embodiments of the present disclosure, the image sensor can be a binocular image sensor. The binocular image sensor can correspond to two camera coordinate systems, and the target pose can be the pose of the tracking object in any one of the two camera coordinate systems. Of course, the target pose can also be the pose in other coordinate systems other than the two camera coordinate systems, as long as the conversion relationship between the other coordinate systems and any one of the two camera coordinate systems is known or can be calculated.
[0052] In the embodiments of the present disclosure, based on the original image of the tracking object and the image to be processed including the tracking object, the target point pairs of the tracking object can be determined. The target point pairs of the tracking object can be used to determine the target pose of the tracking object. In this way, based on the information carried by the image, the pose estimation of the tracking object can be effectively achieved.
[0053] Figure 3 It is a schematic flowchart of a method for pose estimation based on images provided by some exemplary embodiments of the present disclosure. Figure 3The method shown may include step 310, step 320, step 330, and step 340. It should be noted that Figure 3 the method shown and Figure 1 the method shown have some similarities. Regarding these similarities, no further introduction will be given here. For specific details, please refer to the relevant introduction in the above text.
[0054] Step 310: Obtain a to-be-processed image including a tracking object.
[0055] Step 320: Determine key point pairs of the tracking object based on the original image and the to-be-processed image of the tracking object obtained in advance. Among them, the key point pairs include corner points of the tracking object on the original image and two-dimensional key points of the tracking object in the to-be-processed image corresponding to the corner points. The corner points have three-dimensional coordinates in the reference coordinate system where the original image is located.
[0056] In some alternative embodiments of the present disclosure, the corner points of the tracking object on the original image can be manually marked to ensure the accuracy of the corner points of the tracking object. The number of corner points of the tracking object can be multiple, for example, it can be N. In an alternative example, the original image can be as Figure 2-3 shown. The number N of the corner points of the tracking object can be 6, and the 6 corner points can be respectively represented as R1, R2, R3, R4, R5, and R6.
[0057] It can be understood that the corner points of the tracking object on the original image belong to a kind of key points. In some alternative embodiments of the present disclosure, other key points on the tracking object in the original image can be extracted to form key point pairs with the two-dimensional key points of the tracking object in the to-be-processed image.
[0058] In some alternative embodiments of the present disclosure, the two-dimensional key points of the corner points belonging to the tracking object in the to-be-processed image can be determined by a deep learning method. For example, before executing the image-based pose estimation method provided by the embodiments of the present disclosure, a target detection network and a key point detection network can be trained. Through the target detection network, target detection can be performed on the to-be-processed image to identify the two-dimensional bounding boxes and classes of the tracking object from the to-be-processed image. Through the key point detection network, key point detection can be performed on the two-dimensional bounding boxes to determine the two-dimensional key points of the corner points belonging to the tracking object in the to-be-processed image. By combining the corner points of the tracking object on the original image and the two-dimensional key points of the corner points belonging to the tracking object in the to-be-processed image, key point pairs can be obtained. The number of key point pairs can be multiple, for example, it can be N. Each key point pair can be used as a target point pair in the above text.
[0059] In an alternative example, the to-be-processed image can be a binocular image, including Figure 2-1The left-eye image shown and Figure 2-2 the right-eye image shown. From Figure 2-1 and Figure 2-2 it can be seen that the number of tracking objects is three. For the sake of convenience in explanation, here only the tracking object located between the user's two hands can be considered.
[0060] Through the object detection network, performing object detection on the left-eye image, the two-dimensional bounding box 210 shown in Figure 2-4 can be obtained. Through the key-point detection network, performing key-point detection on the two-dimensional bounding box 210, six two-dimensional key points of the corner points belonging to the tracking object in the left-eye image can be determined. Specifically, reference can be made to Figure 2-5 . Assuming that starting from the two-dimensional key point located at the upper left position among these six two-dimensional key points, in clockwise order, these six two-dimensional key points are respectively denoted as C1, C2, C3, C4, C5, C6, then six key-point pairs can be formed between the original image and the left-eye image, which are respectively the key-point pair including R1 and C1, the key-point pair including R2 and C2, the key-point pair including R3 and C3, the key-point pair including R4 and C4, the key-point pair including R5 and C5, and the key-point pair including R6 and C6.
[0061] Similarly, through the object detection network, performing object detection on the right-eye image, the two-dimensional bounding box 220 shown in Figure 2-6 can be obtained. Through the key-point detection network, performing key-point detection on the two-dimensional bounding box 220, six two-dimensional key points of the corner points belonging to the tracking object in the right-eye image can be determined. Specifically, reference can be made to Figure 2-7 . Assuming that these six two-dimensional key points are respectively denoted as D1, D2, D3, D4, D5, D6, then six key-point pairs can be formed between the original image and the right-eye image, which are respectively the key-point pair including R1 and D1, the key-point pair including R2 and D2, the key-point pair including R3 and D3, the key-point pair including R4 and D4, the key-point pair including R5 and D5, and the key-point pair including R6 and D6.
[0062] From the above example, it can be seen that when the number N of the corner points of the tracking object is 6, if the image to be processed is a binocular image, a total of 12 key-point pairs can be obtained in step 320. It can be understood that if the image to be processed is a monocular image, a total of 6 key-point pairs can be obtained in step 320.
[0063] In some alternative embodiments of the present disclosure, the reference coordinate system where the original image is located may be a three-dimensional coordinate system determined based on the original image. In an alternative example, the upper left pixel point of the original image may be used as the origin of the reference coordinate system, the width direction of the original image may be used as the X-axis direction of the reference coordinate system, the height direction of the original image may be used as the Y-axis direction of the reference coordinate system, and the direction perpendicular to both the width direction and the height direction of the original image may be used as the Z-axis direction of the reference coordinate system.
[0064] Of course, the determination method of the reference coordinate system is not limited to this. For example, the upper right pixel point, the lower left pixel point, or the lower right pixel point of the original image may be used as the origin of the reference coordinate system. For another example, the X-axis direction and the Y-axis direction of the reference coordinate system may be interchanged.
[0065] Since the reference coordinate system is determined based on the original image, for any pixel point in the original image, a corresponding three-dimensional point can be found in the reference coordinate system. Then, for any corner point in the original image, a corresponding three-dimensional point can be found in the reference coordinate system, and the three-dimensional coordinates of this three-dimensional point can be used as the three-dimensional coordinates of this corner point in the reference coordinate system.
[0066] Step 330, determining a matching point pair of the tracking object based on the original image and the image to be processed, where the matching point pair includes the feature point of the tracking object on the original image and the two-dimensional feature point of the tracking object in the image to be processed that matches the feature point on the original image, and the feature point on the original image has three-dimensional coordinates in the reference coordinate system.
[0067] In some alternative embodiments of the present disclosure, the feature points of the tracking object on the original image may be determined by a non-deep learning method. For example, traditional methods such as Scale-invariant Feature Transform (SIFT) may be used to determine the feature points of the tracking object on the original image. The number of feature points of the tracking object on the original image may be multiple, for example, it may be M. M may be greater than N. For example, M may be 100 and N may be 6. In an alternative example, Figure 2-3 each corner point of each rectangle located inside the hexagon may be used as a feature point.
[0068] In some alternative embodiments of the present disclosure, after determining the feature points of the tracking object on the original image, an image feature matching algorithm may be used to determine the two-dimensional feature points of the tracking object that match the feature points from the image to be processed. By combining the feature points of the tracking object on the original image and the two-dimensional feature points of the tracking object that match the feature points, a matching point pair can be obtained. The number of matching point pairs may be multiple. Each matching point pair may be used as one of the target point pairs described above.
[0069] In an alternative example, the image to be processed may be a binocular image. The number M of feature points of the tracking object on the original image may be 100, for example, and are respectively represented as P1, P2, P3, ……, P100. The two-dimensional feature points respectively matching P1, P2, P3, ……, P100 on the left-eye image may be represented as G1, G2, G3, ……, G100. The two-dimensional feature points respectively matching P1, P2, P3, ……, P100 on the right-eye image may be represented as S1, S2, S3, ……, S100. In this way, 100 matching point pairs can be formed between the original image and the left-eye image, which are respectively the matching point pair including P1 and G1, the matching point pair including P2 and G2, ……, the matching point pair including P100 and G100. Similarly, 100 matching point pairs can be formed between the original image and the right-eye image, which are respectively the matching point pair including P1 and S1, the matching point pair including P2 and S2, ……, the matching point pair including P100 and S100.
[0070] As can be seen from the above example, when the number M of feature points of the tracking object on the original image is 100, if the image to be processed is a binocular image, 200 matching point pairs can be obtained in step 330. It can be understood that if the image to be processed is a monocular image, 100 key point pairs can be obtained in step 320.
[0071] In some embodiments, even if M is 100 and the image to be processed is a binocular image, the number of matching point pairs obtained in step 330 may not be 200 in total, but may be less than 200. In addition, the number of matching point pairs formed between the original image and the left-eye image and the number of matching point pairs formed between the original image and the right-eye image may be different.
[0072] As introduced above, for any pixel point in the original image, a corresponding three-dimensional point can be found in the reference coordinate system. Then, for any feature point in the original image, a corresponding three-dimensional point can be found in the reference coordinate system, and the three-dimensional coordinates of this three-dimensional point can be used as the three-dimensional coordinates of this feature point in the reference coordinate system.
[0073] Step 340, obtaining the target pose of the tracking object based on the key point pairs and the matching point pairs.
[0074] In some alternative embodiments of the present disclosure, all the key point pairs determined in step 320 and all the matching point pairs determined in step 330 can be used to determine the target pose of the tracking object.
[0075] In some other alternative embodiments of the present disclosure, according to certain rules, some key point pairs can be selected from all the key point pairs determined in step 320, and some matching point pairs can be selected from all the matching point pairs determined in step 330, and only the selected partial key point pairs and the selected partial matching point pairs are used to determine the target pose of the tracking object.
[0076] In the embodiments of the present disclosure, by combining the original image of the tracking object and the image to be processed including the tracking object, the key point pairs of the tracking object and the matching point pairs of the tracking object can be determined. Since the corner points in the key point pairs have three-dimensional coordinates in the reference coordinate system where the original image is located, it can be considered that the corner points in the key point pairs are associated with the reference coordinate system. In addition, the two-dimensional key points in the key point pairs are associated with the camera coordinate system. Then, the key point pairs can associate the camera coordinate system with the reference coordinate system. Since the feature points in the matching point pairs have three-dimensional coordinates in the reference coordinate system, it can be considered that the feature points in the matching point pairs are associated with the reference coordinate system. In addition, the two-dimensional feature points in the matching point pairs are associated with the camera coordinate system. Then, the matching point pairs can associate the camera coordinate system with the reference coordinate system. In this way, based on the key point pairs and the matching point pairs, the conversion relationship between the reference coordinate system and the camera coordinate system can be solved to obtain the pose of the reference coordinate system relative to the camera coordinate system, and this pose can be used as the target pose, or, based on this pose, the target pose can be obtained through coordinate transformation. Therefore, by using the embodiments of the present disclosure, by utilizing the information carried by the image and introducing the reference coordinate system, the pose estimation of the tracking object can be effectively realized by combining deep learning methods and non-deep learning methods.
[0077] In some alternative embodiments of the present disclosure, the three-dimensional coordinates of the corner points on the original image in the reference coordinate system can be determined based on the scale relationship, and the scale relationship can be the relationship between the size of the tracking object in the original image and the true physical size of the tracking object. The scale relationship can be determined in the following manner:
[0078] When the image to be processed is a monocular image, the scale relationship can be a preset value;
[0079] When the image to be processed is a binocular image, the three-dimensional key points of the tracking object can be determined based on the binocular image, and the three-dimensional key points are aligned with the corner points on the original image to obtain the scale relationship.
[0080] It should be noted that the scale relationship can be the ratio relationship between the size of the tracking object in the original image and the true physical size of the tracking object. For example, if the size of the tracking object along the width direction in the original image is a1 and the true physical size of the tracking object along the width direction is a2, then a1 / a2 can be used as the scale relationship, or a2 / a1 can be used as the scale relationship.
[0081] In some alternative embodiments of the present disclosure, the image to be processed may be a monocular image, and the scale relationship may be a preset value. That is to say, the scale relationship may be known. Optionally, the preset value may be 1. That is to say, the size of the tracking object in the original image is the same as the true physical size of the tracking object. In this case, the tracking object in the original image can be regarded as the tracking object in the real physical world. On this basis, the three-dimensional coordinates of any pixel point on the original image in the reference coordinate system can be determined. Of course, the preset value may not be 1, but a value between 0 and 1, or a value greater than 1. In this case, the tracking object in the original image can be regarded as the result of scaling the tracking object in the real physical world by a certain proportion. On this basis, the three-dimensional coordinates of any pixel point on the original image in the reference coordinate system can also be determined.
[0082] In an alternative example, the upper-left pixel point of the original image is used as the origin of the reference coordinate system, the width direction of the original image is used as the X-axis direction of the reference coordinate system, and the height direction of the original image is used as the Y-axis direction of the reference coordinate system. The pixel coordinates of any two-dimensional point in the original image are (x, y). If the preset value is 1, the three-dimensional coordinates of this two-dimensional point in the reference coordinate system can be (x, y, 0). If the preset value is k (k is between 0 and 1), and a2 / a1 is used as the scale relationship, the three-dimensional coordinates of this two-dimensional point in the reference coordinate system can be (k*x, k*y, 0).
[0083] In some other alternative embodiments of the present disclosure, the image to be processed may be a binocular image, and the scale relationship may be unknown. The scale relationship can be obtained by calculation. Assume that by performing step 320, the two-dimensional key points of the corners belonging to the tracking object in the left-eye image are determined, and the two-dimensional key points of the corners belonging to the tracking object in the right-eye image are determined. Based on the two-dimensional key points in the left-eye image and the two-dimensional key points in the right-eye image, through triangulation, the three-dimensional key points of the tracking object in the camera coordinate system corresponding to the left-eye image can be determined. It should be noted that if the scale relationship is used, the three-dimensional key points of the tracking object in the reference coordinate system can be determined from the corners in the original image. On this basis, the three-dimensional key points of the tracking object in the reference coordinate system and the three-dimensional key points of the tracking object in the camera coordinate system corresponding to the left-eye image can be point-cloud aligned using the point-cloud alignment algorithm. In this way, based on the three-dimensional key points of the tracking object in the camera coordinate system corresponding to the left-eye image and the six corners in the original image, the scale relationship can be inversely deduced by introducing the point-cloud alignment algorithm.
[0084] In an alternative example, the point-cloud alignment algorithm may include the Umeyama algorithm. The objective function of the Umeyama algorithm can be seen in the following formula:
[0085]
[0086] Among them, n can be 6, q i can represent the i-th three-dimensional key point of the tracking object in the camera coordinate system corresponding to the left-eye image, c can represent the scale relationship, R can represent the relative rotation between the reference coordinate system and the camera coordinate system, p i can represent the three-dimensional coordinates in the reference coordinate system directly obtained from the corner points of the tracking object in the original image without introducing the scale relationship, and t can represent the relative translation between the reference coordinate system and the camera coordinate system. By minimizing the objective function, c as the scale relationship can be obtained.
[0087] The above describes the situation of aligning the three-dimensional key points of the tracking object in the camera coordinate system corresponding to the left-eye image with the corner points on the original image to obtain the scale relationship. Based on this, the three-dimensional coordinates of each corner point on the original image in the reference coordinate system can be obtained. Using a similar method, the three-dimensional key points of the tracking object in the camera coordinate system corresponding to the right-eye image can be aligned with the corner points on the original image to obtain the scale relationship, and based on this, the three-dimensional coordinates of each corner point on the original image in the reference coordinate system can also be obtained.
[0088] In the embodiments of the present disclosure, when the image to be processed is a monocular image, based on the known scale relationship, the three-dimensional coordinates of the corner points on the original image in the reference coordinate system can be efficiently and reliably determined, so as to provide an effective reference for the pose estimation of the tracking object. When the image to be processed is a binocular image, the scale relationship may be unknown. Through triangulation processing, the scale relationship can be calculated. Based on the calculated scale relationship, the three-dimensional coordinates of the corner points on the original image in the reference coordinate system can be efficiently and reliably determined, so as to provide an effective reference for the pose estimation of the tracking object. Therefore, in the embodiments of the present disclosure, the target pose of the tracking object can be estimated from a monocular image, and the target pose of the tracking object can also be estimated from a binocular image, and the applicable range is very wide.
[0089] See Figure 4 , which is a schematic flowchart of a method for determining matching point pairs of a tracking object provided by some exemplary embodiments of the present disclosure. Figure 4 The method shown may include step 410 and step 420. Optionally, step 410 may be executed before step 330. Step 420 may be an alternative implementation of step 330 of the present disclosure.
[0090] Step 410, estimate the pose estimation value of the tracking object according to the key point pairs.
[0091] It should be noted that the pose estimation value can be similar to the meaning of the target pose in the above text. The main difference is that the pose estimation value can be determined only based on key point pairs, with low accuracy, while the target pose can be determined based on both key point pairs and matching point pairs, with high accuracy.
[0092] In some alternative embodiments of the present disclosure, as Figure 5 shown, step 410 may include step 4101, step 4103, and step 4105.
[0093] Step 4101: Based on the two-dimensional key points of the tracked object in the binocular image, determine the three-dimensional key points of the two-dimensional key points in the camera coordinate system.
[0094] As introduced above, based on the two-dimensional key points in the left-eye image and the two-dimensional key points in the right-eye image, the three-dimensional key points of the tracked object in the camera coordinate system corresponding to the left-eye image and the three-dimensional key points of the tracked object in the camera coordinate system corresponding to the right-eye image can be determined through triangulation.
[0095] Step 4103: Align the three-dimensional key points in the camera coordinate system with the corner points on the original image to obtain the conversion relationship between the camera coordinate system and the reference coordinate system.
[0096] As introduced above, based on the three-dimensional key points of the tracked object in the camera coordinate system corresponding to the left-eye image and the corner points on the original image, the relative rotation R and relative translation t between the reference coordinate system and the camera coordinate system can be obtained through the application of the Umeyama algorithm. R and t can form the conversion relationship between the camera coordinate system and the reference coordinate system. Similarly, based on the three-dimensional key points of the tracked object in the camera coordinate system corresponding to the right-eye image and the corner points on the original image, the conversion relationship between the camera coordinate system and the reference coordinate system can also be obtained through the application of the Umeyama algorithm.
[0097] Step 4105: Estimate the pose estimation value of the tracked object according to the conversion relationship.
[0098] In some alternative embodiments of the present disclosure, two transformation relationships can be obtained by performing step 4103, namely, the transformation relationship determined based on the left-eye image and the transformation relationship determined based on the right-eye image. Optionally, both transformation relationships can be in the form of poses, which are used to represent the pose of the reference coordinate system relative to the corresponding camera coordinate system. Then, in step 4105, the transformation relationship determined based on the left-eye image can be determined as the pose estimation value of the tracking object, and the transformation relationship determined based on the right-eye image can be determined as another pose estimation value of the tracking object. Alternatively, the transformation relationship determined based on the left-eye image can be optimized by the least squares method, and the optimization result can be used as the pose estimation value of the tracking object. Similarly, the transformation relationship determined based on the right-eye image can be optimized by the least squares method, and the optimization result can be used as the pose estimation value of the tracking object.
[0099] Figure 5 In the illustrated embodiment, for the case where the image to be processed is a binocular image, by combining triangulation processing and point cloud alignment processing, the pose estimation value of the tracking object can be determined efficiently and reliably, so as to provide an effective reference for determining the matching point pairs.
[0100] For the case where the image to be processed is a monocular image, the three-dimensional coordinates of the corner points on the original image in the reference coordinate system can be directly determined based on the known scale relationship. The three-dimensional key points with the three-dimensional coordinates and the two-dimensional key points in the key point pair can form 3D-2D point pairs. On the premise of the known 3D-2D point pairs, the transformation relationship between the camera coordinate system and the reference coordinate system can be obtained through the Perspective-n-Point (PnP) algorithm. According to the obtained transformation relationship, the pose estimation value of the tracking object can be determined.
[0101] Step 420: Based on the pose estimation value and the scale relationship, project the feature points on the original image onto the image to be processed to obtain the projected points in the image to be processed, optimize the positions of the projected points, obtain the two-dimensional feature points of the tracking object in the image to be processed that match the feature points on the original image, and use the feature points on the original image and the two-dimensional feature points as the matching point pairs of the tracking object.
[0102] In some alternative embodiments of the present disclosure, the image to be processed may be a monocular image. According to known scale relationships, the three-dimensional coordinates of feature points on the original image in the reference coordinate system can be determined. Using the pose estimation value, the three-dimensional feature points with the three-dimensional coordinates in the reference coordinate system can be projected onto the image to be processed to obtain the projection points in the image to be processed. Next, a template matching algorithm can be used to optimize the positions of the projection points, and the pixel points at the optimized positions can be used as the two-dimensional feature points of the tracking object in the image to be processed that match the feature points on the original image. Optionally, the template matching algorithm may be a Normalized Cross Correlation (NCC) template matching algorithm. In this way, for the feature points on the original image, a small area can be intercepted centered on the feature point, and based on the pixel values of this small area, a position more accurate than the position of the projection point can be searched near the projection point in the image to be processed, thereby realizing the optimization of the position of the projection point. It can be understood that the NCC template matching algorithm is a template matching algorithm based on pixel values. In specific implementation, other types of template matching algorithms may also be used, and the present disclosure does not limit this. By combining the feature points on the original image with the corresponding two-dimensional feature points, the matching point pairs can be obtained.
[0103] The method for determining the matching point pairs when the image to be processed is a monocular image is introduced in the above paragraph. When the image to be processed is a binocular image, the method for determining the matching point pairs is similar. The main difference is that matching point pairs need to be determined between the original image and the left-eye image, and matching point pairs also need to be determined between the original image and the right-eye image.
[0104] Figure 4 In the illustrated embodiment, the key point pairs can provide a very effective reference for the determination of the pose estimation value of the tracking object. The determined pose estimation value can be used together with the scale relationship for the projection of the feature points on the original image. Then, through the optimization of the positions of the projection points, the two-dimensional feature points of the tracking object in the image to be processed that match the feature points on the original image can be found more accurately. Thus, the matching point pairs can be reliably determined for the pose estimation of the tracking object.
[0105] In some embodiments, when the image to be processed is a binocular image, the scale relationship may also be known. In this case, point cloud alignment is not required, and the known scale relationship can be directly utilized.
[0106] See Figure 6 , which is a schematic flowchart of a method for determining the key point pairs of a tracking object provided by some exemplary embodiments of the present disclosure. Figure 6 The method shown may include step 610, step 620, step 630, and step 640. Optionally, Figure 6The method shown can be applied to the case where the image to be processed is a binocular image. The combination of steps 610 to 640 can be used as an alternative implementation for determining the key point pairs of the tracking object based on the original image and the image to be processed of the tracking object obtained in advance in step 320 of the present disclosure.
[0107] Step 610: Based on the binocular image, respectively determine the two-dimensional key points of the tracking object in the binocular image.
[0108] In some alternative implementations of the present disclosure, the two-dimensional key points of the tracking object in the left-eye image and the two-dimensional key points of the tracking object in the right-eye image can be determined by a deep learning method.
[0109] Step 620: Use the corner points on the original image and the two-dimensional key points of the tracking object in one of the binocular images corresponding to the corner points as the first key point pairs.
[0110] In some alternative implementations of the present disclosure, the corner points on the original image and the two-dimensional key points in the left-eye image can be combined to form the first key point pairs.
[0111] Step 630: Use the corner points on the original image and the two-dimensional key points of the tracking object in the other of the binocular images corresponding to the corner points as the second key point pairs.
[0112] In some alternative implementations of the present disclosure, the corner points on the original image and the two-dimensional key points in the right-eye image can be combined to form the second key point pairs.
[0113] Step 640: Use the first key point pairs and the second key point pairs as the key point pairs of the tracking object.
[0114] In step 640, the first key point pairs and the second key point pairs can jointly form the key point pairs of the tracking object. For example, if the corner points on the original image and the two-dimensional key points in the left-eye image form 6 first key point pairs, and the corner points on the original image and the two-dimensional key points in the right-eye image form 6 second key point pairs, then the total number of key point pairs of the tracking object can be 12.
[0115] In the embodiments of the present disclosure, key point pairs can be determined based on the left-eye image and the right-eye image respectively, and the determined key point pairs can all be used as the key point pairs of the tracking object. In this way, both the left-eye image and the right-eye image can provide constraints for the pose estimation process, which is beneficial to improving the accuracy and reliability of the finally obtained target pose.
[0116] In some embodiments, only the first key point pairs can be used as the key point pairs of the tracking object for pose estimation of the tracking object, or only the second key point pairs can be used as the key point pairs of the tracking object for pose estimation of the tracking object.
[0117] See Figure 7 , which is a schematic flowchart of a method for obtaining a target pose provided by some exemplary embodiments of the present disclosure. Figure 7 The method shown may include step 710 and step 720. Optionally, the combination of step 710 and step 720 may be an alternative implementation of step 340 of the present disclosure.
[0118] Step 710: Based on the key point pairs and the matching point pairs, determine the initial pose of the tracking object, and screen out the point pairs that meet the preset retention conditions from the key point pairs and the matching point pairs.
[0119] It should be noted that the key point pairs and the matching point pairs can form a first set of point pairs. Since the corner points and feature points of the tracking object on the original image both have three-dimensional coordinates in the reference coordinate system, a 3D-2D point pair can be obtained from each point pair in the first set of point pairs, and thus a second set of point pairs including multiple 3D-2D point pairs can be obtained. Among them, the 3D point in each 3D-2D point pair represents a point in the reference coordinate system, and the 2D point in each 3D-2D point pair represents a point in the image to be processed.
[0120] In some alternative embodiments of the present disclosure, three point pairs can be randomly selected from the first set of point pairs. According to the three 3D-2D point pairs corresponding one-to-one to the three point pairs selected from the second set of point pairs, and in combination with the PnP algorithm, the pose of the tracking object in the camera coordinate system can be estimated. According to the estimated pose, it can be distinguished which of the remaining point pairs in the first set of point pairs except for the at least three point pairs selected belong to outlier point pairs and which belong to inlier point pairs. Then, three point pairs can be reselected from the first set of point pairs, and the above operations can be repeatedly executed.
[0121] Through the operations in the above paragraph, it can be determined which three point pairs are selected from the first set of point pairs so that the number of inlier point pairs in the first set of point pairs is as large as possible. The pose obtained from the three 3D-2D point pairs corresponding one-to-one to the three point pairs selected and the PnP algorithm can be used as the initial pose. In this case, the inlier point pairs in the first set of point pairs and the three point pairs selected can both be used as the point pairs that meet the preset retention conditions.
[0122] Step 720: Optimize the initial pose based on the point pairs that meet the preset retention conditions to obtain the target pose of the tracking object.
[0123] In some alternative embodiments of the present disclosure, 3D-2D point pairs corresponding to the point pairs in the second point pair set that meet the preset retention conditions can be determined. Based on these 3D-2D point pairs, a predetermined optimization algorithm can be used to optimize the initial pose value by minimizing the reprojection error to obtain the target pose. Optionally, the predetermined optimization algorithm can be either a linear optimization method or a non-linear optimization method. In an alternative example, the predetermined optimization algorithm can be the least squares method.
[0124] In the embodiments of the present disclosure, the random sample consensus algorithm and the PnP algorithm can be combined first to determine the initial pose value of the tracking object and screen the point pairs that meet the preset retention conditions, and then, based on the initial pose value and the point pairs that meet the preset retention conditions, the initial pose value can be further optimized. In this way, the accuracy and reliability of the point pairs used to obtain the final target pose can be better guaranteed, which is conducive to improving the accuracy and reliability of the finally obtained target pose. Moreover, the target pose is not obtained in one step, but is roughly estimated first and then accurately estimated, which is conducive to improving the accuracy and reliability of the point pairs of the final target pose.
[0125] In some embodiments, the initial pose value of the tracking object can also be determined only based on the matching point pairs, and the point pairs that meet the preset retention conditions can be screened from the matching point pairs, and based on the point pairs that meet the preset retention conditions, the initial pose value can be optimized to obtain the target pose of the tracking object.
[0126] See Figure 8 , which is a schematic flow chart of a method for rendering a virtual image provided by some exemplary embodiments of the present disclosure. Figure 8 The method shown may include step 810, step 820, and step 830.
[0127] Step 810, determine the scale relationship.
[0128] As introduced above, when the image to be processed is a monocular image, the scale relationship can be a preset value. When the image to be processed is a binocular image, the scale relationship can be obtained by calculation. For example, the three-dimensional key points determined based on the binocular image can be aligned with the corner points on the original image to obtain the scale relationship. In some embodiments, when optimizing the initial pose value based on the point pairs that meet the preset retention conditions in step 720, the scale relationship obtained by alignment can also be optimized, and the optimization result can be used as the scale relationship in step 810.
[0129] Step 820, determine the target pose.
[0130] Optionally, the target pose can be determined either by Figure 1 the method shown or by Figure 3 the method shown.
[0131] Step 830, render a virtual image based on the scale relationship and the target pose.
[0132] In some alternative embodiments of the present disclosure, the image sensor for collecting the image to be processed may be disposed on the head-mounted display device. The size of the tracking object in the original image can be considered known. Combining with the scale relationship, the true physical size of the tracking object can be determined. Additionally, according to the target pose, the position and orientation of the tracking object relative to the head-mounted display device can be determined. Combining the scale relationship and the target pose, a suitable virtual object can be rendered through the head-mounted display device. For example, a cup that fits the size of the tracking object and has the same orientation can be rendered directly above the tracking object. For another example, a toy that fits the size of the tracking object and has the opposite orientation can be rendered directly above the tracking object. Therefore, adopting Figure 8 the shown embodiment is beneficial to ensuring the display effect of the head-mounted display device, thereby enhancing the user experience.
[0133] Refer to Figure 9 , which is a flowchart showing a method for determining feature points on an original image provided by some exemplary embodiments of the present disclosure. Figure 9 The method shown may include step 910, step 920, step 930, step 940, and step 950. Optionally, Figure 9 the method shown may be executed before step 330 of the present disclosure.
[0134] Step 910, perform a scaling process on the original image to obtain multiple scaled images with different image sizes.
[0135] In some alternative embodiments of the present disclosure, the original image may be subjected to several rounds of scaling processes, and each round of scaling process may obtain a scaled image. In this way, multiple scaled images can be obtained.
[0136] In an optional example, if the width of the original image is W and the height is H, then the original image may first be subjected to a first round of scaling process to obtain a first scaled image with a width of W / 2 and a height of H / 2, then a second round of scaling process to obtain a second scaled image with a width of W / 4 and a height of H / 4, then a third round of scaling process to obtain a third scaled image with a width of W / 8 and a height of H / 8, and so on. Details are not described herein again.
[0137] Step 920, perform feature point extraction on the original image to obtain a first set of feature points.
[0138] In some alternative embodiments of the present disclosure, traditional methods such as SIFT may be used to perform feature point extraction on the original image to obtain a first set of feature points.
[0139] Step 930: Extract feature points from multiple scaled images to obtain a second set of feature points.
[0140] In some alternative embodiments of the present disclosure, traditional methods such as SIFT can be used to extract feature points from multiple scaled images to obtain a second set of feature points. The second set of feature points may include multiple feature points in each scaled image.
[0141] Step 940: Map the second set of feature points to the original image to obtain a third set of feature points.
[0142] It should be noted that there is a certain proportional relationship between the scale of each scaled image and the scale of the original image. According to the proportional relationship, multiple feature points in each scaled image can be mapped to the original image. In this way, a third set of feature points composed of the mapped points corresponding to multiple feature points in each scaled image can be obtained.
[0143] Step 950: Determine the feature points on the original image based on the first set of feature points and the third set of feature points.
[0144] In some alternative embodiments of the present disclosure, if the first set of feature points and the third set of feature points have no intersection, each feature point in the set of feature points composed of the first set of feature points and the third set of feature points can be directly used as the feature points on the original image.
[0145] In some other alternative embodiments of the present disclosure, if the first set of feature points and the third set of feature points have an intersection, the union of the first set of feature points and the third set of feature points can be determined, and each feature point in the union can be used as the feature points on the original image.
[0146] In the embodiments of the present disclosure, not only can feature extraction be performed on the original image, but also feature extraction can be performed on the scaled images obtained by scaling the original image. In this way, it is equivalent to extracting multiple layers of features through pyramid layering of the original image, and using the multiple layers of features for pose estimation of the tracked object, which is beneficial to improving the accuracy and reliability of the finally obtained target pose.
[0147] In some embodiments, pyramid layering may not be performed on the original image, but feature points may be directly extracted from the original image to obtain a first set of feature points, and each feature point in the first set of feature points can be used as the feature points on the original image.
[0148] In some alternative embodiments of the present disclosure, as Figure 10 shown, the pose estimation of the tracked object can be implemented through the following module:
[0149] Bounding box detection module: Using an object detection network (such as YOLOv5), detect the position and category of the tracking object in the left and right eye images respectively.
[0150] Key point detection module: For each detected tracking object, use a key point detection network (such as SimCC) to identify the 2D key points of the tracking object in the left and right eye images.
[0151] Feature point extraction module: Perform pyramid layering on the original image, extract multi-layer features, and obtain the feature points on the original image.
[0152] Matched point pair extraction module: First, triangulate using the 2D key points in the left and right eye images to estimate the position of the 3D key points of the tracking object in the camera coordinate system corresponding to the left eye image. Then, align the estimated position of the 3D key points with the corner point position of the tracking object on the original image (the umeyama algorithm can be used) to estimate the rough pose (equivalent to the pose estimation value in the above text) and scale relationship of the tracking object. The feature points on the original image can be reprojected using this rough pose (projected into the camera coordinate system corresponding to the left eye image) to obtain the initial position of the points matching the feature points on the original image in the left eye image. Finally, through the NCC template matching algorithm, obtain the accurate position of the matched points on the left eye image. In a similar manner, the accurate position of the matched points on the right eye image can be obtained.
[0153] Pose estimation module: Based on the matched point pairs and key point pairs, first use the random sample consensus algorithm and the PnP algorithm to estimate the initial pose value of the tracking object, and filter out the point pairs belonging to outliers. Then, through the least squares optimization algorithm, for the point pairs belonging to inliers, minimize the reprojection error to estimate the accurate pose (equivalent to the target pose in the above text) and accurate scale relationship of the tracking object.
[0154] In summary, by adopting the embodiments of the present disclosure, it is possible to effectively implement the pose estimation of the tracking object by combining deep learning methods and non-deep learning methods. Even if the area occupied by the tracking object in the image to be processed is small and the observation angle is large, a good pose estimation effect can still be achieved.
[0155] Any of the image-based pose estimation methods provided by the embodiments of the present disclosure can be executed by any suitable device with data processing capabilities, including but not limited to: terminal devices and servers, etc. Or, any of the image-based pose estimation methods provided by the embodiments of the present disclosure can be executed by a processor. For example, the processor executes any of the image-based pose estimation methods mentioned in the embodiments of the present disclosure by calling the corresponding instructions stored in the memory. This will not be elaborated further below.
[0156] Exemplary Device
[0157] Figure 11 It is a schematic structural diagram of an image-based pose estimation device provided by some exemplary embodiments of the present disclosure. Figure 11 The illustrated device includes: a first acquisition module 1110 for acquiring a to-be-processed image including a tracking object; a first determination module 1120 for determining key point pairs of the tracking object based on a pre-acquired original image and the to-be-processed image of the tracking object, where the key point pairs include corner points of the tracking object on the original image and two-dimensional key points of the tracking object in the to-be-processed image corresponding to the corner points, and the corner points have three-dimensional coordinates in the reference coordinate system where the original image is located; a second determination module 1130 for determining matching point pairs of the tracking object based on the original image and the to-be-processed image, where the matching point pairs include feature points of the tracking object on the original image and two-dimensional feature points of the tracking object in the to-be-processed image that match the feature points on the original image, and the feature points on the original image have three-dimensional coordinates in the reference coordinate system; a second acquisition module 1140 for obtaining the target pose of the tracking object based on the key point pairs and the matching point pairs.
[0158] In some alternative embodiments of the present disclosure, the three-dimensional coordinates of the corner points of the tracking object on the original image in the reference coordinate system are determined based on a scale relationship, and the scale relationship is the relationship between the size of the tracking object in the original image and the true physical size of the tracking object. The scale relationship is determined in the following manner: when the to-be-processed image is a monocular image, the scale relationship is a preset value; when the to-be-processed image is a binocular image, three-dimensional key points of the tracking object are determined based on the binocular image, and the three-dimensional key points are aligned with the corner points on the original image to obtain the scale relationship.
[0159] It can be understood that the corner points of the tracking object on the original image belong to a type of key points. In some alternative embodiments of the present disclosure, other key points on the tracking object in the original image can be extracted to form key point pairs with the two-dimensional key points of the tracking object in the to-be-processed image.
[0160] In some alternative embodiments of the present disclosure, as Figure 12 shown, the device provided by the embodiments of the present disclosure further includes: an estimation module 1210 for estimating a pose estimation value of the tracking object according to the key point pairs before determining the matching point pairs of the tracking object based on the original image and the to-be-processed image; a second determination module 1130 for projecting the feature points on the original image to the to-be-processed image based on the pose estimation value and the scale relationship to obtain projection points in the to-be-processed image, optimizing the positions of the projection points to obtain two-dimensional feature points of the tracking object in the to-be-processed image that match the feature points on the original image, and using the feature points on the original image and the two-dimensional feature points as the matching point pairs of the tracking object.
[0161] In some alternative embodiments of the present disclosure, asFigure 13 As shown in Figure 13 , when the image to be processed is a binocular image, the first determination module 1120 includes: a first determination sub-module 1310, configured to respectively determine two-dimensional key points of a tracking object in the binocular image based on the binocular image; a second determination sub-module 1320, configured to use the corner points on the original image and the two-dimensional key points of the tracking object in one of the binocular images corresponding to the corner points as the first key point pair; a third determination sub-module 1330, configured to use the corner points on the original image and the two-dimensional key points of the tracking object in the other of the binocular images corresponding to the corner points as the second key point pair; a fourth determination sub-module 1340, configured to use the first key point pair and the second key point pair as the key point pair of the tracking object.
[0162] In some alternative embodiments of the present disclosure, as Figure 14 As shown in Figure 14 , the estimation module 1210 includes: a fifth determination sub-module 1410, configured to determine three-dimensional key points of the two-dimensional key points in the camera coordinate system based on the two-dimensional key points of the tracking object in the binocular image; an alignment sub-module 1420, configured to align the three-dimensional key points in the camera coordinate system with the corner points on the original image to obtain the conversion relationship between the camera coordinate system and the reference coordinate system; an estimation sub-module 1430, configured to estimate the pose estimation value of the tracking object according to the conversion relationship.
[0163] In some alternative embodiments of the present disclosure, as Figure 15 As shown in Figure 15 , the second acquisition module 1140 includes: a processing sub-module 1510, configured to determine an initial pose value of the tracking object based on the key point pair and the matching point pair, and screen out the point pairs that meet the preset retention conditions from the key point pair and the matching point pair; an optimization sub-module 1520, configured to optimize the initial pose value based on the point pairs that meet the preset retention conditions to obtain the target pose of the tracking object.
[0164] In some alternative embodiments of the present disclosure, as Figure 16 As shown in Figure 16 , the device provided by the embodiments of the present disclosure further includes: a rendering module 1610, configured to render a virtual image based on the scale relationship and the target pose.
[0165] In some alternative embodiments of the present disclosure, as Figure 17As shown, the apparatus provided by an embodiment of the present disclosure further includes: a scaling module 1710, configured to perform scaling processing on an original image to obtain multiple scaled images with different image sizes before determining a pair of matching points of a tracking object based on a pose estimation value and an image to be processed; a first feature point extraction module 1720, configured to extract feature points from the original image to obtain a first set of feature points; a second feature point extraction module 1730, configured to extract feature points from the multiple scaled images to obtain a second set of feature points; a mapping module 1740, configured to map the second set of feature points to the original image to obtain a third set of feature points; and a third determination module 1750, configured to determine feature points on the original image based on the first set of feature points and the third set of feature points.
[0166] In the apparatus of the present disclosure, the various optional embodiments, optional implementations, and optional examples disclosed above can be flexibly selected and combined as needed to achieve corresponding functions and effects, and the present disclosure does not list them one by one.
[0167] Exemplary Electronic Device
[0168] Figure 18 The block diagram of an electronic device according to an embodiment of the present disclosure is illustrated. The electronic device 1800 includes one or more processors 1810 and a memory 1820.
[0169] The processor 1810 may be a central processing unit (CPU) or other form of processing unit having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 1800 to perform desired functions.
[0170] The memory 1820 may include one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage media, and the processor 1810 may run the one or more computer program instructions to implement the methods of the various embodiments of the present disclosure described above and / or other desired functions.
[0171] In one example, the electronic device 1800 may further include: an input device 1830 and an output device 1840, and these components are interconnected through a bus system and / or other forms of connection mechanisms (not shown).
[0172] The input device 1830 may further include, for example, a keyboard, a mouse, and so on.
[0173] The output device 1840 can output various information to the outside, which can include, for example, a display, a speaker, a printer, a communication network, and remote output devices connected thereto, and so on.
[0174] Of course, for simplicity, Figure 18 only some of the components related to the present disclosure in the electronic device 1800 are shown, and components such as a bus, an input / output interface, and so on are omitted. In addition, according to specific application scenarios, the electronic device 1800 may further include any other appropriate components.
[0175] Exemplary Computer Program Product and Computer Readable Storage Medium
[0176] In addition to the above methods and devices, embodiments of the present disclosure may also be a computer program product, which includes computer program instructions that, when run by a processor, cause the processor to execute the steps in the methods according to various embodiments of the present disclosure described in the above "Exemplary Method" section of this specification.
[0177] The computer program product can be written in any combination of one or more programming languages to write program code for performing the operations of the embodiments of the present disclosure. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, executed as an independent software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0178] Furthermore, embodiments of the present disclosure may also be a computer-readable storage medium having computer program instructions stored thereon, and the computer program instructions, when run by a processor, cause the processor to execute the steps in the methods according to various embodiments of the present disclosure described in the above "Exemplary Method" section of this specification.
[0179] The computer-readable storage medium may employ any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may include, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0180] The basic principles of the present disclosure have been described in conjunction with specific embodiments. However, the advantages, benefits, effects, etc. mentioned in the present disclosure are only examples and not limitations. It cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present disclosure. The specific details disclosed above are only for the purposes of illustration and easy understanding, rather than limitations. These details do not limit the present disclosure to necessarily adopt the above specific details for implementation.
[0181] Those skilled in the art can make various changes and modifications to the present disclosure without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present disclosure and their equivalent technologies, the present disclosure also intends to include these changes and modifications.
Claims
1. An image-based pose estimation method, comprising: Obtaining a to-be-processed image including a tracking object; Determining key point pairs of the tracking object based on a pre-obtained original image of the tracking object and the to-be-processed image, wherein the key point pairs include corner points of the tracking object on the original image and two-dimensional key points of the tracking object in the to-be-processed image corresponding to the corner points, and the corner points have three-dimensional coordinates in a reference coordinate system where the original image is located; Determining matching point pairs of the tracking object based on the original image and the to-be-processed image, wherein the matching point pairs include feature points of the tracking object on the original image and two-dimensional feature points of the tracking object in the to-be-processed image matching the feature points on the original image, and the feature points on the original image have three-dimensional coordinates in the reference coordinate system; Obtaining the target pose of the tracking object based on the key point pairs and the matching point pairs.
2. The method according to claim 1, wherein, The three-dimensional coordinates of the corner points of the tracking object on the original image in the reference coordinate system are determined based on a scale relationship, where the scale relationship is the relationship between the size of the tracking object in the original image and the true physical size of the tracking object, and the scale relationship is determined by the following method: When the to-be-processed image is a monocular image, the scale relationship is a preset value; When the to-be-processed image is a binocular image, determining three-dimensional key points of the tracking object based on the binocular image, aligning the three-dimensional key points with the corner points on the original image, and obtaining the scale relationship.
3. The method according to any one of claims 1 or 2, wherein Before the step of determining the matching point pairs of the tracking object based on the original image and the to-be-processed image, the method further includes: Estimating a pose estimation value of the tracking object according to the key point pairs; The determining the matching point pairs of the tracking object based on the original image and the to-be-processed image includes: Based on the pose estimation value and the scale relationship, projecting the feature points on the original image onto the to-be-processed image to obtain projection points in the to-be-processed image, optimizing the positions of the projection points, obtaining two-dimensional feature points of the tracking object in the to-be-processed image matching the feature points on the original image, and using the feature points on the original image and the two-dimensional feature points as the matching point pairs of the tracking object.
4. The method according to any one of claims 2 or 3, wherein When the to-be-processed image is a binocular image, the determining the key point pairs of the tracking object based on a pre-obtained original image of the tracking object and the to-be-processed image includes: Based on the binocular image, respectively determining two-dimensional key points of the tracking object in the binocular image; Using the corner points on the original image and the two-dimensional key points of the tracking object in one of the binocular images corresponding to the corner points as a first key point pair; Using the corner points on the original image and the two-dimensional key points of the tracking object in the other of the binocular images corresponding to the corner points as a second key point pair; Using the first key point pair and the second key point pair as the key point pairs of the tracking object.
5. The method according to claim 3 or 4, wherein The estimating a pose estimation value of the tracking object according to the key point pairs includes: Determine the three-dimensional key points of the two-dimensional key points of the tracked object in the binocular image in the camera coordinate system; Align the three-dimensional key points in the camera coordinate system with the corner points on the original image to obtain the conversion relationship between the camera coordinate system and the reference coordinate system; Estimate the pose estimate of the tracked object according to the conversion relationship.
6. The method according to any one of claims 2-5, wherein, The method further includes: rendering a virtual image based on the scale relationship and the target pose.
7. The method according to any one of claims 1-6, wherein, The obtaining the target pose of the tracked object based on the key point pairs and the matching point pairs includes: Based on the key point pairs and the matching point pairs, determine the initial pose value of the tracked object, and screen out the point pairs that meet the preset retention conditions from the key point pairs and the matching point pairs; Based on the point pairs that meet the preset retention conditions, optimize the initial pose value to obtain the target pose of the tracked object.
8. According to the method as claimed in any one of claims 1-7, wherein, Before the step of determining the matching point pairs of the tracked object based on the pose estimate value and the image to be processed, the method further includes: Perform a scaling process on the original image to obtain multiple scaled images with different image sizes; Extract feature points from the original image to obtain a first set of feature points; Extract feature points from the multiple scaled images to obtain a second set of feature points; Map the second set of feature points to the original image to obtain a third set of feature points; Based on the first set of feature points and the third set of feature points, determine the feature points on the original image.
9. An image-based pose estimation device, comprising: A first acquisition module, configured to acquire an image to be processed including a tracked object; A first determination module, configured to determine key point pairs of the tracked object based on the original image and the image to be processed of the tracked object acquired in advance, where the key point pairs include corner points of the tracked object on the original image and two-dimensional key points of the tracked object in the image to be processed corresponding to the corner points, and the corner points have three-dimensional coordinates in the reference coordinate system where the original image is located; A second determination module, configured to determine matching point pairs of the tracked object based on the original image and the image to be processed, where the matching point pairs include feature points of the tracked object on the original image and two-dimensional feature points of the tracked object in the image to be processed that match the feature points on the original image, and the feature points on the original image have three-dimensional coordinates in the reference coordinate system; A second acquisition module, configured to obtain the target pose of the tracked object based on the key point pairs and the matching point pairs.
10. An electronic device, comprising: A memory, configured to store a computer program product; A processor, configured to execute the computer program product stored in the memory, and when the computer program product is executed, implement the image-based pose estimation method according to any one of claims 1 to 8 above.
11. A computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, implement the image-based pose estimation method according to any one of claims 1 to 8 above.