Object pose estimation method and apparatus, electronic device, and computer program product

By using the golden angle spherical sampling method to render 3D models and generate multiple uniform viewpoint images, the problem of low accuracy and efficiency in existing 6D pose estimation is solved, achieving efficient and accurate object pose estimation and improving user experience.

CN122265382APending Publication Date: 2026-06-23UBTECH ROBOTICS CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610570050.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-27
Publication Date
2026-06-23

AI Technical Summary

Technical Problem

Existing 6D pose estimation techniques suffer from poor accuracy, especially in 6D pose estimation based on unknown objects. The accuracy of 3D reconstruction is low, the pose generation efficiency is low, and it cannot cover most of the object's viewpoint, resulting in a poor user experience.

Method used

A spherical sampling method with a golden angle is used to obtain spherical pose. Multiple uniformly distributed viewpoint images are generated by rendering a 3D model and input into the object pose estimation model for processing, thereby improving the accuracy and efficiency of pose estimation.

Benefits of technology

It achieves efficient and accurate object pose estimation, covering various perspectives of the target object, thus improving the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122265382A_ABST
    Figure CN122265382A_ABST
Patent Text Reader

Abstract

This application relates to the field of image processing technology, and particularly to object pose estimation methods, apparatuses, electronic devices, and computer program products. In this method, after acquiring a first image containing a target object, the electronic device can determine a first 3D model corresponding to the target object, and render the first 3D model based on a first spherical pose to obtain a second image corresponding to the first 3D model. The first and second images are then input into an object pose estimation model for processing to obtain the object pose corresponding to the target object. The first spherical pose is obtained based on a spherical sampling method using the golden angle, which enables efficient acquisition of the spherical pose, ensuring uniform and non-clustered spherical pose. This allows the image rendered based on the spherical pose to cover all viewpoints of the target object, enabling the object pose estimation model to efficiently and accurately determine the object pose corresponding to the target object based on the rendered image covering all viewpoints, thus improving the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of image processing technology, and in particular relates to an object pose estimation method, device, electronic device and computer program product. Background Technology

[0002] With the rapid development of artificial intelligence (AI) and robotics technologies, six-degree-of-freedom (6DOF) pose estimation has become a key task connecting visual perception and physical manipulation. The goal of 6D pose estimation is to estimate the position of an object in three-dimensional space. x , y , z Spatial information, including rotational pose (roll, pitch, yaw, or rotation matrix), can provide spatial information for robot grasping, augmented reality (AR), autonomous navigation, automated assembly, and visual measurement. However, current 6D pose estimation suffers from poor accuracy. Summary of the Invention

[0003] This application provides an object pose estimation method, apparatus, electronic device, and computer program product, which can improve the accuracy of object pose estimation.

[0004] In a first aspect, embodiments of this application provide an object pose estimation method, including: Obtain the first image containing the target object; Determine the first three-dimensional model corresponding to the target object; The first 3D model is rendered based on the first spherical pose to obtain the second image corresponding to the first 3D model; the first spherical pose is a pose generated based on spherical points, which are obtained by sampling on a unit sphere using a spherical sampling method based on the golden angle. The spherical points include multiple points. When sampling on the unit sphere using the spherical sampling method based on the golden angle, a preset vertex of the unit sphere is determined as the first spherical point, and the next spherical point is obtained by rotating the golden angle each time. The first image and the second image are input into the object pose estimation model for processing to obtain the object pose corresponding to the target object.

[0005] In the object pose estimation method provided above, the electronic device can acquire a first image containing the target object and determine the first 3D model corresponding to the target object. Subsequently, the electronic device can render the first 3D model based on the first spherical pose to obtain a second image corresponding to the first 3D model. The first and second images are then input into the object pose estimation model for processing to obtain the object pose corresponding to the target object. The first spherical pose can be obtained based on a spherical sampling method using the golden angle, enabling efficient acquisition of the spherical pose. This ensures that the sampled spherical pose is uniform and non-clustered, allowing the rendered image based on the spherical pose to cover all viewpoints of the target object. Therefore, the object pose estimation model can efficiently and accurately determine the object pose corresponding to the target object based on the rendered image covering all viewpoints of the target object, thus improving the user experience.

[0006] In some embodiments, before rendering the 3D model based on the first spherical pose, the object pose estimation method further includes: Determine the number of points corresponding to the spherical surface; Based on the unit sphere and the number of corresponding spherical points, determine the corresponding number of each spherical point. y coordinate; For each of the spherical points, based on the unit sphere and the corresponding spherical point... y The coordinates are used to determine the target radius corresponding to the spherical point, and the circumference of the spherical point is determined according to the golden angle. y The target angle of axis rotation is used to determine the corresponding point on the sphere based on the target radius and the target angle. x coordinates and z coordinate.

[0007] For example, the step of determining the number of points corresponding to each spherical point based on the unit sphere and the number of points corresponding to the spherical surface. y Coordinates, including: The corresponding point on each sphere is determined according to the following formula. y coordinate: y i = R - ; in, y i For the first i The corresponding spherical point y coordinate, R The radius corresponding to a unit sphere. R =1, N The number of points corresponding to the sphere. i It is a positive integer, and 0 ≤ i ≤N-1 .

[0008] For example, the step of determining the spherical point around the golden angle... y The target angle of axis rotation includes: The target angle of rotation of the spherical point about the y-axis is determined using the following formula: ; in, For the first i The target angle corresponding to each spherical point This refers to the golden angle.

[0009] In one embodiment, each of the first spherical poses is a 4×4 homogeneous transformation matrix, the homogeneous transformation matrix including a rotation matrix and a translation vector, the rotation matrix including... x The column vectors corresponding to the axes y The column vectors corresponding to the axes z The column vector corresponding to the axis.

[0010] For example, the object pose estimation method further includes: For each of the spherical points, according to the spherical point corresponding to x coordinate, y coordinates and z Coordinates are used to determine the translation vector in the homogeneous transformation matrix; Based on the spherical points, determine a first unit vector pointing to the origin of the unit sphere, and define the first unit vector as the value in the rotation matrix. z The column vectors corresponding to the axes; The preset direction vector is cross-producted with the first unit vector to obtain a second unit vector, which is then defined as the second unit vector in the rotation matrix. x The column vectors corresponding to the axes; Perform a cross product between the first unit vector and the second unit vector to obtain a third unit vector, and define the third unit vector as a component in the rotation matrix. y The column vector corresponding to the axis.

[0011] In some embodiments, the object pose estimation model is trained based on a third image and a fourth image. The third image is an image containing the training object, and the fourth image is an image generated by rendering a second three-dimensional model based on a second spherical pose. The second spherical pose is a pose obtained by a spherical sampling method based on the golden angle, and the second three-dimensional model is a three-dimensional model corresponding to the training object.

[0012] Secondly, embodiments of this application provide an object pose estimation device, which may include: The image acquisition module is used to acquire a first image containing the target object; A 3D model determination module is used to determine the first 3D model corresponding to the target object; The image rendering module is used to render the first three-dimensional model according to the first spherical pose to obtain a second image corresponding to the first three-dimensional model; the first spherical pose is a pose generated based on spherical points, the spherical points are obtained by sampling on a unit sphere based on the golden angle spherical sampling method, the spherical points include multiple points, when sampling on the unit sphere based on the golden angle spherical sampling method, the preset vertex of the unit sphere is determined as the first spherical point, and the next spherical point is obtained by rotating the golden angle each time; The pose estimation module is used to input the first image and the second image into the object pose estimation model for processing to obtain the object pose corresponding to the target object.

[0013] In some embodiments, the object pose estimation device may further include: A spherical pose generation module is used to determine the number of points corresponding to the spherical surface; and to determine the number of points corresponding to each spherical surface based on the unit sphere and the number of points corresponding to the spherical surface. y Coordinates; for each of the spherical points, based on the unit sphere and the coordinates corresponding to the spherical point. y The coordinates are used to determine the target radius corresponding to the spherical point, and the circumference of the spherical point is determined according to the golden angle. y The target angle of axis rotation is used to determine the corresponding point on the sphere based on the target radius and the target angle. x coordinates and z coordinate.

[0014] For example, the spherical pose generation module is further configured to determine the corresponding point on each spherical surface according to the following formula. y coordinate: y i = R - ; in, y i For the first i The corresponding spherical point y coordinate, R The radius corresponding to a unit sphere. R =1, N The number of points corresponding to the sphere. i It is a positive integer, and 0 ≤ i ≤ N-1 .

[0015] For example, the spherical pose generation module is further configured to determine the spherical point around the sphere according to the following formula. y Target angle of axis rotation: ; in, For the first i The target angle corresponding to each spherical point This refers to the golden angle.

[0016] In one embodiment, each of the first spherical poses is a 4×4 homogeneous transformation matrix, the homogeneous transformation matrix including a rotation matrix and a translation vector, the rotation matrix including... x The column vectors corresponding to the axes y The column vectors corresponding to the axes z The column vector corresponding to the axis.

[0017] For example, the spherical pose generation module is further configured to, for each spherical point, determine the position of the spherical point according to the position of the spherical point. x coordinate, y coordinates and z The coordinates are used to determine the translation vector in the homogeneous transformation matrix; based on the spherical point, a first unit vector pointing to the origin of the unit sphere is determined, and this first unit vector is used as the vector in the rotation matrix. z The column vector corresponding to the axis; the cross product of the preset direction vector and the first unit vector is taken to obtain the second unit vector, and the second unit vector is determined as the column vector in the rotation matrix. x The column vector corresponding to the axis; the cross product of the first unit vector and the second unit vector is taken to obtain the third unit vector, and the third unit vector is determined as the column vector in the rotation matrix. y The column vector corresponding to the axis.

[0018] In some embodiments, the object pose estimation model is trained based on a third image and a fourth image. The third image is an image containing the training object, and the fourth image is an image generated by rendering a second three-dimensional model based on a second spherical pose. The second spherical pose is a pose obtained by a spherical sampling method based on the golden angle, and the second three-dimensional model is a three-dimensional model corresponding to the training object.

[0019] Thirdly, embodiments of this application provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the electronic device implements the object pose estimation method described in any one of the first aspects above.

[0020] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by an electronic device, causes the electronic device to implement the object pose estimation method described in any one of the first aspects above.

[0021] Fifthly, embodiments of this application provide a computer program product, which includes a computer program that, when executed by an electronic device, causes the electronic device to implement the object pose estimation method described in any one of the first aspects.

[0022] It is understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 This is a flowchart illustrating the object pose estimation method provided in the embodiments of this application; Figure 2 This is a schematic diagram of the object pose estimation device provided in the embodiments of this application; Figure 3 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0025] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0026] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0027] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0028] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."

[0029] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0030] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0031] With the rapid development of AI and robotics technologies, 6D pose estimation has become a key task connecting visual perception and physical manipulation. The goal of 6D pose estimation is to estimate the position of an object in three-dimensional space. x, y, z Spatial information, along with rotational attitude (roll, pitch, yaw, or rotation matrix), can provide spatial information for robot grasping, augmented reality (AR), autonomous navigation, automated assembly, and visual measurement.

[0032] 6D pose estimation can be divided into 6D pose estimation based on known objects and 6D pose estimation based on unknown objects. 6D pose estimation based on known objects refers to performing 6D pose estimation based on the object's 3D model, such as a computer-aided design (CAD) model. In this method, several poses need to be generated, and the CAD model of the object is rendered based on these generated poses to obtain rendered images of the object's CAD model in different poses. These rendered images are then used as input to the pose estimation model, resulting in higher accuracy. In contrast, 6D pose estimation based on unknown objects, lacking a CAD model of the object, typically requires 3D reconstruction, resulting in lower accuracy.

[0033] In 6D pose estimation based on known objects, the CAD model of the object needs to be rendered based on several generated poses to obtain a rendered image, which serves as input to the pose estimation model. If the pose generation efficiency is low and cannot cover most of the object's viewpoint, it will reduce the speed and accuracy of 6D pose estimation. For example, pose generation can be based on a sampling method that subdivides triangles into a regular icosahedron. However, this method requires continuously subdividing the triangle faces to obtain new sampling points, which is inefficient. Furthermore, the order of subdividing the triangle faces is difficult to control, making it difficult to arbitrarily sample a number of points that are uniformly distributed on the sphere.

[0034] In summary, the above pose sampling methods suffer from poor efficiency and accuracy, resulting in poor speed and accuracy of 6D pose estimation and a poor user experience.

[0035] To address the aforementioned issues, this application provides an object pose estimation method, apparatus, electronic device, and computer program product. In this object pose estimation method, the electronic device can acquire a first image containing a target object and determine a first 3D model corresponding to the target object. Subsequently, the electronic device can render the first 3D model based on a first spherical pose to obtain a second image corresponding to the first 3D model. The first and second images are then input into an object pose estimation model for processing to obtain the object pose corresponding to the target object. The first spherical pose can be obtained using a spherical sampling method based on the golden angle, enabling efficient acquisition of the spherical pose. This ensures that the sampled spherical pose is uniform and non-clustered, allowing the rendered image based on the spherical pose to cover all viewpoints of the target object. Therefore, the object pose estimation model can efficiently and accurately determine the object pose corresponding to the target object based on the rendered image covering all viewpoints, improving user experience and demonstrating strong usability and practicality.

[0036] The object pose estimation method provided in this application can be applied to electronic devices such as robots, mobile phones, tablets, wearable devices, vehicle-mounted devices, augmented reality (AR) / virtual reality (VR) devices, laptops, ultra-mobile personal computers (UMPCs), netbooks, personal digital assistants (PDAs), and cloud servers. This application does not impose any restrictions on the specific type of electronic device.

[0037] Please see Figure 1 , Figure 1 A schematic flowchart illustrating an object pose estimation method provided in an embodiment of this application is shown. This object pose estimation method can be applied to any of the electronic devices described above. Figure 1 As shown, the object pose estimation method may include: S101, The electronic device acquires a first image containing the target object.

[0038] It should be noted that the target object can be any object. The first image can be a real image containing the target object captured by a camera. For example, in a scenario where an electronic device (e.g., a robot) grasps an object, the target object can be the object that the robot needs to grasp. For instance, the robot may include a camera, and when the robot needs to grasp the target object, it can capture a real image containing the target object through the camera, i.e., the first image. Alternatively, the robot may not include a camera, but it can communicate with device 1, and device 1 may include a camera. When the robot needs to grasp the target object, it can send an instruction to device 1 to capture an image. Device 1 can then control its camera to capture the first image containing the target object based on this instruction, and can send the captured first image to the robot.

[0039] S102. The electronic device determines the first three-dimensional model corresponding to the target object.

[0040] For example, after acquiring the first image, the electronic device can analyze the first image to determine the target object contained in the first image. After determining the target object contained in the first image, the electronic device can acquire a first three-dimensional model corresponding to the target object.

[0041] It should be noted that the electronic device or the device 2 connected in communication with the electronic device may store three-dimensional models corresponding to each object. After determining the target object contained in the first image, the electronic device can determine the first three-dimensional model corresponding to the target object from the three-dimensional models stored in the electronic device or the three-dimensional models stored in the device 2 connected in communication with the electronic device. For example, the three-dimensional models corresponding to each object can be CAD models, that is, the first three-dimensional model corresponding to the target object can be the CAD model corresponding to the target object. It should be understood that the CAD models corresponding to each object can be constructed according to the actual scene, and the embodiments of this application do not limit the construction process of the CAD models corresponding to each object.

[0042] S103. The electronic device renders the first three-dimensional model according to the first spherical pose to obtain the second image corresponding to the first three-dimensional model. The first spherical pose is a pose generated based on spherical points. The spherical points are obtained by sampling on a unit sphere using a spherical sampling method based on the golden angle. There are multiple spherical points. When sampling on a unit sphere using the spherical sampling method based on the golden angle, the preset vertex of the unit sphere is determined as the first spherical point. The next spherical point is obtained by rotating the golden angle each time.

[0043] For example, a first spherical pose may include multiple poses. The electronic device can render the first 3D model according to each first spherical pose to obtain a second image of the first 3D model under each first spherical pose. A second image can be a rendered image generated by projecting the first 3D model onto a 2D image plane under a first spherical pose. Each second image may include an RGB image, or may include an RGB image and a depth image. That is, a first spherical pose can represent a viewpoint and can be used to characterize a viewing angle. A second image can refer to an image of a target object presented from a viewing angle.

[0044] It should be understood that the embodiments of this application do not limit the specific method by which the electronic device renders the first three-dimensional model according to the first spherical pose to obtain the second image corresponding to the first three-dimensional model. Rendering can be performed according to any existing method to obtain the second image.

[0045] In some embodiments, the electronic device can acquire spherical points based on a unit sphere and generate a first spherical pose based on these points. The first spherical pose can be a pose obtained using Fibonacci spherical sampling. For example, a preset vertex of the unit sphere (e.g., the "North Pole" vertex) can be used as the initial point (i.e., the first spherical point). The device rotates by a golden angle each time to acquire the next sampling point (i.e., the spherical point), generating multiple spherical points evenly distributed on the sphere. This allows for the generation of multiple uniformly distributed, non-clustered first spherical poses. In other words, this embodiment of the application can efficiently and uniformly acquire multiple first spherical poses using the Fibonacci spherical sampling method. This ensures that the rendered image based on the multiple first spherical poses covers all viewpoints of the target object, enabling the object pose estimation model to efficiently and accurately determine the object pose corresponding to the target object based on the rendered image covering all viewpoints. Furthermore, the Fibonacci spherical sampling method can adaptively adjust based on the number of spherical poses, enabling the generation of N uniformly distributed spherical points on the sphere from an input of N spherical points. Among these, the golden angle... That is, the golden angle can be 137.5 degrees.

[0046] In one embodiment, when acquiring spherical points based on a unit sphere, the electronic device can determine the number of spherical points to determine the coordinates of each spherical point. After determining the coordinates of a certain spherical point, the electronic device can generate a first spherical pose based on those coordinates. It should be understood that the number of spherical points can be user-defined or set by default by the electronic device. This embodiment does not limit this and can be determined according to the actual scenario. For example, the user can determine the number of spherical points based on the actual computing power of the electronic device, or the electronic device can adaptively adjust the number of spherical points based on its actual computing power. This can achieve a balance between the determination speed and accuracy of the object pose estimation model, improving both the determination speed and the accuracy of pose estimation.

[0047] The following is the first i Taking a spherical point as an example, the process of an electronic device determining the coordinates of the spherical point and generating the first spherical pose corresponding to the spherical point is illustrated.

[0048] In one embodiment, after determining the number of points corresponding to the spherical surface, the electronic device can determine the number of points corresponding to the unit sphere based on the number of points corresponding to the spherical surface. i The corresponding spherical point y Coordinates. Where, the first... i The corresponding spherical point y coordinatey i = R - . y i For the first i The corresponding spherical point y coordinate, R The radius corresponding to a unit sphere, i.e. R =1, N The number of points corresponding to the sphere. i It is a positive integer, and 0 ≤ i ≤ N-1 .

[0049] In determining the first i The corresponding spherical point y After coordinates, the electronic device can determine the coordinates based on the unit sphere and the first... i The corresponding spherical point y Coordinates, determine the first i The target radius corresponding to each spherical point is determined based on the golden angle. i A spherical point around y The target angle of rotation of the axis. It should be understood that the first... i The target radius corresponding to a point on a sphere can refer to the radius of a unit sphere. y = y i The radius of the circle intersecting the horizontal planes they lie on. Wherein, the... i The target radius corresponding to each spherical point . No. i A spherical point around y Target angle of axis rotation . It is the golden angle.

[0050] In obtaining the i After determining the target radius and target angle corresponding to the first spherical point, the electronic device can then... i The target radius and target angle corresponding to the nth spherical point are used to determine the nth... i The corresponding spherical point x coordinates and z Coordinates. Where, the first... i The corresponding spherical point x coordinate . No. i The corresponding spherical point z coordinate Therefore, electronic devices can determine the first... i The coordinates of each point on the sphere .

[0051] In one embodiment, each first spherical pose can be a 4×4 homogeneous transformation matrix. The homogeneous transformation matrix can include a rotation matrix and a translation vector. The rotation matrix can include... x The column vectors corresponding to the axes y The column vectors corresponding to the axes z The column vectors corresponding to the axes. Wherein, x The column vectors corresponding to the axes y The column vectors corresponding to the axes z The column vectors corresponding to the axes are all 3×1 column vectors, and are orthogonal unit vectors. For example, the first... i First spherical pose . For the first i In the rotation matrix corresponding to the first spherical pose x The column vectors corresponding to the axes, For the first i In the rotation matrix corresponding to the first spherical pose y The column vectors corresponding to the axes, For the first i In the rotation matrix corresponding to the first spherical pose z The column vectors corresponding to the axes, For the first i Translation vector in the first spherical pose.

[0052] For example, in the rotation matrix z The axial direction can be a unit vector pointing from each point on the unit sphere to the origin (i.e., the center of the unit sphere, for example, (0, 0, 0)). x The axis direction can be connected to the direction vector (0, 0, 1) and... z The cross product of the axes is obtained. y The axial direction can be passed z shaft and x The cross product of axes is obtained. That is, in determining the first... i The coordinates of each point on the sphere Afterwards, the electronic device can be based on the first i The coordinates of each point on the sphere Determine the first unit vector pointing to the origin of the unit sphere, and the first unit vector can be defined as the... i In the rotation matrix corresponding to the first spherical pose z The column vector corresponding to the axis. Additionally, the electronic device can perform a cross product between a preset direction vector and a first unit vector to obtain a second unit vector, and can define the second unit vector as the first... i In the rotation matrix corresponding to the first spherical pose xThe column vectors corresponding to the axes. The preset direction vector can be (0, 0, 1). After obtaining the second unit vector, the electronic device can perform a cross product between the first and second unit vectors to obtain a third unit vector, and can then define the third unit vector as the first... i In the rotation matrix corresponding to the first spherical pose y The column vector corresponding to the axis can be used to obtain the first... i Each first spherical pose corresponds to a rotation matrix. The orientations of each first spherical pose are consistent and can all point to the first 3D model.

[0053] For example, the translation vector is also a 3×1 column vector. Wherein, the... i Translation vector in the first spherical pose That is, the electronic device determines the first... i The coordinates of each point on the sphere After that, the first i The coordinates of each point on the sphere , determined as the i Translation vector in the first spherical pose.

[0054] S104. The electronic device inputs the first image and the second image into the object pose estimation model for processing to obtain the object pose corresponding to the target object.

[0055] In this embodiment, after acquiring a first image captured by a camera and second images generated based on the poses of each first spherical surface, the electronic device can input the first image and each second image into an object pose estimation model. After acquiring the first image and each second image, the object pose estimation model can estimate the object pose corresponding to the target object based on the first image and each second image. It should be understood that this embodiment does not limit the specific method by which the object pose estimation model estimates the object pose corresponding to the target object based on the first image and each second image; the specific method for determining the object pose can be determined by referring to existing object pose estimation methods based on known objects.

[0056] It should be noted that the object pose estimation model can be a trained model.

[0057] For example, when training the object pose estimation model, device 3 can acquire a real image containing the training object (hereinafter referred to as the third image). Alternatively, device 3 can obtain M second spherical poses based on the golden angle spherical sampling method, and render the second 3D model corresponding to the training object based on the M second spherical poses to obtain images of the second 3D model under each second spherical pose (hereinafter referred to as the fourth image). Subsequently, device 3 can use the third and fourth images to train the object pose estimation model to obtain the trained object pose estimation model. It should be understood that device 3 can be the aforementioned electronic device or other devices; this application embodiment does not limit this. M and the aforementioned N can be the same or different; this application embodiment does not limit this. For example, M can be determined based on the actual computing power of device 3.

[0058] In this embodiment of the application, when training the object pose estimation model, the Fibonacci spherical sampling method can be used to efficiently acquire the second spherical pose. This can improve the generation efficiency of the second spherical pose, make the sampled second spherical pose uniform and non-clustered, and ensure that the rendered image based on the second spherical pose can cover all viewpoints of the training object. Thus, the object pose estimation model can be trained using rendered images covering all viewpoints of the training object, which can improve the convergence speed and accuracy of the object pose estimation model. As a result, the trained object pose estimation model can efficiently and accurately determine the object pose corresponding to the target object, thereby improving the user experience.

[0059] In this embodiment, the electronic device can acquire a first image containing a target object and determine a first 3D model corresponding to the target object. Subsequently, the electronic device can render the first 3D model based on a first spherical pose to obtain a second image corresponding to the first 3D model. The first and second images can then be input into an object pose estimation model for processing to obtain the object pose corresponding to the target object. The first spherical pose can be obtained using a spherical sampling method based on the golden angle, enabling efficient acquisition of the spherical pose. This ensures that the sampled spherical pose is uniform and non-clustered, allowing the rendered image based on the spherical pose to cover all viewpoints of the target object. Therefore, the object pose estimation model can efficiently and accurately determine the object pose corresponding to the target object based on the rendered image covering all viewpoints of the target object, thus improving the user experience.

[0060] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0061] Corresponding to the object pose estimation method described in the above embodiments, Figure 2 A structural block diagram of the object pose estimation device provided in the embodiments of this application is shown. For ease of explanation, only the parts related to the embodiments of this application are shown.

[0062] like Figure 2 As shown, the object pose estimation device may include: Image acquisition module 201 is used to acquire a first image containing the target object; The 3D model determination module 202 is used to determine the first 3D model corresponding to the target object; The image rendering module 203 is used to render the first three-dimensional model according to the first spherical pose to obtain a second image corresponding to the first three-dimensional model; the first spherical pose is a pose generated based on spherical points, the spherical points are obtained by sampling on a unit sphere based on the golden angle spherical sampling method, the spherical points include multiple points, when sampling on the unit sphere based on the golden angle spherical sampling method, the preset vertex of the unit sphere is determined as the first spherical point, and the next spherical point is obtained by rotating the golden angle each time; The pose estimation module 204 is used to input the first image and the second image into the object pose estimation model for processing to obtain the object pose corresponding to the target object.

[0063] In some embodiments, the object pose estimation device may further include: The spherical pose generation module is specifically used to determine the number of points corresponding to the spherical surface; and to determine the number of points corresponding to each spherical surface based on the unit sphere and the number of points corresponding to the spherical surface. y Coordinates; for each of the spherical points, based on the unit sphere and the coordinates corresponding to the spherical point. y The coordinates are used to determine the target radius corresponding to the spherical point, and the circumference of the spherical point is determined according to the golden angle. y The target angle of axis rotation is used to determine the corresponding point on the sphere based on the target radius and the target angle. x coordinates and z coordinate.

[0064] For example, the spherical pose generation module is further configured to determine the corresponding point on each spherical surface according to the following formula. y coordinate: y i = R- ; in, y i For the first i The corresponding spherical point ycoordinate, R The radius corresponding to a unit sphere. R =1, N The number of points corresponding to the sphere. i It is a positive integer, and 0 ≤ i ≤ N-1 .

[0065] For example, the spherical pose generation module is further configured to determine the target angle of rotation of the spherical point about the y-axis according to the following formula: ; in, For the first i The target angle corresponding to each spherical point This refers to the golden angle.

[0066] For example, each of the first spherical poses is a 4×4 homogeneous transformation matrix, the homogeneous transformation matrix including a rotation matrix and a translation vector, the rotation matrix including... x The column vectors corresponding to the axes y The column vectors corresponding to the axes z The column vector corresponding to the axis.

[0067] The spherical pose generation module is further configured to, for each spherical point, determine the pose of the spherical point based on the corresponding position of the spherical point. x coordinate, y coordinates and z The coordinates are used to determine the translation vector in the homogeneous transformation matrix; based on the spherical point, a first unit vector pointing to the origin of the unit sphere is determined, and this first unit vector is used as the vector in the rotation matrix. z The column vector corresponding to the axis; the cross product of the preset direction vector and the first unit vector is taken to obtain the second unit vector, and the second unit vector is determined as the column vector in the rotation matrix. x The column vector corresponding to the axis; the cross product of the first unit vector and the second unit vector is taken to obtain the third unit vector, and the third unit vector is determined as the column vector in the rotation matrix. y The column vector corresponding to the axis.

[0068] In some embodiments, the object pose estimation model is trained based on a third image and a fourth image. The third image is an image containing the training object, and the fourth image is an image generated by rendering a second three-dimensional model based on a second spherical pose. The second spherical pose is a pose obtained by a spherical sampling method based on the golden angle, and the second three-dimensional model is a three-dimensional model corresponding to the training object.

[0069] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.

[0070] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0071] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 3 As shown, the electronic device 3 of this embodiment includes: at least one processor 30 ( Figure 3 (Only one is shown in the diagram), memory 31, and computer program 32 stored in the memory 31 and executable on the at least one processor 30, wherein the processor 30 executes the computer program 32 to implement the steps in any of the above-described embodiments of the object pose estimation method.

[0072] The electronic device 3 can be a robot, mobile phone, tablet computer, wearable device, in-vehicle device, augmented reality (AR) / virtual reality (VR) device, laptop computer, UMPC, netbook, PDA, cloud server, or other computing device. The electronic device 3 may include, but is not limited to, a processor 30 and a memory 31. Those skilled in the art will understand that... Figure 3 This is merely an example of electronic device 3 and does not constitute a limitation on electronic device 3. It may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, it may also include input / output devices, network access devices, etc.

[0073] The processor 30 can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.

[0074] In some embodiments, the memory 31 may be an internal storage unit of the electronic device 3, such as a hard disk or memory of the electronic device 3. In other embodiments, the memory 31 may be an external storage device of the electronic device 3, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the electronic device 3. Furthermore, the memory 31 may include both internal and external storage units of the electronic device 3. The memory 31 is used to store the operating system, applications, bootloader, data, and other programs, such as the program code of the computer program. The memory 31 can also be used to temporarily store data that has been output or will be output.

[0075] This application also provides a computer-readable storage medium storing a computer program that, when executed by an electronic device, causes the electronic device to implement the steps in the above-described embodiments of the object pose estimation methods.

[0076] This application provides a computer program product, which includes a computer program. When the computer program is executed by an electronic device, the electronic device performs the steps described in the above-described embodiments of the object pose estimation methods.

[0077] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or some intermediate form. The computer-readable storage medium can include at least: any entity or device capable of carrying computer program code to a device / electronic device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks.

[0078] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0079] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0080] In the embodiments provided in this application, it should be understood that the disclosed devices / electronic devices and methods can be implemented in other ways. For example, the device / electronic device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual couplings or direct couplings or communication connections may be through some interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.

[0081] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0082] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A method for estimating the pose of an object, characterized in that, include: Obtain the first image containing the target object; Determine the first three-dimensional model corresponding to the target object; The first 3D model is rendered based on the first spherical pose to obtain the second image corresponding to the first 3D model; the first spherical pose is a pose generated based on spherical points, which are obtained by sampling on a unit sphere using a spherical sampling method based on the golden angle. The spherical points include multiple points. When sampling on the unit sphere using the spherical sampling method based on the golden angle, a preset vertex of the unit sphere is determined as the first spherical point, and the next spherical point is obtained by rotating the golden angle each time. The first image and the second image are input into the object pose estimation model for processing to obtain the object pose corresponding to the target object.

2. The object pose estimation method according to claim 1, characterized in that, Before rendering the first 3D model based on the first spherical pose, the object pose estimation method further includes: Determine the number of points corresponding to the spherical surface; Based on the unit sphere and the number of corresponding spherical points, determine the corresponding number of each spherical point. y coordinate; For each of the spherical points, based on the unit sphere and the corresponding spherical point... y The coordinates are used to determine the target radius corresponding to the spherical point, and the circumference of the spherical point is determined according to the golden angle. y The target angle of axis rotation is used to determine the corresponding point on the sphere based on the target radius and the target angle. x coordinates and z coordinate.

3. The object pose estimation method according to claim 2, characterized in that, The method involves determining the number of points corresponding to each spherical surface based on the unit sphere and the corresponding number of spherical points. y Coordinates, including: The corresponding point on each sphere is determined according to the following formula. y coordinate: y i = R- ; in, y i For the first i The corresponding spherical point y coordinate, R The radius corresponding to a unit sphere. R =1, N The number of points corresponding to the sphere. i It is a positive integer, and 0 ≤ i ≤ N-1 .

4. The object pose estimation method according to claim 2 or 3, characterized in that, The spherical point is determined according to the golden angle. y The target angle of axis rotation includes: The spherical point is determined according to the following formula. y Target angle of axis rotation: ; in, For the first i The target angle corresponding to each spherical point This refers to the golden angle.

5. The object pose estimation method according to claim 2 or 3, characterized in that, Each of the first spherical poses is a 4×4 homogeneous transformation matrix, which includes a rotation matrix and a translation vector. The rotation matrix includes... x The column vectors corresponding to the axes y The column vectors corresponding to the axes z The column vectors corresponding to the axes; The object pose estimation method further includes: For each of the spherical points, according to the spherical point corresponding to x coordinate, y coordinates and z Coordinates are used to determine the translation vector in the homogeneous transformation matrix; Based on the spherical points, determine a first unit vector pointing to the origin of the unit sphere, and define the first unit vector as the value in the rotation matrix. z The column vectors corresponding to the axes; The preset direction vector is cross-producted with the first unit vector to obtain a second unit vector, which is then defined as the second unit vector in the rotation matrix. x The column vectors corresponding to the axes; Perform a cross product between the first unit vector and the second unit vector to obtain a third unit vector, and define the third unit vector as a component in the rotation matrix. y The column vector corresponding to the axis.

6. The object pose estimation method according to any one of claims 1 to 3, characterized in that, The object pose estimation model is trained based on the third image and the fourth image. The third image is an image containing the training object, and the fourth image is an image generated by rendering the second three-dimensional model based on the second spherical pose. The second spherical pose is the pose obtained by the spherical sampling method based on the golden angle, and the second three-dimensional model is the three-dimensional model corresponding to the training object.

7. An object pose estimation device, characterized in that, include: The image acquisition module is used to acquire a first image containing the target object; A 3D model determination module is used to determine the first 3D model corresponding to the target object; The image rendering module is used to render the first three-dimensional model according to the first spherical pose to obtain a second image corresponding to the first three-dimensional model; the first spherical pose is a pose generated based on spherical points, the spherical points are obtained by sampling on a unit sphere based on the golden angle spherical sampling method, the spherical points include multiple points, when sampling on the unit sphere based on the golden angle spherical sampling method, the preset vertex of the unit sphere is determined as the first spherical point, and the next spherical point is obtained by rotating the golden angle each time; The pose estimation module is used to input the first image and the second image into the object pose estimation model for processing, so as to obtain the object pose corresponding to the target object.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it causes the electronic device to implement the object pose estimation method as described in any one of claims 1 to 6.

9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the electronic device, the electronic device implements the object pose estimation method as described in any one of claims 1 to 6.

10. A computer program product, the computer program product comprising a computer program, characterized in that, When the computer program is executed by the electronic device, the electronic device implements the object pose estimation method as described in any one of claims 1 to 6.