Pose estimation method and apparatus, electronic device, and storage medium

By using pose estimation models trained using three-dimensional images and deterministic annealing algorithm, the objective function is constructed and multiple geometric features are matched, which solves the problem of insufficient pose estimation accuracy in the prior art and achieves higher pose estimation accuracy.

WO2025123367A1PCT designated stage expired Publication Date: 2025-06-19SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI

Patent Information

Application Number
PCT/CN2023/139291
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-15
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

The existing pose estimation method is based on a monocular camera. It is difficult to accurately determine the matching relationship between image data and model through geometric feature matching algorithms, resulting in insufficient accuracy of pose estimation.

Method used

Three-dimensional images are used for pose estimation. By constructing the objective function and calling the pose estimation model trained by deterministic annealing algorithm, pose estimation is used to perform pose estimation using a variety of geometric features.

Benefits of technology

It improves the accuracy of pose estimation, can effectively handle a variety of geometric features, and enhances the information dimension and matching accuracy of pose recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2023139291_19062025_PF_FP_ABST
    Figure CN2023139291_19062025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of image processing, and provides a pose estimation method and apparatus, an electronic device, and a storage medium. The method comprises: acquiring an image to be recognized, said image being a three-dimensional image obtained by photographing a target to be recognized; constructing a target function on the basis of geometric features of said target in the image to be recognized; and calling a pose estimation model, and, on the basis of the target function, performing pose estimation on the image to be recognized, to obtain a pose estimation result corresponding to said target. The present application solves the problem in the related art of low pose estimation accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Position estimation method, device, electronic device and storage medium Technical Field

[0001] The present application relates to the field of image processing technology. Specifically, the present application relates to a posture estimation method, device, electronic device and storage medium. Background Art

[0002] Methods for estimating relative object pose have numerous applications in aerospace, industrial assembly, and other fields. For example, in aerospace, estimating the relative pose between two spacecraft is a key issue in space rendezvous and docking, as well as target tracking. In industry, pose estimation methods are used to estimate the pose of an object relative to a camera to enable automated assembly tasks.

[0003] Existing pose estimation methods rely on capturing a certain type of geometric features from a monocular camera's pose image of the target. These features are then used to estimate the pose of the target by matching the image data with a model. However, algorithms that match only a single type of geometric feature in an image cannot accurately determine the matching relationship between the image data and the model. Furthermore, basic geometric feature matching algorithms based on monocular cameras can only process single straight line segments or circular features, resulting in insufficient robustness and pose estimation accuracy.

[0004] From the above, we can see that how to improve the accuracy of pose estimation still needs to be solved.

[0005] Summary of the Invention

[0006] This application provides a method, device, electronic device, and storage medium for posture estimation, which can solve the problem of low posture estimation accuracy in related technologies. The technical solution is as follows:

[0007] According to one aspect of the present application, a pose estimation method is characterized in that it includes: acquiring an image to be identified; the image to be identified is a three-dimensional image obtained by photographing the target to be identified; constructing an objective function based on the geometric features of the target to be identified in the image to be identified; calling a pose estimation model, performing pose estimation on the image to be identified according to the objective function, and obtaining a pose estimation result corresponding to the target to be identified; the pose estimation model is a machine learning model that has been trained by a deterministic annealing algorithm and has the function of performing pose estimation on the target to be identified.

[0008] According to one aspect of the present application, a posture estimation device is characterized in that it includes: an image acquisition module for acquiring an image to be identified; the image to be identified is a three-dimensional image obtained by shooting the target to be identified; a function construction module for constructing an objective function according to the geometric features of the target to be identified in the image to be identified; a posture estimation module for calling a posture estimation model, performing posture estimation on the image to be identified according to the objective function, and obtaining a posture estimation result corresponding to the target to be identified; the posture estimation model is a machine learning model that has been trained by a deterministic annealing algorithm and has the function of performing posture estimation on the target to be identified.

[0009] In an exemplary embodiment, the function construction module includes: an image reconstruction unit, used to reconstruct the image to be identified to obtain the geometric features of the target to be identified; a sampling unit, used to sample the rotation parameter space of the image to be identified to obtain at least one rotation parameter; and a function construction unit, used to construct an objective function based on each of the rotation parameters and the geometric features.

[0010] In an exemplary embodiment, the image reconstruction unit includes: an image reconstruction subunit, used to perform image reconstruction on at least one geometric dimension of the image to be identified, and obtain a direction vector corresponding to each geometric dimension, each of the geometric dimensions including a straight line dimension and a circle dimension; a feature generation subunit, used to generate geometric features of the target to be identified based on the direction vector corresponding to each of the geometric dimensions.

[0011] In an exemplary embodiment, the pose estimation module includes: a screening unit, used to screen each of the rotation parameters according to the objective function to obtain at least one target rotation parameter that meets a preset threshold; a pose estimation unit, used to input the geometric features of the target to be identified and each of the target rotation parameters into the pose estimation model, perform pose estimation on the image to be identified, and obtain preliminary estimation results corresponding to each of the target rotation parameters; a result evaluation unit, used to perform matching evaluation on each of the preliminary estimation results, and determine one of the preliminary evaluation results as the pose estimation result according to the evaluation result.

[0012] In an exemplary embodiment, the screening unit includes: an output determination subunit, which is used to synchronously compare the objective functions generated by each rotation parameter to determine the output value of the objective function under the same parameter conditions; and a parameter determination subunit, which is used to determine the rotation parameter corresponding to the objective function whose output value is within a preset threshold as the target rotation parameter.

[0013] In an exemplary embodiment, the preliminary estimation result includes a displacement vector and a feature matching relationship corresponding to each target rotation parameter; the result evaluation unit includes: a result generation subunit, which is used to perform matching calculations based on each displacement vector and each target rotation parameter, determine the evaluation score of each feature matching relationship, and generate a pose estimation result based on the feature matching relationship with the smallest evaluation score.

[0014] In an exemplary embodiment, the apparatus further comprises:

[0015] A training image acquisition module, configured to acquire a plurality of training images, wherein the training images include a three-dimensional model of an object to be identified in the training image;

[0016] A training result generation module, configured to input the current training image into the pose estimation model and perform pose estimation using a deterministic annealing algorithm to obtain a pose estimation result;

[0017] An updating module is used to determine whether the pose estimation model meets the training completion condition based on the difference between the pose estimation result and the three-dimensional model in the current training image; if not, the pose estimation model parameters are updated and updated according to the training completion condition until the pose estimation model meets the training completion condition, thereby obtaining the pose estimation model that has completed training.

[0018] According to one aspect of the present application, an electronic device includes at least one processor and at least one memory, wherein the memory stores computer-readable instructions; the computer-readable instructions are executed by one or more of the processors, so that the electronic device implements the pose estimation method as described above.

[0019] According to one aspect of the present application, a storage medium stores computer-readable instructions thereon, and the computer-readable instructions are executed by one or more processors to implement the pose estimation method as described above.

[0020] The beneficial effects of the technical solution provided by this application are:

[0021] In the above technical solution, a three-dimensional image to be identified is obtained; then a target function is constructed according to the geometric features of the target to be identified in the image to be identified; a pose estimation model trained by an annealing algorithm is used to perform pose estimation according to multiple geometric features in the image to be identified by the target function, so as to obtain a more accurate pose estimation result, thereby effectively solving the problem of low pose estimation accuracy existing in related technologies. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for describing the embodiments of the present application. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without inventive efforts.

[0023] FIG1 is a schematic diagram of an implementation environment according to the present application;

[0024] FIG2 is a flow chart of a method for posture estimation according to an exemplary embodiment;

[0025] FIG3 is a flow chart of step 230 in one embodiment of the embodiment corresponding to FIG2 ;

[0026] FIG4 is a flow chart of step 231 in one embodiment of the embodiment corresponding to FIG3 ;

[0027] FIG4 a is a schematic diagram showing a geometric feature according to an exemplary embodiment;

[0028] FIG5 is a flow chart of step 250 in one embodiment of the embodiment corresponding to FIG2 ;

[0029] FIG6 is a flow chart of step 251 in one embodiment of the embodiment corresponding to FIG5 ;

[0030] FIG7 is a flow chart of step 255 in one embodiment of the embodiment corresponding to FIG5 ;

[0031] FIG8 is a flowchart showing training of a pose estimation model according to an exemplary embodiment;

[0032] FIG9 a is a schematic diagram showing a specific implementation of a pose estimation method in an application scenario;

[0033] FIG9 b is a schematic diagram of a specific implementation of a pose estimation method in an application scenario;

[0034] FIG10 is a structural block diagram of a posture estimation apparatus according to an exemplary embodiment;

[0035] Fig. 11 is a structural block diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION

[0036] The following describes in detail embodiments of the present application, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and are not to be construed as limiting the present application.

[0037] It will be understood by those skilled in the art that, unless expressly stated otherwise, the singular forms "a", "an", "said" and "the" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present disclosure refers to the presence of the features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or there may be intermediate elements. In addition, "connected" or "coupled" as used herein may include wireless connections or wireless couplings. The term "and / or" used herein includes all or any units and all combinations of one or more associated listed items.

[0038] As previously mentioned, existing pose estimation methods rely on capturing a certain type of geometric features from a monocular camera's pose image of the target. These features are then used to estimate the pose of the target by matching the image data with a model. However, algorithms matching only a single type of geometric feature in an image cannot effectively determine the matching relationship between the image data and the model. Furthermore, basic geometric feature matching algorithms based on monocular cameras can only process single straight line segments or circular features, resulting in insufficient robustness and pose estimation accuracy.

[0039] From the above, it can be seen that the relevant technology still has the defect of low accuracy of pose estimation.

[0040] To this end, the pose estimation method provided in the present application can effectively improve the accuracy of pose estimation. Accordingly, the pose estimation method is applicable to a pose estimation device, which can be deployed in an electronic device. The electronic device can be a computer device configured with a von Neumann architecture, for example, the computer device includes a desktop computer, a laptop computer, a server, etc.; the electronic device can also be an electronic device with a central control function, for example, the electronic device includes a gateway, etc.; the electronic device can also refer to a portable and mobile electronic device, for example, the electronic device includes a smart phone, a tablet computer, etc.

[0041] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0042] Figure 1 is a schematic diagram of an implementation environment involved in a pose estimation method. It should be noted that this implementation environment is only an example adapted to the present invention and should not be considered as providing any limitation on the scope of application of the present invention.

[0043] The implementation environment includes a collection end 110 and a service end 130 .

[0044] Specifically, the acquisition end 110 can also be considered as an image acquisition device, including but not limited to electronic devices with shooting functions such as cameras, cameras, and camcorders. For example, the acquisition end 110 is a stereo camera.

[0045] Server 130 can be an electronic device such as a desktop computer, laptop computer, or server, or a computer cluster consisting of multiple servers, or even a cloud computing center consisting of multiple servers. Server 130 is used to provide background services, such as, but not limited to, pose estimation services.

[0046] The server 130 and the acquisition terminal 110 establish a network communication connection in advance through wired or wireless means, and data transmission between the server 130 and the acquisition terminal 110 is achieved through the network communication connection. The transmitted data includes but is not limited to: the image of the object to be identified, etc.

[0047] In one application scenario, through the interaction between the acquisition terminal 110 and the server 130, the acquisition terminal 110 shoots and acquires the image to be identified for the object to be identified, and uploads the image to be identified to the server 130 to request the server 130 to provide a pose estimation service.

[0048] For the server 130, after receiving the image of the object to be identified uploaded by the acquisition end 110, it calls the pose estimation service to perform pose estimation on the image of the object to be identified and obtain the pose estimation result of the object to be identified, so as to solve the problem of poor effect existing in the related technology.

[0049] Please refer to FIG. 2 . An embodiment of the present application provides a posture estimation method. The method is applicable to an electronic device, which may be the server 130 in the implementation environment shown in FIG. 1 .

[0050] In the following method embodiments, for ease of description, the execution subject of each step of the method is taken as an electronic device as an example for illustration, but this does not constitute a specific limitation.

[0051] As shown in FIG2 , the method may include the following steps:

[0052] Step 210: Obtain an image to be recognized.

[0053] The image to be identified is a three-dimensional image obtained by photographing the target to be identified.

[0054] In a possible implementation, a stereo camera is used to photograph the target to be identified at different positions to generate left and right images of the target to be identified as the image to be identified.

[0055] Step 230: constructing an objective function based on the geometric features of the target to be identified in the image to be identified.

[0056] The objective function is used to specifically describe the state of the target during pose estimation using its geometric features within the image. As will be appreciated, the geometric features in each image are different, and therefore the generated objective function is also different. This means that the objective function uniquely corresponds to the image being identified.

[0057] In a possible implementation, the spatial straight line segments and circles in the image to be identified are reconstructed to generate the direction vector of the spatial straight line, the direction vector of the spatial circle rotation axis, and the rotation parameters to generate the corresponding objective function.

[0058] Step 250 , calling the pose estimation model, performing pose estimation on the image to be identified according to the objective function, and obtaining a pose estimation result corresponding to the target to be identified.

[0059] Among them, the pose estimation model is a machine learning model that has been trained with a deterministic annealing algorithm and has the ability to estimate the pose of the target to be identified.

[0060] It should be noted that before using image features to generate pose estimation results, it is necessary to match the data obtained in the image to be identified with the model of the target to be identified. In this application, the pose estimation model is matched using a deterministic annealing algorithm to determine the matching relationship between the data in the image to be identified and the model of the target to be identified, and to find the pose estimation result that is closest to the true value of the image to be identified, such as the matching information of the six-degree-of-freedom pose information of the target to be identified and the model of the target to be identified.

[0061] Through the above process, the objective function is generated by using multiple geometric features of the image to be identified, and pose estimation is performed according to the pose estimation model based on the deterministic annealing algorithm. Multiple features in the image to be identified can be processed simultaneously, and there is no limit on the number of features, which improves the information dimension that can be obtained by pose estimation and improves the accuracy of pose estimation. At the same time, matching processing is performed through the deterministic annealing algorithm, which improves the matching accuracy between the geometric features in the three-dimensional image and the target to be identified, and improves the accuracy of pose estimation.

[0062] In an exemplary embodiment, as shown in FIG3 , step 230 may include the following steps:

[0063] Step 231 : reconstruct the image to be identified to obtain the geometric features of the target to be identified.

[0064] In an exemplary embodiment, as shown in FIG4 , step 231 may include the following steps:

[0065] Step 2311: Reconstruct at least one geometric dimension of the image to be identified to obtain a direction vector corresponding to each geometric dimension.

[0066] Among them, the geometric dimension includes the straight line segment dimension and the circle dimension.

[0067] Step 2313: Generate geometric features of the target to be identified based on the direction vectors corresponding to each geometric dimension.

[0068] By reconstructing the image to be identified in the line segment and circle dimensions and extracting features, we can obtain reconstructed line and reconstructed spatial circle data in the image to be identified. For example, as shown in Figure 4a, the figure shows the geometric features generated based on the line segment and circle dimensions. In the upper figure, the dots represent sampling points on the reconstructed line, and the solid line segments represent the sampled model line segments. In the lower figure, the dots represent sampling points on the reconstructed spatial circle and its rotation axis, and the solid circle represents the sampled model circle.

[0069] Step 233: Sampling the rotation parameter space of the image to be identified to obtain at least one rotation parameter.

[0070] In a possible implementation, the rotation parameter is obtained by uniformly sampling the rotation parameter space of the image to be identified using the Thomson method.

[0071] In one possible implementation, the rotation parameters are obtained in the form of a rotation matrix, and the K nearest neighbor method (K-NN) is used to estimate the K R A rotation matrix R.

[0072] Among them, the metric function formula used by the K-nearest neighbor method is:

[0073] in,

[0074] Among them, N L and M L Respectively represent the number of fitted lines on the image to be identified and the number of straight lines on the target model to be identified. C and M C They represent the number of fitted ellipses on the image to be identified and the number of circles on the target model to be identified. Li and m Lk They represent the unit direction vector of the i-th reconstructed straight line segment and the unit direction vector of the k-th model straight line, respectively. Ci and m Ck Represents the unit direction vector of the rotation axis of the i-th reconstructed space circle and the The unit direction vector of the rotation axis of the model circle. β1 and α1 represent a set constant.

[0075] It should be noted that The physical meaning of same. Function and The function is a function that indicates the degree of matching between the rotation matrix obtained from the image to be recognized and the true value. Medium, larger and This will make the exponential function exp() very small. The smallest K R R is the rotation matrix estimate closest to the true value. For example, when the rotation matrix R is closer to the true value, n Li and Rm Lk Under the correct matching relationship, it should be close to the collinear state. Approximately close to 0. However, R with incorrect matching or far from the true value will Larger. α1 is a set value, representing The value if the rebuilt feature has no matching model feature.

[0076] Step 235: construct an objective function based on the rotation parameters and geometric features.

[0077] In one possible implementation, the three-dimensional image is a left and right image, and the objective function formula constructed based on the left and right images is as follows:

[0078] in,

[0079] in, Indicates the matching relationship between the reconstructed straight line in the image to be identified and the straight line of the target model to be identified. For example, if Indicates that the i-th image reconstruction line to be identified matches the k-th model line if Indicates that the i-th image reconstructed line to be identified does not match the k-th model line.

[0080] Represents the matching relationship between the reconstructed circle of the image to be identified and the circle of the target model to be identified. For example, if Indicates that the i-th image reconstructed circle to be identified matches the k-th model circle if Indicates that the reconstructed circle of the i-th image to be identified does not match the k-th model circle.

[0081] also, K l and K r They represent the internal parameter matrix of the camera shooting the recognition target, L liIndicates the i-th fitting line in the left image, L ri represents the i-th fitted line in the right image. Represent the planes formed by the straight line segments in the left and right images and the optical center of the camera respectively.

[0082] n sl1 represents the total number of sampling points on the kth model line, P kj Indicates the jth sampling point on the kth model line. C i and C k They represent the center coordinates of the reconstructed circle of the i-th image to be identified (relative to the left camera) and the center coordinates of the k-th model circle (relative to the model coordinate system).

[0083] in, m ck represents the unit direction vector of the rotation axis of the k-th model circle, and Q k is the projection matrix. For example, for any vector v, Q k is the projection of vector v on the rotation axis of the kth model circle. sc1 represents the number of sampling points on the rotation axis of the i-th image reconstruction circle to be identified, Represents the jth sampling point on the rotation axis of the i-th image reconstruction circle to be identified. sc2 represents the number of sampling points on the circumference of the reconstructed circle of the i-th image to be identified, Represents the jth sampling point on the circumference of the i-th image to be identified and reconstructed. i represents the radius of the reconstructed circle of the i-th image to be identified, r k represents the radius of the kth model circle.

[0084] In addition, the objective function The formula is as follows:

[0085] in, The function indicates whether the rotation matrix R and the displacement vector t are close to the true value and whether the matching relationship is correct. It is: the distance from the sampling point on the model line to the spatial plane formed by the camera optical center and the image fitting line. When the rotation matrix R and the displacement vector t are close to the true value and the matching relationship is correct, the model line will be very close to the spatial plane formed by the camera optical center and the image fitting line. Will be very close to 0. Under the wrong rotation matrix R and displacement vector t or the wrong matching relationship, It will be very big.

[0086] In addition, the objective function The formula is as follows:

[0087] in, The function consists of the following four parts:

[0088] (1) The distance from the center of the reconstructed circle of the image to be identified to the center of the model circle.

[0089] (2) The distance between the sampling point on the rotation axis of the space circle of the image to be identified and the rotation axis of the model circle.

[0090] (3) The distance from the sampling point on the circumference of the image reconstruction space circle to the support plane of the model circle.

[0091] (4) The difference between the radius of the reconstructed space circle of the image to be identified and the radius of the model circle.

[0092] pass The value of the function indicates whether the rotation matrix R and the displacement vector t are close to the true value. For example, when the rotation matrix R and the displacement vector t are close to the true value and the matching relationship is correct, the physical quantities represented by (1), (2), (3) and (4) will be very close to 0. The function is also close to 0. However, under the wrong rotation matrix R and displacement vector t or the wrong matching relationship, at least one of the physical quantities represented by (1), (2), (3) and (4) will be very large. The function is also very large.

[0093] It should be noted that and In the example, the upper limit of k is M L +1 and M C +1. When k=M L +1(or M C +1) indicates that the linear feature (or circular feature) reconstructed by the i-th image to be identified has no matching model feature. In this case, the value range of k is set between correct matching and incorrect matching. (or )between.

[0094] It should be noted that in the process of constructing the objective function, the objective function can be simplified, and the simplification process is as follows:

[0095] Again Find the derivative with respect to t:

[0096] in

[0097] At this time, The linear equation for the displacement vector t is as follows:

[0098] remember:

[0099] Then the linear equation about the displacement vector t can be expressed as: t=B -1 A..

[0100] Through the above process, the objective function corresponding to the image to be identified is constructed, and various basic geometric features of the image to be identified can be processed, thereby improving the accuracy and robustness of the objective function generation, ensuring the stability of the posture recognition process, and improving the accuracy of posture recognition.

[0101] In an exemplary embodiment, as shown in FIG5 , step 250 may include the following steps:

[0102] Step 251 : Screen the rotation parameters according to the target function to obtain at least one target rotation parameter that meets a preset threshold.

[0103] In an exemplary embodiment, as shown in FIG6 , step 251 may include the following steps:

[0104] Step 2511 , synchronously compare the objective functions generated by the rotation parameters to determine the output values ​​of the objective functions under the same parameter conditions.

[0105] Step 2513: Determine the rotation parameter corresponding to the objective function whose output value is within a preset threshold as the target rotation parameter.

[0106] Among them, by inputting all rotation parameters into the target function, the function output values ​​are arranged in order of size, and the rotation parameter with the smallest function output value is determined as the target rotation parameter.

[0107] In one possible implementation, the k-nearest neighbor method is used to obtain the rotation parameters, so it is necessary to obtain K R target rotation parameters.

[0108] The rotation parameters are collected in the form of a rotation matrix R, where the Rodriguez formula of R is expressed as follows: n=[n1 n2 n3] T =[cosαsinβ sinαsinβ cosβ] T ,,

[0109] Among them, 0≤α<2π, 0≤β<π, 0≤θ<π, α, β and θ are the parameters of the image to be identified obtained by the Thomson sampling method, and finally N can be generated by collecting the image to be identified.R A rotation matrix R, which is N R Substitute R into the metric function formula to obtain N R indivual The smallest K R indivual The corresponding R is the target rotation parameter.

[0110] In step 253 , the geometric features of the target to be identified and the rotation parameters of each target are input into a pose estimation model, and pose estimation is performed on the image to be identified to obtain preliminary estimation results of the rotation parameters of each target.

[0111] The preliminary estimation results include the displacement vectors and feature matching relationships corresponding to the rotation parameters of each target. R The rotation parameters are input into the pose estimation model to generate K R unique vectors and K R Group feature matching relationship.

[0112] Step 255 , performing a matching evaluation on each preliminary estimation result, and determining one of the preliminary estimation results as the pose estimation result according to the evaluation result.

[0113] In an exemplary embodiment, as shown in FIG7 , step 255 may include the following steps:

[0114] Step 2551: Perform matching calculation based on each displacement vector and each target rotation parameter to determine the evaluation score of each feature matching relationship.

[0115] Among them, the evaluation score is a parameter to measure the matching of rotation parameters and displacement vectors. The smaller the evaluation score, the more the rotation parameters and displacement vectors generated by the pose estimation model match the true values, and the better the matching of the feature matching relationship.

[0116] Step 2553: Generate a pose estimation result based on the feature matching relationship with the smallest evaluation score.

[0117] In one possible implementation, the feature matching relationship is screened through an objective function to obtain the optimal feature matching relationship and posture estimation result.

[0118] Among them, K is obtained by the K nearest neighbor algorithm R Rotation matrix R, and calculate K respectively R A displacement vector t.

[0119] The rotation matrix R and the displacement vector t are input into the objective function for calculation. At this time, in the objective function, the loss function formula of the straight line is as follows:

[0120] in

[0121] P ij represents the jth sampling point on the i-th reconstructed spatial line, P k1 and P k2 Respectively represent the two endpoints on the k-th model line. median() represents the median. α4 represents the set threshold. n sl2 Represents the number of points sampled on the straight line segment in the reconstructed space, P ij Represents the jth sampling point on the i-th reconstructed space line segment.

[0122] The loss function formula for the circle is as follows:

[0123] The metric function of the deterministic annealing algorithm for matching relationships is as follows:

[0124] Where z = Z (Z represents the number of temperatures set in the deterministic annealing algorithm), (R j , t j ) represents the K obtained by the K nearest neighbor algorithm R The jth group of postures in the group of postures.

[0125] By obtaining K R indivual The feature matching relationship corresponding to the minimum value in and the posture (R, t) are used to generate the pose estimation result.

[0126] Under the effect of the above embodiment, the feature matching relationship is screened through the objective function to obtain the optimal feature matching relationship and posture estimation result, thereby improving the accuracy of posture recognition.

[0127] In an exemplary embodiment, as shown in FIG8 , the pose estimation model is a machine learning model trained with a deterministic annealing algorithm and capable of performing pose estimation on the target to be identified. The training process includes:

[0128] Step 810: Acquire multiple training images and initialize the parameters of the pose estimation model.

[0129] Step 830: Input the current training image into the pose estimation model, and perform pose estimation using a deterministic annealing algorithm to obtain a pose estimation result.

[0130] In one possible implementation, the process of the deterministic annealing algorithm is as follows: set a series of temperatures β z , z=1,2,....,Z; According to the linear equation of the generated displacement vector t, calculate the initial value of t (denoted as ),in By traversing all the temperatures β z Get the displacement vector t.

[0131] Specifically, use initialization According to the linear equation of displacement vector t and The formula is calculated to get and Update according to the simplified linear equation about the displacement vector t Repeat this process until Converges, and the t obtained at this time is the final output displacement vector t.

[0132] Step 850, based on the difference between the pose estimation result and the current training image, determine whether the pose estimation model meets the training completion conditions; if not, update the pose estimation model parameters until the pose estimation model meets the training completion conditions, and obtain a pose estimation model that has completed training.

[0133] By updating the parameters of the pose estimation model, the pose estimation model can be trained in the direction of improving the pose estimation capability until the pose estimation capability meets the user requirements.

[0134] Under the effect of the above embodiment, a pose estimation model is generated that can perform pose estimation according to the deterministic annealing algorithm, thereby improving the accuracy of pose estimation.

[0135] Figure 9a is a schematic diagram of a specific implementation of a pose estimation method in an application scenario. In this application scenario, the target to be identified is a spacecraft. As shown in Figure 9a, a three-dimensional image of the spacecraft is acquired using multiple cameras. The circled line in the figure represents the projection of the spacecraft CAD model onto the three-dimensional image and the pose estimation result generated by the pose estimation model, calculated using the Levenberg-Marquardt optimization algorithm.

[0136] Firstly, the Thomson method is used to uniformly sample the rotation parameter space to obtain the rotation parameters, and the geometric features of the target surface are reconstructed through the left and right images of the stereo camera. The objective function is constructed using the direction vector of the spatial straight line, the direction vector of the spatial circle rotation axis and the rotation parameters.

[0137] From all collected rotation parameters, find K that minimizes the objective function R The target rotation parameters and the reconstructed basic geometric features are input into the deterministic annealing algorithm, and the K corresponding to the rotation parameters is generated under the constraint of the objective function. R displacement vectors, and K R Group feature matching relationship.

[0138] The target rotation parameters and K R The displacement vector is used as the input of the objective function, and the target rotation parameter and displacement vector that minimize the objective function are found as the object's pose estimation. The feature matching relationship corresponding to the target rotation parameter and the displacement vector is used as the final feature matching relationship to generate the six-degree-of-freedom pose of the target relative to the stereo camera.

[0139] The pose estimation results generated by the pose estimation model are shown in Figures 9a and 9b.

[0140] The circled line in FIG9a is the projection in the three-dimensional image generated by the CAD model according to the pose estimation result generated by the pose estimation model.

[0141] In Figure 9b, the dots in the upper figure represent sampling points on the reconstructed straight line, and the solid line segments represent model straight line segments; in the lower figure, the dots represent sampling points on the reconstructed spatial circle and its rotation axis, and the solid circle represents the model circle; the experimental noise is set to 2, 8 lines, and 8 circles; the matching results of the data obtained in the image to be identified and the target model to be identified are completely correct; the displacement vector error in the pose estimation result is 0.6083%, and the rotation matrix error is 0.0505 degrees.

[0142] The results show that the matching results and pose estimation of the present invention are quite accurate and can provide prior knowledge for downstream tasks.

[0143] The following are embodiments of the apparatus of the present application, which can be used to perform the pose estimation method involved in the present application. For details not disclosed in the apparatus embodiments of the present application, please refer to the method embodiments of the pose estimation method involved in the present application.

[0144] Please refer to Figure 10. An embodiment of the present application provides a posture estimation device 900, including but not limited to: an image acquisition module 910, a function construction module 930, and a posture estimation module 950.

[0145] The image acquisition module 910 is used to acquire an image to be identified, which is a three-dimensional image obtained by photographing the target to be identified.

[0146] The function construction module 930 is used to construct an objective function according to the geometric features of the target to be identified in the image to be identified.

[0147] The pose estimation module 950 is used to call a pose estimation model to perform pose estimation on the image to be identified based on the objective function, and obtain a pose estimation result corresponding to the target to be identified. The pose estimation model is a machine learning model that uses a deterministic annealing algorithm and has the ability to estimate the pose of the target to be identified.

[0148] It should be noted that, when the posture estimation device provided in the above embodiment performs posture estimation, the division of the above-mentioned functional modules is only used as an example. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the posture estimation device will be divided into different functional modules to complete all or part of the functions described above.

[0149] In addition, the posture estimation device and posture estimation method provided in the above embodiments belong to the same concept, and the specific manner in which each module performs operations has been described in detail in the method embodiments and will not be repeated here.

[0150] Please refer to FIG. 11 . An electronic device 4000 is provided in an embodiment of the present application. The electronic device 4000 may include a desktop computer, a laptop computer, a server, etc.

[0151] In FIG. 11 , the electronic device 4000 includes at least one processor 4001 and at least one memory 4003 .

[0152] Data exchange between processor 4001 and memory 4003 can be achieved via at least one communication bus 4002. Communication bus 4002 may include a path for transmitting data between processor 4001 and memory 4003. Communication bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, for example. Communication bus 4002 may be divided into an address bus, a data bus, a control bus, and the like. For ease of illustration, FIG11 shows only one thick line, but this does not indicate that there is only one bus or only one type of bus.

[0153] Optionally, the electronic device 4000 may further include a transceiver 4004, which may be used for data exchange between the electronic device and other electronic devices, such as data transmission and / or data reception. It should be noted that in actual applications, the number of transceivers 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of the present application.

[0154] Processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 4001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and the like.

[0155] The memory 4003 can be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program instructions or codes in the form of instructions or data structures and can be accessed by the electronic device 400, but is not limited to these.

[0156] Computer-readable instructions are stored in the memory 4003 , and the processor 4001 can read the computer-readable instructions stored in the memory 4003 through the communication bus 4002 .

[0157] The computer-readable instructions are executed by one or more processors 4001 to implement the pose estimation method in the above-mentioned embodiments.

[0158] In addition, an embodiment of the present application provides a storage medium on which computer-readable instructions are stored. The computer-readable instructions are executed by one or more processors to implement the above-mentioned posture estimation method.

[0159] In an embodiment of the present application, a computer program product is provided, which includes computer-readable instructions. The computer-readable instructions are stored in a storage medium. One or more processors of an electronic device read the computer-readable instructions from the storage medium, load and execute the computer-readable instructions, so that the electronic device implements the above-mentioned posture estimation method.

[0160] Compared with related technologies, this method generates an objective function based on multiple geometric features of the image to be identified and performs pose estimation based on a pose estimation model based on a deterministic annealing algorithm. This allows for simultaneous processing of multiple features in the image to be identified, with no limit on the number of features. This increases the dimensionality of information available for pose estimation and improves pose estimation accuracy. Simultaneously, matching processing using a deterministic annealing algorithm improves the matching accuracy between geometric features in the 3D image and the object to be identified, thereby improving pose estimation accuracy. Constructing an objective function corresponding to the image to be identified processes multiple basic geometric features of the image to be identified, improving the accuracy and robustness of the generated objective function, ensuring the stability of the pose recognition process and enhancing pose recognition accuracy. The objective function filters feature matching relationships to obtain the optimal feature matching relationship and pose estimation result, thereby improving pose recognition accuracy. A pose estimation model is generated that can perform pose estimation based on a deterministic annealing algorithm, thereby improving pose estimation accuracy.

[0161] It should be understood that although the steps in the flowcharts of the accompanying drawings are shown in sequence as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the flowcharts of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.

[0162] The above are only some of the implementation methods of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A pose estimation method, characterized in that, Including: Obtain an image to be recognized; The image to be recognized is a three-dimensional image obtained by photographing a target to be recognized; Construct an objective function according to the geometric features of the target to be recognized in the image to be recognized; Call a pose estimation model, and perform pose estimation on the image to be recognized according to the objective function to obtain a pose estimation result corresponding to the target to be recognized; The pose estimation model is a machine learning model trained by a deterministic annealing algorithm and capable of performing pose estimation on the target to be recognized.

2. The method according to claim 1, characterized in that, The constructing an objective function according to the geometric features of the target to be recognized in the image to be recognized includes: Perform image reconstruction on the image to be recognized to obtain the geometric features of the target to be recognized; Sample the rotation parameter space of the image to be recognized to obtain at least one rotation parameter; Construct an objective function according to each rotation parameter and the geometric features.

3. The method according to claim 2, characterized in that, The performing image reconstruction on the image to be recognized to obtain the geometric features of the target to be recognized includes: Perform image reconstruction on at least one geometric dimension of the image to be recognized to obtain a direction vector corresponding to each geometric dimension, and each geometric dimension includes a straight-line segment dimension and a circle dimension; Generate the geometric features of the target to be recognized according to the direction vectors corresponding to each geometric dimension.

4. The method according to claim 2, characterized in that, The calling a pose estimation model, and performing pose estimation on the image to be recognized according to the objective function to obtain a pose estimation result corresponding to the target to be recognized includes: Screen each rotation parameter according to the objective function to obtain at least one target rotation parameter that meets a preset threshold; Input the geometric features of the target to be recognized and each target rotation parameter into the pose estimation model, and perform pose estimation on the image to be recognized to obtain a preliminary estimation result corresponding to each target rotation parameter; Perform a matching evaluation on each preliminary estimation result, and determine one of the preliminary estimation results as the pose estimation result according to the evaluation result.

5. The method according to claim 4, characterized in that, The screening each rotation parameter according to the objective function to obtain at least one target rotation parameter that meets a preset threshold includes: Synchronously compare the objective functions generated by each rotation parameter to determine the output value of the objective function under the condition of the same parameter; Determine the rotation parameter corresponding to the objective function whose output value is within the preset threshold as the target rotation parameter.

6. The method according to claim 4, characterized in that, The preliminary estimation result includes a displacement vector and a feature matching relationship corresponding to each target rotation parameter; The performing a matching evaluation on each preliminary estimation result, and determining one of the preliminary estimation results as the pose estimation result according to the evaluation result includes: Perform a matching calculation according to each displacement vector and each target rotation parameter to determine the evaluation score of each feature matching relationship; Generate a pose estimation result according to the feature matching relationship with the smallest evaluation score.

7. The method according to claim 1, characterized in that, The training process of the pose estimation model includes: Obtain a plurality of training images, and the training images include a three-dimensional model of the target to be recognized in the training images; Input the current training image into the pose estimation model, and perform pose estimation through a deterministic annealing algorithm to obtain a pose estimation result; Determine whether the pose estimation model meets the training completion condition according to the difference between the pose estimation result and the 3D model in the current training image; if not, update the pose estimation model parameters and so on until the pose estimation model meets the training completion condition, and obtain the pose estimation model that has completed training.

8. A pose estimation device, characterized in that, Including: An image acquisition module for acquiring an image to be recognized; the image to be recognized is a 3D image obtained by photographing a target to be recognized. A function construction module for constructing an objective function according to the geometric features of the target to be recognized in the image to be recognized. A pose estimation module for calling a pose estimation model and performing pose estimation on the image to be recognized according to the objective function to obtain a pose estimation result corresponding to the target to be recognized. The pose estimation model is a machine learning model trained by a deterministic algorithm and capable of performing pose estimation on the target to be recognized.

9. An electronic device, characterized in that, Including: At least one processor and at least one memory, wherein, The memory stores computer-readable instructions; The computer-readable instructions are executed by one or more of the processors, so that the electronic device implements the pose estimation method according to any one of claims 1 to 7.

10. A storage medium having computer-readable instructions stored thereon, characterized in that, The computer-readable instructions are executed by one or more processors to implement the pose estimation method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Image scene depth estimation method and device thereof, terminal equipment and storage medium

    CN113160294A

  • Pose estimation method and device, terminal equipment and storage medium

    CN114820779A

  • A method for estimating the pose of a camera in the frame of reference of a three-dimensional scene, device, augmented reality system and computer program therefor

    US20210174539A1

  • Method for estimating pose, associated device, system and computer program

    WO2018185104A1

Cited By

  • Sequence image relative pose estimation method, device, equipment and medium

    CN121437621A