Position estimation method, device, terminal equipment and storage medium
By identifying the model to identify feature points in the target image and matching them with preset spatial points, the problem of low pose estimation accuracy in the prior art is solved, and higher feature point recognition accuracy and pose estimation accuracy are achieved.
Patent Information
- Application Number
- CN202210270950.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-18
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2042-03-18
AI Technical Summary
The prior art has low accuracy when estimating the position of an object, especially when the non-geometric corner points of the object in the image are collected as feature points, the positioning is not accurate enough, resulting in inaccurate prediction of the position.
By acquiring the target image, using the recognition model to identify the target image, obtain multiple feature points of the target object in the target image, and match these feature points with the preset spatial points of the target object to calculate the position of the target object.
The accuracy of feature point recognition of target objects is improved, the accuracy of pose estimation is enhanced, and the limitations of geometric corner point recognition of objects are avoided.
Smart Images

Figure CN114820779B_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the technical field of posture estimation, and in particular, relates to a posture estimation method, apparatus, terminal device and storage medium. Background Art
[0002] In the prior art, deep learning is used to estimate the pose of an object. The principle is to use the powerful feature extraction capability of a neural network to extract the features of an object in an image, and then use the features to estimate the pose of the object. That is, by fitting and training a large amount of data, the entire neural network can better predict the pose of an object in an image in one embodiment.
[0003] However, when the above neural network directly predicts the pose of an object, it uses the geometric corner points of the object in the captured image as feature points for subsequent processing. For example, the intersection or vertex of the object's lines. Therefore, the existing direct use of neural networks to predict the pose of an object has strong limitations. If the non-geometric corner points of the object in the captured image are used as feature points for subsequent processing, the positioning of the non-geometric corner points in the image is not accurate enough. When the inaccurately positioned feature points are used for subsequent processing, the predicted pose is also relatively inaccurate. Summary of the invention
[0004] The embodiments of the present application provide a posture estimation method, apparatus, terminal device and storage medium, which can solve the problem of low accuracy when predicting the posture of a target object in an image.
[0005] In a first aspect, an embodiment of the present application provides a posture estimation method, the method comprising:
[0006] Acquire a target image, where the target image is obtained by photographing the target object;
[0007] The target image is recognized by using the recognition model to obtain multiple feature points of the target object in the target image;
[0008] The position and posture of the target object are determined based on multiple feature points and preset spatial points of the target object.
[0009] In a second aspect, an embodiment of the present application provides a posture estimation device, the device comprising:
[0010] An acquisition module is used to acquire a target image, where the target image is obtained by photographing the target object;
[0011] A recognition module is used to recognize a target image using a recognition model to obtain a plurality of feature points of a target object in the target image;
[0012] The determination module is used to determine the position and posture of the target object based on multiple feature points and preset spatial points of the target object.
[0013] In a third aspect, an embodiment of the present application provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method of the first aspect described above when executing the computer program.
[0014] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the method of the first aspect described above is implemented.
[0015] In a fifth aspect, an embodiment of the present application provides a computer program product, which, when executed on a terminal device, enables the terminal device to execute the method of the first aspect.
[0016] Compared with the prior art, the embodiments of the present application have the following beneficial effects: the target image is directly identified by a recognition model to obtain a plurality of feature points of the target object in the target image. Then, the feature points are matched with the preset spatial points of the target object respectively to calculate the position and posture of the target object. At this time, the recognition model only outputs the feature points of the target object in the target image, and does not directly output the position and posture of the target object in the target image. Therefore, when the recognition model identifies and locates the non-geometric corner points of the target object, it can also improve the accuracy of identifying the feature points of the target object. In this way, when determining the position and posture of the target object, it is not necessary for the recognition model to be limited to identifying the geometric corner points of the object. In addition, the estimation accuracy of the position and posture of the target object can also be improved based on the feature points with higher recognition accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0018] Figure 1 is a flowchart of an implementation method of a posture estimation method provided in an embodiment of the present application;
[0019] Figure 2 This is a schematic diagram of an implementation method for determining an actual offset in a posture estimation method provided in an embodiment of the present application;
[0020] Figure 3 This is a schematic diagram of the network structure of the recognition model in one embodiment of the present application;
[0021] Figure 4 It is a schematic diagram of an implementation method for determining the pose of a target object in a pose estimation method provided in an embodiment of the present application;
[0022] Figure 5 This is a schematic diagram of an application scenario of a recognition model for identifying feature points in an embodiment of the present application;
[0023] Figure 6 This is a schematic diagram of an application scenario for determining a posture in an embodiment of the present application;
[0024] Figure 7 is a structural schematic diagram of a posture estimation device provided by an embodiment of the present application;
[0025] Figure 8 It is a structural diagram of a terminal device provided in one embodiment of the present application. DETAILED DESCRIPTION
[0026] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.
[0027] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or combinations thereof.
[0028] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.
[0029] The posture estimation method provided in the embodiment of the present application can be applied to terminal devices such as tablet computers, laptop computers, ultra-mobile personal computers (UMPCs), netbooks, etc. The embodiment of the present application does not impose any restrictions on the specific type of terminal devices.
[0030] As described in the background technology, in the prior art, when estimating the pose of an object, deep learning is used to achieve this. Specifically, the powerful feature extraction capability of a neural network is used to extract the features of an object in an image, and then the features are used to directly estimate the pose of the object.
[0031] However, when using this method to estimate the pose of an object, there are high requirements for the shape of the object. Specifically, the shape of an object usually needs to contain more geometric corner points. The geometric corner points are usually easily identified by a neural network as feature points, and then the position of the feature points in the image is accurately determined to perform pose estimation. Based on this, in order to enable the terminal device to estimate the pose of an object without being constrained by the geometric shape of the object, and to have a certain estimation accuracy when using the non-geometric corner points of the object as feature points for pose estimation, the terminal device can estimate the pose of the target object in the following manner S101-S103.
[0032] The following is an exemplary description of a method for predicting the interaction relationship of a drug pair provided in the present application in conjunction with specific embodiments.
[0033] See also Figure 1 , Figure 1 A flowchart of a method for estimating a posture provided by an embodiment of the present application is shown, and the method comprises the following steps:
[0034] S101. A terminal device acquires a target image, where the target image is obtained by photographing a target object.
[0035] In one embodiment, the target image is an image containing a target object, and the terminal device needs to process the target image to estimate the position and posture of the target object, wherein the target object may be an aircraft, a vehicle, a pedestrian, etc., without limitation.
[0036] In one embodiment, when photographing the target object, the three-dimensional model of the target object may be photographed, or the target object in the actual scene may be photographed, without limitation.
[0037] In one embodiment, the target image may contain one or more target objects, which is not limited to one or more. It should be noted that when there are multiple target objects, the position and posture of each target object needs to be estimated based on its feature points.
[0038] S102: The terminal device recognizes the target image using a recognition model to obtain a plurality of feature points of the target object in the target image.
[0039] In one embodiment, the recognition model is a model for recognizing feature points of a target object in a target image. At the same time, after recognizing the feature points, the recognition model can also determine the predicted position of the feature points in the target image.
[0040] It is understandable that the target image is composed of multiple pixels, and the feature point is a pixel in the target image. Based on this, the terminal device can pre-determine any position in the target image as the coordinate origin and construct a two-dimensional coordinate system to determine the predicted position of the feature point in the target image.
[0041] Specifically, the recognition model is obtained by training the preset initial model through a preset training set, a preset first loss function and a preset second loss function. The preset training set includes multiple sample images, the actual position and actual offset of the training feature points of the target object in each sample image; the first loss function is used to characterize the loss between the actual position of each training feature point and the predicted position predicted by the initial model; the second loss function is used to characterize the loss between the actual offset of each training feature point and the predicted offset predicted by the initial model.
[0042] In one embodiment, the preset training set is a training set for training a recognition model. The preset training set includes a plurality of sample images. It should be noted that when the target object is an aircraft, the sample images are images of aircraft of different positions and types. The initial model can be obtained by pre-training a neural network using a public geometric figure data set. In this way, the initial model has basic image recognition capabilities, and thus better model training can be performed.
[0043] It should be added that when the recognition model identifies feature points and their locations, it does not directly output the predicted location, but first outputs the predicted location of the feature point and the predicted offset. Then, the initial predicted location is corrected according to the predicted offset to obtain the final predicted location. Therefore, when training the recognition model, the preset training set needs to include not only the actual location of the feature point in the image, but also the actual offset corresponding to the feature point.
[0044] It should be noted that the feature map output by the initial model after processing the sample image may be reduced by S times compared to the sample image. That is, one pixel in the feature map corresponds to a pixel area composed of S times the pixels in the sample image. It can be seen that when the initial model predicts the position of the feature point, it may predict that the feature point is one of the S times the pixels. That is, the error of the final prediction result of the initial model may be large. Therefore, if the preset training set only includes the actual position of the feature point in the sample image for model training, the prediction accuracy of the feature point by the initial model is low. Therefore, in this embodiment, the preset training set needs to be specially processed so that it not only includes the actual position of the feature point in the image, but also includes the actual offset corresponding to the feature point for model training.
[0045] Based on this, it can be known that during training, the initial model needs to use the above two loss functions to iterate the model parameters respectively to obtain the recognition model. That is, it is necessary to use: a first loss function that calculates the loss between the actual position of each training feature point and the predicted position predicted by the initial model; and a second loss function that calculates the loss between the actual offset of each training feature point and the predicted offset predicted by the initial model.
[0046] It can be understood that at this time, after the recognition model is generated, the recognition model will output the predicted position of the feature point in the target image and the predicted offset at the time of this prediction. After that, the predicted position and the predicted offset are added after linear transformation, and the corresponding position of the feature point in the target image is obtained. In this way, the accuracy of the recognition model in identifying and locating the feature point and the position of the feature point in the target image can be improved.
[0047] It should be added that when determining the actual offset, the feature points may have offsets in width and / or height. Therefore, the actual offset will consist of the actual height offset and the actual width offset. In addition, if the actual offsets are manually annotated one by one, the annotation cost will increase. Based on this, refer to Figure 2 The terminal device can determine the actual offset of the training feature point in the sample image through the following steps S201-S204, as detailed below:
[0048] S201. The terminal device calculates the scaling factor of the sample image when it is processed by the initial model according to the size of the sample image and the dimensional size of the feature map output by the initial model.
[0049] In one embodiment, the size of the sample image can be determined according to the width and height of the sample image. Exemplarily, the width of the image is defined as D and the height is defined as E. The output feature map is the feature map output after the initial model processes the sample image.
[0050] Based on the above description, it can be known that a pixel point in the feature map corresponds to a pixel area composed of multiple pixel points in the sample image. Therefore, the dimension of the feature map is usually much smaller than the dimension of the sample image. Among them, the scaling factor is a factor determined by both the width and the height. Usually, in the above scaling factors, the scaling factor of the sample image in width should be consistent with the scaling factor of the sample image in height, and there is no limitation on this.
[0051] S202: The terminal device obtains the actual width and actual height of the training feature point in the sample image.
[0052] In one embodiment, the training feature point is usually a pixel in the sample image. However, there are also cases where the training feature point is multiple pixels in the sample image. That is, multiple pixels are combined to represent a training feature point. In this case, the actual width and actual height of the training feature point will not only be the width and height of one pixel.
[0053] It should be added that if the training feature point is a combination of multiple pixel points, the position of the training feature point can also correspond to multiple ones; or, the position of one of the pixel points is used to represent the actual position of the training feature point, which is not limited.
[0054] S203. The terminal device calculates a first ratio of the actual width to the zoom factor, and rounds the first ratio down to obtain a target first ratio; and calculates a second ratio of the actual height to the zoom factor, and rounds the second ratio down to obtain a target second ratio.
[0055] S204. The terminal device determines the difference between the first ratio and the target first ratio as the actual width offset; and determines the difference between the second ratio and the target second ratio as the actual height offset.
[0056] In one embodiment, the terminal device may specifically calculate the actual height offset and the actual width offset corresponding to each training feature point by using the following formula (1) and formula (2):
[0057]
[0058]
[0059] Wherein, d and e represent the actual width and height of the training feature points respectively; S represents the scaling factor; denote the first ratio and the second ratio respectively; d' and e' denote the actual width offset and the actual height offset respectively. Indicates the floor symbol.
[0060] Based on this, the calculated actual offset will be slightly adjusted based on the actual position, so that the recognition model trained based on the actual offset and the actual position can improve the accuracy of identifying and locating the feature points.
[0061] In a specific embodiment, the recognition model includes an encoder and a decoder; the encoder includes a multi-layer residual network, and the output layer of the residual network is an adaptive averaging layer; the decoder includes a positioning decoder for predicting the pixel area where the feature point is located, and an offset decoder for predicting the offset of the feature point in the pixel area.
[0062] Specifically, refer to Figure 3 , Figure 3 The network structure diagram of the recognition model in one embodiment of the present application; the multi-layer residual network is specifically a 50-layer residual network (ResNet50). In addition, the last fully connected layer (the output layer) in the multi-layer residual network is replaced by an adaptive averaging layer to improve the accuracy of the residual network output result.
[0063] in, Figure 3 The feature point decoder in includes a positioning decoder and an offset decoder. Moreover, the positioning decoder and the offset decoder are composed of 2 hidden layers and 1 output layer. Among them, the first hidden layer is a 1×1 convolution kernel with 512 channels, and the second hidden layer is a 3×3 convolution kernel with 256 channels. Among them, the output layer of the positioning decoder is a single-channel feature map after binarization. Moreover, each pixel point on the feature map maps a 16×16 pixel area in the target image. And, the probability that each pixel point in the corresponding pixel area belongs to a feature point is output. The output layer of the offset decoder is used to output a 2-channel feature map, the first channel represents the predicted offset of the width of the corresponding area of the detected feature point in the target image, and the second channel represents the predicted offset of the height of the corresponding area of the detected feature point in the target image.
[0064] The settings corresponding to the entire initial model are as follows:.
[0065] a) The network parameters of the initial model are initialized using normal distribution, and the activation function is the ReLU function;
[0066] b) The optimizer uses the Adam optimizer, the learning rate is adjusted to 0.001, and the first-order decay rate β1 and the second-order decay rate β2 of the optimizer are set to 0.9 and 0.009 respectively;
[0067] c) The mini-batch size for training is set to 4;
[0068] d) The iteration period of pre-training is set to 6, and the iteration period of transfer learning is set to 4.
[0069] Afterwards, the preset initial model is trained using a preset training set, a preset first loss function, and a preset second loss function to obtain a recognition model.
[0070] In addition, during the training phase, the first loss function and the second loss function in the initial model are defined as follows:
[0071] Define the channel where the pixel point on the feature map is located as k, the width is d, and the height is e. On the first loss function, it can be shown as formula (3):
[0072]
[0073] Among them, d and e respectively indicate that the predicted feature point is in a pixel area with a width of d and a length of e, and x a,b Indicates the probability value of the pixel at the predicted feature point position ab; y a,b Indicates the actual value of the pixel at position ab; that is, if the pixel at position ab is a feature point, the actual value is set to 1; otherwise, it is set to 0 for calculation. It should be added that the above 1 and 0 are only one setting method in this embodiment, and in other embodiments, other values can also be set for calculation.
[0074] On the second loss function, it can be shown as formula (4):
[0075]
[0076] For the above two loss functions, they can also be processed by the following formula (5) to obtain the final loss function:
[0077] L total =2L position +L offset (5)
[0078] The above is the network structure and training method of the recognition model. The terminal device can use the above recognition model to recognize the target image and accurately obtain multiple feature points of the target object in the target image without being limited to the geometric corner points of the target object. Afterwards, the terminal device can perform the following S103 step:
[0079] S103: The terminal device determines the position and posture of the target object based on the multiple feature points and the preset spatial point of the target object.
[0080] In one embodiment, the above-mentioned preset spatial points of the target object usually have multiple points, all of which are spatial points that are pre-marked and determined on the model of the target object. It should be noted that the number of feature points identified by the recognition model may be the same as the number of spatial points, or may be different. Moreover, the identified feature points may match the above-mentioned spatial points one by one, or may not match the above-mentioned spatial points.
[0081] In one embodiment, the terminal device can specifically use a random sampling consensus algorithm (RANSAC) to first establish a matching relationship between feature points and spatial points to form feature point pairs, and then use an efficient multi-point perspective imaging algorithm (EPnP) to solve the posture, and use the Gauss-Newton iteration method to optimize the posture.
[0082] In this embodiment, the target image is directly identified by using a recognition model to obtain a plurality of feature points of the target object in the target image. Then, the image feature points are matched with preset spatial points of the target object respectively to calculate the position and posture of the target object. At this time, the recognition model only outputs the feature points of the target object in the target image, and does not directly output the position and posture of the target object in the target image. Therefore, when the recognition model identifies and locates the non-geometric corner points of the target object, it can also improve the accuracy of identifying the feature points of the target object. In this way, when determining the position and posture of the target object, not only is it not necessary for the recognition model to be limited to identifying the geometric corner points of the object, but the estimation accuracy of the position and posture of the target object can also be improved based on the feature points with higher recognition accuracy.
[0083] In a specific embodiment, referring to Figure 4 The terminal device can determine the position and posture of the target object through the following steps S401-S406, as detailed below:
[0084] S401. The terminal device randomly selects a preset number of feature points and matches them with a preset number of spatial points to obtain an initial matching relationship between the feature points and the spatial points.
[0085] In one embodiment, the preset number may be a number set by the terminal device according to actual conditions, for example, the number may be 4. Usually, the number of feature points will be much larger than the preset number, and therefore, a plurality of feature points may be randomly selected. However, when the number of feature points is less than or equal to the preset number, all feature points need to be selected and randomly matched with the same number of spatial points.
[0086] It should be noted that the initial matching relationship at this time is only the matching relationship between a preset number of feature points and a preset number of spatial points; however, whether other feature points and other spatial points also have this initial matching relationship needs to be determined through subsequent steps.
[0087] S402: The terminal device maps the spatial points to the target image according to the initial matching relationship to obtain the projected feature points of the target object.
[0088] In one embodiment, the position of the above-mentioned spatial point is usually a three-dimensional coordinate, and the position of the feature point is a two-dimensional coordinate. Therefore, it is necessary to map each spatial point to the target image according to the above-mentioned initial matching relationship, and obtain a plurality of corresponding projected feature points. At this time, the position of the projected feature point is a two-dimensional coordinate.
[0089] S403: The terminal device calculates the geometric distances between the projected feature points and the corresponding feature points respectively.
[0090] In one embodiment, the geometric distance is the geometric distance between the projected feature point and the corresponding feature point. The geometric distance can be calculated based on the two-dimensional coordinates of the projected feature point and the two-dimensional coordinates of the feature point, which is not limited to this description.
[0091] In one embodiment, the feature points corresponding to each projection feature point may be determined by a random sampling consistency algorithm. It should be noted that, in an initial matching relationship, the number of geometric distances between the projection feature point and the corresponding feature point should be 1.
[0092] It can be understood that the geometric distance can represent the matching degree between the projected feature point and the corresponding feature point under the initial matching relationship. Generally, the larger the geometric distance, the lower the matching degree.
[0093] S404. The terminal device repeatedly executes S401-S404 to obtain multiple geometric distances.
[0094] S405. The terminal device determines the feature point corresponding to the minimum value of the geometric distance as the target feature point.
[0095] S406: The terminal device determines the position and posture according to the matching relationship between the target feature point and the space point.
[0096] In one embodiment, if the initial matching relationship is obtained based on the randomly selected feature points only once, the initial matching relationship may not be the optimal matching relationship between the projected feature points and the feature points. That is, the matching degree is not the best. Therefore, the terminal device can repeatedly perform the above steps to obtain multiple geometric distances. For example, repeat the steps N times to obtain N geometric distances. Among them, N can be set according to the actual situation, which is the judgment condition for the terminal device to jump out of the loop.
[0097] It can be understood that the geometric distance can characterize the degree of matching between the projected feature point and the corresponding feature point under the initial matching relationship. Therefore, after obtaining multiple geometric distances, the terminal device can determine the feature point corresponding to the minimum value of the geometric distance as the target feature point. At the same time, the pose is determined based on the initial matching relationship calculated between the target feature point and the spatial point. In this way, the above method can improve the accuracy of estimating the pose of the target object.
[0098] In a specific embodiment, the terminal device can calculate the position and posture of the target object by the following method and formula:
[0099] First, a set of three-dimensional spatial points q on the target object is known. i (i=1,2,..,M) and use the recognition model to obtain a set of two-dimensional feature points p j(j=1,2,…,N); then, four spatial points are randomly sampled to form a spatial point matrix Q and four feature points to form a feature point matrix P, so that Q and P form a 2D-3D matching relationship (if the number of feature points is less than or equal to 4, all are selected). Then, the EPnP algorithm f is used to solve the pose and obtain the initial matching relationship (initial rotation matrix R and translation vector t). Then, R and t are used to transform all other spatial feature points Project it into the target image to get the homogeneous coordinates of the spatial points projected on the image (i.e., the coordinates of the projected feature points); and, according to the intrinsic parameter matrix K when photographing the target object (e.g., the intrinsic parameter matrix of the camera), the projected feature points and the corresponding matching feature points are calculated; then, (8) the geometric distances e between all matching projected feature points and the feature points are calculated; wherein, and Respectively represent the homogeneous coordinates of the spatial points that form the matching relationship and the homogeneous coordinates of the feature points, and W represents the number of projected feature points that form the matching relationship. Then, the above steps are repeated until the number of iterations is met, and multiple set distances are obtained. Finally, (9) returns the initial matching relationship (R' and t') corresponding to the minimum value of the geometric distance as the final estimated pose of the target object.
[0100] The calculation formula is as follows:
[0101] [R,t]=f(Q,P), (3)
[0102]
[0103]
[0104] [R',t']=argmine(R,t) (6)
[0105] The meaning of each character in the above formula has been explained in the above example and will not be explained again.
[0106] Among them, the EPnP algorithm f is used to solve the pose and obtain the initial matching relationship, which can be specifically performed through the following steps and formulas:
[0107] (1) Control point selection:
[0108] The core idea of EPnP is to use a set of non-coplanar virtual control points to represent a space point. In this embodiment, four non-coplanar virtual control points can be used. For example: Represents the non-homogeneous coordinates of the virtual control point in the world coordinate system, Represents the non-homogeneous coordinates of a spatial point in the world coordinate system. The strategy for selecting control points is as follows:
[0109] As (10), find the centroid of the spatial point set as the first virtual control point
[0110]
[0111] According to (11), a matrix A is constructed, and then A is calculated T The eigenvalue λ of A i (i=1,2,3) and the corresponding eigenvector v i , therefore, the remaining three control points can be obtained from (12)
[0112]
[0113]
[0114] (2) Spatial points are represented by virtual control points:
[0115] After obtaining the virtual control point from 0, we can establish (13) and The relationship between
[0116]
[0117] where α ji yes use represents the weight parameter. Therefore, we can obtain α ji for:
[0118]
[0119] (3) Representation of virtual control points in the camera coordinate system:
[0120] In this embodiment, it is possible to Represents the non-homogeneous coordinates of the spatial point in the camera coordinates, and the non-homogeneous coordinates of the feature point matched with it are represented as u j =[w j h j ] T , Represents the non-homogeneous coordinates of the virtual control point in the camera coordinates. Thus, the relationship between the feature point and the virtual control point is obtained:
[0121]
[0122] Therefore, each feature point u j This corresponds to a homogeneous system of equations:
[0123]
[0124] After that, we can use U = [u 1 … u N ] T Indicates that all feature points form a 2N×1 feature point vector, using It means that all virtual control points in the camera coordinate system form a 12×1 virtual control point vector. Therefore, the relationship between U and C is expressed as
[0125]
[0126]
[0127] Among them, M is a 2N×12 matrix. Next, calculate M T The N zero eigenvalues βk (k = 1, 2, ..., N) of M and the corresponding eigenvectors v k However, in practical applications, due to the error relationship, M T M may not have a zero eigenvalue, so the 4 smallest eigenvalues β can be selected k (k=1,2,3,4) and the corresponding eigenvector v k Therefore, C can be expressed as
[0128]
[0129] (4) Solving the posture
[0130] Because of the above β k is an undetermined parameter, so we need to k Solution. Ideally, the Euclidean distance (the geometric distance mentioned above) between virtual control points in the world coordinate system should be equal to the Euclidean distance in the camera coordinate system. For the convenience of calculation, the following equivalent relationship can be obtained:
[0131]
[0132] Therefore, 4 points can get more than 6 forms of equations. And, for ease of calculation, we can use β mn Variable substitution β m β n Therefore, there will be 10 undetermined parameters βmn. After that, we can use β{β mn}(m≤n) represents the 10×1 unknown vector formed by the unknown parameters, and the remaining items on the left side of the equal sign form a 6×10 matrix A. B represents the value vector of the equation system, so the equation system can be written in the following form:
[0133] Aβ=B (16)
[0134] Because A is a 6×10 matrix and rank(A)=6, for A TPerform QR decomposition to obtain a set of standard orthogonal bases Q of the A column vector space 10×6 and the upper triangular matrix R 6×6 Therefore, β can be obtained by (21):
[0135] β=QR -T B (17)
[0136] Then, the coordinates of all spatial points under the camera are obtained through (14) and (22):
[0137]
[0138] And, (23) calculate The center of gravity and the matrix B are constructed as (24):
[0139]
[0140]
[0141] Then, the matrix H is calculated using the matrices A and B obtained from (8) (20)
[0142] H=B T A (21)
[0143] Then perform singular value decomposition on the above H:
[0144] H=U∑V T (twenty two)
[0145] Afterwards, the rotation matrix R can be obtained from U and V:
[0146] R=UV T (twenty three)
[0147] And, if |R|<0, let R(2,:)=-R(2,:), and then calculate the translation vector t:
[0148]
[0149] Based on this, the pose estimation of the target object can be achieved. Finally, the above algorithm can be implemented and tested on a desktop computer with a CPU of 2.2GHz, 10 cores, and RAM of 32GB. The feature point and space point matching algorithm can be implemented through the pytorch framework, and the above matching algorithm and pose estimation algorithm can also be implemented through the MATLAB program. When the effectiveness of the above algorithm is verified through simulation experiments, the attitude angle error in the above pose error is about 1.96°, and the relative translation error is about 3%, which has a better effect on the pose estimation of the target object.
[0150] For example, refer to Figure 5 , Figure 5 Schematic diagram of an application scenario of a recognition model for identifying feature points in one embodiment of the present application. Figure 5 The target object in the image is an aircraft, which includes target images generated by photographing aircraft models in different positions. The dots in the image are the feature points determined when the desktop computer recognizes the target image through the recognition model.
[0151] And, refer to Figure 6 , Figure 6 This is a schematic diagram of an application scenario for determining a posture in an embodiment of the present application. Figure 6 The outline generated by the white line in the middle is the estimated position and posture of the aircraft, which has a high degree of overlap with the actual position and posture of the aircraft.
[0152] See also Figure 7 , Figure 7 is a structural block diagram of a posture estimation device provided in an embodiment of the present application. The posture estimation device in this embodiment includes modules for executing Figure 1 , Figure 2 , Figure 4 For details, please refer to the steps in the corresponding embodiment. Figure 1 , Figure 2 , Figure 4 as well as Figure 1 , Figure 2 , Figure 4 For the convenience of explanation, only the parts related to this embodiment are shown. Figure 7 , the posture estimation device 700 may include: an acquisition module 710, an identification module 720 and a determination module 740, wherein:
[0153] The acquisition module 710 is used to acquire a target image, where the target image is obtained by photographing the target object.
[0154] The recognition module 720 is used to recognize the target image using the recognition model to obtain multiple feature points of the target object in the target image.
[0155] The determination module 730 is used to determine the position and posture of the target object according to the multiple feature points and the preset spatial points of the target object.
[0156] In one embodiment, the recognition model is obtained by training a preset initial model through a preset training set, a preset first loss function and a preset second loss function; wherein the preset training set includes multiple sample images, the actual position and actual offset of the training feature points of the target object in each sample image; the first loss function is used to characterize the loss between the actual position of each training feature point and the predicted position predicted by the initial model; the second loss function is used to characterize the loss between the actual offset of each training feature point and the predicted offset predicted by the initial model.
[0157] In one embodiment, the pose estimation device 700 further includes the following device to determine the actual offset of the training feature point:
[0158] The scaling factor calculation module is used to calculate the scaling factor of the sample image when it is processed by the initial model according to the size of the sample image and the dimension size of the feature map output by the initial model.
[0159] The actual offset calculation module is used to calculate the actual offset of the training feature point in the sample image according to the zoom factor and the actual position.
[0160] In one embodiment, the actual offset includes an actual height offset and an actual width offset; the actual offset calculation module is further used to:
[0161] Obtain the actual width and actual height of the training feature point in the sample image; calculate a first ratio of the actual width to the zoom factor, and round the first ratio down to obtain a target first ratio; and calculate a second ratio of the actual height to the zoom factor, and round the second ratio down to obtain a target second ratio; determine the difference between the first ratio and the target first ratio as the actual width offset; and determine the difference between the second ratio and the target second ratio as the actual height offset.
[0162] In one embodiment, the recognition model includes an encoder and a decoder; the encoder includes a multi-layer residual network, and the output layer of the residual network is an adaptive averaging layer; the decoder includes a positioning decoder for predicting the pixel area where the feature point is located, and an offset decoder for predicting the offset of the feature point in the pixel area.
[0163] In one embodiment, the determination module 730 is further configured to:
[0164] Determine the target feature point that matches the spatial point from multiple feature points; determine the position and posture based on the matching relationship between the target feature point and the spatial point.
[0165] In one embodiment, the determination module 730 is further configured to:
[0166] S1. Randomly select a preset number of feature points and match them with a preset number of spatial points to obtain an initial matching relationship between the feature points and the spatial points; S2. According to the initial matching relationship, map the spatial points to the target image to obtain the projected feature points of the target object; S3. Calculate the geometric distances between the projected feature points and the corresponding feature points; S4. Repeat S1-S4 to obtain multiple geometric distances; S5. Determine the feature point corresponding to the minimum value of the geometric distance as the target feature point.
[0167] When it is understood that Figure 7 In the structural block diagram of the posture estimation device shown, each module is used to perform Figure 1 , Figure 2 , Figure 4 The steps in the corresponding embodiments, and for Figure 1 , Figure 2 , Figure 4 Each step in the corresponding embodiment has been explained in detail in the above embodiment. Figure 1 , Figure 2 , Figure 4 as well as Figure 1 , Figure 2 , Figure 4 The relevant descriptions in the corresponding embodiments are not repeated here.
[0168] Figure 8 1 is a block diagram of a terminal device provided by an embodiment of the present application. Figure 8 As shown, the terminal device 800 of this embodiment includes: a processor 810, a memory 820, and a computer program 830 stored in the memory 820 and executable on the processor 810, such as a program of the pose estimation method. When the processor 810 executes the computer program 830, the steps in each embodiment of the pose estimation method described above are implemented, such as Figure 1 Alternatively, the processor 810 implements the above when executing the computer program 830 Figure 7 The functions of each module in the corresponding embodiment are, for example, Figure 7 For details on the functions of modules 710 to 730, please refer to Figure 7 Related description in the corresponding embodiment.
[0169] Exemplarily, the computer program 830 may be divided into one or more modules, one or more modules are stored in the memory 820, and are executed by the processor 810 to implement the pose estimation method provided in the embodiment of the present application. One or more modules may be a series of computer program instruction segments capable of completing specific functions, and the instruction segments are used to describe the execution process of the computer program 830 in the terminal device 800. For example, the computer program 830 may implement the pose estimation method provided in the embodiment of the present application.
[0170] The terminal device 800 may include, but is not limited to, a processor 810 and a memory 820. Those skilled in the art will appreciate that Figure 8 It is only an example of the terminal device 800 and does not constitute a limitation of the terminal device 800. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the terminal device may also include input and output devices, network access devices, buses, etc.
[0171] The processor 810 may be a central processing unit, or other general-purpose processors, digital signal processors, application-specific integrated circuits, off-the-shelf programmable gate arrays or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.
[0172] The memory 820 may be an internal storage unit of the terminal device 800, such as a hard disk or memory of the terminal device 800. The memory 820 may also be an external storage device of the terminal device 800, such as a plug-in hard disk, a smart memory card, a flash memory card, etc. equipped on the terminal device 800. Further, the memory 820 may also include both an internal storage unit of the terminal device 800 and an external storage device.
[0173] An embodiment of the present application provides a computer-readable storage medium, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, a posture estimation method as described in the above-mentioned embodiments is implemented.
[0174] An embodiment of the present application provides a computer program product. When the computer program product is run on a terminal device, the terminal device executes the posture estimation method in each of the above embodiments.
[0175] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A pose estimation method, It is characterized in that The method comprises: Acquire a target image, where the target image is obtained by photographing the target object; A recognition model is used to recognize the target image to obtain multiple feature points of the target object in the target image; the recognition model is obtained by training a preset initial model through a preset training set, a preset first loss function and a preset second loss function; the preset training set includes multiple sample images, and the actual position and actual offset of the training feature points of the target object in each of the sample images; the first loss function is used to characterize the loss between the actual position of each of the training feature points and the predicted position predicted by the initial model; the second loss function is used to characterize the loss between the actual offset of each of the training feature points and the predicted offset predicted by the initial model; the sum of twice the loss corresponding to the first loss function and the loss corresponding to the second loss function is used to train the initial model; Determining the position and posture of the target object according to the multiple feature points and the preset spatial point of the target object; The method for determining the actual offset of the training feature point is: Calculate, according to the size of the sample image and the dimension size of the feature map output by the initial model, a scaling factor of the sample image when being processed by the initial model; The actual offset of the training feature point in the sample image is calculated according to the scaling factor and the actual position.
2. The method according to claim 1, It is characterized in that The actual offset includes an actual height offset and an actual width offset; The calculating the actual offset of the training feature point in the sample image according to the scaling factor and the actual position includes: Obtaining the actual width and actual height of the training feature point in the sample image; Calculating a first ratio of the actual width to the zoom factor, and rounding the first ratio down to obtain a target first ratio; and calculating a second ratio of the actual height to the zoom factor, and rounding the second ratio down to obtain a target second ratio; The difference between the first ratio and the target first ratio is determined as the actual width offset; and the difference between the second ratio and the target second ratio is determined as the actual height offset.
3. The method according to claim 1, It is characterized in that The recognition model includes an encoder and a decoder; the encoder includes a multi-layer residual network, and the output layer of the residual network is an adaptive averaging layer; the decoder includes a positioning decoder for predicting the pixel area where the feature point is located, and an offset decoder for predicting the predicted offset of the feature point in the pixel area.
4. The method according to any one of claims 1 to 3, It is characterized in that The step of determining the position and posture of the target object according to the plurality of feature points and a preset spatial point of the target object comprises: Determine a target feature point matching the spatial point from the multiple feature points; The position and posture are determined according to the matching relationship between the target feature point and the spatial point.
5. The method according to claim 4, It is characterized in that The determining, from the plurality of feature points, a target feature point that matches the spatial point comprises: S1. Randomly select a preset number of feature points and match them with the preset number of space points to obtain an initial matching relationship between the feature points and the space points; S2. Mapping the spatial points to the target image respectively according to the initial matching relationship to obtain projection feature points of the target object; S3, respectively calculating the geometric distances between the projected feature points and the corresponding feature points; S4, repeatedly execute S1-S4 to obtain multiple geometric distances; S5. Determine the feature point corresponding to the minimum value of the geometric distance as the target feature point.
6. A posture estimation device, It is characterized in that The device comprises: An acquisition module, used to acquire a target image, wherein the target image is obtained by photographing a target object; A recognition module, used for recognizing the target image by using a recognition model to obtain a plurality of feature points of the target object in the target image; the recognition model is obtained by training a preset initial model through a preset training set, a preset first loss function and a preset second loss function; the preset training set includes a plurality of sample images, and the actual position and actual offset of the training feature points of the target object in each of the sample images; the first loss function is used to characterize the loss between the actual position of each of the training feature points and the predicted position predicted by the initial model; the second loss function is used to characterize the loss between the actual offset of each of the training feature points and the predicted offset predicted by the initial model; the sum of twice the loss corresponding to the first loss function and the loss corresponding to the second loss function is used to train the initial model; A determination module, used to determine the position and posture of the target object according to the multiple feature points and the preset spatial point of the target object; A scaling factor calculation module, used to calculate the scaling factor of the sample image when it is processed by the initial model according to the size of the sample image and the dimension size of the feature map output by the initial model; The actual offset calculation module is used to calculate the actual offset of the training feature point in the sample image according to the zoom factor and the actual position.
7. A terminal device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, It is characterized in that When the processor executes the computer program, the method according to any one of claims 1 to 5 is implemented.
8. A computer-readable storage medium storing a computer program. It is characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Attitude estimation method and device, medium and equipment
CN112149477A
Fingertip position detection method and device, equipment and storage medium
CN112558810A
Object attitude tracking method and device, terminal equipment and storage medium
CN113298870A