Three-dimensional coordinate determination method and device, electronic equipment and storage medium

By acquiring two-dimensional coordinates and depth information from joint images and using a neural network model for staged prediction, the problem of determining three-dimensional coordinates in joint images is solved, improving the accuracy and speed of coordinate determination.

CN115345937BActive Publication Date: 2025-11-21NANCHANG VIRTUAL REALITY RES INST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210933320.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-04
Publication Date
2025-11-21
Estimated Expiration
2042-08-04

AI Technical Summary

Technical Problem

In existing technologies, it is impossible to determine the coordinates of a joint in three-dimensional space using joint images.

Method used

By acquiring the two-dimensional coordinates of key joint points, the actual joint length, and the relative depth in the joint image, the first neural network model is used to predict the length value, and the second neural network model is used to iteratively predict the change in length value, thereby calculating the three-dimensional coordinates of the key joint points.

Benefits of technology

This method enables the phased determination of the three-dimensional coordinates of key joint points using joint images and joint lengths, improving the accuracy and speed of coordinate determination.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115345937B_ABST
    Figure CN115345937B_ABST
Patent Text Reader

Abstract

The application provides a three-dimensional coordinate determination method and device, electronic equipment and storage medium, including: obtaining a joint image of a joint; determining first input information; calculating a unit vector of a joint key point; predicting a predicted length value from the first input information by a first neural network model; calculating a predicted three-dimensional coordinate of the joint key point; calculating a predicted joint length of the joint; calculating a predicted joint length loss; determining second input information; predicting a predicted length value change from the second input information by a second neural network model; determining a total length value; determining whether an iteration end condition is reached according to the predicted joint length and the actual joint length; if so, calculating a three-dimensional coordinate of the joint key point according to the total length value and the unit vector of the joint key point. The application can determine the three-dimensional coordinate of the joint key point based on the joint image and the joint length.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, and more particularly, to a three-dimensional coordinate determination method and device, electronic equipment and storage medium. BACKGROUND

[0002] With the rapid development of computer vision technology, deep learning technology is increasingly widely used in the field of image processing. In the related art, although joint images of joints can be collected by a camera or the like, the coordinates of the joints in three-dimensional space cannot be determined according to the joint images. Therefore, how to inversely determine the coordinates of joint key points in three-dimensional space from the joint images of the joints is a technical problem to be solved in the related art. SUMMARY

[0003] In view of the above problems, the embodiments of the present application provide a three-dimensional coordinate determination method and device, electronic equipment and storage medium to improve the above problems.

[0004] According to an aspect of the embodiments of the present application, a three-dimensional coordinate determination method is provided, including: acquiring a joint image of a joint; determining first input information according to two-dimensional coordinates of a joint key point in the joint image, an actual joint length of the joint and a relative depth of the joint key point in the joint image; the first input information includes a unit vector of the joint key point; the unit vector is calculated according to the two-dimensional coordinates of the joint key point in the joint image and camera intrinsic parameters of a camera from which the joint image is derived; performing length value prediction by a first neural network model according to the first input information to obtain a predicted length value, the length value being a distance value between the three-dimensional coordinates of the joint key point and an origin; calculating a predicted three-dimensional coordinate of the joint key point according to the predicted length value and the unit vector of the joint key point; calculating a predicted joint length of the joint according to the predicted three-dimensional coordinate of the joint key point; calculating a predicted joint length loss according to the predicted joint length of the joint and the actual joint length of the joint; determining second input information according to the predicted joint length loss and the first input information; performing length value change amount prediction by a second neural network model according to the second input information to obtain a predicted length value change amount; adding the predicted length value change amount and the predicted length value to obtain a total length value; determining whether an iteration end condition is reached according to the predicted joint length and the actual joint length; if it is determined that the iteration end condition is reached, calculating the three-dimensional coordinates of the joint key point according to the total length value and the unit vector of the joint key point.

[0005] According to an aspect of some embodiments of the present application, a device for determining a three-dimensional coordinate is provided. The device comprises: an obtaining module configured to obtain a joint image of a joint; a first input information determining module configured to determine first input information according to a two-dimensional coordinate of a joint key point in the joint image, an actual joint length of the joint, and a relative depth of the joint key point in the joint image; the first input information comprises a unit vector of the joint key point; the unit vector is calculated according to the two-dimensional coordinate of the joint key point in the joint image and a camera intrinsic parameter of a camera from which the joint image is derived; a first predicting module configured to perform length value prediction by a first neural network model according to the first input information to obtain a predicted length value; the length value refers to a distance value between a three-dimensional coordinate of a joint key point and an origin; a predicted three-dimensional coordinate calculating module configured to calculate a predicted three-dimensional coordinate of the joint key point according to the predicted length value and the unit vector of the joint key point; a predicted joint length calculating module configured to calculate a predicted joint length of the joint according to the predicted three-dimensional coordinate of the joint key point; a predicted joint length loss calculating module configured to calculate a predicted joint length loss according to the predicted joint length of the joint and the actual joint length of the joint; a second input information calculating module configured to determine second input information according to the predicted joint length loss and the first input information; a second predicting module configured to perform length value change amount prediction by a second neural network model according to the second input information to obtain a length value change amount; a total length value determining module configured to add the predicted length value change amount and the predicted length value to obtain a total length value; a judging module configured to judge whether an iteration end condition is reached according to the predicted joint length and the actual joint length; and a three-dimensional coordinate calculating module configured to calculate the three-dimensional coordinate of the joint key point according to the total length value and the unit vector of the joint key point if it is determined that the iteration end condition is reached.

[0006] In some embodiments, the judging module comprises: a ratio calculating unit configured to calculate a ratio of the predicted joint length and the actual joint length; a loss value calculating unit configured to determine a loss value of a loss function according to the ratio; a judging unit configured to determine that the iteration end condition is reached if the loss value of the loss function is less than a loss threshold; and determine that the iteration end condition is not reached if the loss value of the loss function is not less than the loss threshold.

[0007] In some embodiments, the first input information determination module comprises: an intermediate three-dimensional coordinate calculation unit configured to calculate intermediate three-dimensional coordinates of the joint key points according to two-dimensional coordinates of the joint key points in the joint image and camera intrinsic parameters of a camera from which the joint image is derived; a unit vector calculation unit configured to normalize the intermediate three-dimensional coordinates to obtain unit vectors of the joint key points; a relative depth calculation unit configured to calculate relative depths of the joint key points in the joint image according to depth information of each pixel in the joint image and a depth value of the joint key point in the joint image; and a first input information determination unit configured to determine the first input information according to the unit vectors, an actual joint length of the joint, and the relative depths of the joint key points in the joint image.

[0008] In some embodiments, the first input information determination unit comprises: a first preprocessing subunit configured to input the unit vectors into a first preprocessing neural network to obtain unit vector features of the unit vectors; a second preprocessing subunit configured to input the actual joint length into a second preprocessing neural network to obtain joint length features of the actual joint length; a third preprocessing subunit configured to input the relative depths into a third preprocessing neural network to obtain relative depth features of the relative depths; a reshaping subunit configured to reshape the unit vector features to obtain two-dimensional unit vector features, reshape the joint length features to obtain two-dimensional joint length features, and reshape the relative depth features to obtain two-dimensional relative depth features; and a first input information determination subunit configured to combine the two-dimensional unit vector features, the two-dimensional joint length features, and the two-dimensional relative depth features to obtain the first input information.

[0009] In some embodiments, the three-dimensional coordinate determination apparatus further comprises: a predicted relative depth calculation module configured to calculate predicted relative depths of the joint key points according to predicted three-dimensional coordinates of the joint key points; a relative depth loss calculation module configured to calculate a relative depth loss according to the predicted relative depths of the joint key points and the relative depths of the joint key points; a fourth preprocessing module configured to input the relative depth loss into a fourth preprocessing neural network to obtain a relative depth loss feature; a fifth preprocessing module configured to input the predicted length value into a fifth preprocessing neural network for preprocessing to obtain a predicted length value feature; and an adding module configured to add the relative depth loss feature and the predicted length value feature to the second input information.

[0010] In some embodiments, the second input information determination module comprises: a sixth preprocessing unit, configured to preprocess the predicted joint length loss by a sixth preprocessing neural network to obtain predicted joint length loss features; and a second input information determination unit, configured to combine the predicted joint length loss features with the first input information to obtain the second input information.

[0011] In some embodiments, the joint key points of the joint comprise a first joint key point indicating one end of the joint and a second joint key point indicating another end of the joint.

[0012] In some embodiments, the predicted joint length calculation module comprises: an Euclidean distance calculation unit, configured to calculate an Euclidean distance between the first joint key point and the second joint key point according to the predicted three-dimensional coordinates of the first joint key point and the predicted three-dimensional coordinates of the second joint key point; and a predicted joint length determination unit, configured to determine the calculated Euclidean distance as the predicted joint length of the joint. According to an aspect of an embodiment of the present application, there is provided an electronic device, comprising: a processor; a memory, the memory having computer readable instructions stored thereon, the computer readable instructions being executed by the processor to implement the three-dimensional coordinate determination method as described above.

[0013] According to an aspect of an embodiment of the present application, there is provided a computer readable storage medium having computer readable instructions stored thereon, the computer readable instructions being executed by a processor to implement the three-dimensional coordinate determination method as described above.

[0014] According to an aspect of an embodiment of the present application, there is provided a computer program product comprising computer instructions, the computer instructions being executed by a processor to implement the three-dimensional coordinate determination method as described above.

[0015] In the present application, the length value is determined in two stages, i.e., after the first input information is determined according to the two-dimensional coordinates of the joint key points in the joint image, the first neural network is used to predict the length value according to the first input information to obtain a predicted length value, then the second neural network is used to iteratively predict the length value change amount according to the second input information to obtain a predicted length value change amount, then the total length value is obtained by adding the predicted length value and the predicted length value change amount, and the three-dimensional coordinates of the joint key points are determined based on the total length value, which realizes the determination of the three-dimensional coordinates of the joint key points in stages by using the joint image and the joint length.

[0016] Further, in the scheme of the present application, the second input information comprises a predicted joint length loss, which is calculated based on an actual joint length and a predicted joint length calculated based on the predicted length value, so as to use the characteristic that the joint length of a joint is constant as supervision information for determining the length value variation, thereby ensuring the accuracy of the predicted length value variation determined, and further ensuring the accuracy of the three-dimensional coordinates of the joint key points determined subsequently.

[0017] It should be understood that the general description above and the detailed description below are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF DRAWINGS

[0018] The drawings incorporated in the specification and forming a part thereof illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the application. It is clear that the drawings described below are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.

[0019] Figure 1 A flow chart of the method for determining three-dimensional coordinates according to an embodiment of the present application.

[0020] Figure 2 A specific step flow chart of step 102 according to an embodiment of the present application.

[0021] Figure 3 A specific step flow chart of step 240 according to an embodiment of the present application.

[0022] Figure 4 A specific step flow chart of step 110 according to an embodiment of the present application.

[0023] Figure 5 A specific step flow chart after step 107 according to an embodiment of the present application.

[0024] Figure 6 A schematic diagram of the process for determining three-dimensional coordinates according to an embodiment of the present application.

[0025] Figure 7 A block diagram of the device for determining three-dimensional coordinates according to an embodiment of the present application.

[0026] Figure 8 A structural schematic diagram of a computer system of an electronic device suitable for implementing embodiments of the present application is shown. DETAILED DESCRIPTION

[0027] Example implementations will now be described more fully with reference to the accompanying drawings. Example implementations may, however, be implemented in many different forms and should not be construed as limited to the examples set forth herein; rather, these implementations are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the example implementations to those skilled in the art.

[0028] Moreover, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of embodiments of the application. One skilled in the relevant art will recognize, however, that the

[0029] The block diagrams in the drawings show only the functionality of the example implementations and do not imply any particular physical or architectural arrangement of the example implementations. For example, functions shown as discrete blocks in the example implementations can be provided in one or more modules, components, or integrated circuits. The example implementations can be implemented in hardware, software, or a combination of both, and can be implemented in one or more devices.

[0030] The flow diagrams depicted in the drawings show the functionality of example implementations and do not imply any particular physical or architectural arrangement of the example implementations. For example, the flow diagrams can be implemented in hardware, software, or a combination of both. Further, the flow diagrams can be implemented in one or more devices, such as a server, a cloud server, or other electronic devices.

[0031] Figure 1 is a flow diagram of a method for determining a three-dimensional coordinate according to an embodiment of the present application. The method of the present application can be executed by an electronic device with processing capability, such as a server, a cloud server, etc., which is not specifically limited herein. As shown in Figure 1 the method includes:

[0032] At step 101, an image of a joint is obtained.

[0033] In the present embodiment, the image of the joint can be a grayscale image (gray scale image), in which each pixel in the image can be represented by an intensity value from 0 (black) to 255 (white). Different gray levels are represented between 0-255. In other embodiments, the image of the joint can also be an RGB image.

[0034] At step 102, first input information is determined according to the two-dimensional coordinates of the joint key point in the joint image, the actual joint length of the joint, and the relative depth of the joint key point in the joint image; the first input information includes a unit vector of the joint key point; the unit vector is calculated according to the two-dimensional coordinates of the joint key point in the joint image and the camera intrinsic parameter of the camera from which the joint image is derived.

[0035] The joint key point refers to a point on a joint that has a joint identification function, which can be a joint point, a bone point, etc., and of course can also be other points defined by a user. It can be understood that for the same joint, it can include multiple joint key points, and in the scheme of the present application, since the joint length of the joint is involved, the joint key point in the present application at least includes a first joint key point located at one end of the joint and a second joint key point at the other end of the joint, and of course, in specific embodiments, it can also include other joint key points in addition to the first joint key point and the second joint key point.

[0036] For example, the joint key point can be: a point at the tip position of a finger, a point at the base of the phalanx of the distal phalanx joint of the finger, a point at the end of the phalanx joint of the finger, a point at the metacarpophalangeal joint of the attachment point of the finger and the palm, or a point at the wrist position of the attachment point of the palm and the human forearm, etc. The number and position of the specific joint key points can be set according to actual needs, which are not specifically limited here.

[0037] In some embodiments, a joint key point set can be pre-set for a joint, so that in step 102, the joint key point in the joint key point set corresponding to the joint is located in the joint image. It can be understood that the joint key points included in the joint key point set set for different joints are different.

[0038] The two-dimensional coordinates of the joint key point in the joint image refer to the coordinates of the joint key point in the image coordinate system of the joint image. Therefore, after locating the pixel where the joint key point is located in the joint image, the two-dimensional coordinates of the joint key point in the joint image can be correspondingly obtained.

[0039] Since in practice, there can be problems such as positioning errors of joint key points due to unclear images, in order to ensure the accuracy of the three-dimensional coordinates of the joint key points determined subsequently, abnormal pixel points can be filtered first, specifically, an Isolation Forest algorithm or a Local Outlier Factor (LOF) algorithm can be used to exclude abnormal pixel points.

[0040] The relative depth of the joint key point in the joint image refers to the depth value of the pixel where the joint key point is located relative to a reference pixel in the joint image.

[0041] wherein the unit vector of the joint key point refers to the unit vector of the ray between the three-dimensional coordinates of the joint key point and the coordinate origin, i.e., the unit vector between the origin (0, 0, 0). Optionally, the unit vector of a joint key point can be expressed as wherein in the three-dimensional space rectangular coordinate system, is a set of bases representing a space vector, and the coordinates are expressed as: If the unit vector of a joint key point is preceded by a negative sign, it means that the vector with the negative sign in the joint key point is located on the negative half-axis of the corresponding coordinate axis, for example, if the unit vector of a joint key point is wherein in the three-dimensional space rectangular coordinate system, the coordinates corresponding to the x-axis are (-1, 0, 0), and the coordinates corresponding to the z-axis are (0, 0, -1).

[0042] In some embodiments, as shown in Figure 2 step 102 includes:

[0043] Step 210: calculating the intermediate three-dimensional coordinates of the joint key point according to the two-dimensional coordinates of the joint key point in the joint image and the camera intrinsic parameters of the camera from which the joint image is derived.

[0044] The camera intrinsic parameters include the focal length of the camera and the coordinates of the optical center of the camera. The camera intrinsic parameter matrix is a matrix constructed according to the focal length of the camera and the coordinates of the optical center of the camera. This matrix is called the camera intrinsic parameter matrix, denoted as K, and K is:

[0045]

[0046] wherein (f x , f y ) is the focal length of the camera, and (c x , c y ) is the coordinates of the optical center of the camera.

[0047] For a point P in the world coordinate system, assuming its coordinates in the world coordinate system are (X, Y, C), the coordinates of the point P in the camera coordinate system of the camera are (X C , Y C , Z C ), based on the camera, the point is perspective projected to the image coordinate system, and the coordinates of the point P are (u, v). The specific coordinate transformation process can be described according to the following formula:

[0048]

[0049] From the above, we can get:

[0050]

[0051] Also, because, It can be regarded as the ray vector of the joint key point in the camera coordinate system, that is,

[0052] If point P is a joint key point in the present application, the above coordinates (X C , Y C , Z C ) are regarded as the intermediate three-dimensional coordinates of the joint key point, It can be regarded as the ray vector of the joint key point in the camera coordinate system.

[0053] Optionally, in order to calculate the intermediate three-dimensional coordinates of the joint key point, one dimension is added to the two-dimensional coordinates of the joint key point to obtain Then left multiply the inverse matrix of the camera intrinsic matrix to obtain the intermediate three-dimensional coordinates, denoted as P 2dz :

[0054] P 2dz = K -1 ×P′ 2d (Formula 3)

[0055] Where K -1 represents the inverse matrix of the camera intrinsic matrix of the camera.

[0056] Step 220, normalizing the intermediate three-dimensional coordinates to obtain the unit vector of the joint key point.

[0057] In some embodiments, the distance between the intermediate three-dimensional coordinates and the origin (0, 0, 0) can be calculated first, that is, the length of the intermediate three-dimensional coordinate vector, and then the unit vector of the joint key point is calculated according to the distance. For example, the intermediate three-dimensional coordinates are (X C , Y C , Z c ), and the distance between the origin (0, 0, 0) is: The corresponding unit vector of the joint key point is:

[0058] Step 230, according to the depth information of each pixel in the joint image and the depth value of the joint key point in the joint image, the relative depth of the joint key point in the joint image is calculated.

[0059] The depth information of each pixel in the joint image is used to indicate the depth value of the corresponding pixel in the joint image.

[0060] In some embodiments, during the image acquisition towards the joint, the joint image of the joint is acquired simultaneously, and the depth image of the joint is acquired, wherein the value of each pixel point in the depth image is the depth value of the corresponding pixel point. Similarly, in order to determine the depth value corresponding to the joint key point, the joint key point is positioned in the depth image, so that the depth value of the joint key point is correspondingly acquired.

[0061] On this basis, after obtaining the depth value of each pixel, one of the pixels is selected as a reference pixel point, and the relative depth of each joint key point in the joint image is correspondingly calculated. Among them, the reference pixel point can be the pixel point with the smallest depth value in the joint image, or the pixel point with the median depth value in the joint image, and of course it can also be other pixel points, which can be set according to actual needs.

[0062] When the pixel point with the smallest depth value in the joint image is selected as the reference pixel point, the relative depth of the joint key point can be calculated according to the following formula (5):

[0063]

[0064] In some embodiments, in order to unify the dimension, avoid the difference between the depth values of the pixel points in the joint image being too large, leading to the range of the determined relative depth value being too large, the relative depth calculated according to formula (4) can also be divided by the difference between the maximum depth value and the minimum depth value in the joint image, and the quotient obtained by the division is taken as the final relative depth of the joint key point, that is, the relative depth of each joint key point is calculated according to the following formula (5):

[0065]

[0066] In the above formulas (4) and (5), P 0.5d is the relative depth of the joint key point in the joint image; is the depth value of the joint key point; is the minimum depth value in the joint image; is the maximum depth value in the joint image.

[0067] Step 240, determining the first input information according to the unit vector, the actual joint length of the joint and the relative depth of the joint key point in the joint image.

[0068] In some embodiments, the unit vector of the joint key point, the actual joint length of the joint and the relative depth of the joint key point in the joint image are combined to obtain the first input information.

[0069] In some embodiments, as shown in Figure 3 Step 240 includes:

[0070] Step 241, input the unit vector into the first pre-processing neural network to obtain the unit vector feature of the unit vector.

[0071] Step 242, input the actual joint length into the second pre-processing neural network to obtain the joint length feature of the actual joint length.

[0072] Step 243, input the relative depth into the third pre-processing neural network to obtain the relative depth feature of the relative depth.

[0073] In some embodiments, the first pre-processing neural network, the second pre-processing neural network, and the third pre-processing neural network can be constructed by a fully connected neural network, or can be constructed by other neural networks. Optionally, using the first pre-processing neural network, the second pre-processing neural network, and the third pre-processing neural network to extract the features of the corresponding different parameters can improve the accuracy of determining the three-dimensional coordinates of the joint key points.

[0074] Step 244, reshape the unit vector feature to obtain a two-dimensional unit vector feature; reshape the joint length feature to obtain a two-dimensional joint length feature; and reshape the relative depth feature to obtain a two-dimensional relative depth feature.

[0075] In some embodiments, the reshaping processing refers to changing the dimension of a one-dimensional feature group or a multi-dimensional feature group to obtain a new dimension feature group, and the number of feature groups and the features do not change in the reshaping processing. Optionally, the reshape function can be used for reshaping processing. If the unit vector feature is a 1*16 feature group, where 1 represents that the feature group is a one-dimensional feature group, and 16 represents that the feature group has 16 features, the reshape function can be used for reshaping the feature group, input B = reshape(A, m, m), where A represents the original feature group, m represents the row and column number of the features in the returned feature group, and B represents the reshaped feature group, i.e., an m*m feature matrix B is obtained, and the elements in B are obtained from A by column. If the number of elements in A is not m*m, an error will be caused. For example, A is a 1*16 feature group, and by B = reshape(A, 4, 4), a 4*4 feature matrix is obtained, and all the features in B come from A. Reshaping the features can reduce the data amount of each parameter feature, thereby reducing the calculation pressure and storage pressure of the neural network.

[0076] Step 245, combine the two-dimensional unit vector feature, the two-dimensional joint length feature, and the two-dimensional relative depth feature to obtain the first input information.

[0077] In some embodiments, the two-dimensional unit vector feature, the two-dimensional joint length feature, and the two-dimensional relative depth feature can be spliced or combined to obtain the first input information.

[0078] Please continue to refer to Figure 1 Step 103, performing length value prediction according to the first input information by the first neural network model to obtain a predicted length value, wherein the length value refers to the distance value between the three-dimensional coordinates of the joint key point and the origin.

[0079] In this application, the length value predicted by the first neural network for the joint key point is referred to as the predicted length value.

[0080] In some embodiments, the first neural network model can be constructed by a convolutional neural network, and if the first input information is a two-dimensional array or a more-dimensional array, the length value prediction can be performed by the first neural network model constructed by the convolutional neural network.

[0081] In other embodiments, the first neural network model can be set as a neural network model with multiple network structures, the parameters of different network structures are different, and the diversity of the first neural network in predicting the predicted length value can be improved by changing the network structure of the first neural network model.

[0082] Of course, in other embodiments, the first neural network model can also be constructed by other neural networks, which are not specifically limited here.

[0083] As described above, for the joint key point in the world coordinate system, the intermediate three-dimensional coordinates of the joint key point can be determined through the perspective transformation of the camera, and then the coordinates of the joint key point in the image can be determined, and vice versa, the intermediate three-dimensional coordinates of the joint key point can also be calculated reversely according to the coordinates of the joint key point in the image and the length value, and the specific process can be described as:

[0084] P 2d =p roj (K, P 3d ) (Formula 6)

[0085] Wherein, P 2d is the two-dimensional coordinates of the joint key point, p roj represents a function of calculating the intermediate three-dimensional coordinates of the joint key point according to the three-dimensional coordinates of the joint key point and the camera intrinsic parameter matrix of the camera.

[0086] Step 104, calculating the predicted three-dimensional coordinates of the joint key point according to the predicted length value and the unit vector of the joint key point.

[0087] The predicted three-dimensional coordinates of the joint key point refer to the three-dimensional coordinates calculated according to the predicted length value and the unit vector of the joint key point.

[0088] In the embodiment, the specific method for calculating the predicted three-dimensional coordinates of the joint key point is: multiplying the predicted length value and the unit vector of the joint key point to obtain the predicted three-dimensional coordinates of the joint key point.

[0089] Specifically, the predicted three-dimensional coordinates of the joint key point are calculated according to the following formula:

[0090] P′ 3d = L p × V p (Formula 7)

[0091] wherein P′ 3d is the predicted three-dimensional coordinates of the joint key point; L p is the predicted length value; and V p is the unit vector of the joint key point.

[0092] Step 105, according to the predicted three-dimensional coordinates of the joint key point, the predicted joint length of the joint is calculated.

[0093] As described above, the joint key points of the joint include a first joint key point indicating one end of the joint and a second joint key point indicating the other end of the joint; according to the above process, the predicted three-dimensional coordinates of the first joint key point and the predicted three-dimensional coordinates of the second joint key point can be determined. On this basis, the joint length can be calculated according to the following process: according to the predicted three-dimensional coordinates of the first joint key point and the predicted three-dimensional coordinates of the second joint key point, the Euclidean distance between the first joint key point and the second joint key point is calculated; the calculated Euclidean distance is taken as the predicted joint length of the joint.

[0094] The calculation formula of the Euclidean distance is as follows:

[0095]

[0096] wherein the predicted three-dimensional coordinates of the first joint key point are (x1, y1, z1); the predicted three-dimensional coordinates of the second joint key point are (x2, y2, z2); and p is the Euclidean distance between the first joint key point and the second joint key point.

[0097] Step 106, according to the predicted joint length of the joint and the actual joint length of the joint, the predicted joint length loss is calculated.

[0098] In some embodiments, the actual joint length of the joint and the first predicted joint length are subtracted to obtain the joint length loss.

[0099] Step 107, according to the predicted joint length loss and the first input information, the second input information is determined.

[0100] In some embodiments, the predicted joint length loss and the first input information can be combined to obtain second input information.

[0101] At step 108, a second neural network model is used to predict a length value change amount according to the second input information, to obtain a predicted length value change amount.

[0102] In the present embodiment, the second neural network model can be constructed by a convolutional neural network. Of course, in other embodiments, the second neural network model can also be constructed by other neural networks, which are not specifically limited herein.

[0103] In some embodiments, the second neural network model can be set as a neural network model with multiple network structures, and the parameters of different network structures are different. The value of the predicted length value change amount can be enriched by setting the network structure of the second neural network model.

[0104] At step 109, the predicted length value change amount and the predicted length value are added to obtain a total length value.

[0105] At step 110, whether an iteration end condition is reached is determined according to the predicted joint length and the actual joint length.

[0106] In some embodiments, determining whether the iteration end condition is reached according to the predicted joint length and the actual joint length can make the predicted joint length closer to the actual joint length, and thus the total length value predicted according to the predicted joint length is more accurate, and the three-dimensional coordinates of the joint key points obtained are more accurate.

[0107] In some embodiments, as shown in FIG. 11, step 110 includes: Figure 4

[0108] At step 310, a ratio of the predicted joint length and the actual joint length is calculated.

[0109] In step 110, whether the iteration end condition is reached can be determined according to the process shown in FIG. 12. Figure 4

[0110] At step 320, a loss value of a loss function is determined according to the ratio.

[0111] In some embodiments, the ratio can be logarithmically calculated, and the logarithm calculation is a logarithm calculation with base 10. Optionally, the logarithm calculation can be wherein B'1 is the predicted joint length, and B1 is the actual joint length. Optionally, after the ratio is logarithmically calculated, the absolute value of the logarithmically calculated value is taken, and finally the average value of the absolute value is taken. Specifically, the loss function is specifically represented as: ​​Wherein, the abs function represents the absolute value function, the mean function represents the mean value function, and the final value is the loss value.

[0112] Step 330, if the loss value of the loss function is less than the loss threshold, it is determined that the iteration end condition is reached.

[0113] The loss threshold can be set according to actual needs, and is not specifically limited here.

[0114] Step 340, if the loss value of the loss function is not less than the loss threshold, it is determined that the iteration end condition is not reached.

[0115] In some embodiments, after step 340, the method further comprises: if it is determined that the iteration end condition is not reached, taking the total length value as the predicted length value in the next iteration process, and returning to execute step 105.

[0116] When it is determined that the iteration end condition is not reached, it is necessary to continue to use the first neural network model and the second neural network model for prediction. In the next iteration, the total length value in the last iteration is taken as the predicted length value in step 105, and steps 105 and subsequent steps are re-executed according to the replaced predicted length value, until the iteration end condition is reached according to the second prediction obtained from the joint length loss. The last round of predicted length value refers to the total length value obtained in the last iteration process relative to the current iteration round.

[0117] In some embodiments, when returning to execute step 105, the fourth preprocessing neural network, the fifth preprocessing neural network and the sixth preprocessing neural network with different results or parameters from the last time are used, so as to ensure the use of different preprocessing networks in each iteration and improve the accuracy of the prediction of the length value change amount. Through experiments, it is found that generally 3 iterations can determine that the iteration end condition is reached.

[0118] Please continue to refer to Figure 1 Step 111, if it is determined that the iteration end condition is reached, the three-dimensional coordinates of the joint key point are calculated according to the total length value and the unit vector of the joint key point.

[0119] The three-dimensional coordinates of the joint key point refer to the coordinates of the joint key point in the world coordinate system. When it is determined that the iteration end condition is reached, the three-dimensional coordinates of the joint key point required can be calculated according to the total length value predicted by the second neural network model and the unit vector of the joint key point.

[0120] The formula used to calculate the three-dimensional coordinates of the joint key point is specifically:

[0121] P′ 3d = L′ p x Vp (Formula 9)

[0122] wherein L' is the total length value. p is the total length value.

[0123] In the scheme of the present application, the length value is determined in two stages, that is, after the first input information is determined according to the two-dimensional coordinates of the joint key points in the joint image, the first neural network is used to predict the length value according to the first input information to obtain the predicted length value, and then the second neural network is used to iteratively predict the length value change amount according to the second input information to obtain the predicted length value change amount. Then, the total length value is obtained by adding the predicted length value and the predicted length value change amount, and the three-dimensional coordinates of the joint key points are determined based on the total length value, which realizes the determination of the three-dimensional coordinates of the joint key points in stages by using the joint image and the joint length.

[0124] Moreover, in the scheme of the present application, the second input information includes a predicted joint length loss, which is calculated based on the actual length of the joint and the predicted joint length calculated based on the predicted length value. Therefore, by using the characteristic that the joint length of the joint is constant, the joint length is used as the supervision information for determining the length value change amount, so as to ensure the accuracy of the predicted length value change amount determined, and further ensure the accuracy of the three-dimensional coordinates of the joint key points subsequently determined.

[0125] Further, in the scheme of the present application, in the case where it is judged based on the total length value that the iteration end condition is not reached, step 105 is returned to be executed again to predict the length value change amount, without the need to repeatedly predict the length value. Compared with the prediction process of the length value, the calculation amount in the prediction process of the length value change amount is smaller, and the iteration rate is faster. Therefore, the present scheme can improve the speed of determining the three-dimensional coordinates of the joint key points. In practice, the scheme of the present application is tested, and generally, three times of repeated iteration can reach the iteration end condition.

[0126] Further, the scheme of the present application can be applied to the online application stage and the training stage of the second neural network. It can be understood that in the training stage, if it is judged that the iteration end condition is not reached, the parameters of the second neural network model need to be adjusted to predict the length value change amount again by the second neural network after the parameters are adjusted based on the new second input information. Therefore, the scheme of the present application has a wide range of application.

[0127] In some embodiments, step 107 comprises: preprocessing the predicted joint length loss by a sixth preprocessing neural network to obtain predicted joint length loss features. The predicted joint length loss features are combined with the first input information to obtain the second input information.

[0128] In some embodiments, the preprocessing of the predicted joint length loss by the third sixth preprocessing neural network can include feature extraction of the predicted joint length loss by the sixth preprocessing neural network to obtain one-dimensional predicted joint length loss features, reshaping of the one-dimensional predicted joint length loss features by a reshape function to obtain two-dimensional predicted joint length loss features with equal width and height, and combination of the reshaped two-dimensional predicted joint length loss features with the first input information to obtain the second input information.

[0129] In some embodiments, the sixth preprocessing neural network can be constructed by a fully connected network, which can include multiple layers of fully connected network layers. Of course, in other embodiments, the sixth preprocessing neural network can also be other neural networks, which are not specifically limited here.

[0130] In some embodiments, as shown in FIG. 7, after step 107, the method further includes: Figure 5

[0131] Step 410: calculating predicted relative depths of the joint key points according to the predicted three-dimensional coordinates of the joint key points.

[0132] The predicted relative depth of the joint key point refers to the relative depth of the joint key point calculated according to the predicted three-dimensional coordinates of the joint key point. The predicted relative depth of the joint key point can be calculated according to formula (5) or formula (6). In the calculation process, the depth value corresponding to the three-dimensional coordinates of the joint key point is replaced by the depth value corresponding to the predicted three-dimensional coordinates of the joint key point. The Z-axis coordinate value in the predicted three-dimensional coordinates of the joint key point is the corresponding depth value.

[0133] Step 420: calculating a relative depth loss according to the predicted relative depths of the joint key points and the relative depths of the joint key points.

[0134] The relative depth loss of the joint key point is obtained by subtracting the relative depth of the joint key point calculated in step 220 from the predicted relative depth of the joint key point.

[0135] In this embodiment, the relative depth loss is also used as the data basis for the length value change amount prediction by the second neural network, thereby providing more data for the length value change amount prediction, which can improve the accuracy of the predicted length value change amount.

[0136] Step 430: inputting the relative depth loss into a fourth preprocessing neural network to obtain relative depth loss features.

[0137] ​In some embodiments, the fourth preprocessing neural network may be constructed from a fully connected network, which may include multiple layers of fully connected network. Of course, in other embodiments, the fourth preprocessing neural network may also be other neural networks, which are not specifically limited here.

[0138] Step 440: Input the predicted length value into the fifth preprocessing neural network for preprocessing to obtain the predicted length value feature.

[0139] In some embodiments, the fifth preprocessing neural network may be constructed from a fully connected network, which may include multiple layers of fully connected network. Of course, in other embodiments, the fifth preprocessing neural network may also be other neural networks, and no specific limitation is made here.

[0140] Step 450: Add the relative depth loss feature and the predicted length value feature to the second input information.

[0141] In some embodiments, during the process of adding the relative depth loss features and the predicted length value features to the second input information, the relative depth loss features and the predicted length value features can also be reshaped to obtain two-dimensional relative depth loss features and two-dimensional predicted length value features. Then, the two-dimensional relative depth loss features and the two-dimensional predicted length value features are added to the second input information to reduce the amount of feature data, thereby reducing the computational and storage pressure of the second neural network. Figure 6 This is a schematic diagram illustrating the process of determining three-dimensional coordinates according to an embodiment of this application, as shown below. Figure 6 As shown, the process of confirming the 3D coordinates of key joint points is divided into two stages: a first stage and a second stage. In the first stage, a first neural network model is used to predict the length value. In the second stage, a second neural network model is used to predict the change in length value.

[0142] The specific process of the first stage is as follows: Obtain the two-dimensional coordinates P of the joint key points in the joint image. 2d Then, based on the two-dimensional coordinates P of the joint key points in the joint image... 2d The inverse matrix K of the camera intrinsic parameter matrix of the camera from which the joint image originates. -1 Calculate the unit vector V of the joint key points p Then, the first preprocessing network Ln is used. Vp For unit vector V p Preprocessing is performed to obtain the unit vector V. p The one-dimensional unit vector feature is obtained by reshaping the one-dimensional unit vector feature to obtain the two-dimensional unit vector feature I′. V The relative depth value P of the joint key point is calculated based on the depth information of each pixel in the joint image and the depth value of the joint key point in the joint image.0.5d Then, the second preprocessing network Ln is used. P0.5d For relative depth value P 0.5d Preprocessing is performed to obtain the relative depth value P. 0.5d The one-dimensional relative depth feature is obtained, and then the one-dimensional relative depth feature is reshaped to obtain the two-dimensional relative depth feature I′. 0.5 ;Utilizing the third preprocessing network Ln B1 The actual joint length B1 corresponding to the joint key point is preprocessed to obtain a one-dimensional joint length feature of the actual joint length B1. Then, the one-dimensional joint length feature is reconstructed to obtain a two-dimensional joint length feature I′. B1 ; The two-dimensional joint length feature I′ corresponding to the joint key points. B1 The two-dimensional relative depth value feature I′ of the key joint. 0.5 And the two-dimensional unit vector feature I′ of the joint key point. V The first input information I′ is formed by splicing and combining the information, and the first neural network C is formed by combining the first input information I′. onv1 Based on the first input information I′, the length value is predicted to be L. p .

[0143] The second stage process specifically involves: based on the predicted length value L... p V and the unit vector of the joint key points p Calculate the predicted 3D coordinates P′ of the joint key points 3d Then, based on the predicted three-dimensional coordinates P′ of the joint key points 3d Calculate the predicted joint length B′1; subtract the predicted joint length B′1 from the actual joint length B1 to obtain the predicted joint length loss ΔB′1. Then, based on the predicted joint length loss ΔB′1 and the first input information I′, determine the second input information I″, which is then processed by the second neural network C. onv2 Based on the second input information I″, the change in length value is predicted, and the predicted change in length value ΔL is obtained. p .

[0144] In some embodiments, a sixth preprocessing neural network Ln can be utilized. ΔB The predicted joint length loss ΔB′1 is preprocessed to obtain the predicted joint length loss feature, and then the predicted joint length loss feature is reconstructed to obtain the two-dimensional predicted joint length loss feature I′. ΔB The two-dimensional predicted joint length loss feature I′ ΔB Combined with the first input information I′, the second input information I″ is obtained.

[0145] In some embodiments, the predicted relative depth loss ΔP' of the joint key point can be further calculated according to the predicted three-dimensional coordinates of the joint key point and the depth information of each pixel in the joint image 0.5d , and the fourth preprocessing network Ln ΔP0.5d is used again 0.5d to preprocess the relative depth loss ΔP' of the joint key point, to obtain a relative depth loss feature, and then the relative depth loss feature is reshaped to obtain a two-dimensional relative depth loss feature I' ΔP0.5d ; the fifth preprocessing network Ln Lp is used to preprocess the predicted length value L p of the joint key point, to obtain a predicted length value feature, and then the predicted length value feature is reshaped to obtain a two-dimensional predicted length value feature I' Lp ; the two-dimensional relative depth loss feature I' ΔP0.5d and the two-dimensional predicted length value feature I' Lp are added to the second input information I".

[0146] Then, the predicted length value of the first stage is added to the predicted length value change of the second stage to obtain a total length value. Then, a loss value of a loss function is calculated according to the predicted joint length and the actual joint length , and when the loss value of the loss function is less than a loss threshold value, the iteration is ended, the total length value of the last iteration is output, and the total length value is multiplied by the unit vector of the joint key point to calculate the three-dimensional coordinates of the joint key point.

[0147] The device embodiments of the present application are described below, which can be used to execute the methods in the above-mentioned embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the above-mentioned method embodiments of the present application.

[0148] Figure 7 is a block diagram of a three-dimensional coordinate determination device according to an embodiment of the present application, as shown in Figure 7As shown, the three-dimensional coordinate determination apparatus 600 comprises: an acquisition module 601 configured to acquire a joint image of a joint; a first input information determination module 602 configured to determine first input information according to a two-dimensional coordinate of a joint key point in the joint image, an actual joint length of the joint, and a relative depth of the joint key point in the joint image; the first input information comprises a unit vector of the joint key point; the unit vector is calculated according to the two-dimensional coordinate of the joint key point in the joint image and camera intrinsic parameters of a camera from which the joint image is derived; a first prediction module 603 configured to perform length value prediction by a first neural network model according to the first input information to obtain a predicted length value, the length value being a distance value between a three-dimensional coordinate of a joint key point and an origin; a predicted three-dimensional coordinate calculation module 604 configured to calculate a predicted three-dimensional coordinate of the joint key point according to the predicted length value and the unit vector of the joint key point; a predicted joint length calculation module 605 configured to calculate a predicted joint length of the joint according to the predicted three-dimensional coordinate of the joint key point; a predicted joint length loss calculation module 606 configured to calculate a predicted joint length loss according to the predicted joint length of the joint and the actual joint length of the joint; a second input information calculation module 607 configured to determine second input information according to the predicted joint length loss and the first input information; a second prediction module 608 configured to perform length value change amount prediction by a second neural network model according to the second input information to obtain a length value change amount; a total length value determination module 609 configured to add the predicted length value change amount and the predicted length value to obtain a total length value; a judgment module 610 configured to judge whether an iteration end condition is reached according to the predicted joint length and the actual joint length; a three-dimensional coordinate calculation module 611 configured to calculate the three-dimensional coordinate of the joint key point according to the total length value and the unit vector of the joint key point if it is determined that the iteration end condition is reached.

[0149] In some embodiments, the judgment module 610 comprises: a ratio calculation unit configured to calculate a ratio of the predicted joint length and the actual joint length; a loss value calculation unit configured to determine a loss value of a loss function according to the ratio; a judgment unit configured to determine that the iteration end condition is reached if the loss value of the loss function is less than a loss threshold; and determine that the iteration end condition is not reached if the loss value of the loss function is not less than the loss threshold.

[0150] In some embodiments, the first input information determination module 602 comprises: an intermediate three-dimensional coordinate calculation unit configured to calculate intermediate three-dimensional coordinates of the joint key points according to two-dimensional coordinates of the joint key points in the joint image and camera intrinsic parameters of a camera from which the joint image is derived; a unit vector calculation unit configured to normalize the intermediate three-dimensional coordinates to obtain unit vectors of the joint key points; a relative depth calculation unit configured to calculate relative depths of the joint key points in the joint image according to depth information of each pixel in the joint image and a depth value of the joint key point in the joint image; and a first input information determination unit configured to determine the first input information according to the unit vectors, an actual joint length of the joint, and the relative depths of the joint key points in the joint image.

[0151] In some embodiments, the first input information determination unit comprises: a first preprocessing subunit configured to input the unit vectors into a first preprocessing neural network to obtain unit vector features of the unit vectors; a second preprocessing subunit configured to input the actual joint length into a second preprocessing neural network to obtain joint length features of the actual joint length; a third preprocessing subunit configured to input the relative depths into a third preprocessing neural network to obtain relative depth features of the relative depths; a reshaping subunit configured to reshape the unit vector features to obtain two-dimensional unit vector features, reshape the joint length features to obtain two-dimensional joint length features, and reshape the relative depth features to obtain two-dimensional relative depth features; and a first input information determination subunit configured to combine the two-dimensional unit vector features, the two-dimensional joint length features, and the two-dimensional relative depth features to obtain the first input information.

[0152] In some embodiments, the three-dimensional coordinate determination apparatus 600 further comprises: a predicted relative depth calculation module configured to calculate predicted relative depths of the joint key points according to predicted three-dimensional coordinates of the joint key points; a relative depth loss calculation module configured to calculate a relative depth loss according to the predicted relative depths of the joint key points and the relative depths of the joint key points; a fourth preprocessing module configured to input the relative depth loss into a fourth preprocessing neural network to obtain a relative depth loss feature; a fifth preprocessing module configured to input the predicted length value into a fifth preprocessing neural network for preprocessing to obtain a predicted length value feature; and an adding module configured to add the relative depth loss feature and the predicted length value feature to the second input information.

[0153] In some embodiments, the second input information determination module 607 includes: a sixth preprocessing unit, configured to pre-process the predicted joint length loss by a sixth preprocessing neural network to obtain a predicted joint length loss feature; and a second input information determination unit, configured to combine the predicted joint length loss feature and the first input information to obtain the second input information.

[0154] In some embodiments, the joint key point of the joint includes a first joint key point indicating one end of the joint and a second joint key point indicating another end of the joint.

[0155] In some embodiments, the predicted joint length calculation module 605 includes: an Euclidean distance calculation unit, configured to calculate an Euclidean distance between the first joint key point and the second joint key point according to the predicted three-dimensional coordinates of the first joint key point and the predicted three-dimensional coordinates of the second joint key point; and a predicted joint length determination unit, configured to determine the calculated Euclidean distance as the predicted joint length of the joint.

[0156] In some embodiments, the predicted three-dimensional coordinate calculation module 604 includes: a predicted three-dimensional coordinate calculation unit, configured to multiply the predicted length value and the intermediate three-dimensional coordinates of the joint key point to obtain the predicted three-dimensional coordinates of the joint key point.

[0157] According to an aspect of some embodiments of the present application, an electronic device is also provided. As shown in the figure, the electronic device 700 includes a processor 710 and one or more memories 720, the one or more memories 720 are configured to store program instructions executed by the processor 710, and the processor 710 executes the program instructions to implement the object recognition method described above. Figure 8

[0158] ​Further, the processor 710 can include one or more processing cores. The processor 710 runs or executes instructions, programs, code sets or instruction sets stored in the memory 720, and calls data stored in the memory 720. Alternatively, the processor 710 can be implemented in at least one of a digital signal processing (DSP), a field-programmable gate array (FPGA), and a programmable logic array (PLA). The processor 710 can be integrated with a combination of one or more of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. Among them, the CPU is mainly used to process operating systems, user interfaces, and application programs; the GPU is used to be responsible for rendering and drawing display content; and the modem is used to process wireless communication. It can be understood that the above-mentioned modem can also not be integrated into the processor, but be realized by a separate communication chip.

[0159] According to an aspect of the present application, the present application also provides a computer readable storage medium, which can be contained in the electronic device described in the above embodiments, or can exist separately and not be assembled into the electronic device. The above computer readable storage medium carries computer readable instructions, which, when executed by a processor, implement the method in any of the above embodiments.

[0160] It should be noted that the computer readable medium shown in the embodiments of the present application can be a computer readable signal medium or a computer readable storage medium or any combination of the two. The computer readable storage medium may, for example, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any combination of the above. More specific examples of the computer readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, the computer readable storage medium can be any tangible medium that contains or stores a program used by or in connection with an instruction execution system, apparatus or device. In this application, the computer readable signal medium can include a data signal carried in a baseband or as part of a carrier wave, in which the computer readable program code is carried. Such a propagated data signal can take many forms, including but not limited to, an electromagnetic signal, an optical signal or any suitable combination of the above. The computer readable signal medium can also be any computer readable medium that is not a computer readable storage medium and that can communicate, propagate or transport a program for use by or in connection with an instruction execution system, apparatus or device. The program code contained on the computer readable medium can be transmitted by any suitable medium, including but not limited to wireless, wired, optical, etc., or any suitable combination of the above.

[0161] It should be noted that although several modules or units of the device for action execution are mentioned in the foregoing detailed description, such a division is not mandatory. Indeed, according to an embodiment of the application, the features and functionalities of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functionalities of one module or unit described above can be further divided into embodied by a plurality of modules or units.

[0162] The computer program product of the present application can be a computer program product comprising a computer readable storage medium and a computer program mechanism embedded in the computer readable storage medium. The computer readable storage medium is not to be construed as a transitory signal per se. The computer readable storage medium is a tangible (non-transitory) computer-readable medium having computer readable code embodied thereon that, when executed by a computer, causes the computer to perform methods described herein. The computer readable medium can be a computer program product, a memory device, a memory component, a memory module, a memory bank, a memory circuit, a data storage system, a data storage container, or any suitable combination of the foregoing. A non-exhaustive list of tangible (non-transitory) computer-readable media includes compact discs (CDs), digital versatile discs (DVDs), Blu-Ray discs, floppy disks, hard disks, magnetic tape, magnetic-based storage, optical-based storage, flash memory, volatile memory, non-volatile memory, static random access memory (SRAM), dynamic random access memory (DRAM), volatile memory modules, non-volatile memory modules, flash memory modules, programmable logic devices, field-programmable gate arrays (FPGAs), eeprom, and other memory devices. The above memory devices are examples (and not limitation) of computer readable storage media. The computer readable storage medium can have stored thereon this computer program mechanism.

[0163] Other embodiments of the present application will be apparent to those skilled in the art from consideration of the specification and practice of the features disclosed herein. It is intended that the present application cover any and all variations of the application that come within the scope of the claims and a concept of equivalents. It is intended that the present application not be limited to the specific examples disclosed, but will include all implementations that are within the scope of the present application as defined by the appended claims, including full equivalents.

[0164] It is to be understood that the application is not limited to the precise construction described and as shown in the attached drawings, and that various modifications and changes can be effected thereon without departing from the scope of the application. The scope of the application is to be limited only by the appended claims.

Claims

1. A method of determining a three-dimensional coordinate, characterized by, The method comprises: obtaining a joint image of a joint; determining first input information according to two-dimensional coordinates of a joint key point in the joint image, an actual joint length of the joint and a relative depth of the joint key point in the joint image; the first input information comprises a unit vector of the joint key point; the unit vector is calculated according to the two-dimensional coordinates of the joint key point in the joint image and camera intrinsic parameters of a camera from which the joint image is derived; performing length value prediction by a first neural network model according to the first input information to obtain a predicted length value, wherein the length value refers to a distance value between three-dimensional coordinates of a joint key point and an origin; calculating predicted three-dimensional coordinates of the joint key point according to the predicted length value and the unit vector of the joint key point; calculating a predicted joint length of the joint according to the predicted three-dimensional coordinates of the joint key point; calculating a predicted joint length loss according to the predicted joint length of the joint and the actual joint length of the joint; determining second input information according to the predicted joint length loss and the first input information; performing length value change amount prediction by a second neural network model according to the second input information to obtain a predicted length value change amount; adding the predicted length value change amount and the predicted length value to obtain a total length value; judging whether an iteration end condition is reached according to the predicted joint length and the actual joint length; if it is determined that the iteration end condition is reached, calculating three-dimensional coordinates of the joint key point according to the total length value and the unit vector of the joint key point.

2. The method of claim 1, wherein, The judging whether the iteration end condition is reached according to the predicted joint length and the actual joint length comprises: calculating a ratio of the predicted joint length and the actual joint length; determining a loss value of a loss function according to the ratio; if the loss value of the loss function is less than a loss threshold value, it is determined that the iteration end condition is reached; if the loss value of the loss function is not less than the loss threshold value, it is determined that the iteration end condition is not reached.

3. The method of claim 1, wherein, The determining first input information according to two-dimensional coordinates of a joint key point in the joint image, an actual joint length of the joint and a relative depth of the joint key point in the joint image comprises: calculating intermediate three-dimensional coordinates of the joint key point according to the two-dimensional coordinates of the joint key point in the joint image and camera intrinsic parameters of a camera from which the joint image is derived; performing regularization on the intermediate three-dimensional coordinates to obtain a unit vector of the joint key point; calculating the relative depth of the joint key point in the joint image according to depth information of each pixel in the joint image and a depth value of the joint key point in the joint image; determining the first input information according to the unit vector, the actual joint length of the joint and the relative depth of the joint key point in the joint image.

4. The method of claim 3, wherein, The determining first input information according to the unit vector, the actual joint length of the joint and the relative depth of the joint key point in the joint image comprises: input the unit vector into a first preprocessing neural network to obtain a unit vector feature of the unit vector; input the actual joint length into a second preprocessing neural network to obtain a joint length feature of the actual joint length; input the relative depth into a third preprocessing neural network to obtain a relative depth feature of the relative depth; perform reshaping processing on the unit vector feature to obtain a two-dimensional unit vector feature, perform reshaping processing on the joint length feature to obtain a two-dimensional joint length feature, and perform reshaping processing on the relative depth feature to obtain a two-dimensional relative depth feature; merge the two-dimensional unit vector feature, the two-dimensional joint length feature, and the two-dimensional relative depth feature to obtain first input information.

5. The method of claim 1, wherein, After determining the second input information according to the predicted joint length loss and the first input information, the method further includes: calculating a predicted relative depth of the joint key point according to the predicted three-dimensional coordinates of the joint key point; calculating a relative depth loss according to the predicted relative depth of the joint key point and the relative depth of the joint key point; inputting the relative depth loss into a fourth preprocessing neural network to obtain a relative depth loss feature; inputting the predicted length value into a fifth preprocessing neural network for preprocessing to obtain a predicted length value feature; adding the relative depth loss feature and the predicted length value feature to the second input information.

6. The method of claim 1, wherein, The determining of the second input information according to the predicted joint length loss and the first input information includes: preprocessing the predicted joint length loss by a sixth preprocessing neural network to obtain a predicted joint length loss feature; combining the predicted joint length loss feature and the first input information to obtain the second input information.

7. The method of claim 1, wherein, The joint key points of the joint include a first joint key point indicating one end of the joint and a second joint key point indicating the other end of the joint; The calculating of the predicted joint length of the joint according to the predicted three-dimensional coordinates of the joint key point includes: calculating the Euclidean distance between the first joint key point and the second joint key point according to the predicted three-dimensional coordinates of the first joint key point and the predicted three-dimensional coordinates of the second joint key point; taking the calculated Euclidean distance as the predicted joint length of the joint.

8. A device for determining a three-dimensional coordinate, characterized by The device includes: an acquisition module configured to acquire a joint image of a joint; a first input information determination module configured to determine first input information according to two-dimensional coordinates of a joint key point in the joint image, an actual joint length of the joint, and a relative depth of the joint key point in the joint image; the first input information includes a unit vector of the joint key point; the unit vector is calculated according to two-dimensional coordinates of the joint key point in the joint image and camera intrinsic parameters of a camera from which the joint image is derived; The first prediction module is configured to perform length value prediction on the first input information by using a first neural network model to obtain a predicted length value, the length value being a distance value between a three-dimensional coordinate of a joint key point and an origin; The predicted three-dimensional coordinate calculation module is configured to calculate a predicted three-dimensional coordinate of the joint key point according to the predicted length value and a unit vector of the joint key point; The predicted joint length calculation module is configured to calculate a predicted joint length of the joint according to the predicted three-dimensional coordinate of the joint key point; The predicted joint length loss calculation module is configured to calculate a predicted joint length loss according to the predicted joint length of the joint and an actual joint length of the joint; The second input information calculation module is configured to determine second input information according to the predicted joint length loss and the first input information; The second prediction module is configured to perform length value change amount prediction on the second input information by using a second neural network model to obtain a predicted length value change amount; The total length value determination module is configured to add the predicted length value change amount and the predicted length value to obtain a total length value; The judgment module is configured to determine whether an iteration end condition is reached according to the predicted joint length and the actual joint length; The three-dimensional coordinate calculation module is configured to calculate a three-dimensional coordinate of the joint key point according to the total length value and the unit vector of the joint key point if it is determined that the iteration end condition is reached.

9. An electronic device, comprising: The electronic device comprises: one or more processors; a memory electrically connected to the one or more processors; one or more application programs, wherein the one or more application programs are stored in the memory and configured to be executed by the one or more processors, and the one or more application programs are configured to perform the method of any one of claims 1-7.

10. A computer readable storage medium, characterized in that, The computer-readable storage medium stores program codes, and the program codes can be called and executed by a processor to perform the method of any one of claims 1-7.

Citation Information

Patent Citations

  • Direction angle information-based three-dimensional target detection method

    CN108597009A

  • Hand posture three-dimensional reconstruction method and device and storage medium

    CN113362452A