Three-dimensional pose estimation method and system, and computer-readable storage medium
By employing a heat map and verification code network to calculate pixel confidence scores, the method addresses noise issues in Pix2Pose networks, improving the precision of three-dimensional pose estimation through selective point selection.
Patent Information
- Application Number
- CN202210617194.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-01
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-06-01
AI Technical Summary
The three-dimensional pose estimation algorithm based on a monocular RGB camera in the prior art has a large error in the RANSAC and PnP algorithms due to the large number of noise points, which affects the accuracy of the pose estimation results.
By obtaining the target two-dimensional image, input it to the heat map network and the verification code network, calculate the confidence of each pixel point, use the preset selection algorithm to select two-dimensional points with higher confidence, and calculate the three-dimensional pose with Pix2Pose network and RANSAC and PnP algorithms.
The accuracy of the target three-dimensional pose is improved, and the accuracy of pose estimation is improved by screening out two-dimensional points with higher confidence, reducing the influence of noise points.
Smart Images

Figure CN115147573B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of three-dimensional pose estimation, and particularly relates to a three-dimensional pose estimation method, system, and computer-readable storage medium. Background Art
[0002] Target three-dimensional pose estimation is the cornerstone of many perception-related applications. Especially in fields such as virtual reality and robot control, the accuracy of target three-dimensional pose estimation directly determines the intuitive interaction effect. The target three-dimensional pose estimation algorithm based on a monocular RGB camera is widely used in virtual reality applications due to its low requirements for hardware devices. Pix2Pose designs an autoencoder architecture to estimate the three-dimensional coordinates of each pixel in a two-dimensional image, forming a two-dimensional to three-dimensional correspondence, and then uses the RANSAC and PnP algorithms to calculate the target three-dimensional pose, effectively handling occlusion and symmetry problems.
[0003] However, when there are many noise points with estimation errors in the three-dimensional coordinates obtained by the Pix2Pose network, the errors of the RANSAC and PnP algorithms that rely on two-dimensional to three-dimensional point pairs for calculation also increase significantly, and have a serious adverse impact on the pose estimation result. Summary of the Invention
[0004] The present application aims to at least solve one of the technical problems existing in the prior art. For this reason, the present application provides a three-dimensional pose estimation method, system, and computer-readable storage medium, which can select two-dimensional points with higher confidence, thereby improving the accuracy of the target three-dimensional pose.
[0005] In a first aspect, the present application provides a three-dimensional pose estimation method, including:
[0006] Obtain a target two-dimensional image;
[0007] Input the target two-dimensional image into a heatmap network to obtain a target heatmap image;
[0008] Input the target two-dimensional image into a check code network to obtain a target check image;
[0009] Calculate the confidence of each pixel point on the target two-dimensional image according to the target heatmap image and the target check image to obtain a confidence set;
[0010] Determine at least one target two-dimensional point from the confidence set according to a preset confidence selection algorithm;
[0011] Calculate the target three-dimensional pose according to the three-dimensional points corresponding to the target two-dimensional points.
[0012] The three-dimensional pose estimation method according to the embodiments of the first aspect of the present application has at least the following beneficial effects: obtaining a target two-dimensional image, inputting the target two-dimensional image into a heatmap network to obtain a target heatmap image, and at the same time, inputting the same target two-dimensional image into a check code network to obtain a target check image, calculating the confidence of each pixel point on the target two-dimensional image by the target heatmap image and the target check image to obtain a confidence set, selecting at least one target two-dimensional point in the two-dimensional image from the confidence set according to a preset selection algorithm, estimating the three-dimensional coordinates of the target two-dimensional point through a Pix2Pose network, and then calculating the target three-dimensional pose using RANSAC and PnP algorithms. By calculating the confidence of the generated target check image and the existing target heatmap image, two-dimensional points with higher confidence are screened out, thereby improving the accuracy of the target three-dimensional pose.
[0013] According to some embodiments of the first aspect of the present application, the check code network is trained through the following steps: obtaining training two-dimensional images rendered from a training three-dimensional map at different angles and the vertex coordinates and center coordinates of the training three-dimensional map; rendering a self-supervised check image through a vertex shader according to the training two-dimensional image, the vertex coordinates, and the center coordinates; inputting the training two-dimensional image into an incompletely trained check code network to obtain a predicted check image; and adjusting the parameters of the check code network according to the predicted check image and the self-supervised check image.
[0014] According to some embodiments of the first aspect of the present application, the obtaining of the training two-dimensional images rendered from a training three-dimensional map at different angles includes: presetting a rotation angle and calculating the minimum number of rotation times according to the rotation angle; obtaining the training two-dimensional images rendered from the training three-dimensional map at different angles according to the rotation angle and the training number of rotation times; where the training number of rotation times is an integer less than or equal to the minimum number of rotation times.
[0015] According to some embodiments of the first aspect of the present application, after the step of rendering a self-supervised check image through a vertex shader according to the training two-dimensional image, the vertex coordinates, and the center coordinates, it includes: creating a training data set; where the training data set includes multiple training data groups; and recording the training two-dimensional image corresponding to the training three-dimensional map at the same angle and the self-supervised check image into the same training data group.
[0016] According to some embodiments of the first aspect of the present application, the training step of the check code network further includes: rendering a self-supervised heatmap image through a vertex shader according to the training two-dimensional image, the vertex coordinates, and the center coordinates; and recording the self-supervised heatmap image into the training data group corresponding to the training three-dimensional map at the same angle.
[0017] According to some embodiments of the first aspect of the present application, rendering the self-supervised verification image through a vertex shader according to the training two-dimensional image, the vertex coordinates, and the center coordinates includes: calculating the maximum distance between the vertex coordinates and the center coordinates and the first distance from each vertex coordinate to the center coordinate according to the center coordinate and all the vertex coordinates; calculating the vertex gray value corresponding to each vertex coordinate according to the maximum distance and the first distance; and rendering the self-supervised verification image through the vertex shader according to the vertex gray value.
[0018] According to some embodiments of the first aspect of the present application, adjusting the parameters of the verification code network according to the predicted verification image and the self-supervised verification image includes: obtaining the total number of two-dimensional points and the number of two-dimensional points of the target mask on the training two-dimensional image; calculating the loss value of the verification code network according to the predicted verification image, the self-supervised verification image, the total number of two-dimensional points, and the number of two-dimensional points of the target mask; and adjusting the parameters of the verification code network according to the loss value.
[0019] According to some embodiments of the first aspect of the present application, calculating the confidence of each pixel point on the target two-dimensional image according to the target heat map image and the target verification image to obtain a confidence set includes: calculating the color value of each two-dimensional point according to the target heat map image; calculating the gray value of each two-dimensional point according to the target verification image; and calculating the confidence of the corresponding two-dimensional point according to the color value and the gray value of the same two-dimensional point, so as to obtain the confidence set.
[0020] In a second aspect, the present application provides a three-dimensional pose estimation system, including: at least one memory; at least one processor; at least one program; the program is stored in the memory, and the processor executes at least one program to implement the three-dimensional pose estimation method according to any one of the embodiments of the first aspect.
[0021] The three-dimensional pose estimation system according to the second aspect embodiment of the present application has at least the following beneficial effects: obtaining a target two-dimensional image, inputting the target two-dimensional image into a heat map network to obtain a target heat map image, and at the same time, inputting the same target two-dimensional image into a check code network to obtain a target check image, calculating the confidence of each pixel point on the target two-dimensional image from the target heat map image and the target check image to obtain a confidence set, selecting at least one target two-dimensional point in the two-dimensional image from the confidence set according to a preset selection algorithm, estimating the three-dimensional coordinates of the target two-dimensional point through a Pix2Pose network, and then using the RANSAC and PnP algorithms to calculate the target three-dimensional pose. By calculating the confidence between the generated target check image and the existing target heat map image, two-dimensional points with higher confidence are screened out, thereby improving the accuracy of the target three-dimensional pose.
[0022] In a third aspect, the present application also provides a computer-readable storage medium, which stores computer-executable signals for executing the three-dimensional pose estimation method according to any one of the embodiments in the first aspect.
[0023] Additional aspects and advantages of the present application will be given in part in the following description, become apparent in part from the following description, or be understood through the practice of the present application. Description of the Drawings
[0024] The additional aspects and advantages of the present application will become apparent and be easily understood in conjunction with the description of the embodiments with the following drawings, where:
[0025] Figure 1 is a flowchart of a three-dimensional pose estimation method according to an embodiment of the present application;
[0026] Figure 2 is a flowchart of a three-dimensional pose estimation method according to another embodiment of the present application;
[0027] Figure 3 is a flowchart of a three-dimensional pose estimation method according to another embodiment of the present application;
[0028] Figure 4 is a flowchart of a three-dimensional pose estimation method according to another embodiment of the present application;
[0029] Figure 5 is a flowchart of a three-dimensional pose estimation method according to another embodiment of the present application;
[0030] Figure 6 is a flowchart of a three-dimensional pose estimation method according to another embodiment of the present application;
[0031] Figure 7 is a flowchart of a three-dimensional pose estimation method according to another embodiment of the present application;
[0032] Figure 8 Flow chart of the three - dimensional pose estimation method for another embodiment of the present application;
[0033] Figure 9 Schematic diagram of the training two - dimensional image and the self - supervised verification image for an embodiment of the present application;
[0034] Figure 10 Schematic diagram of the training two - dimensional image and the self - supervised heat map for an embodiment of the present application. Detailed implementation manners
[0035] The embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary only for explaining the present application and should not be construed as limiting the present application.
[0036] In the description of the present application, it should be understood that for the orientation description, such as the orientation or positional relationship indicated by up, down, front, back, left, right, etc., is based on the orientation or positional relationship shown in the accompanying drawings. It is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as limiting the present application.
[0037] In the description of the present application, if the first and second are described only for the purpose of distinguishing technical features, they should not be construed as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or the sequence relationship of the indicated technical features.
[0038] In the description of the present application, unless otherwise clearly defined, words such as setting, installation, connection, etc. should be understood in a broad sense, and those skilled in the art can reasonably determine the specific meanings of the above words in the present application in combination with the specific content of the technical solution.
[0039] In the first aspect, referring to Figure 1 , the present application provides a three - dimensional pose estimation method, including but not limited to the following steps:
[0040] Step S110: Obtain a target two - dimensional image;
[0041] Step S120: Input the target two - dimensional image into a heat map network to obtain a target heat map;
[0042] Step S130: Input the target two - dimensional image into a verification code network to obtain a target verification image;
[0043] Step S140: Calculate the confidence of each pixel on the target two-dimensional image according to the target thermal image and the target verification image to obtain a confidence set;
[0044] Step S150: determining at least one target two-dimensional point from the confidence set according to a preset confidence selection algorithm;
[0045] Step S160: Calculate the target three-dimensional posture according to the three-dimensional point corresponding to the target two-dimensional point.
[0046] Get the target two-dimensional image I t , and input the target two-dimensional image into the thermal map network to obtain the target thermal image C t At the same time, the same target two-dimensional image is input into the verification code network to obtain the target verification image G t , the target thermal image C t and target verification image G t The confidence of each pixel on the target two-dimensional image is calculated to obtain a confidence set. According to the preset selection algorithm, at least one target two-dimensional point is selected from the two-dimensional image and the three-dimensional coordinates of the target two-dimensional point are estimated through the Pix2Pose network. The RANSAC and PnP algorithms are then used to calculate the target three-dimensional posture. The confidence calculation is performed by comparing the generated target verification image with the existing target thermal image to screen out two-dimensional points with higher confidence, thereby improving the accuracy of the target three-dimensional posture.
[0047] Reference Figure 2 It can be understood that the verification code network in step S130 is obtained by training the following steps:
[0048] Step S210: obtaining vertex coordinates and center coordinates of the training two-dimensional image and the training three-dimensional image rendered from the training three-dimensional image at different angles;
[0049] Step S220: Rendering a self-supervised verification image through a vertex shader according to the training two-dimensional image, vertex coordinates and center coordinates;
[0050] Step S230: inputting the training two-dimensional image into the untrained verification code network to obtain a predicted verification image;
[0051] Step S240: Adjust the parameters of the verification code network according to the predicted verification image and the self-supervised verification image.
[0052] Reference Figure 9 , Figure 9 Four different angles are provided to obtain the training 2D images of the oil tank. i (like Figure 9 As shown in the left figure in the figure), and obtain the vertex coordinates v of the training three-dimensional graphi =(x i ,y i ,z i ), i ∈ [0, N) and the central coordinate v center =(x center ,y center ,z center ), where N represents the number of vertex coordinates of the training three-dimensional graph. According to the training two-dimensional image I i , the corresponding vertex coordinate v i , and the central coordinate v center , through the vertex shader for rendering, four self-supervised verification images G i of the oil tank corresponding to four different angles are obtained respectively, which represent the true images of the training two-dimensional image I i of the oil tank at the corresponding angle (as shown in the right figure in Figure 9 ). The training two-dimensional image I i is input into the incomplete training verification code network to obtain the predicted verification image G i '. According to the correct self-supervised verification image G i and the predicted verification image G i ' obtained from the verification code network, it is judged whether the verification code network needs to adjust its parameters.
[0053] Referring to Figure 3 , in step S210, the steps of obtaining the training two-dimensional images rendered from the training three-dimensional graph at different angles include but are not limited to the following steps:
[0054] Step S310: Preset the rotation angle and calculate the minimum number of rotation times according to the rotation angle;
[0055] Step S320: Obtain the training two-dimensional images rendered from the training three-dimensional graph at different angles according to the rotation angle and the training number of rotation times; where the training number of rotation times is an integer less than or equal to the minimum number of rotation times.
[0056] In this application, if the preset rotation angle is θ, then the minimum number of rotation times is For the training three-dimensional graph, its three-dimensional rotation angle is q = (θα, θβ, θγ), where α, β, γ represent the training number of rotation times in different directions, that is, the training three-dimensional Figure 1 needs to be rotated times in total, that is, corresponding to generating training two-dimensional images. Each training two-dimensional image needs to be rendered through the vertex shader to obtain a corresponding self-supervised verification image. For example, if the rotation angle θ is 5°, then the minimum number of rotation times is 72 times, that is, the training three-dimensional Figure 1A total of 373,248 rotations are required. By rotating the training 3D image multiple times according to the preset rotation angle, the training 2D images rendered at different angles are obtained, and more sets of data are used for training to improve the reliability of the verification code network training process.
[0057] Reference Figure 4 It is understandable that after step S220, the following steps are also included but not limited to:
[0058] Step S410: creating a training data collection; wherein the training data collection includes multiple training data groups;
[0059] Step S420: Record the training two-dimensional image and the self-supervised verification image corresponding to the training three-dimensional image at the same angle into the same training data group.
[0060] By creating a training data collection, the corresponding training two-dimensional images and self-built supervised verification images are stored in the same training array, which is convenient for the subsequent input of the training two-dimensional images corresponding to the same angle training three-dimensional images into the verification code network to obtain the predicted verification image and the corresponding self-built supervised verification image for comparison and discrimination, so as to avoid too much training data causing messy data storage, that is, the training data set D i ∈(( i ,G i ).
[0061] Reference Figure 5 , the training steps of the verification code network also include but are not limited to the following steps:
[0062] Step S510: Rendering a self-supervised thermal image through a vertex shader according to the training two-dimensional image, vertex coordinates and center coordinates;
[0063] Step S520: Record the self-supervised thermal image into the training data group corresponding to the three-dimensional training image at the same angle.
[0064] Reference Figure 10 , Figure 10 Four different angles are provided to obtain the training 2D images of the oil tank. i (like Figure 10 As shown in the left figure), each training two-dimensional image I i The corresponding self-supervised thermal image C can be obtained by calculation i (like Figure 10 As shown in the right figure), the specific self-supervised thermal image C i The calculation process is to calculate each vertex v in the catenary two-dimensional image by the following formula i =(x i ,y i ,z i )’s RGB color value:
[0065]
[0066] Among them, x max , y max , z max , x min , y min , z min respectively represent the maximum and minimum values of all vertex coordinates in the x, y, and z coordinate axes. At the same time, the self-supervised thermal image is recorded into the training data group corresponding to the three-dimensional map trained at the same angle, that is, the training data group D i ∈(( i , C i , G i ), that is, through the self-supervised thermal image C i the reference function for training the thermal map network can be achieved, and it can also be used as a control for the self-supervised verification image G i . This is not limited in this application.
[0067] With reference to Figure 6 , it can be understood that in step S220, it may include but is not limited to the following steps:
[0068] Step S610: Calculate the maximum distance between the vertex coordinates and the center coordinates and the first distance from each vertex coordinate to the center coordinates according to the center coordinates and all vertex coordinates;
[0069] Step S620: Calculate the vertex gray value corresponding to each vertex coordinate according to the maximum distance and the first distance;
[0070] Step S630: Render the self-supervised verification image through the vertex shader according to the vertex gray value.
[0071] For the specific calculation process of the self-supervised verification image G i , the gray values of all vertices v i are calculated through the following formula:
[0072] d i =||v i -v center ||,
[0073] Among them, v i represents the vertex coordinate, v center represents the center coordinate, d i represents the first distance between the i-th vertex coordinate and the center coordinate, d max represents the maximum distance between all vertex coordinates and the center coordinate, and g i represents the gray value of the i-th vertex.
[0074] Referring to Figure 7 , it can be understood that in step S240, it may include but is not limited to the following steps:
[0075] Step S710: Obtain the total number of 2D points and the number of 2D points of the target mask on the training 2D image;
[0076] Step S720: Calculate the loss value of the verification code network according to the predicted verification image, self-supervised verification image, total number of 2D points, and number of 2D points of the target mask;
[0077] Step S730: Adjust the parameters of the verification code network according to the loss value.
[0078] The total number of 2D points on the training 2D image in step S710 represents all the 2D points on the entire training 2D image, that is, it includes the 2D points of the oil tank part and the 2D points of the non-oil tank part such as Figure 9 or Figure 10 , while the number of 2D points of the target mask represents all the 2D points of the target part in the training 2D image, that is, it includes the 2D points of the oil tank part such as Figure 9 or Figure 10 . And the loss value of the verification code network is calculated through the following loss function formula:
[0079]
[0080] where M is the total number of 2D points on the training 2D image, Mmask is the number of 2D points of the target mask on the training 2D image, g j ’∈G’, g j ’∈G’, j ∈ α0, M), σ is an adjustable weight value, ranging from 0 to 1. The accuracy of the verification code network is measured by the loss value.
[0081] Referring to Figure 8 , it can be understood that in step S140, it may include but is not limited to the following steps:
[0082] Step S810: Calculate the color value of each 2D point according to the target thermal image;
[0083] Step S820: Calculate the gray value of each 2D point according to the target verification image;
[0084] Step S830: Calculate the confidence of the corresponding 2D point according to the color value and gray value of the same 2D point, so as to obtain a confidence set.
[0085] Calculate the confidence level by comparing the generated target verification image with the existing target heat map image, and select the two-dimensional points with higher confidence level, so as to improve the accuracy of the target three-dimensional pose. The confidence level calculation is obtained through the following formula:
[0086] p i = 1 - g i + ||c i ||;
[0087] where c i represents the color value of the two-dimensional point, and g i represents the grayscale value of the two-dimensional point.
[0088] In a second aspect, the present application also provides a three-dimensional pose estimation system, including: at least one memory, at least one processor, and at least one program. The program is stored in the memory, and the processor executes one or more programs to implement the above three-dimensional pose estimation method.
[0089] Obtain the target two-dimensional image I t , and input the target two-dimensional image into the heat map network to obtain the target heat map image C t . At the same time, input the same target two-dimensional image into the verification code network to obtain the target verification image G t . Calculate the confidence level of each pixel point on the target two-dimensional image for the target heat map image C t and the target verification image G t to obtain a confidence level set. According to a preset selection algorithm, select at least one target two-dimensional point in the two-dimensional image from the confidence level set, and estimate the three-dimensional coordinates of the target two-dimensional point through the Pix2Pose network. Then use the RANSAC and PnP algorithms to calculate the target three-dimensional pose. By calculating the confidence level by comparing the generated target verification image with the existing target heat map image, select the two-dimensional points with higher confidence level, so as to improve the accuracy of the target three-dimensional pose.
[0090] The processor and the memory can be connected by a bus or other means.
[0091] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs, non-transitory computer-executable programs, and signals, such as the program instructions / signals corresponding to the processing module in the embodiments of the present application. The processor executes various functional applications and data processing by running the non-transitory software programs, instructions, and signals stored in the memory, that is, implements the three-dimensional pose estimation method in the above method embodiments.
[0092] The memory may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store relevant data of the above three-dimensional pose estimation method, etc. In addition, the memory may include high-speed random access memory and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely provided relative to the processor, and these remote memories can be connected to the processing module through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0093] One or more signals are stored in the memory and, when executed by one or more processors, perform the three-dimensional pose estimation method in any of the above method embodiments. For example, perform the method steps S110 to S160 described above Figure 1 in, Figure 2 the method steps S210 to S240 in, Figure 3 the method steps S310 to S320 in, Figure 4 the method steps S410 to S420 in, Figure 5 the method steps S510 to S520 in, Figure 6 the method steps S610 to S630 in, Figure 7 the method steps S710 to S730 in and Figure 8 the method steps S810 to S830 in.
[0094] In a third aspect, embodiments of the present application provide a computer-readable storage medium. The computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by one or more processors, the above one or more processors can be caused to execute the three-dimensional pose estimation method in the above method embodiments. For example, perform the method steps S110 to S160 described above Figure 1 in, Figure 2 the method steps S210 to S240 in, Figure 3 the method steps S310 to S320 in, Figure 4 the method steps S410 to S420 in, Figure 5 the method steps S510 to S520 in, Figure 6 the method steps S610 to S630 in, Figure 7 the method steps S710 to S730 in and Figure 8 the method steps S810 to S830 in.
[0095] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0096] Through the description of the above embodiments, those of ordinary skill in the art can understand that all or some of the steps and systems disclosed in the above methods can be implemented as software, firmware, hardware, and their appropriate combinations. Some physical components or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable signals, data structures, program modules, or other data. Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassette, tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those of ordinary skill in the art, a communication medium typically includes computer-readable signals, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium.
[0097] In the description of this specification, the descriptions referring to the terms "one embodiment", "some embodiments", "illustrative embodiments", "examples", "specifically", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0098] The embodiments of the present application have been described in detail above in conjunction with the accompanying drawings. However, the present application is not limited to the above embodiments. Within the scope of knowledge possessed by those of ordinary skill in the art to which the present application pertains, various changes can be made without departing from the purpose of the present application.
Claims
1. A three-dimensional pose estimation method, characterized in that, Including: Obtain a target two-dimensional image; Input the target two-dimensional image into a heatmap network to obtain a target heatmap image; Input the target two-dimensional image into a check code network to obtain a target check image; According to the target heatmap image and the target check image, calculate the confidence of each pixel point on the target two-dimensional image to obtain a confidence set; Determine at least one target two-dimensional point from the confidence set according to a preset confidence selection algorithm; Calculate the target three-dimensional pose according to the three-dimensional points corresponding to the target two-dimensional points; Wherein, the check code network is obtained through training by the following steps: Obtain training two-dimensional images rendered from a training three-dimensional map at different angles, and the vertex coordinates and center coordinates of the training three-dimensional map; According to the training two-dimensional image, the vertex coordinates and the center coordinates, render through a vertex shader to obtain a self-supervised check image; Input the training two-dimensional image into an incompletely trained check code network to obtain a predicted check image; Adjust the parameters of the check code network according to the predicted check image and the self-supervised check image; Wherein, the step of rendering through a vertex shader according to the training two-dimensional image, the vertex coordinates and the center coordinates to obtain a self-supervised check image includes: According to the center coordinates and all the vertex coordinates, calculate the maximum distance between the vertex coordinates and the center coordinates and the first distance from each vertex coordinate to the center coordinates; According to the maximum distance and the first distance, calculate the vertex gray value corresponding to each vertex coordinate; According to the vertex gray value, render through a vertex shader to obtain a self-supervised check image; Wherein, the step of adjusting the parameters of the check code network according to the predicted check image and the self-supervised check image includes: Obtain the total number of two-dimensional points and the number of target mask two-dimensional points on the training two-dimensional image; Calculate the loss value of the check code network according to the predicted check image, the self-supervised check image, the total number of two-dimensional points and the number of target mask two-dimensional points; Adjust the parameters of the check code network according to the loss value; Wherein, the step of calculating the confidence of each pixel point on the target two-dimensional image according to the target heatmap image and the target check image to obtain a confidence set includes: Calculate the color value of each two-dimensional point according to the target heatmap image; Calculate the gray value of each two-dimensional point according to the target check image; According to the color value and the gray value of the same two-dimensional point, calculate the confidence of the corresponding two-dimensional point, and calculate the confidence of the corresponding two-dimensional point, so as to obtain a confidence set.
2. The three-dimensional attitude estimation method according to claim 1, wherein The step of obtaining training two-dimensional images rendered from a training three-dimensional map at different angles includes: Preset a rotation angle and calculate the minimum number of rotation times according to the rotation angle; Obtain training two-dimensional images rendered from a training three-dimensional map at different angles according to the rotation angle and the training number of rotation times; wherein, the training number of rotation times is an integer less than or equal to the minimum number of rotation times.
3. The three-dimensional attitude estimation method according to claim 2, characterized in that, After the step of rendering the self-supervised verification image through a vertex shader according to the training two-dimensional image, the vertex coordinates, and the center coordinates, it includes: Create a training data set; wherein, the training data set includes multiple training data groups; Record the training two-dimensional image corresponding to the training three-dimensional map at the same angle and the self-supervised verification image into the same training data group.
4. The three-dimensional attitude estimation method according to claim 3, characterized in that The training step of the verification code network further includes: Render a self-supervised heat map through a vertex shader according to the training two-dimensional image, the vertex coordinates, and the center coordinates; Record the self-supervised heat map into the training data group corresponding to the training three-dimensional map at the same angle.
5. A three-dimensional attitude estimation system, characterized in that, It includes: At least one memory; At least one processor; At least one program; The program is stored in the memory, and the processor executes at least one program to implement the three-dimensional pose estimation method according to any one of claims 1 to 4.
6. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable signals for executing the three-dimensional pose estimation method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Object 6D attitude estimation method and device based on 2D detection and computer equipment
CN112614184A
Two-dimensional and three-dimensional multi-person posture estimation system and method
CN112651316A