A three-dimensional human body pose estimation method, system, storage medium and electronic device
Through the ring arrangement of camera columns and optimized calibration algorithm, combined with two-dimensional and three-dimensional data processing, low-latency synchronization and accurate three-dimensional human posture estimation of multi-camera systems are achieved, solving the problems of high synchronization delay, large calibration error and poor robustness in the prior art, and improving the accuracy and efficiency of three-dimensional human posture reconstruction.
Patent Information
- Application Number
- CN202210974354.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-15
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-08-15
AI Technical Summary
In the prior art, multi-camera has high synchronization delay, low data processing efficiency, cumbersome camera calibration and large errors, and poor human posture estimation, especially at high frame rates, it is difficult to achieve accurate three-dimensional human posture reconstruction.
The camera column is arranged in a circular shape, and the two-dimensional posture estimation model is used to obtain two-dimensional human posture data. The three-dimensional human posture reconstruction is carried out by combining the camera's internal and external parameters. The camera parameters are optimized using a monocular and dual-objective algorithm, and the three-dimensional skeleton joint length checksum fusion is performed, and the final estimation is performed using a parametric human model.
It realizes low-latency synchronization of multi-camera and accurate three-dimensional human posture estimation, solves the calibration error problem, and improves robustness and reconstruction accuracy.
Smart Images

Figure CN115457594B_ABST
Abstract
Description
Background Art
[0002] Current three-dimensional human pose estimation methods generally can be divided into two steps. In the first step, two-dimensional human poses in all viewpoint cameras are generated. Two-dimensional human pose estimation can be divided into bottom-up methods and top-down methods. Generally speaking, benefiting from human instance information, top-down methods show higher average accuracy. Bottom-up methods assemble low-level features through the positions of low-level features, but it is challenging to assemble them accurately and quickly. The well-known algorithm OpenPose in this field introduces the Part Affinity Fields (PAF) to help parse low-level key points on limbs, thus obtaining high-precision real-time performance. In the second step, based on the coordinate transformation relationship between different viewpoints and combined with two-dimensional key points of viewpoints, the three-dimensional human key point coordinates are calculated through the binocular stereo vision principle. In this step, the pose relationship between different viewpoints and the camera internal parameters need to be known. Therefore, camera calibration needs to be carried out in advance when restoring the three-dimensional key point coordinates, and the calibration accuracy will have a great impact on the accuracy of the three-dimensional key point coordinates.
[0003] However, in the prior art, there are defects such as relatively high multi-camera synchronization delay, low efficiency of camera data writing and processing in high frame rate situations, high cumulative error in the external parameter calibration of multi-camera systems, cumbersome calibration steps, and relatively low robustness of human pose estimation to limb occlusion recognition. Therefore, there is an urgent need to provide a technical solution to solve the problems existing in the prior art. Summary of the Invention
[0004] To solve the above technical problems, the present invention provides a three-dimensional human pose estimation method, system, storage medium, and electronic device.
[0005] The technical solution of a three-dimensional human pose estimation method of the present invention is as follows:
[0006] Control each camera placed on each camera column to synchronously collect the original image data of the human body to be measured, and use the trained two-dimensional pose estimation model to obtain the two-dimensional human pose data corresponding to each original image data; wherein, all camera columns are arranged in a ring;
[0007] According to the target camera internal parameters, target camera external parameters, and two-dimensional human pose data of all cameras on each camera column, obtain the original three-dimensional human pose data of the human body to be measured under the viewpoints of each camera column respectively;
[0008] Perform three-dimensional skeleton joint length verification on each original three-dimensional human pose data to obtain multiple target three-dimensional human pose data, and perform three-dimensional human pose fusion on all target three-dimensional human pose data to obtain the three-dimensional human pose estimation result of the human body to be measured.
[0009] The beneficial effects of a three-dimensional human pose estimation method of the present invention are as follows:
[0010] The method of the present invention solves problems such as low-latency synchronization of multiple cameras, automatic and accurate calibration of multiple cameras, and accurate three-dimensional reconstruction of human poses, and can achieve accurate estimation of three-dimensional human poses.
[0011] On the basis of the above solution, a three-dimensional human pose estimation method of the present invention can also be improved as follows.
[0012] Further, the process of obtaining the target camera internal parameters of each camera is as follows:
[0013] Control each camera to respectively perform multiple synchronous acquisitions on the calibration board arranged at the center of the target area, and obtain a plurality of original camera calibration data collected by each camera; wherein, the target area is formed by circularly arranging all camera columns;
[0014] Use the checkerboard corner detection algorithm to detect all the original camera calibration data, determine each original camera calibration data that meets the preset conditions as the target camera calibration data, and obtain the corner position information corresponding to each target camera calibration data;
[0015] According to all the corner position information corresponding to any camera, perform monocular camera internal parameter calibration on the any camera to obtain the target camera internal parameters of the any camera, until the target camera internal parameters of each camera are obtained.
[0016] Further, the process of obtaining the target camera external parameters of each camera is as follows:
[0017] According to the preset calibration conditions, divide all cameras into multiple calibration groups, and determine the origin camera of each calibration group;
[0018] Perform binocular external parameter calibration on each pair of adjacent cameras in each calibration group to obtain the original camera external parameters of each camera relative to the origin camera of the corresponding calibration group, and use the LM algorithm to respectively optimize the minimum reprojection error of the original camera external parameters of each camera to obtain the first optimized camera external parameters of each camera;
[0019] Use local bundle adjustment to iteratively optimize the first optimized camera external parameters of each camera in each calibration group to obtain the intra-group optimized camera external parameters of each camera relative to the corresponding calibration group;
[0020] Among them, the process of local bundle adjustment is: use the least squares optimization of the cost function to accumulate errors to iteratively optimize the first optimized camera external parameters within each calibration group;
[0021] The cost function is: h(T i ,vij ) = K i T ipj , v ij is the coordinate of the j-th corner point in the i-th camera pixel coordinate system, T i is the transformation relationship between the i-th camera and the origin camera coordinate system, K i is the target camera internal parameter of the i-th camera, m is the number of all cameras in the corresponding calibration group, n is the number of co-viewing points in the corresponding calibration group, p j is the three-dimensional coordinate of the j-th corner point in the origin camera coordinate system;
[0022] Convert the camera coordinates of all cameras in each calibration group to the coordinates relative to the origin camera of the corresponding calibration group, and determine the origin camera of the first calibration group as the global origin camera;
[0023] Based on the first preset formula, convert the intra-group optimized camera external parameter of each camera to the target camera external parameter relative to the global origin camera; wherein, the first preset formula is:
[0024] is the first camera of the first calibration group, is the first camera of the fifth calibration group, is the first camera of the k-th calibration group, is the i-th camera of the k-th calibration group, represents the rigid body transformation relationship of camera C2 relative to the coordinate system of camera C1.
[0025] Furthermore, the monocular camera internal parameter calibration of any camera according to all corner point position information corresponding to the any camera to obtain the target camera internal parameter of the any camera includes:
[0026] Perform monocular camera internal parameter calibration on the any camera according to Zhang Zhengyou's planar calibration method and all corner point position information corresponding to the any camera to obtain the original camera internal parameter of the any camera;
[0027] Based on the PnP algorithm, calculate the first conversion relationship between the checkerboard and the camera coordinate system in each target camera calibration data corresponding to the any camera, and obtain the converted camera internal parameter of the any camera according to the first conversion relationship and the original camera internal parameter of the any camera;
[0028] Use the LM algorithm to optimize the minimized reprojection error of the converted camera internal parameter of the any camera to obtain the target camera internal parameter of the any camera.
[0029] Furthermore, the original three-dimensional human body pose data includes: three-dimensional coordinates of multiple joint points and visibility information of multiple joint points;
[0030] Obtaining the original three-dimensional human body pose data of the human body to be measured under the viewpoints of each camera column according to the target camera intrinsic parameters, target camera extrinsic parameters, and two-dimensional human body pose data of all cameras on each camera column, including:
[0031] According to the target camera intrinsic parameters, target camera extrinsic parameters, and two-dimensional human body pose data of all cameras on any one camera column, controlling all cameras on the any one camera column to perform triangulation to obtain the three-dimensional coordinates of multiple joint points and the visibility information of multiple joint points of the human body to be measured under the viewpoints of the any one camera column until the three-dimensional coordinates of multiple joint points and the visibility information of multiple joint points of the human body to be measured under the viewpoints of each camera column are obtained.
[0032] Furthermore, performing three-dimensional human body pose fusion on all the target three-dimensional human body pose data to obtain the three-dimensional human body pose estimation result of the human body to be measured, including:
[0033] According to the fusion skeleton calculation formula, the target camera extrinsic parameters corresponding to each camera, and all the target three-dimensional human body pose data, obtaining the fusion skeleton data of each joint point of the human body to be measured; wherein, the fusion skeleton calculation formula is: is the fusion skeleton data of the s-th joint point, is the s-th joint coordinate of the three-dimensional skeleton corresponding to the i-th viewpoint, is the corresponding weight, , θ i ∈ (0, 90), is the joint point visibility information, z i is the depth value, θ i is the joint angle;
[0034] According to the fusion skeleton data of all joint points of the human body to be measured and the parametric human body model, obtaining the three-dimensional human body pose estimation result of the human body to be measured; wherein, the parametric human body model is: E fused (o, β) = ω pro E pro + ω shape E shape + ω geo E geo , E pro represents aligning the two-dimensional projections on each view to the three-dimensional joints, E shape represents preprocessing the human body shape, E geo represents constraining the joint points by multi-view geometric consistency, ω pro is the first balance weight corresponding to E pro ω shape is the second balance weight corresponding to E shape ωgeo For E geo The corresponding third balance weight, θ is used to control the bone length, and β is used to control the posture of each joint.
[0035] The technical solution of a three-dimensional human pose estimation system of the present invention is as follows:
[0036] It includes: a processing module, a generation module, and an estimation module;
[0037] The processing module is used to: control each camera placed on each camera column to synchronously collect the original image data of the human body to be measured, and use the trained two-dimensional pose estimation model to obtain the two-dimensional human pose data corresponding to each original image data; wherein, all the camera columns are arranged in a ring;
[0038] The generation module is used to: obtain the original three-dimensional human pose data of the human body to be measured under the viewpoints of each camera column respectively according to the target camera internal parameters, target camera external parameters and two-dimensional human pose data of all the cameras on each camera column;
[0039] The estimation module is used to: perform three-dimensional skeleton joint length verification on each original three-dimensional human pose data to obtain multiple target three-dimensional human pose data, and perform three-dimensional human pose fusion on all the target three-dimensional human pose data to obtain the three-dimensional human pose estimation result of the human body to be measured.
[0040] The beneficial effects of a three-dimensional human pose estimation system of the present invention are as follows:
[0041] The system of the present invention solves problems such as multi-camera low-latency synchronization, automatic and accurate calibration of multi-cameras, and accurate three-dimensional reconstruction of human poses, and can achieve accurate estimation of three-dimensional human poses.
[0042] On the basis of the above solution, a three-dimensional human pose estimation system of the present invention can also be improved as follows.
[0043] Further, the process of obtaining the target camera internal parameters of each camera is as follows:
[0044] Control each camera to respectively perform multiple synchronous acquisitions on the calibration board arranged at the center of the target area to obtain multiple original camera calibration data collected by each camera; wherein, the target area is formed by the ring arrangement of all the camera columns;
[0045] Use the checkerboard corner detection algorithm to detect all the original camera calibration data, determine each original camera calibration data that meets the preset conditions as the target camera calibration data, and obtain the corner position information corresponding to each target camera calibration data;
[0046] Based on the position information of all corner points corresponding to any camera, perform monocular camera internal parameter calibration on the any camera to obtain the target camera internal parameters of the any camera, until the target camera internal parameters of each camera are obtained.
[0047] The technical solution of a storage medium of the present invention is as follows:
[0048] Instructions are stored in the storage medium. When the computer reads the instructions, the computer is made to execute the steps of a three-dimensional human pose estimation method of the present invention.
[0049] The technical solution of an electronic device of the present invention is as follows:
[0050] It includes a memory, a processor, and a computer program stored on the memory and executable on the processor. It is characterized in that when the processor executes the computer program, the computer is made to execute the steps of a three-dimensional human pose estimation method of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 It is a schematic flow chart of a three-dimensional human pose estimation method according to an embodiment of the present invention;
[0052] Figure 2 It is a schematic diagram of a camera acquisition area in a three-dimensional human pose estimation method according to an embodiment of the present invention;
[0053] Figure 3 It is a schematic diagram of a hardware system in a three-dimensional human pose estimation method according to an embodiment of the present invention;
[0054] Figure 4 It is a schematic diagram of camera group calibration in a three-dimensional human pose estimation method according to an embodiment of the present invention;
[0055] Figure 5 It is a schematic structural diagram of a three-dimensional human pose estimation system according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0056] As Figure 1 shown, a three-dimensional human pose estimation method according to an embodiment of the present invention includes the following steps:
[0057] S1. Control each camera placed on each camera column to synchronously collect the original image data of the human body to be measured, and use the trained two-dimensional pose estimation model to obtain the two-dimensional human pose data corresponding to each original image data.
[0058] Among them, as Figure 2As shown, in this embodiment, 32 high-frame-rate industrial cameras are adopted. Every two cameras are installed on a camera column, and every two camera columns form a camera array. Each camera array is connected to a server through an active hybrid fiber data cable, and 8 servers are networked through an Ethernet switch. One server is connected to each server (a total of eight servers (the first server 31 to the eighth server 38), one of which is the main server and the other seven are slave servers). All the camera columns are arranged in a ring to form a target area (shooting area). Specifically, as Figure 3 shown, the entire system is configured with two customized synchronization signal transmitting devices (the first synchronization device 10 and the second synchronization device 20). Each synchronization device is connected to 16 cameras through special trigger cables. At the same time, the two synchronization devices are cascaded to achieve synchronous triggering of 32 cameras (the first camera 101 to the thirty-second camera 132). The acquisition area of the entire system is covered by a special curtain, and a lighting system providing multiple light sources is configured.
[0059] Among them, the human body to be measured is located at the center of the target area. The original image data is the image data of the human body to be measured captured by the camera. The trained two-dimensional pose estimation model is a pre-trained human pose estimation neural network model used to estimate the two-dimensional human pose of the human body to be measured from the viewpoints of each camera.
[0060] It should be noted that for the two-dimensional pose estimation model, in this embodiment, the backbone neural network structure of the open-source human pose estimation framework OpenPose is optimized and trained using the constructed dataset. In actual use, the trained two-dimensional pose estimation model is used for multi-camera human pose estimation, which has better performance than the default OpenPose pre-trained model.
[0061] S2. According to the target camera internal parameters, target camera external parameters, and two-dimensional human pose data of all cameras on each camera column, obtain the original three-dimensional human pose data of the human body to be measured from the viewpoints of each camera column respectively.
[0062] Among them, the target camera internal parameter is the camera internal parameter obtained after multi-camera calibration. The target camera external parameter is the camera external parameter obtained after multi-camera calibration. The original three-dimensional human pose is the three-dimensional human pose generated according to the camera internal parameters, camera external parameters, and two-dimensional human pose data of all cameras on the camera column.
[0063] S3. Perform three-dimensional skeleton joint length verification on each original three-dimensional human pose data to obtain multiple target three-dimensional human pose data, and perform three-dimensional human pose fusion on all the target three-dimensional human pose data to obtain the three-dimensional human pose estimation result of the human body to be measured.
[0064] Among them, the process of performing three-dimensional skeleton joint length verification on each piece of original three-dimensional human pose data is as follows: Use the skeleton prior length template to detect the three-dimensional skeleton joint length for each perspective, exclude the joint points with abnormal coordinates, and if an abnormal joint point is detected, set the visibility of this joint point to 0.
[0065] Among them, the three-dimensional human pose estimation result is: the visualized three-dimensional human pose of the human body to be measured.
[0066] It should be noted that before performing pose estimation, the user can create corresponding templates according to the actual situation of the human body to be measured.
[0067] Preferably, the process of obtaining the target camera internal parameters of each camera is as follows:
[0068] Control each camera to perform multiple synchronous acquisitions on the calibration board set at the center of the target area respectively, and obtain multiple pieces of original camera calibration data collected by each camera; among them, the target area is formed by the annular arrangement of all camera columns.
[0069] Specifically, a planar checkerboard is used as the calibration tool. After all viewpoint cameras are triggered synchronously, the calibration personnel hold the checkerboard calibration board and rotate it in a circle at the center of the camera array (target area). At the same time, during the rotation process, the calibration board is translated up and down in a small range, tilted and rotated at a small angle. The pictures collected by each camera are named with "camera number + frame number". Each time the trigger emits a square wave signal, all cameras take pictures synchronously, and the frame number increases simultaneously. The calibration software ensures the synchronization of the calibration data of different cameras by the same frame number of the collected pictures.
[0070] It should be noted that the number of times of controlling all cameras to synchronously acquire calibration images is not limited and can be set according to user requirements.
[0071] Use the checkerboard corner detection algorithm to detect all the original camera calibration data, determine each piece of original camera calibration data that meets the preset conditions as the target camera calibration data, and obtain the corner position information corresponding to each target camera calibration data.
[0072] Among them, use the checkerboard corner detection algorithm provided by the open-source computer vision algorithm library OpenCV to detect all the original camera calibration data. The specific detection process is the prior art and will not be elaborated here.
[0073] Among them, the preset conditions are set by the corresponding checkerboard corner detection algorithm and are used to screen out valid camera calibration data.
[0074] Based on the position information of all corner points corresponding to any camera, perform monocular camera internal parameter calibration on the any camera to obtain the target camera internal parameters of the any camera until the target camera internal parameters of each camera are obtained.
[0075] Specifically, based on Zhang Zhengyou's planar calibration method and the position information of all corner points corresponding to any camera, perform monocular camera internal parameter calibration on the any camera to obtain the original camera internal parameters of the any camera.
[0076] Among them, when calibrating the internal parameters using Zhang Zhengyou's planar calibration method, it is assumed that there is no distortion.
[0077] Based on the PnP algorithm, calculate the first transformation relationship between the checkerboard and the camera coordinate system in each target camera calibration data corresponding to the any camera, and obtain the transformed camera internal parameters of the any camera according to the first transformation relationship and the original camera internal parameters of the any camera.
[0078] Among them, the first transformation relationship is: the rigid body transformation relationship between the checkerboard and the camera coordinate system.
[0079] Use the LM algorithm to optimize the minimized reprojection error of the transformed camera internal parameters of the any camera to obtain the target camera internal parameters of the any camera.
[0080] It should be noted that Zhang Zhengyou's planar calibration method, the PnP algorithm, and the LM algorithm are all existing technologies and will not be elaborated here.
[0081] Preferably, the process of obtaining the target camera external parameters of each camera is as follows:
[0082] According to the preset calibration conditions, divide all cameras into multiple calibration groups and determine the origin camera of each calibration group.
[0083] Among them, the preset calibration conditions are: divide all cameras into 5 calibration groups, with 8 cameras in each calibration group. The specific grouping schematic diagram is as Figure 4 shown.
[0084] Perform binocular external parameter calibration on each pair of adjacent cameras in each calibration group to obtain the original camera external parameters of each camera relative to the origin camera of the corresponding calibration group, and use the LM algorithm to optimize the minimized reprojection error of the original camera external parameters of each camera respectively to obtain the first optimized camera external parameters of each camera.
[0085] Specifically, for each pair of adjacent cameras in each calibration group, binocular extrinsic parameter calibration is performed respectively, and the original camera extrinsic parameters of each camera in the calibration group relative to the origin camera in the group are obtained by recursion. Among them, the process of binocular extrinsic parameter calibration is as follows: matching the corner point pairs with the same frame number of two cameras, calculating the coordinates of the undistorted corner points using the trial position method, solving the essential matrix E using the Ransanc algorithm, decomposing the E matrix to obtain the initial extrinsic parameters, and using the LM algorithm to minimize the reprojection error to optimize the original camera extrinsic parameters to obtain the first optimized camera extrinsic parameters of each camera.
[0086] Using local bundle adjustment to iteratively optimize the first optimized camera extrinsic parameters of each camera in each calibration group to obtain the intra-group optimized camera extrinsic parameters of each camera relative to the corresponding calibration group.
[0087] Among them, the process of local bundle adjustment is as follows: using the least squares optimization of the cost function to accumulate the error to iteratively optimize the first optimized camera extrinsic parameters within each calibration group.
[0088] The cost function is: h(T i , v ij ) = K i T i p j , v ij is the coordinate of the jth corner point in the pixel coordinate system of the ith camera, T i is the transformation relationship between the ith camera and the origin camera coordinate system, K i is the target camera intrinsic parameter of the ith camera, m is the number of all cameras in the corresponding calibration group, n is the number of co-visible points in the corresponding calibration group, and p j is the three-dimensional coordinate of the jth corner point in the origin camera coordinate system.
[0089] It should be noted that the co-visible points refer to the same corner points recognized by different viewpoints at the same time point. In this embodiment, the number of co-visible points is the comprehensive number of co-visible points in all calibration frames with co-visible points among the cameras in a calibration group. In addition, when solving, Schur elimination is used to accelerate the calculation. After the iteration ends, the intra-group optimized camera extrinsic parameters of each camera relative to the corresponding calibration group are obtained.
[0090] Convert the camera coordinates of all cameras in each calibration group to the coordinates relative to the origin camera of the corresponding calibration group, and determine the origin camera of the first calibration group as the global origin camera.
[0091] Among them, the origin camera of the first calibration group is the first camera.
[0092] Based on the first preset formula, the intra-group optimized camera extrinsic parameters of each camera are converted into target camera extrinsic parameters relative to the global origin camera. The first preset formula is as follows:
[0093] is the first camera of the first calibration group, is the first camera of the fifth calibration group, is the first camera of the k-th calibration group, is the i-th camera of the k-th calibration group; T(C1, C2) represents the rigid body transformation relationship of camera C2 relative to the coordinate system of camera C1.
[0094] Preferably, the original three-dimensional human pose data includes: three-dimensional coordinates of multiple joint points and visibility information of multiple joint points.
[0095] The obtaining of the original three-dimensional human pose data of the human body to be measured at the viewpoints of each camera column according to the target camera intrinsic parameters, target camera extrinsic parameters and two-dimensional human pose data of all cameras on each camera column includes:
[0096] According to the target camera intrinsic parameters, target camera extrinsic parameters and two-dimensional human pose data of all cameras on any camera column, controlling all cameras of the any camera column to perform triangulation to obtain three-dimensional coordinates of multiple joint points and visibility information of multiple joint points of the human body to be measured at the viewpoints of the any camera column until three-dimensional coordinates of multiple joint points and visibility information of multiple joint points of the human body to be measured at the viewpoints of each camera column are obtained.
[0097] Among them, the visibility information of the detected human joint points is 1, and the visibility information of the undetected human joint points is 0.
[0098] Preferably, the performing of three-dimensional human pose fusion on all the target three-dimensional human pose data to obtain the three-dimensional human pose estimation result of the human body to be measured includes:
[0099] According to the fusion skeleton calculation formula, the target camera extrinsic parameters corresponding to each camera and all the target three-dimensional human pose data, obtaining the fusion skeleton data of each joint point of the human body to be measured; among them, the fusion skeleton calculation formula is: is the fusion skeleton data of the s-th joint point, is the s-th joint coordinate of the three-dimensional skeleton corresponding to the i-th viewpoint, is the corresponding weight, , θ i ∈(0, 90), is the joint visibility information, z i is the depth value, θ iis the joint angle.
[0100] It should be noted that the depth value is the distance from the joint point to the camera, and the joint angle is the angle between the joint surface and the main optical axis of the viewing point. Among them, T i is the transformation relationship between the i-th camera and the origin camera coordinate system; i ∈ [1, 16]. Since the trained model contains 25 human joint points, s ∈ [1, 25].
[0101] According to the fused skeleton data and the parametric human model of all joint points of the human body to be measured, obtain the three-dimensional human body pose estimation result of the human body to be measured; among them, the parametric human model is: E fused (θ, β) = ω pro E pro + ω shape E shape + ω geo E geo , E pro represents aligning the two-dimensional projection on each view to the three-dimensional joint, E shape represents preprocessing the human body shape, E geo represents constraining the joint points by multi-view geometric consistency, ω pro is the first balance weight corresponding to E pro , ω shape is the second balance weight corresponding to E shape , ω geo is the third balance weight corresponding to E geo , θ is used to control the bone length, and β is used to control the pose of each joint.
[0102] The technical solution of this embodiment solves problems such as multi-camera low-latency synchronization, automatic and accurate calibration of multi-cameras, and accurate three-dimensional reconstruction of human body poses, and can achieve precise estimation of three-dimensional human body poses.
[0103] As Figure 2 shown, a three-dimensional human body pose estimation system 200 according to an embodiment of the present invention includes: a processing module 210, a generation module 220, and an estimation module 230;
[0104] The processing module 210 is configured to: control each camera placed on each camera column to synchronously collect the original image data of the human body to be measured, and use the trained two-dimensional pose estimation model to obtain the two-dimensional human body pose data corresponding to each original image data; among them, all the camera columns are arranged in a ring;
[0105] The generation module 220 is configured to: obtain the original three-dimensional human body pose data of the human body to be measured under the viewing points of each camera column respectively according to the target camera internal parameters, target camera external parameters, and two-dimensional human body pose data of all the cameras on each camera column;
[0106] The estimation module 230 is configured to: perform three-dimensional skeleton joint length verification on each piece of original three-dimensional human body pose data to obtain a plurality of target three-dimensional human body pose data, and perform three-dimensional human body pose fusion on all the target three-dimensional human body pose data to obtain the three-dimensional human body pose estimation result of the human body to be measured.
[0107] Preferably, the process of obtaining the target camera internal parameters of each camera is as follows:
[0108] Control each camera to perform multiple synchronous acquisitions on a calibration board set at the center of the target area respectively to obtain a plurality of original camera calibration data collected by each camera; wherein, the target area is formed by arranging all the camera columns in a ring;
[0109] Use the checkerboard corner detection algorithm to detect all the original camera calibration data, determine each piece of original camera calibration data that meets the preset conditions as the target camera calibration data, and obtain the corner position information corresponding to each target camera calibration data;
[0110] According to all the corner position information corresponding to any one camera, perform monocular camera internal parameter calibration on the any one camera to obtain the target camera internal parameters of the any one camera until the target camera internal parameters of each camera are obtained.
[0111] The technical solution of this embodiment solves problems such as multi-camera low-latency synchronization, automatic and accurate calibration of multi-cameras, and accurate three-dimensional reconstruction of human body poses, and can achieve accurate estimation of three-dimensional human body poses.
[0112] For the parameters and the steps for each module in the three-dimensional human body pose estimation system 200 of this embodiment to implement corresponding functions, reference can be made to the parameters and steps in the embodiment of the three-dimensional human body pose estimation method in the above text, which will not be elaborated here.
[0113] A storage medium provided by an embodiment of the present invention includes: instructions are stored in the storage medium, and when a computer reads the instructions, the computer is made to execute the steps of a three-dimensional human body pose estimation method. Specifically, reference can be made to the parameters and steps in the embodiment of the three-dimensional human body pose estimation method in the above text, which will not be elaborated here.
[0114] Computer storage media such as: USB flash drives, external hard drives, etc.
[0115] An electronic device provided by an embodiment of the present invention includes a memory, a processor, and a computer program stored on the memory and executable on the processor. It is characterized in that when the processor executes the computer program, the computer is made to execute the steps of a three-dimensional human pose estimation method. Specifically, reference can be made to the various parameters and steps in the embodiments of the three-dimensional human pose estimation method described above, which will not be elaborated here.
[0116] Those skilled in the art of the present technology know that the present invention can be implemented as a method, a system, a storage medium, and an electronic device.
[0117] Therefore, the present invention can be specifically implemented in the following forms: it can be entirely hardware, entirely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software, generally referred to as "circuit", "module", or "system" in this document. In addition, in some embodiments, the present invention can also be implemented in the form of a computer program product in one or more computer-readable media, which contain computer-readable program code. Any combination of one or more computer-readable media can be adopted. The computer-readable media can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium can be any tangible medium that contains or stores a program, which can be used by or in combination with an instruction execution system, device, or component. Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. A three-dimensional human body pose estimation method, characterized in that Including: Controlling each camera placed on each camera column to synchronously collect the original image data of the human body to be measured, and using the trained two-dimensional pose estimation model to obtain the two-dimensional human body pose data corresponding to each original image data; wherein, all the camera columns are arranged in a ring; According to the target camera intrinsic parameters, target camera extrinsic parameters and two-dimensional human body pose data of all the cameras on each camera column, obtaining the original three-dimensional human body pose data of the human body to be measured under the viewpoints of each camera column respectively; Performing three-dimensional skeleton joint length verification on each original three-dimensional human body pose data to obtain a plurality of target three-dimensional human body pose data, and performing three-dimensional human body pose fusion on all the target three-dimensional human body pose data to obtain the three-dimensional human body pose estimation result of the human body to be measured; The process of performing three-dimensional human body pose fusion on all the target three-dimensional human body pose data to obtain the three-dimensional human body pose estimation result of the human body to be measured includes: According to the fused skeleton calculation formula, the target camera extrinsic parameters corresponding to each camera, and all the target three-dimensional human pose data, the fused skeleton data of each joint point of the human body to be measured is obtained; wherein, the fused skeleton calculation formula is: is the fused skeleton data of the s-th joint point, is the s-th joint coordinate of the three-dimensional skeleton corresponding to the i-th viewpoint, is the corresponding weight, θ i ∈(0, 90), is the joint visibility information, z i is the depth value, θ i is the joint angle; Based on the fused skeleton data and the parametric human model of all joint points of the human body to be measured, obtain the three-dimensional human body pose estimation result of the human body to be measured; wherein, the parametric human model is: E fused (θ, β) = ω pro E pro + ω shape E shape + ω geo E geo , E pro represents aligning the two-dimensional projection on each view to the three-dimensional joints, E shape represents preprocessing the human body form, E geo represents constraining the joint points by multi-view geometric consistency, ω pro is the first balance weight corresponding to E pro , ω shape is the second balance weight corresponding to E shape , ω geo is the third balance weight corresponding to E geo , θ is used to control the bone length, β is used to control the pose of each joint, T i is the transformation relationship between the i-th camera and the origin camera coordinate system.
2. The three-dimensional human body pose estimation method according to claim 1, wherein, The process of obtaining the target camera intrinsic parameters of each camera is as follows: Controlling each camera to perform multiple synchronous acquisitions on the calibration board set at the center of the target area respectively, to obtain a plurality of original camera calibration data collected by each camera; wherein, the target area is formed by the ring arrangement of all the camera columns; Using the checkerboard corner detection algorithm to detect all the original camera calibration data, determining each original camera calibration data that meets the preset conditions as the target camera calibration data, and obtaining the corner position information corresponding to each target camera calibration data; According to all the corner position information corresponding to any one camera, performing monocular camera intrinsic parameter calibration on the any one camera to obtain the target camera intrinsic parameter of the any one camera, until the target camera intrinsic parameters of each camera are obtained.
3. The three-dimensional human body pose estimation method according to claim 2, wherein, The process of obtaining the target camera extrinsic parameters of each camera is as follows: According to the preset calibration conditions, dividing all the cameras into multiple calibration groups, and determining the origin camera of each calibration group; Performing binocular extrinsic parameter calibration on each pair of adjacent cameras in each calibration group respectively to obtain the original camera extrinsic parameters of each camera relative to the origin camera of the corresponding calibration group, and using the LM algorithm to optimize the minimum reprojection error of the original camera extrinsic parameters of each camera respectively to obtain the first optimized camera extrinsic parameters of each camera; Using local bundle adjustment to iteratively optimize the first optimized camera extrinsic parameters of each camera in each calibration group to obtain the intra-group optimized camera extrinsic parameters of each camera relative to the corresponding calibration group; Wherein, the process of local bundle adjustment is: using the least squares optimization cumulative error of the cost function to iteratively optimize the first optimized camera extrinsic parameters within each calibration group; The cost function is as follows: h(T i ,v ij ) = K i T i p j ,v ij is the coordinate of the j-th corner point in the i-th camera pixel coordinate system, T i is the transformation relationship between the i-th camera and the origin camera coordinate system, K i is the target camera internal parameter of the i-th camera, m is the number of all cameras in the corresponding calibration group, n is the number of co-viewpoint points in the corresponding calibration group, p j is the three-dimensional coordinate of the j-th corner point in the origin camera coordinate system; Converting the camera coordinates of all the cameras in each calibration group to the coordinates relative to the origin camera of the corresponding calibration group, and determining the origin camera of the first calibration group as the global origin camera; Based on the first preset formula, converting the intra-group optimized camera extrinsic parameters of each camera to the target camera extrinsic parameters relative to the global origin camera; wherein, the first preset formula is: is the first camera of the first calibration group, is the first camera of the fifth calibration group, is the first camera of the k-th calibration group, is the i-th camera of the k-th calibration group; T(C1, C2) represents the rigid body transformation relationship of camera C2 relative to the coordinate system of camera C1.
4. A three-dimensional human body pose estimation method according to claim 3, characterized in that, The process of performing monocular camera intrinsic parameter calibration on the any one camera according to all the corner position information corresponding to the any one camera to obtain the target camera intrinsic parameter of the any one camera includes: According to the Zhang-Zhengyou planar calibration method and the position information of all corner points corresponding to any one of the cameras, perform monocular camera intrinsic parameter calibration on any one of the cameras to obtain the original camera intrinsic parameters of any one of the cameras; Based on the PnP algorithm, calculate the first conversion relationship between the checkerboard and the camera coordinate system in each target camera calibration data corresponding to any one of the cameras, and obtain the converted camera intrinsic parameters of any one of the cameras according to the first conversion relationship and the original camera intrinsic parameters of any one of the cameras; Use the LM algorithm to optimize the minimized reprojection error of the converted camera intrinsic parameters of any one of the cameras to obtain the target camera intrinsic parameters of any one of the cameras.
5. A three-dimensional human body pose estimation method according to any one of claims 1-4, characterized in that, The original three-dimensional human pose data includes: three-dimensional coordinates of multiple joint points and visibility information of multiple joint points; The process of obtaining the original three-dimensional human pose data of the human body to be measured at the viewpoints of each camera column according to the target camera intrinsic parameters, target camera extrinsic parameters, and two-dimensional human pose data of all cameras on each camera column includes: According to the target camera intrinsic parameters, target camera extrinsic parameters, and two-dimensional human pose data of all cameras on any one camera column, control all cameras on any one camera column to perform triangulation to obtain the three-dimensional coordinates of multiple joint points and visibility information of multiple joint points of the human body to be measured at the viewpoints of any one camera column respectively, until the three-dimensional coordinates of multiple joint points and visibility information of multiple joint points of the human body to be measured at the viewpoints of each camera column are obtained.
6. A three-dimensional human body pose estimation system, characterized in that, Including: A processing module, a generation module, and an estimation module; The processing module is used to: control each camera placed on each camera column to synchronously collect the original image data of the human body to be measured, and use the trained two-dimensional pose estimation model to obtain the two-dimensional human pose data corresponding to each original image data; wherein, all camera columns are arranged in a ring; The generation module is used to: obtain the original three-dimensional human pose data of the human body to be measured at the viewpoints of each camera column according to the target camera intrinsic parameters, target camera extrinsic parameters, and two-dimensional human pose data of all cameras on each camera column; The estimation module is used to: perform three-dimensional skeleton joint length verification on each original three-dimensional human pose data to obtain multiple target three-dimensional human pose data, and perform three-dimensional human pose fusion on all target three-dimensional human pose data to obtain the three-dimensional human pose estimation result of the human body to be measured; The process of performing three-dimensional human pose fusion on all target three-dimensional human pose data in the estimation module to obtain the three-dimensional human pose estimation result of the human body to be measured includes: According to the fused skeleton calculation formula, the extrinsic parameters of the target camera corresponding to each camera, and all the target three-dimensional human pose data, the fused skeleton data of each joint point of the human body to be measured is obtained; wherein, the fused skeleton calculation formula is: is the fused skeleton data of the s-th joint point, is the s-th joint coordinate of the three-dimensional skeleton corresponding to the i-th viewpoint, is the corresponding weight, θ i ∈(0, 90), is the joint visibility information, z i is the depth value, θ i is the joint angle; Based on the fused skeleton data and the parametric human model of all joint points of the human body to be measured, obtain the three-dimensional human body pose estimation result of the human body to be measured; wherein, the parametric human model is: E fused (θ, β) = ω pro E pro + ω shape E shape + ω geo E geo , E pro represents aligning the two-dimensional projection on each view to the three-dimensional joints, E shape represents preprocessing the human body shape, E geo represents constraining the joint points by multi-view geometric consistency, ω pro is the first balance weight corresponding to E pro , ω shape is the second balance weight corresponding to E shape , ω geo is the third balance weight corresponding to E geo , θ is to control the bone length, β is to control the pose of each joint, T i is the transformation relationship between the i-th camera and the origin camera coordinate system.
7. A three-dimensional human body pose estimation system according to claim 6, wherein The process of obtaining the target camera intrinsic parameters of each camera is: Control each camera to perform multiple synchronous acquisitions on the calibration board set at the center of the target area respectively to obtain multiple original camera calibration data collected by each camera; wherein, the target area is formed by the circular arrangement of all camera columns; Use the checkerboard corner point detection algorithm to detect all the original camera calibration data, determine each original camera calibration data that meets the preset conditions as the target camera calibration data, and obtain the corner point position information corresponding to each target camera calibration data. According to all the corner point position information corresponding to any one camera, perform monocular camera internal parameter calibration on the any one camera to obtain the target camera internal parameters of the any one camera until the target camera internal parameters of each camera are obtained.
8. A storage medium, characterized in that, Instructions are stored in the storage medium, and when the computer reads the instructions, the computer is caused to execute a three-dimensional human pose estimation method according to any one of claims 1 to 5.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, the computer is caused to execute a three-dimensional human pose estimation method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Human body three-dimensional posture estimation method based on multi-view fusion
CN114529605A
Method for reconstructing three-dimensional human body by using mirror
WO2022165635A1