Feedback-based 3D human body posture reconstruction method and lighting system

By associating the predicted 3D position with the actual 2D position and combining multi-view geometry parameter optimization and Kalman filter, the pose determination error caused by a limited number of cameras or occlusion is resolved, achieving more accurate and stable 3D human pose reconstruction.

CN120672934APending Publication Date: 2025-09-19GUANGZHOU HAOYANG ELECTRONICS CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510575147.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

When using cameras to shoot the human body for continuous posture determination, due to the limited number of cameras or the occlusion of key points of the human body, large deviations may occur in the human posture determination, affecting the accuracy.

Method used

By back-projecting the predicted 3D position at time t back to the camera image coordinate system and associating it with the actual 2D position, a 2D joint point set is formed. The multi-view geometric parameter optimization method and Kalman filter are used for fusion and optimization to obtain a more accurate 3D human key point position.

Benefits of technology

The robustness and accuracy of human pose reconstruction are improved, the error under occlusion problem is reduced, and the position of human key points can be estimated stably and accurately even when the number of cameras is small.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120672934A_ABST
    Figure CN120672934A_ABST
Patent Text Reader

Abstract

The invention discloses a 3D human body posture reconstruction method and a lighting system based on feedback, and the 3D human body posture reconstruction method comprises the steps: carrying out the inverse projection of a predicted 3D position of a human body key point at a t moment, associating the predicted 3D position with an actual 2D position at the t moment to form a 2D joint point set of the human body key point, and calculating a plurality of 3D joint points corresponding to the human body key point according to the 2D joint point set; and fusing the 3D joint point clouds by using a fusion algorithm to obtain fused 3D positions of the key points of the human body, and optimizing the fused 3D positions by using a multi-view geometric parameter optimization method to obtain estimated 3D positions of the key points of the human body at the t moment. The 3D position estimation is more accurate and more reliable and is close to the real situation, the shielding problem can be relieved to a certain degree, and a more stable and more accurate estimation effect can be generated when the number of cameras is small.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of human body posture determination, and more particularly to a feedback-based 3D human body posture reconstruction method and lighting system. Background Art

[0002] When continuously determining the pose of a person using a camera to capture the human body, the state information of key points at the previous moment is often used to predict the state information of key points at the current moment. This is done to narrow the recognition range of key points in the current shot or to evaluate the accuracy of the positions of key points captured and identified. However, due to the limited number of cameras or occlusion of key points, continuous pose determination can lead to gradually large deviations from the actual state, resulting in inaccurate pose determination. Summary of the Invention

[0003] In order to overcome at least one of the defects described in the above-mentioned prior art, the present invention provides a feedback-based 3D human posture reconstruction method, which reverse-projects the predicted 3D position at time t and associates it with the actual 2D position, and obtains multiple 3D joint points corresponding to the key points of the human body. After fusing and optimizing them, the estimated 3D position of the key points of the human body at time t is obtained. The estimated 3D position is more accurate, more reliable, and closer to the actual situation.

[0004] To solve the above technical problems, the present invention adopts a technical solution: a feedback-based 3D human posture reconstruction method, comprising the following steps: S1, based on the estimated 3D position and motion information of the human body key point at time t-1, predict the predicted 3D position of the human body key point at time t; S2, associating the multiple actual 2D positions of the key points of the human body captured by the camera at time t with the predicted 3D positions by back-projecting them back into the image coordinate system of each camera to obtain multiple predicted 2D positions, to form a 2D joint point set of the key points of the human body at time t; S3, according to the 2D joint point set of the key point of the human body at time t, obtain the 3D joint point cloud of the key point of the human body at time t; S4, fusing the 3D joint point cloud using a fusion algorithm to obtain the fused 3D position of the key points of the human body; S5. Using the multi-view geometric parameter optimization method, the fused 3D position is optimized by combining multiple actual 2D positions of the key points of the human body captured by the camera at time t to obtain the estimated 3D position of the key points of the human body at time t.

[0005] The feedback-based 3D human posture reconstruction method obtains more possible positions of 2D joint points corresponding to a human key point by reverse-projecting the predicted 3D position at time t back into the image coordinate system of each camera and correlating it with the actual 2D position. This feedback alleviates the reconstruction error caused by the failure to recognize the actual 2D position of the human key point at the current moment to a certain extent, thereby improving the overall robustness; more 2D joint points can make the 3D joint points denser, and the fused 3D position after fusion is also more accurate, and the fused 3D position is globally optimized using a multi-view geometric parameter optimization method to obtain an estimated 3D position, further enhancing the accuracy and stability of human key point estimation; this scheme can alleviate the occlusion problem to a certain extent, and produce a more stable and accurate estimation effect when the number of cameras is small.

[0006] Furthermore, between steps S2 and S3, a verification step is included: using the epipolar RANAC algorithm to verify two 2D joint points from any two image coordinate systems in the 2D joint point set, and excluding pairs of 2D joint points with significant deviations. Due to noise, occlusion, or other factors, the acquired positions of human key points may contain errors, resulting in inaccurate associations. Using the epipolar RANAC algorithm to exclude pairs of 2D joint points with significant deviations not only reduces the computational effort required to reconstruct 3D joint points, but also reduces the noise in the 3D joint point cloud, thereby providing a more accurate initial value estimate for subsequent reconstruction.

[0007] Furthermore, in the verification step, the Sampson error is used to measure the degree of deviation of two 2D joint points of any two image coordinate systems under the epipolar geometry constraint. , where F is a 3×3 basic matrix representing the projection relationship between two camera perspectives, x1 and x2 correspond to the homogeneous coordinates of two 2D joint points in any two image coordinate systems, and it is considered that When the error is less than a predetermined threshold, the deviation is acceptable. The Sampson error measures the degree of deviation between matching point pairs x1 and x2 under epipolar geometry constraints. This metric takes into account the geometric error caused by the fundamental matrix F and is computationally more efficient and more stable than other error metrics, especially when processing large-scale data.

[0008] Furthermore, the predetermined threshold is γ 2 , γ is three times the mean of the standard deviations of the image noise of all cameras. That is, after determining the image noise of each camera, the standard deviation of the noise of each image is calculated, ultimately obtaining the mean of all standard deviations. The predetermined threshold is the square of three times this mean.

[0009] Furthermore, step S3 specifically involves combining two 2D joint points from any two image coordinate systems in the 2D joint point set to form a 3D joint point cloud. This method can generate as many 3D joint points as possible, making the 3D joint point cloud denser.

[0010] Furthermore, in step S4, the 3D joint point clouds are fused using a gain Kalman filter to obtain the fused 3D positions of the body's key points. This method has the advantage of leveraging state information at time t-1 to provide a priori values ​​for estimating the positions of the body's key points at time t, thereby enabling dynamic prediction of the positions of the body's key points at time t. Furthermore, the Kalman filter effectively models noise in both the system and observation processes, and adaptively adjusts the filter gain to suppress measurement errors and noise interference. It also corrects for potential outliers, preventing them from negatively impacting the pose estimation results.

[0011] Furthermore, the S4 step is as follows: record the 3D joint point cloud as ,in represents the coordinates of the vth 3D joint point (1≤v≤N), then the mean of the 3D joint point cloud is , the covariance matrix is , where Var(x v ) 、Var(y v ) and Var(z v ) represent x v 、y v 、z v The variance of represents the measurement noise of the vth 3D joint point coordinate in the corresponding direction, Cov(·,·) represents the covariance between the components of the vth 3D joint point on the corresponding two axes and the components of the other 3D joint points on the corresponding two axes; Based on the state information of the key points of the human body at time t-1, the predicted state at time t is obtained , the predicted covariance , where A is the state transition matrix, is the state information at time t-1, is the covariance matrix at time t-1, Q is the process noise covariance matrix, then the Kalman gain is , where H is the measurement matrix, and its structure is H=[I 3×3 0 3×3 ], I 3×3 is a 3×3 identity matrix, 0 3×3 is a 3×3 zero matrix; and the state information at time t , the covariance matrix , fused 3D position .

[0012] Obtaining the mean of the 3D joint point cloud can achieve data smoothing, thereby reducing random errors; the covariance matrix R is used to characterize the uncertainty of the measurement data, enabling the system to better adapt to observation noise and improve the robustness and accuracy of data processing.

[0013] Furthermore, in step S5, the multi-view geometric parameter optimization method is a bundle adjustment method. The bundle adjustment method optimizes the position estimates of 3D points by minimizing the reprojection error of the fused 3D positions. Specifically, it incorporates the geometric constraints between images captured by different cameras and employs nonlinear optimization techniques. The aim is to find an optimal solution that maintains geometric consistency across multiple views by minimizing the error between viewpoints.

[0014] Furthermore, the S5 step is specifically as follows: optimize the fused 3D position so that the loss function The position at the minimum is used as the estimated 3D position, where x j Represents the coordinates of the corresponding position in the image coordinate system, δ(x j ) is the indicator function, indicating x j The confidence level of detection in the image coordinate system, P j represents the transformation matrix that projects the fused 3D position into the image coordinate system. Due to detection and projection errors, there is a certain deviation between the detected keypoints and the inverse projection of the 3D point. By minimizing this loss function, the bundle adjustment method can further optimize the position estimates of the 3D points while ensuring geometric consistency. This optimization process uses an iterative algorithm, updating the camera parameters and the coordinates of the 3D points through a nonlinear least squares method in each iteration, ultimately converging to an optimal solution and significantly improving the positioning accuracy of the 3D points.

[0015] Furthermore, after step S5, step S6 is further included: combining the estimated 3D positions of the key points of the human body at time t-1 with the motion information of the key points of the human body at time t, so as to facilitate the subsequent prediction of the 3D positions of the key points of the human body at time t+1.

[0016] The present invention also provides a lighting system comprising a camera for capturing images of a human body, a processor for constructing a 3D human pose using any of the aforementioned feedback-based 3D human pose reconstruction methods, and a lamp for projecting light based on the 3D human pose. The lamp can estimate the 3D position of key points of the human body, accurately projecting light, providing efficient interaction and an enhanced audience experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is a schematic diagram of the process structure of the feedback-based 3D human posture reconstruction method of the present invention.

[0018] Figure 23D position estimation process of the human body key points at time t is shown in FIG.

[0019] Figure 3 This is a schematic diagram of the principle of polarity error.

[0020] Figure 4 It is a schematic diagram comparing the 3D joint point cloud and the actual 3D position of the key points of the human body at time t.

[0021] Figure 5 This is a schematic diagram of the principle of bundle adjustment method.

[0022] Figure 6 It is a structural schematic diagram of the lighting system of the present invention. DETAILED DESCRIPTION

[0023] The drawings are for illustrative purposes only and should not be construed as limiting this patent. To better illustrate the embodiments, some components in the drawings may be omitted, enlarged, or reduced in size, and do not represent actual product dimensions. Those skilled in the art will understand that some well-known structures and their descriptions may be omitted from the drawings. The positional relationships depicted in the drawings are for illustrative purposes only and should not be construed as limiting this patent.

[0024] like Figures 1 to 2 The present invention provides a feedback-based 3D human posture reconstruction method, comprising the following steps: S1, based on the estimated 3D position and motion information of the human body key point at time t-1, predict the predicted 3D position of the human body key point at time t; S2, associating the multiple actual 2D positions of the key points of the human body captured by the camera at time t with the predicted 3D positions by back-projecting them back into the image coordinate system of each camera to obtain multiple predicted 2D positions, to form a 2D joint point set of the key points of the human body at time t; S3, according to the 2D joint point set of the key point of the human body at time t, obtain the 3D joint point cloud of the key point of the human body at time t; S4, fusing the 3D joint point cloud using a fusion algorithm to obtain the fused 3D position of the key points of the human body; S5. Using the multi-view geometric parameter optimization method, the fused 3D position is optimized by combining multiple actual 2D positions of the key points of the human body captured by the camera at time t to obtain the estimated 3D position of the key points of the human body at time t.

[0025] The feedback-based 3D human posture reconstruction method obtains more possible positions of 2D joint points corresponding to a human key point by reverse-projecting the predicted 3D position at time t back into the image coordinate system of each camera and correlating it with the actual 2D position. This feedback alleviates the reconstruction error caused by the failure to recognize the actual 2D position of the human key point at the current moment to a certain extent, thereby improving the overall robustness; more 2D joint points can make the 3D joint points denser and the fused 3D position more accurate, and the multi-view geometric parameter optimization method is used to globally optimize the fused 3D position to obtain the estimated 3D position, further enhancing the accuracy and stability of the human key point estimation; this scheme can alleviate the occlusion problem to a certain extent, and produce a more stable and accurate estimation effect when the number of cameras is small.

[0026] In this embodiment, the motion information includes speed and direction.

[0027] Steps S1 and S2 are as follows: Assume that the key points of the individual body at time t-1 , corresponding to multiple actual 2D positions at time t, using the extended Kalman filter In the camera's image coordinate system, the predicted 3D position inverse projection set at time t is obtained , and the set of multiple actual 2D positions , according to the human body key point labels, the corresponding human body key points are extracted, and the 2D joint point set can be expressed as , by merging the data in the set, the reconstructed set can be expressed as .

[0028] It should be noted that the predicted 3D position and motion information at the third moment is calculated based on the actual 3D position of the human body's key points at the first and second moments. The predicted 3D position and motion information at the third moment is used as its estimated 3D position and motion information, and subsequent calculations are performed. Therefore, only at the fourth moment will its true estimated 3D position and motion information be obtained, so t≥5 in this application.

[0029] In a preferred embodiment of the present invention, a verification step is further included between step S2 and step S3: using the epipolar RANAC algorithm (i.e., epipolar random sampling consensus algorithm) to verify two 2D joint points from any two image coordinate systems in the 2D joint point set, and excluding a pair of 2D joint points with a large degree of deviation. Due to noise, occlusion or other factors, the positions of the key points of the human body may have errors, resulting in inaccurate association relationships, such as Figure 3, intuitively demonstrating the presence of epipolar error when matching points deviate from their ideal corresponding positions. Here, l1 and l2 are the epipolar lines of x1 and x2, respectively. Using the epipolar RANAC algorithm to exclude pairs of 2D joint points with significant deviations (however, any of these 2D joint points can be rematched and verified with the other 2D joint points) not only reduces the computational effort when reconstructing 3D joint points, but also reduces the noise in the 3D joint point cloud, thereby providing a more accurate initial value estimate for subsequent reconstructions.

[0030] Figure 3 In [1], the epipolar geometry between two views describes how the projections of the same 3D point on the image plane from the two camera perspectives constrain each other. The epipolar geometry is usually expressed through a fundamental matrix F.

[0031] Specifically, the fundamental matrix F is a 3×3 matrix that describes the projection relationship between the two camera views. Ideally, for points x1 and x2 in two given images (corresponding to the points in the first and second views, respectively), they must satisfy the following relationship: ; Where x1 and x2 are their secondary coordinates on the image plane. This equation states that the relationship between a point x2 in the second view and a point x1 in the first view is restricted, i.e., for each pair of matching points, the fundamental matrix provides a geometric constraint.

[0032] The geometric significance of this relationship lies in the fact that the epipolar point is the intersection of the line of sight of a point in one view and the camera optical axis in another view. All corresponding points in the second view must lie on an epipolar line. The fundamental matrix describes the correspondence between these points and lines.

[0033] Due to the existence of noise and measurement errors in practical applications, the above equation Therefore, it is necessary to introduce an error metric to quantify the degree of deviation of matching points under the epipolar geometry constraint.

[0034] In a preferred embodiment of the present invention, in the verification step, the Sampson error is used to measure the degree of deviation of two 2D joint points of any two image coordinate systems under the epipolar geometry constraint. , where F is a 3×3 basic matrix representing the projection relationship between two camera perspectives, x1 and x2 correspond to the homogeneous coordinates of two 2D joint points in any two image coordinate systems, and it is considered that When the error is less than a predetermined threshold, the deviation is acceptable. The Sampson error measures the degree of deviation between matching point pairs x1 and x2 under epipolar geometry constraints. This metric takes into account the geometric error caused by the fundamental matrix F and is computationally more efficient and more stable than other error metrics, especially when processing large-scale data.

[0035] In a preferred embodiment of the present invention, the predetermined threshold is γ 2 , γ is three times the mean of the image noise standard deviations of all camera images. Specifically, after determining the image noise for each camera image, the standard deviation of each image noise is calculated, ultimately averaging all standard deviations. The predetermined threshold is the square of three times this mean. If the Sampson error between two 2D joint points in any two image coordinate systems exceeds the predetermined threshold, that pair will not be used in subsequent 3D joint point calculations.

[0036] In a preferred embodiment of the present invention, step S3 specifically comprises combining two 2D joint points from any two image coordinate systems in the 2D joint point set to form a 3D joint point cloud of all the inferred 3D joint points. Using triangulation, two 2D joint points from any two image coordinate systems can be used to obtain corresponding 3D joint points. This method maximizes the number of 3D joint points, making the 3D joint point cloud denser and facilitating the precise determination of key points on the human body.

[0037] like Figure 4 After the predicted 3D position at time t-1 is fed back, the obtained 2D joint point set is used to generate the 3D joint point cloud at time t. Compared with the actual 3D position obtained by direct camera shooting and recognition at time t (calculated from multiple actual 2D positions), the number is significantly denser, thereby improving the accuracy of the position of key points of the human body.

[0038] In a preferred embodiment of the present invention, in step S4, a gain Kalman filter is used to fuse the 3D joint point clouds to obtain the fused 3D positions of the body's key points. This method has the advantage of leveraging state information at time t-1 to provide a priori values ​​for estimating the positions of the body's key points at time t, thereby enabling dynamic prediction of the positions of the body's key points at time t. Furthermore, the Kalman filter can effectively model noise in both the system and observation processes, and adaptively adjust the filter gain to suppress measurement errors and noise interference. It can also correct for potential outliers, preventing them from negatively impacting the pose estimation results.

[0039] The state information includes the estimated 3D position and motion information of key points of the human body, and the motion information includes speed and direction.

[0040] In a preferred embodiment of the present invention, step S4 is specifically as follows: record the 3D joint point cloud as ,in represents the coordinates of the vth 3D joint point (1≤v≤N), then the mean of the 3D joint point cloud is , the covariance matrix is , where Var(x v ) 、Var(y v ) and Var(z v ) represent x v 、y v 、z v The variance of represents the measurement noise of the vth 3D joint point coordinate in the corresponding direction, Cov(·,·) represents the covariance between the components of the vth 3D joint point on the corresponding two axes and the components of the other 3D joint points on the corresponding two axes; Based on the state information of the key points of the human body at time t-1, the predicted state at time t is obtained , the predicted covariance , where A is the state transition matrix, is the state information at time t-1, is the covariance matrix at time t-1, Q is the process noise covariance matrix, then the Kalman gain is , where H is the measurement matrix, and its structure is H=[I 3×3 0 3×3 ], I 3×3 is a 3×3 identity matrix, 0 3×3 is a 3×3 zero matrix; and the state information at time t , the covariance matrix , fused 3D position .

[0041] Cov(x v ,y v )、Cov(y v ,x v ) represents the covariance between the components of the vth 3D joint point on the y-axis and x-axis and the corresponding components of other 3D joint points, Cov(x v ,z v )、Cov(z v ,x v ) represents the covariance between the components of the vth 3D joint point on the z-axis and x-axis and the corresponding components of other 3D joint points, Cov(z v ,y v )、Cov(y v ,z v ) represents the covariance between the components of the vth 3D joint point on the y-axis and the z-axis and the corresponding components of other 3D joint points, Cov(x v,y v )、Cov(y v ,x v )、Cov(x v ,z v )、Cov(z v ,x v )、Cov(z v ,y v )、Cov(y v ,z v ) together represent the correlation between the components of the vth 3D joint point on the x, y, and z axes and the corresponding components of other 3D joint points.

[0042] Obtaining the mean of the 3D joint point cloud can achieve data smoothing, thereby reducing random errors; the covariance matrix R is used to characterize the uncertainty of the measurement data, enabling the system to better adapt to observation noise and improve the robustness and accuracy of data processing.

[0043] In a preferred embodiment of the present invention, in step S5, the multi-view geometric parameter optimization method is a bundle adjustment method. The bundle adjustment method optimizes the position estimates of 3D points by minimizing the reprojection error of the fused 3D positions. Specifically, it incorporates the geometric constraints between images captured by different cameras and employs nonlinear optimization techniques to find an optimal solution that maintains geometric consistency across multiple views by minimizing the error between viewpoints.

[0044] like Figure 5 This demonstrates the principle of bundle adjustment (BA). In the view, dots represent keypoints identified by the detector, while triangulated points represent the positions of the reconstructed keypoints projected onto the view plane. Due to detection and projection errors, there is a certain deviation between the detected keypoints and the projected points.

[0045] In a preferred embodiment of the present invention, step S5 is specifically as follows: optimizing the fused 3D position so that the loss function The position at the minimum is used as the estimated 3D position, where x j Indicates the corresponding image coordinate system The coordinates of the position, δ(x j ) is the indicator function, indicating x j The confidence level of detection in the image coordinate system, P jRepresents the transformation matrix that projects the fused 3D position into the image coordinate system. Due to detection and projection errors, there is a certain deviation between the detected keypoints and the inverse projection of the 3D point. By minimizing the loss function, which primarily measures the sum of the reprojection errors in each view, the bundle adjustment method can further optimize the position estimate of the 3D point while ensuring geometric consistency. This optimization process uses an iterative algorithm, updating the camera parameters and the coordinates of the 3D point through a nonlinear least squares method in each iteration, ultimately converging to an optimal solution and significantly improving the positioning accuracy of the 3D point.

[0046] It should be noted that when minimizing the loss function, the true position of the human key points on the image is usually unknown (the detected ones are not necessarily true). This application uses the confidence x on the image as j Human key points with a value greater than 0.3 are considered as true values ​​to solve this problem.

[0047] In a preferred embodiment of the present invention, after step S5, step S6 is further included: combining the estimated 3D positions of the key points of the human body at time t-1 with the motion information of the key points of the human body at time t, so as to facilitate the subsequent prediction of the 3D positions of the key points of the human body at time t+1.

[0048] like Figure 5 The present invention also provides a lighting system comprising a camera for capturing images of a human body, a processor for constructing a 3D human pose using any of the aforementioned feedback-based 3D human pose reconstruction methods, and a lamp for projecting light based on the 3D human pose. The lamp can estimate the 3D position of key points of the human body, accurately projecting light, providing efficient interaction and an enhanced audience experience.

[0049] Obviously, the above embodiments of the present invention are merely examples for the purpose of clearly illustrating the present invention, and are not intended to limit the embodiments of the present invention. Those skilled in the art will appreciate that other variations or modifications can be made based on the above description. It is not necessary and impossible to enumerate all embodiments here. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the claims of the present invention.

Claims

1. A feedback-based 3D human posture reconstruction method, characterized in that: The following steps are involved: S1, based on the estimated 3D position and motion information of the human body key point at time t-1, predict the predicted 3D position of the human body key point at time t; S2, associating the multiple actual 2D positions of the key points of the human body captured by the camera at time t with the predicted 3D positions by back-projecting them back into the image coordinate system of each camera to obtain multiple predicted 2D positions, to form a 2D joint point set of the key points of the human body at time t; S3, according to the 2D joint point set of the key point of the human body at time t, obtain the 3D joint point cloud of the key point of the human body at time t; S4, fusing the 3D joint point cloud using a fusion algorithm to obtain the fused 3D position of the key points of the human body; S5. Using the multi-view geometric parameter optimization method, the fused 3D position is optimized by combining multiple actual 2D positions of the key points of the human body captured by the camera at time t to obtain the estimated 3D position of the key points of the human body at time t.

2. The feedback-based 3D human posture reconstruction method according to claim 1, characterized in that: A verification step is also included between step S2 and step S3: using the epipolar RANAC algorithm to verify two 2D joint points from any two image coordinate systems in the 2D joint point set, and excluding a pair of 2D joint points with a large degree of deviation.

3. The feedback-based 3D human posture reconstruction method according to claim 2, wherein: In the verification step, the Sampson error is used to measure the degree of deviation of two 2D joint points of any two image coordinate systems under the epipolar geometry constraint. , where F is a 3×3 basic matrix representing the projection relationship between two camera perspectives, x1 and x2 correspond to the homogeneous coordinates of two 2D joint points in any two image coordinate systems, and it is considered that When the deviation is less than a predetermined threshold, the deviation is acceptable.

4. The feedback-based 3D human posture reconstruction method according to claim 3, characterized in that: The predetermined threshold is γ 2 , γ is three times the mean of the image noise standard deviation of all cameras.

5. The feedback-based 3D human posture reconstruction method according to claim 1, characterized in that: Step S3 is specifically as follows: combining two 2D joint points from any two image coordinate systems in the 2D joint point set, and forming a 3D joint point cloud from all the inferred 3D joint points.

6. The feedback-based 3D human posture reconstruction method according to claim 1, characterized in that: In step S4, the 3D joint point cloud is fused using the gain Kalman filter to obtain the fused 3D position of the key points of the human body.

7. The feedback-based 3D human posture reconstruction method according to claim 6, wherein: The specific steps of S4 are: record the 3D joint point cloud as ,in represents the coordinates of the vth 3D joint point (1≤v≤N), then the mean of the 3D joint point cloud is , the covariance matrix is , where Var(x v )、Var(y v ) and Var(z v ) represent x v 、y v 、z v The variance of represents the measurement noise of the v-th 3D joint point coordinate in the corresponding direction, and Cov(·,·) represents the covariance between two points. Based on the state information of the key points of the human body at time t-1, the predicted state at time t is obtained. , the predicted covariance , where A is the state transition matrix, is the state information at time t-1, is the covariance matrix at time t-1, Q is the process noise covariance matrix, then the Kalman gain is , where H is the measurement matrix, and its structure is H=[I 3×3 0 3×3 ], I 3×3 is a 3×3 identity matrix, 0 3×3 is a 3×3 zero matrix; and the state information at time t , the covariance matrix , fused 3D position .

8. The feedback-based 3D human posture reconstruction method according to claim 1, characterized in that: In step S5, the multi-view geometric parameter optimization method is the bundle adjustment method.

9. The feedback-based 3D human posture reconstruction method according to claim 8, characterized in that: The specific steps of S5 are: optimize the fused 3D position so that the loss function The position at the minimum is used as the estimated 3D position, where x j Indicates the corresponding image coordinate system The coordinates of the position, δ(x j ) is the indicator function, indicating x j The confidence level of detection in the image coordinate system, P j Represents the transformation matrix that projects the fused 3D position into the image coordinate system.

10. The feedback-based 3D human posture reconstruction method according to claim 1, characterized in that: After step S5, the method further includes step S6: obtaining motion information of the key points of the human body at time t by combining the estimated 3D positions of the key points of the human body at time t-1.

11. A lighting system, characterized in that: The invention comprises a camera for capturing images of a human body, a processor for constructing a 3D human body posture using the feedback-based 3D human body posture reconstruction method according to any one of claims 1 to 10, and a lamp for projecting light according to the 3D human body posture.

Citation Information

Cited By

  • Shooting dangerous action monitoring system for human body posture capture and environment perception

    CN121811506A