Shielding state 3D human body posture estimation method based on multistage optimization
By acquiring RGB images from multi-view cameras and combining occlusion estimation and data augmentation techniques, a multi-view occlusion estimation and optimization network was designed. This solved the accuracy and robustness issues of 3D human joint detection in occluded scenarios, achieving high-precision 3D joint detection.
Patent Information
- Application Number
- CN202510994392.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-17
- Publication Date
- 2025-10-21
AI Technical Summary
Existing 3D human joint detection methods struggle to accurately capture the joint positions of occluded parts in occluded scenes, especially in densely populated or complex environments. Furthermore, multi-view methods have significant shortcomings in occlusion dynamic modeling, collaborative fusion between viewpoints, and network robustness.
A 3D human pose estimation method based on multi-level optimization is adopted. RGB images are acquired by multi-view cameras. Combining occlusion estimation and data augmentation techniques, a multi-view occlusion estimation and optimization network is designed. The visibility of joints is dynamically evaluated by the occlusion perception module. High-precision reconstruction is performed through local occlusion estimation, multi-level optimization and pose prior library.
In real-world applications with severe occlusion, this technology enables high-precision and structurally sound 3D human joint detection, improving the model's occlusion robustness and generalization ability. It can output accurate 3D joint coordinates in real time with only multi-view RGB images as input.
Smart Images

Figure CN120823622A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for detecting 3D human joints, and in particular to a method for estimating 3D human posture in an occluded state based on multi-level optimization, and belongs to the fields of computer vision and artificial intelligence. Background Art
[0002] Three-dimensional human joint detection is a core technology in computer vision. By extracting the 3D human skeleton from images or videos, it enables applications such as action recognition, human-computer interaction, robot collaboration, and security monitoring. However, in real-world scenarios, occlusion—both self-occlusion and inter-person occlusion—significantly limits the performance of detection systems. Single-view methods often fail due to information loss when occlusion occurs, especially in crowded or complex environments, where it is difficult to accurately capture the joint locations of occluded parts.
[0003] Early studies mainly performed 2D joint detection based on a single RGB image, and directly predicted the joint coordinates through regression methods. With the development of deep learning, models based on convolutional neural networks (CNN), such as OpenPose, have significantly improved detection accuracy, using heatmaps to predict joint positions and associating joints through affinity fields. However, when these methods are extended to 3D detection, they still face the problem of missing depth information due to occlusion. To this end, researchers introduced multi-view input, using multiple cameras to capture images from different angles to provide information redundancy to alleviate the impact of occlusion. For example, by fusing multi-view 2D joints through triangulation, 3D positions can be inferred, but detection errors in occluded areas are still difficult to avoid.
[0004] On the other hand, some studies have attempted to improve model robustness through data augmentation techniques. For example, synthetic occlusions (such as geometric shapes or objects) are added to the training data to simulate real-world occlusion scenarios. However, single data augmentation is difficult to fully cover complex occlusion patterns, especially in multi-person interactions or dynamic environments. In addition, although multi-view methods can provide more perspective information, how to effectively estimate occlusion and optimize the fusion process remains an unresolved problem. Traditional methods often rely on manual rules or simple weight adjustments, lack a more comprehensive occlusion analysis, and limit detection accuracy and real-time performance.
[0005] Multi-view 3D joint detection has the advantage of natural information redundancy, which means that joints that are occluded in one view may still be clearly observed in other views. However, existing methods still have significant shortcomings in occlusion dynamic modeling, collaborative fusion between views, and network robustness. For example, most methods fail to comprehensively consider the relationship between multiple views when facing self-occlusion or multi-person interactive occlusion, nor do they build a fusion mechanism that can adaptively adjust over time and posture changes. Therefore, a comprehensive method is needed that integrates multi-view geometric constraints, temporal dynamic constraints of character movements, posture constraints, and skeletal structure constraints to achieve high-precision recovery and dynamic reconstruction of 3D joints in complex occlusion scenes.
[0006] The present invention proposes a 3D human posture estimation method in occluded state based on multi-level optimization, which comprehensively introduces key technologies such as data enhancement mechanism, local occlusion estimation, and multi-level optimization. In a multi-view environment, the system dynamically evaluates the visibility of joints through the occlusion perception module, and combines structural consistency with temporal priors to achieve accurate completion and robust reconstruction of occluded areas. During the training phase, various forms of synthetic occlusion data enhancement strategies are introduced to significantly improve the model's adaptability to complex occlusion scenarios. Ultimately, the system can output high-precision, structurally sound 3D joint coordinates of the human body, only requiring the input of multi-view RGB images, which is suitable for practical application scenarios with severe occlusion. Summary of the Invention
[0007] To address the challenge of occlusion problem to 3D human joint detection, this paper proposes a 3D human pose estimation method in occluded state based on multi-level optimization. A multi-view occlusion estimation and optimization network is constructed, which utilizes the information redundancy provided by multi-view cameras, combines occlusion estimation and data enhancement technology, and accurately estimates the positions of 3D human joints.
[0008] The technical solution adopted by the present invention is as follows: a method for estimating 3D human posture in occluded state based on multi-level optimization, characterized in that: this method collects multi-view RGB images through multiple synchronously calibrated cameras to ensure the accuracy and consistency of the data; a multi-view occlusion estimation and optimization network is designed to fuse multi-view information in an occluded environment into a 3D joint point detection model. During the training process, the occlusion estimation and data enhancement techniques are combined to assist the model in predicting the 3D joint points of the human body, so that the network has strong occlusion robustness and generalization ability. During operation, only the multi-view RGB images need to be input to accurately output the coordinates of the 3D joint points of the human body. The specific steps include:
[0009] Step 1: Data collection and 2D joint detection model training
[0010] Multiple synchronously calibrated wide-baseline cameras capture RGB images from various angles, with a resolution of at least 640x480 pixels and a frame rate of at least 25 FPS, ensuring image quality and smooth motion capture. Each image is projected and calibrated using camera intrinsics (such as focal length and principal point position), aligning the image center to the center of the human body bounding box. This effectively reduces distortion introduced by perspective effects and improves the accuracy of subsequent processing.
[0011] The captured RGB image is fed into an object detection network (YOLO11), which outputs the coordinates of a person's bounding box. The image within the bounding box is then cropped and scaled to 256×256 pixels, maintaining the same aspect ratio. To improve the model's robustness to perspective changes, conventional data augmentation operations are introduced, including random rotation (±15°), translation (±10% of the image width), and horizontal flipping. This increases the data's perspective diversity and provides the model with more training samples.
[0012] The standardized RGB image is input into the 2D human joint detection network. The network detects the coordinates of the 2D human joints (including 15 joints such as the head, shoulders, and elbows) and outputs a heat map for each joint. Because the 2D human joint detection network uses its graph-based posture modeling mechanism, even when some joints are occluded, it can use the contextual information of the visible parts (such as the spatial layout and connection relationship of adjacent joints) for indirect inference, thereby achieving a reasonable estimation of the position of the occluded points from a single perspective. The network consists of a feature extraction module and a two-stage joint prediction module, which are specifically defined as follows:
[0013] Feature extraction module: The input image is passed through the EfficientNet network to extract features and output a feature map F with a size of 256×256×18, where 18 is the number of channels.
[0014] Joint point prediction module: divided into two stages: stage 1 and stage 2;
[0015] Stage 1: Feature map F is input into two parallel channels:
[0016] (1) Confidence channel: The network predicts the joint confidence map S, which is used to detect the positions of all joints. It contains two layers of 5×5 convolution (step size 1), one layer of 3×3 convolution (step size 1), and two layers of 1×1 convolution (step size 1). It outputs a multi-channel heat map, where each channel corresponds to the probability distribution of a joint point. The joint point position is determined by the following formula:
[0017] (x k ,y k )=argmax (x,y) S k (x,y)
[0018] Among them S k is the heat map of the k-th joint point, (x k ,y k ) is the pixel coordinate of the kth joint point, and (x, y) is the pixel coordinate of the heat map.
[0019] (2) Association channel: predict the affinity L between joint points through the network affinity , to determine whether multiple joint points belong to the same person, it includes two layers of 5×5 convolution (step size 1), one layer of 3×3 convolution (step size 1) and two layers of 1×1 convolution (step size 1), and outputs the associated feature part affinity field (Part Affinity Fields, PAFs).
[0020] Phase 2: Phase 2 uses the same dual-channel architecture as Phase 1, including a confidence channel and an association channel. The dual-channel features output by Phase 1 are fused with the basic feature map extracted by the backbone network to form a new feature representation, which serves as the input for Phase 2. Finally, the partial affinity fields output by Phase 2 are used to correlate joint points through bipartite graph matching, resulting in a heatmap of joint points for each character.
[0021] Based on the above network architecture, the system can generate stereo heat maps of human joints from multi-view RGB images. To improve the model's robustness to occlusion, the following data augmentation strategies are adopted:
[0022] (1) Geometric occlusion: By adding randomly generated circular or rectangular occlusion areas to the image, we simulate partial occlusion that may occur in reality. The proportion of these occlusions ranges from 0% to 70%, covering both mild and severe occlusions.
[0023] (2) Object Occlusion: We extract object images from the Pascal VOC dataset (common objects such as chairs and tables) and randomly paste them onto the human body region to simulate scenes where foreground objects occlude the human body. For example, we overlay the image of a chair onto the leg or torso region to enable the model to learn how to identify partially occluded joints in a complex background.
[0024] (3) Random Drop: Remove joints according to the Bernoulli distribution. Some joints are randomly removed according to the Bernoulli distribution (p = 0.5) to simulate the situation where some joints are completely invisible due to occlusion or viewing angle restrictions. For example, the head joint may be lost due to side view, or the elbow joint may not be detected due to arm occlusion, thus improving its generalization ability under extreme conditions.
[0025] Step 2: Personnel detection and tracking
[0026] Sub-step 1: Pedestrian detection and ReID feature extraction
[0027] In order to achieve the detection and continuous tracking of human targets from multiple perspectives, a person association and tracking method that integrates the person re-identification (ReID) mechanism is adopted to make full use of the target's appearance information to assist in achieving stable identity association and continuous tracking. The object detection network (YOLO11) is used to obtain the human bounding box and its confidence from all perspectives. Subsequently, the pre-trained pedestrian re-identification model OSNet is used to input the cropped image of each detection box area into the ReID network to obtain a feature vector with a dimension of d = 256.
[0028] Sub-step 2: Cross-perspective character association and identity matching
[0029] At each time step, the system performs feature matching on the detection results of all cameras. For the detection target A from two different perspectives i and B j , extract their ReID feature vector f respectively i and f j , calculate its normalized cosine similarity:
[0030]
[0031] The calculated normalized cosine similarity matrix D ReID The Hungarian Algorithm is executed as the cost matrix to obtain the character associations across perspectives and record the IDs of the associated characters.
[0032] Sub-step 3: Tracking people across time
[0033] After completing the target identity matching across cameras, it is necessary to continuously maintain the identity status of each target in the time dimension to achieve "tracking" of the person. The description of the person is constructed by the person identity ID and ReID feature vector. The person identity ID is obtained by cross-view person association and identity matching, and ReID is obtained by multi-view feature vector calculation. The feature vector of the person in multiple views is {f1,f2,…,f V}, the corresponding human bounding box confidence is {w1,w2,…,w V}, where V is the number of viewing angles. The ReID feature vector f of the person ReID The calculation formula is as follows:
[0034]
[0035] Get the feature vector f corresponding to the character ID ReID Later, when new data arrives, the Hungarian algorithm is used to perform calculations and matching to achieve personnel tracking.
[0036] Step 3: Initial 3D pose estimation
[0037] To estimate 3D human poses from 2D human joints from multiple perspectives, we use voxel space back-projection and soft-argmax techniques to complete the mapping from 2D to 3D. The specific process is as follows:
[0038] For each camera n, use the 2D human joint point detection network to detect the two-dimensional heat map of each joint point:
[0039]
[0040] Where k∈{1,2,…,K} represents the index of the joint point, (u,v) is the pixel coordinate of the image plane, Represents the heat map of the k-th joint point under the n-th camera.
[0041] After obtaining the heatmaps of all joints from all perspectives, a 16×16×16 human body space is defined in three-dimensional space. This space is divided into a voxel grid with equal spacing. For each voxel (x, y, z) in the voxel space, the camera internal and external parameters are used to project it onto the image plane at each perspective:
[0042] (u n ,v n )=Π n (x,y,z)
[0043] where π n Represents the projection function of the nth camera. If the projection is within the image boundary, the value of the pixel in the corresponding heat map can be obtained.
[0044] For each joint point k, the heat values of the voxel position from all perspectives are fused to construct a voxel heat map:
[0045]
[0046] Where V represents the number of viewing angles, and the calculated H k (x, y, z) is the response strength of the voxel predicted as the kth joint point in all viewing angles. The three-dimensional coordinates of the joint point are obtained by soft-argmax:
[0047]
[0048] Among them, (X k ,Y k ,Z k ) represents the coordinates of the kth joint point. This process realizes the preliminary mapping from multi-view 2D information to 3D posture.
[0049] Step 4: Local Occlusion Estimation
[0050] Partial occlusion can generally be divided into two categories: one is occlusion caused by specific body postures within a single viewpoint, known as "self-occlusion"; the other is occlusion caused by different body parts blocking each other or by other objects, known as "viewpoint occlusion." To improve the consistency and interpretability of the analysis, occlusion estimation is divided into self-occlusion and viewpoint occlusion, with each modeling the occlusion impact and calculating occlusion weights.
[0051] Self-occlusion estimation: It is used to capture the occlusion characteristics of the human body's own structure or limbs, and emphasizes the evaluation of the observation quality of joints under a single view. This method combines the joint confidence and view geometry parameters to define the local visibility coefficient of the joint:
[0052] C ij =S ij ·cos(θ ij )
[0053] Among them, S ij is the joint point confidence, θ ij is the angle between the joint point and the camera, reflecting its orientation. ij Below threshold C th , then it is determined that joint point i is self-occluded in view j, and its weight is adjusted exponentially:
[0054]
[0055] Among them, γ is the basic weight under normal circumstances, and β is the attenuation coefficient. Normalize the self-occlusion estimation weight:
[0056]
[0057] Viewpoint occlusion estimation: This is used to identify joint point occlusions caused by spatial positional relationships between multiple perspectives, emphasizing the possibility of mutual compensation between perspectives. By projecting the initial 3D skeleton onto the 2D image planes of all perspectives, depth sorting and overlap detection strategies are used to analyze the occlusion relationships between joint points. When the depth of joint point A is greater than that of joint point B, and its projected area in the image overlaps with that of B, it can be determined that A is occluded by B. At this point, the visibility of joint point A is reduced, and the system dynamically adjusts its viewpoint weight based on the degree of occlusion:
[0058]
[0059] Among them, d ij is the relative depth of joint point i in view j, d this the occlusion threshold, and α is the smoothing coefficient. This mechanism effectively quantifies the intensity of depth-induced occlusion. To improve real-time performance, a Z-buffer algorithm is used to rapidly compare depth maps, quickly generating a viewpoint occlusion map. This mechanism can determine whether occlusion at certain viewpoints can be compensated for by other viewpoints, providing a basis for the subsequent design of fusion weights.
[0060] Weight the view occlusion and self-occlusion weight Perform weighted averaging to obtain the final fusion weight:
[0061]
[0062] Where η∈[0,1] is the fusion factor, which controls the relative influence of view angle and self-occlusion estimation. Fusion weight matrix W final The dimension is (N, V), where N is the number of joints and V is the number of viewpoints.
[0063] Step 5: Optimize 3D pose estimation
[0064] Based on the occlusion weight matrix W generated in step 4 final , performing multi-level optimization and information completion on the initial 3D human pose. The core optimization goal is to use visibility weights to guide the 3D reconstruction process in the presence of multi-view occlusion, thereby enhancing the robustness, temporal continuity, and structural consistency of the overall pose estimation. The entire optimization framework integrates the constraints of visual observation consistency, temporal consistency, pose prior retrieval, and skeleton consistency, providing a globally interpretable solution for 3D pose reconstruction in complex occluded scenes.
[0065] (1) Visual observation consistency constraint: Considering the different degrees of occlusion at different viewpoints, in order to avoid the misleading effect of low-quality viewpoints on 3D reconstruction, the fusion occlusion weight calculated in step 4 is introduced. Establish the weighted projection error minimization objective function:
[0066]
[0067] in, represents the three-dimensional spatial position of joint point i, is the two-dimensional observation of the point in view j, P j is the camera projection matrix of the corresponding viewing angle. The low weight automatically weakens the observation contribution from the occluded perspective. This mechanism can adaptively guide the optimization process to focus on high-confidence perspectives, thereby achieving more stable and reliable fusion.
[0068] (2) Temporal consistency constraints: In actual human motion, the pose changes between adjacent frames are usually continuous and smooth. Therefore, introducing temporal consistency constraints is crucial to improving the stability of 3D pose estimation. In particular, when joints are occluded or misjudged, relying solely on the current frame information can easily cause jitter, sudden changes, or pose distortion. Therefore, a cross-frame dynamic modeling mechanism must be introduced to provide historical reference.
[0069] Specifically, for each joint point in a frame The three-dimensional position of the joint point in the previous frame is defined The Euclidean distance of is used as a penalty term, and the time regularization term is constructed as follows:
[0070]
[0071] The goal of this regularization term is to minimize the spatial displacement of joints between adjacent frames, thereby enhancing the smoothness of the posture evolution over time. For different joints, the influence of this term can be dynamically adjusted according to their occlusion degree. Time consistency weighting mechanism:
[0072]
[0073] in, That is, when a joint point is in an occluded state in all viewing angles ( When the time smoothing term is low, the weight of tends to 1, thereby increasing the proportion of historical frames in the estimation of the point; and when the joint point has good observation quality in the current frame ( If the θ is high, the weight of the temporal constraint is reduced to avoid over-smoothing the natural motion.
[0074] (3) Pose prior retrieval and constraints: When there is complete occlusion, observations of some joints from all perspectives are unavailable (i.e. ), this lack of information will significantly affect the stability and accuracy of pose reconstruction. To improve the model's completion capabilities in such extreme situations, a structure-driven 3D pose prior library is introduced to provide data-driven constraint compensation.
[0075] Pose prior library P = {P 1 ,P 2 ,…,P M} is constructed by aggregating large-scale, high-quality motion capture data, and each sample P m Represents a complete 3D human skeleton pose. The pose prior library has a diverse range of poses, covering typical human behavior paradigms such as stillness, walking, running, one-handed operation, and interactive actions, constructing a reference space with rich structural expression.
[0076] When there is severe occlusion, extract the set of visible joint points in the current frame Construct an incomplete local skeleton topology; for each sample P in the prior library m , and calculate its distance with The structural similarity of the corresponding joint point subset in the , the distance error between the current skeleton topology and the prior library sample is {d1, d2, ..., d M}; Select the top K best samples based on the matching results And construct the completion result of the occluded joint points in a weighted average manner:
[0077]
[0078] Among them, ρ k is the normalized matching similarity score, which is calculated by the softmax transformation of the distance error:
[0079]
[0080] The semantic consistency constraint of the pose library is constructed by completing the occluded joint points of the current joint point and the pose prior library:
[0081]
[0082] The pose completion strategy essentially transforms the full occlusion problem into a structural reasoning task. It leverages prior knowledge of human structure across scenes and individuals, effectively improving the stability and biological consistency of joint recovery without relying on current image features.
[0083] (4) Skeleton consistency constraint: 3D human pose estimation requires not only data-driven observation accuracy but also consistency and rationality in the skeleton structure. To this end, explicit skeleton geometry consistency constraints are introduced during the optimization process to prevent skeleton proportion distortion caused by occlusion or optimization drift.
[0084] We define a standard bone segment set ε, where each pair (i, k) represents a set of anatomically connected joint points. Ideally, the three-dimensional bone segment length should remain constant or within a reasonable range, so the bone length regularization term is constructed as follows:
[0085]
[0086] where l i,k The target bone segment length is obtained by statistically analyzing the real motion dataset and serves as the anatomical reference for the constraint.
[0087] Based on the above mechanism, a unified multi-target pose estimation loss function is constructed as follows:
[0088]
[0089] Among them, λ t ,λ p ,λ b is the weight hyperparameter of each loss term, which is used to balance the different loss terms in the total objective function L total The joint loss framework ensures observation consistency while introducing dynamic and structural regularization mechanisms, guaranteeing both reconstruction accuracy and improving the system's generalization and recovery capabilities under complex occlusion conditions. After completing the initial 3D pose estimation, the system constructs and minimizes a multi-objective pose estimation loss function using 3D joint coordinates as optimization variables. It further optimizes the initial estimation results, ultimately outputting a more stable, reasonable, and robust 3D human pose.
[0090] Advantages and significant effects of this method:
[0091] By introducing an occlusion-aware network and a fusion optimization strategy, this method can finely model occluded areas in multi-view images, fully leveraging the complementary advantages of perspectives to achieve high-precision 3D joint detection even under severe occlusion conditions. At the same time, this method introduces a large number of synthetic occlusion and object occlusion enhancement samples during the training phase, significantly improving the model's occlusion adaptability and generalization performance. During the inference phase, the system relies solely on multi-view RGB images to output accurate 3D human joints in real time. It has lightweight deployment characteristics and is widely applicable to complex, dynamic, and occlusion-prone application scenarios such as motion capture, motion analysis, and interactive systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0092] Figure 1 This is a flow chart of the 3D human joint detection method based on multiple perspectives;
[0093] Figure 2 This is a schematic diagram of the feature extraction and 2D joint point detection module;
[0094] Figure 3 This is a diagram showing how the joints are connected. DETAILED DESCRIPTION
[0095] 3D human joint detection is a key topic in computer vision, crucial for understanding human posture, advancing artificial intelligence applications, and enabling intelligent interaction. By accurately extracting the human 3D skeleton from images or videos, it can be widely applied in areas such as action recognition, robotic collaboration, security monitoring, and virtual reality.
[0096] In the early days of this field, single-view 3D joint detection primarily relied on regression methods to predict the 3D coordinates of joints directly from RGB images. With the development of deep learning technology, researchers have designed feature extraction models based on convolutional neural networks, such as deep residual networks (ResNet) and pose estimation networks (such as OpenPose), which use heat maps to predict joint locations and associate them with the skeleton. However, single-view input has significant limitations when dealing with occluded scenes (such as self-occlusion or occlusion between people). Joints in occluded areas often cannot be accurately estimated due to missing information.
[0097] To address the occlusion problem, researchers have attempted to introduce multi-view approaches, utilizing multiple cameras to capture images from different angles, providing redundant information to compensate for the shortcomings of a single view. Multi-view approaches can mitigate the effects of occlusion to a certain extent by inferring 3D coordinates through triangulation or depth estimation. However, existing methods are still insufficiently robust in complex occlusion scenarios, such as those involving crowded environments or dynamic occlusion, and lack effective occlusion estimation and perspective fusion strategies. Furthermore, traditional multi-view detection systems often suffer from high computational complexity, making them difficult to meet the demands of real-time applications.
[0098] The advantage of multi-view 3D joint detection is that it can utilize complementary information from different perspectives. For example, when a person is occluded from a certain perspective, other perspectives may still provide valid observation data. To this end, researchers have proposed data augmentation technology to improve the model's adaptability to occluded scenes by adding synthetic occlusions (such as geometric shapes or objects) to the training data. Despite this, a single enhancement strategy is difficult to fully cover the diversity of occlusions in reality, and existing methods still have room for improvement in fine-grained analysis and model optimization of occlusion estimation. Therefore, a comprehensive solution is needed that combines multi-view information, occlusion estimation and deep learning technology to improve the robustness and practicality of the detection system.
[0099] To address these issues, this paper proposes a method for estimating 3D human pose in occluded conditions based on multi-level optimization. This method constructs a multi-view occlusion estimation and optimization network. This method leverages the information redundancy provided by multi-view cameras and combines occlusion estimation with data augmentation techniques to accurately estimate the positions of 3D human joints. During training, this method synthesizes occlusion-enhanced data to improve the model's generalization to occluded scenes. At runtime, the method outputs precise 3D joint coordinates using only multi-view RGB images as input.
[0100] The technical solution adopted by the present invention is as follows: a method for estimating 3D human posture in occluded state based on multi-level optimization, characterized in that: this method collects multi-perspective RGB images through multiple synchronously calibrated cameras to ensure the accuracy and consistency of the data; designs a multi-perspective occlusion estimation and optimization network to fuse multi-perspective information in an occluded environment into a 3D joint detection model, and in the training process, combines occlusion estimation and data enhancement technology to assist the model in predicting 3D joints of the human body, so that the network has strong occlusion robustness and generalization ability. During operation, only multi-perspective RGB images need to be input to accurately output the coordinates of the 3D joints of the human body. The flow chart of this method is as follows Figure 1 As shown, the specific steps include:
[0101] Step 1: Data collection and 2D joint detection model training
[0102] Multiple synchronously calibrated wide-baseline cameras capture RGB images from various angles, with a resolution of at least 640x480 pixels and a frame rate of at least 25 FPS, ensuring image quality and smooth motion capture. Each image is projected and calibrated using camera intrinsics (such as focal length and principal point position), aligning the image center to the center of the human body bounding box. This effectively reduces distortion introduced by perspective effects and improves the accuracy of subsequent processing.
[0103] The captured RGB image is fed into an object detection network (YOLO11), which outputs the coordinates of a person's bounding box. The image within the bounding box is then cropped and scaled to 256×256 pixels, maintaining the same aspect ratio. To improve the model's robustness to perspective changes, conventional data augmentation operations are introduced, including random rotation (±15°), translation (±10% of the image width), and horizontal flipping. This increases the data's perspective diversity and provides the model with more training samples.
[0104] The standardized RGB image is input into the 2D human joint detection network. The feature extraction and 2D joint detection module of the 2D human joint detection network are shown in the figure. Figure 2 The network detects the coordinates of 2D joints of the human body (including 15 joints such as head, shoulder, elbow, etc.) and outputs a heat map of each joint. The joints detected by the network are as follows: Figure 3 As shown in Figure 2. The 2D human joint detection network uses its graph-based pose modeling mechanism to indirectly infer the position of occluded points even when some joints are occluded by using contextual information from visible parts (such as the spatial layout and connectivity of adjacent joints). This allows for a reasonable estimation of the position of occluded points from a single viewpoint. The network consists of a feature extraction module and a two-stage joint prediction module, specifically defined as follows:
[0105] Feature extraction module: The input image is passed through the EfficientNet network to extract features and output a feature map F with a size of 256×256×18, where 18 is the number of channels.
[0106] Joint point prediction module: divided into two stages: stage 1 and stage 2;
[0107] Stage 1: Feature map F is input into two parallel channels:
[0108] (1) Confidence channel: The network predicts the joint confidence map S, which is used to detect the positions of all joints. It contains two layers of 5×5 convolution (step size 1), one layer of 3×3 convolution (step size 1), and two layers of 1×1 convolution (step size 1). It outputs a multi-channel heat map, where each channel corresponds to the probability distribution of a joint point. The joint point position is determined by the following formula:
[0109] (x k ,y k )=argmax (x,y) S k (x,y)
[0110] Among them S k is the heat map of the k-th joint point, (x k ,y k ) is the pixel coordinate of the kth joint point, and (x, y) is the pixel coordinate of the heat map.
[0111] (2) Association channel: predict the affinity L between joint points through the network affinity , to determine whether multiple joint points belong to the same person, it includes two layers of 5×5 convolution (step size 1), one layer of 3×3 convolution (step size 1) and two layers of 1×1 convolution (step size 1), and outputs the associated feature part affinity field (Part Affinity Fields, PAFs).
[0112] Phase 2: Phase 2 uses the same dual-channel architecture as Phase 1, including a confidence channel and an association channel. The dual-channel features output by Phase 1 are fused with the basic feature map extracted by the backbone network to form a new feature representation, which serves as the input for Phase 2. Finally, the partial affinity fields output by Phase 2 are used to correlate joint points through bipartite graph matching, resulting in a heatmap of joint points for each character.
[0113] Based on the above network architecture, the system can generate stereo heat maps of human joints from multi-view RGB images. To improve the model's robustness to occlusion, the following data augmentation strategies are adopted:
[0114] (1) Geometric occlusion: By adding randomly generated circular or rectangular occlusion areas to the image, we simulate partial occlusion that may occur in reality. The proportion of these occlusions ranges from 0% to 70%, covering both mild and severe occlusions.
[0115] (2) Object Occlusion: We extract object images from the Pascal VOC dataset (common objects such as chairs and tables) and randomly paste them onto the human body region to simulate scenes where foreground objects occlude the human body. For example, we overlay the image of a chair onto the leg or torso region to enable the model to learn how to identify partially occluded joints in a complex background.
[0116] (3) Random Drop: Remove joints according to the Bernoulli distribution. Some joints are randomly removed according to the Bernoulli distribution (p = 0.5) to simulate the situation where some joints are completely invisible due to occlusion or viewing angle restrictions. For example, the head joint may be lost due to side view, or the elbow joint may not be detected due to arm occlusion, thus improving its generalization ability under extreme conditions.
[0117] Step 2: Personnel detection and tracking
[0118] Sub-step 1: Pedestrian detection and ReID feature extraction
[0119] In order to achieve the detection and continuous tracking of human targets from multiple perspectives, a person association and tracking method that integrates the person re-identification (ReID) mechanism is adopted to make full use of the target's appearance information to assist in achieving stable identity association and continuous tracking. The object detection network (YOLO11) is used to obtain the human bounding box and its confidence from all perspectives. Subsequently, the pre-trained pedestrian re-identification model OSNet is used to input the cropped image of each detection box area into the ReID network to obtain a feature vector with a dimension of d = 256.
[0120] Sub-step 2: Cross-perspective character association and identity matching
[0121] At each time step, the system performs feature matching on the detection results of all cameras. For the detection target A from two different perspectives i and B j , extract their ReID feature vector f respectively i and f j , calculate its normalized cosine similarity:
[0122]
[0123] The calculated normalized cosine similarity matrix D ReIDThe Hungarian Algorithm is executed as the cost matrix to obtain the character associations across perspectives and record the IDs of the associated characters.
[0124] Sub-step 3: Tracking people across time
[0125] After completing the target identity matching across cameras, it is necessary to continuously maintain the identity status of each target in the time dimension to achieve "tracking" of the person. The description of the person is constructed by the person identity ID and ReID feature vector. The person identity ID is obtained by cross-view person association and identity matching, and ReID is obtained by multi-view feature vector calculation. The feature vector of the person in multiple views is {f1,f2,…,f V}, the corresponding human bounding box confidence is {w1,w2,…,w V}, where V is the number of viewing angles. The ReID feature vector f of the person ReID The calculation formula is as follows:
[0126]
[0127] Get the feature vector f corresponding to the character ID ReID Later, when new data arrives, the Hungarian algorithm is used to perform calculations and matching to achieve personnel tracking.
[0128] Step 3: Initial 3D pose estimation
[0129] To estimate 3D human poses from 2D human joints from multiple perspectives, we use voxel space back-projection and soft-argmax techniques to complete the mapping from 2D to 3D. The specific process is as follows:
[0130] For each camera n, use the 2D human joint point detection network to detect the two-dimensional heat map of each joint point:
[0131]
[0132] Where k∈{1,2,…,K} represents the index of the joint point, (u,v) is the pixel coordinate of the image plane, Represents the heat map of the k-th joint point under the n-th camera.
[0133] After obtaining the heatmaps of all joints from all perspectives, a 16×16×16 human body space is defined in three-dimensional space. This space is divided into a voxel grid with equal spacing. For each voxel (x, y, z) in the voxel space, the camera internal and external parameters are used to project it onto the image plane at each perspective:
[0134] (u n ,v n )=Πn (x,y,z)
[0135] where π n Represents the projection function of the nth camera. If the projection is within the image boundary, the value of the pixel in the corresponding heat map can be obtained.
[0136] For each joint point k, the heat values of the voxel position from all perspectives are fused to construct a voxel heat map:
[0137]
[0138] Where V represents the number of viewing angles, and the calculated H k (x, y, z) is the response strength of the voxel predicted as the kth joint point in all viewing angles. The three-dimensional coordinates of the joint point are obtained by soft-argmax:
[0139]
[0140] Among them, (X k ,Y k ,Z k ) represents the coordinates of the kth joint point. This process realizes the preliminary mapping from multi-view 2D information to 3D posture.
[0141] Step 4: Local Occlusion Estimation
[0142] Partial occlusion can generally be divided into two categories: one is occlusion caused by specific body postures within a single viewpoint, known as "self-occlusion"; the other is occlusion caused by different body parts blocking each other or by other objects, known as "viewpoint occlusion." To improve the consistency and interpretability of the analysis, occlusion estimation is divided into self-occlusion and viewpoint occlusion, with each modeling the occlusion impact and calculating occlusion weights.
[0143] Self-occlusion estimation: It is used to capture the occlusion characteristics of the human body's own structure or limbs, and emphasizes the evaluation of the observation quality of joints under a single view. This method combines the joint confidence and view geometry parameters to define the local visibility coefficient of the joint:
[0144] C ij =S ij ·cos(θ ij )
[0145] Among them, S ij is the joint point confidence, θ ij is the angle between the joint point and the camera, reflecting its orientation. ij Below threshold C th , then it is determined that joint point i is self-occluded in view j, and its weight is adjusted exponentially:
[0146]
[0147] Among them, γ is the basic weight under normal circumstances, and β is the attenuation coefficient. Normalize the self-occlusion estimation weight:
[0148]
[0149] Viewpoint occlusion estimation: This is used to identify joint point occlusions caused by spatial positional relationships between multiple perspectives, emphasizing the possibility of mutual compensation between perspectives. By projecting the initial 3D skeleton onto the 2D image planes of all perspectives, depth sorting and overlap detection strategies are used to analyze the occlusion relationships between joint points. When the depth of joint point A is greater than that of joint point B, and its projected area in the image overlaps with that of B, it can be determined that A is occluded by B. At this point, the visibility of joint point A is reduced, and the system dynamically adjusts its viewpoint weight based on the degree of occlusion:
[0150]
[0151] Among them, d ij is the relative depth of joint point i in view j, d th is the occlusion threshold, and α is the smoothing coefficient. This mechanism effectively quantifies the intensity of depth-induced occlusion. To improve real-time performance, a Z-buffer algorithm is used to rapidly compare depth maps, quickly generating a viewpoint occlusion map. This mechanism can determine whether occlusion at certain viewpoints can be compensated for by other viewpoints, providing a basis for the subsequent design of fusion weights.
[0152] Weight the view occlusion and self-occlusion weight Perform weighted averaging to obtain the final fusion weight:
[0153]
[0154] Where η∈[0,1] is the fusion factor, which controls the relative influence of view angle and self-occlusion estimation. Fusion weight matrix W final The dimension is (N, V), where N is the number of joints and V is the number of viewpoints.
[0155] Step 5: Optimize 3D pose estimation
[0156] Based on the occlusion weight matrix W generated in step 4 final, performing multi-level optimization and information completion on the initial 3D human pose. The core optimization goal is to use visibility weights to guide the 3D reconstruction process in the presence of multi-view occlusion, thereby enhancing the robustness, temporal continuity, and structural consistency of the overall pose estimation. The entire optimization framework integrates the constraints of visual observation consistency, temporal consistency, pose prior retrieval, and skeleton consistency, providing a globally interpretable solution for 3D pose reconstruction in complex occluded scenes.
[0157] (1) Visual observation consistency constraint: Considering the different degrees of occlusion at different viewpoints, in order to avoid the misleading effect of low-quality viewpoints on 3D reconstruction, the fusion occlusion weight calculated in step 4 is introduced. Establish the weighted projection error minimization objective function:
[0158]
[0159] in, represents the three-dimensional spatial position of joint point i, is the two-dimensional observation of the point in view j, P j is the camera projection matrix of the corresponding viewing angle. The low weight automatically weakens the observation contribution from the occluded perspective. This mechanism can adaptively guide the optimization process to focus on high-confidence perspectives, thereby achieving more stable and reliable fusion.
[0160] (2) Temporal consistency constraints: In actual human motion, the pose changes between adjacent frames are usually continuous and smooth. Therefore, introducing temporal consistency constraints is crucial to improving the stability of 3D pose estimation. In particular, when joints are occluded or misjudged, relying solely on the current frame information can easily cause jitter, sudden changes, or pose distortion. Therefore, a cross-frame dynamic modeling mechanism must be introduced to provide historical reference.
[0161] Specifically, for each joint point in a frame The three-dimensional position of the joint point in the previous frame is defined The Euclidean distance of is used as a penalty term, and the time regularization term is constructed as follows:
[0162]
[0163] The goal of this regularization term is to minimize the spatial displacement of joints between adjacent frames, thereby enhancing the smoothness of the posture evolution over time. For different joints, the influence of this term can be dynamically adjusted according to their occlusion degree. Time consistency weighting mechanism:
[0164]
[0165] in, That is, when a joint point is in an occluded state in all viewing angles ( When the time smoothing term is low, the weight of tends to 1, thereby increasing the proportion of historical frames in the estimation of the point; and when the joint point has good observation quality in the current frame ( If the θ is high, the weight of the temporal constraint is reduced to avoid over-smoothing the natural motion.
[0166] (3) Pose prior retrieval and constraints: When there is complete occlusion, observations of some joints from all perspectives are unavailable (i.e. ), this lack of information will significantly affect the stability and accuracy of pose reconstruction. To improve the model's completion capabilities in such extreme situations, a structure-driven 3D pose prior library is introduced to provide data-driven constraint compensation.
[0167] Pose prior library P = {P 1 ,P 2 ,…,P M} is constructed by aggregating large-scale, high-quality motion capture data, and each sample P m Represents a complete 3D human skeleton pose. The pose prior library has a diverse range of poses, covering typical human behavior paradigms such as stillness, walking, running, one-handed operation, and interactive actions, constructing a reference space with rich structural expression.
[0168] When there is severe occlusion, extract the set of visible joint points in the current frame Construct an incomplete local skeleton topology; for each sample P in the prior library m , and calculate its distance with The structural similarity of the corresponding joint point subset in the , the distance error between the current skeleton topology and the prior library sample is {d1, d2, ..., d M}; Select the top K best samples based on the matching results And construct the completion result of the occluded joint points in a weighted average manner:
[0169]
[0170] Among them, ρ k is the normalized matching similarity score, which is calculated by the softmax transformation of the distance error:
[0171]
[0172] The semantic consistency constraint of the pose library is constructed by completing the occluded joint points of the current joint point and the pose prior library:
[0173]
[0174] The pose completion strategy essentially transforms the full occlusion problem into a structural reasoning task. It leverages prior knowledge of human structure across scenes and individuals, effectively improving the stability and biological consistency of joint recovery without relying on current image features.
[0175] (4) Skeleton consistency constraint: 3D human pose estimation requires not only data-driven observation accuracy but also consistency and rationality in the skeleton structure. To this end, explicit skeleton geometry consistency constraints are introduced during the optimization process to prevent skeleton proportion distortion caused by occlusion or optimization drift.
[0176] We define a standard bone segment set ε, where each pair (i, k) represents a set of anatomically connected joint points. Ideally, the three-dimensional bone segment length should remain constant or within a reasonable range, so the bone length regularization term is constructed as follows:
[0177]
[0178] where l i,k The target bone segment length is obtained by statistically analyzing the real motion dataset and serves as the anatomical reference for the constraint.
[0179] Based on the above mechanism, a unified multi-target pose estimation loss function is constructed as follows:
[0180]
[0181] Among them, λ t ,λ p ,λ b is the weight hyperparameter of each loss term, which is used to balance the different loss terms in the total objective function L total The joint loss framework ensures observation consistency while introducing dynamic and structural regularization mechanisms, guaranteeing both reconstruction accuracy and improving the system's generalization and recovery capabilities under complex occlusion conditions. After completing the initial 3D pose estimation, the system constructs and minimizes a multi-objective pose estimation loss function using 3D joint coordinates as optimization variables. It further optimizes the initial estimation results, ultimately outputting a more stable, reasonable, and robust 3D human pose.
[0182] This method designs a multi-view occlusion estimation and optimization network, which integrates multi-view information in an occluded environment into a 3D human joint detection model. During the training process, multi-view RGB images and synthetic occlusion data are used as input to train the network. When the network is running, it only needs to input multi-view RGB images to accurately output the 3D joint coordinates of the human body.
Claims
1. A method for estimating 3D human pose in occluded state based on multi-level optimization, characterized by: Data is collected synchronously through multi-view cameras. A multi-view occlusion estimation and optimization network is designed to integrate multi-view information into a 3D joint detection model in occluded environments. During training, a small amount of multi-view data is used as input, combined with occlusion estimation and data augmentation techniques to assist the model in predicting 3D human joints. At runtime, the model only needs to input multi-view RGB images to accurately output the coordinates of 3D human joints, making it suitable for complex occlusion scenarios. The specific steps include: Step 1: Data collection and 2D joint detection model training Multiple synchronously calibrated wide-baseline cameras are used to capture RGB images from different angles, with a resolution of at least 640x480 pixels and a frame rate of at least 25 FPS. Each image is projected and calibrated using the camera's intrinsic parameters, aligning the image center to the center of the human body bounding box. The captured RGB image is fed into the object detection network YOLO11, which outputs the coordinates of the human body bounding box. The image within the bounding box is cropped and scaled to 256×256 pixels, maintaining the aspect ratio. Conventional data augmentation operations, including random rotation, translation, and horizontal flipping, are introduced to enhance the perspective diversity of the data. The normalized RGB image is input into the 2D human joint detection network to detect the coordinates of the 2D human joints and output a heat map of each joint. The network consists of a feature extraction module and a two-stage joint prediction module, which are specifically defined as follows: Feature extraction module: The input image is passed through the EfficientNet network to extract features and output a feature map F with a size of 256×256×18, where 18 is the number of channels; Joint point prediction module: divided into two stages: stage 1 and stage 2; Stage 1: Feature map F is input into two parallel channels: (1) Confidence channel: The network predicts the joint confidence map S, which is used to detect the positions of all joints. It contains two layers of 5×5 convolution, one layer of 3×3 convolution, and two layers of 1×1 convolution, and outputs a multi-channel heat map. Each channel corresponds to the probability distribution of a joint point. The joint point position is determined by the following formula: (x k ,y k )=argmax (x,y) S k (x,y) Among them S k is the heat map of the k-th joint point, (x k ,y k ) is the pixel coordinate of the k-th joint point, (x, y) is the pixel coordinate of the heat map; (2) Association channel: predict the affinity L between joint points through the network affinity ,To determine whether multiple joint points belong to the same person, it includes two layers of 5×5 convolution, one layer of 3×3 convolution and two layers of 1×1 convolution, and outputs the associated feature part affinity field. Phase 2: Continuing the same dual-channel structure as Phase 1, including a confidence channel and an association channel, the dual-channel features output from Phase 1 are fused with the basic feature map extracted by the backbone network to form a new feature representation, which serves as the input for Phase 2. Finally, the affinity fields output from Phase 2 are used to associate joint points through bipartite graph matching to obtain joint heatmaps for each character. Based on the above network architecture, we can generate stereo heatmaps of human joints from multi-view RGB images. To improve the model's robustness to occlusion, we use the following data augmentation strategies: (1) Geometric occlusion: By adding randomly generated circular or rectangular occlusion areas to the image, partial occlusion that may occur in reality is simulated; the proportion of these occlusions ranges randomly from 0% to 70%, covering both mild occlusion and severe occlusion; (2) Object occlusion: Object images are extracted from the PascalVOC dataset and randomly pasted onto the human body area to simulate the scene of foreground objects being occluded in the real environment; (3) Random drop: remove joints according to Bernoulli distribution; randomly remove some joints according to Bernoulli distribution p = 0.5 to simulate the situation where some joints are completely invisible due to occlusion or viewing angle limitation; Step 2: Personnel detection and tracking Sub-step 1: Pedestrian detection and ReID feature extraction In order to achieve the detection and continuous tracking of human targets under multiple perspectives, a person association and tracking method integrating pedestrian re-identification mechanism is adopted, which makes full use of the appearance information of the target to assist in achieving stable identity association and continuous tracking; the target detection network YOLO11 is used to obtain the human bounding box and its confidence from all perspectives; then, the pre-trained pedestrian re-identification model OSNet is used to input the cropped image of each detection box area into the ReID network to obtain a feature vector with a dimension of d=256. Sub-step 2: Cross-perspective character association and identity matching At each time step, feature matching is performed on the detection results of all cameras; for the detection target A from two different perspectives i and B j , extract their ReID feature vector f respectively i and f j , calculate its normalized cosine similarity: The calculated normalized cosine similarity matrix D ReID Execute the Hungarian algorithm as the cost matrix to obtain the character associations across perspectives and record the IDs of the associated characters; Sub-step 3: Tracking people across time After completing the target identity matching across cameras, the identity status of each target is continuously maintained in the time dimension to achieve "tracking" of the person; the description of the person is constructed through the person identity ID and ReID feature vector, where the person identity ID is obtained through cross-view person association and identity matching, and ReID is obtained through multi-view feature vector calculation. The feature vector of the person in multiple views is {f1,f2,…,f V }, the corresponding human bounding box confidence is {w1,w2,…,w V }, where V is the number of viewing angles; the ReID feature vector f of the person ReID The calculation formula is as follows: Get the feature vector f corresponding to the character ID ReID Finally, when new data arrives, the Hungarian algorithm is used to calculate and match to achieve personnel tracking; Step 3: Initial 3D pose estimation To estimate 3D human pose from 2D human joints from multiple perspectives, we use voxel space back-projection and soft-argmax techniques to complete the mapping from 2D to 3D. The specific process is as follows: For each camera n, use the 2D human joint point detection network to detect the two-dimensional heat map of each joint point: Where k∈{1,2,…,K} represents the index of the joint point, (u,v) is the pixel coordinate of the image plane, Represents the heat map of the kth joint point under the nth camera; After obtaining the heatmaps of all joints from all perspectives, a 16×16×16 human body space is defined in three-dimensional space. This space is divided into a voxel grid with equal spacing. For each voxel (x, y, z) in the voxel space, the camera internal and external parameters are used to project it onto the image plane at each perspective: (u n ,v n )=Π n (x,y,z) where π n Represents the projection function of the nth camera. If the projection is within the image boundary, the value of the pixel in the corresponding heat map is obtained; For each joint point k, the heat values of the voxel position from all perspectives are fused to construct a voxel heat map: Where V represents the number of viewing angles, and the calculated H k (x, y, z) is the response strength of the voxel predicted to be the kth joint point in all viewing angles; the three-dimensional coordinates of the joint point are obtained by soft-argmax: Among them, (X k ,Y k ,Z k ) represents the coordinates of the kth joint point. This process realizes the preliminary mapping from multi-view 2D information to 3D posture; Step 4: Local Occlusion Estimation Partial occlusion is divided into two categories: one is occlusion caused by a specific body posture under a single viewpoint, called "self-occlusion"; the other is occlusion caused by different body parts blocking each other or being blocked by other objects, called "viewpoint occlusion". To improve the consistency and interpretability of the analysis, occlusion estimation is divided into self-occlusion and viewpoint occlusion, and the occlusion effects are modeled and occlusion weights are calculated for each. Self-occlusion estimation: It is used to capture the occlusion characteristics of the human body's own structure or limbs, and emphasizes the evaluation of the observation quality of joints under a single perspective. It combines the joint confidence and the perspective geometry parameters to define the local visibility coefficient of the joint: C ij =S ij ·cos(θ ij ) Among them, S ij is the joint point confidence, θ ij is the angle between the joint point and the camera, reflecting its orientation; if C ij Below threshold C th , then it is determined that joint point i is self-occluded in view j, and its weight is adjusted exponentially: Among them, γ is the basic weight under normal circumstances, β is the attenuation coefficient; the self-occlusion estimation weight is normalized: Viewpoint occlusion estimation: This method is used to identify joint occlusions caused by spatial relationships between multiple viewpoints, emphasizing the possibility of mutual compensation between viewpoints. The initial 3D skeleton is projected onto the 2D image planes of all viewpoints, and depth sorting and overlap detection strategies are used to analyze the occlusion relationships between joints. When the depth of joint A is greater than that of joint B, and its projected area in the image overlaps with that of B, A is considered to be occluded by B. At this point, the visibility of joint A is reduced, and its viewpoint weight is dynamically adjusted based on the degree of occlusion: Among them, d ij is the relative depth of joint point i in view j, d th is the occlusion threshold, and α is the smoothing coefficient. To improve real-time performance, the Z-buffer algorithm is used to quickly compare the depth map, thereby quickly generating a perspective occlusion map. This mechanism can determine whether occlusion at certain perspectives can be compensated by other perspectives, providing a basis for the design of subsequent fusion weights. Weight the view occlusion and self-occlusion weight Perform weighted averaging to obtain the final fusion weight: Where η∈[0,1] is the fusion factor, which controls the relative influence of view angle and self-occlusion estimation. Fusion weight matrix W final The dimension is (N, V), where N is the number of joints and V is the number of viewpoints. Step 5: Optimize 3D pose estimation Based on the occlusion weight matrix W generated in step 4 final , multi-level optimization and information completion are performed on the initial 3D human pose. The core goal of the optimization is to use visibility weights to guide the 3D reconstruction process under the premise of multi-view occlusion, thereby enhancing the robustness, temporal continuity and structural consistency of the overall pose estimation. The entire optimization framework integrates the constraints of visual observation consistency, temporal consistency, pose prior retrieval and skeleton consistency, providing a global interpretability solution for 3D pose reconstruction in complex occlusion scenes. (1) Visual observation consistency constraint: Considering the different degrees of occlusion at different viewpoints, in order to avoid the misleading effect of low-quality viewpoints on 3D reconstruction, the fusion occlusion weight calculated in step 4 is introduced. Establish the weighted projection error minimization objective function: in, represents the three-dimensional spatial position of joint point i, is the two-dimensional observation of the point in view j, P j is the camera projection matrix of the corresponding viewing angle; since Low weights automatically weaken the contribution of observations from occluded perspectives. This mechanism can adaptively guide the optimization process to focus on high-confidence perspectives, thereby achieving more stable and reliable fusion. (2) Temporal consistency constraints: In the actual human motion process, a cross-frame dynamic modeling mechanism is introduced to provide historical reference; Specifically, for each joint point in a frame The three-dimensional position of the joint point in the previous frame is defined The Euclidean distance of is used as a penalty term, and the time regularization term is constructed as follows: The goal of this regularization term is to minimize the spatial displacement change of joints between adjacent frames, thereby enhancing the smoothness of the posture evolution over time; for different joints, the influence of this regularization term can be dynamically adjusted according to the degree of occlusion; an occlusion weight based regularization term is introduced. Time consistency weighting mechanism: in, That is, when a joint point is in an occluded state in all viewing angles, the weight of its temporal smoothing term is Tends to 1, thereby increasing the proportion of historical frames in the estimation of this point; when the observation quality of the joint point in the current frame is good, the weight of the time constraint is reduced to avoid excessive smoothing of natural movements; (3) Posture prior retrieval and constraints: When there is complete occlusion, the observation of some joints from all perspectives is unavailable, that is, This lack of information significantly impacts the stability and accuracy of pose reconstruction. To improve the model's ability to complete poses in such extreme situations, a structure-driven 3D pose prior library is introduced to provide data-driven constraint compensation. Pose prior library P = {P 1 ,P 2 ,…,P M } is constructed by aggregating large-scale, high-quality motion capture data, and each sample P m Represents a complete 3D human skeleton posture; the posture prior library has posture diversity, covers typical human behavior paradigms, and constructs a reference space with rich structural expression; When there is severe occlusion, extract the set of visible joint points in the current frame Construct an incomplete local skeleton topology; for each sample P in the prior library m , and calculate its distance with The structural similarity of the corresponding joint point subset in the , the distance error between the current skeleton topology and the prior library sample is {d1, d2, ..., d M }; Select the top K best samples based on the matching results And construct the completion result of the occluded joint points in a weighted average manner: Among them, ρ k is the normalized matching similarity score, which is calculated by the softmax transformation of the distance error: The semantic consistency constraint of the pose library is constructed by completing the occluded joint points of the current joint point and the pose prior library: The pose completion strategy essentially transforms the full occlusion problem into a structure reasoning task; (4) Skeleton consistency constraint: 3D human pose estimation requires not only data-driven observation accuracy but also consistency and rationality in the skeleton structure. To this end, explicit skeleton geometry consistency constraints are introduced during the optimization process to prevent skeleton proportion distortion caused by occlusion or optimization drift. Define a standard bone segment set ε, where each pair (i, k) represents a set of anatomically connected joint points; the three-dimensional bone segment length remains constant or within a reasonable range, so the bone length regularization term is constructed as follows: where l i,k The target bone segment length is obtained by statistically analyzing the real motion dataset and serves as the anatomical reference for the constraint; Based on the above mechanism, a unified multi-target pose estimation loss function is constructed as follows: Among them, λ t ,λ p ,λ b is the weight hyperparameter of each loss term, which is used to balance the different loss terms in the total objective function L total After completing the initial 3D pose estimation, the 3D joint coordinates are used as optimization variables to construct and minimize the multi-objective pose estimation loss function, and the initial estimation results are further optimized to finally output a more stable, reasonable and robust 3D human body pose.
Citation Information
Cited By
Dynamic image behavior analysis system and method based on time sequence association
CN121214557A
Three-dimensional attitude estimation method and device, storage medium and electronic equipment
CN121353407A
Rehabilitation scene-oriented multi-view cross-modal three-dimensional human body posture estimation method and system
CN121583006A
Three-dimensional human body posture estimation method and device based on multiple modes and occlusion perception
CN121616663A
Human body posture estimation method with shielding detection correction and kinematics constraint optimization functions
CN121789269A