Human body skeleton point positioning identification method and system based on OpenPose

Through the OpenPose-based human bone point positioning recognition method, combined with the dual-domain multi-path self-supervised diffusion model and signal-to-noise ratio evaluation technology, bone point detection and three-dimensional transformation are carried out, and bone point detection problems under complex postures and occlusion are solved, achieving high-precision and stable 3D bone point reconstruction.

CN120411545APending Publication Date: 2025-08-01GUIZHOU EDUCATION UNIV
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510574294.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

The prior art has insufficient stability and accuracy in handling complex poses, severe occlusion and 2D to 3D conversion. It is difficult to maintain the accuracy of bone point detection in highly dynamic scenarios such as dance, and traditional methods are complex in calculations, making it difficult to meet the needs of real-time applications.

Method used

The human bone point positioning recognition method based on OpenPose is adopted. After preprocessing the input image, bone point detection is carried out by combining the dual-domain multi-path self-supervised diffusion model and convolutional neural network for bone point detection, and convolutional neural network enhanced with signal-to-noise ratio evaluation and bone topology weight mapping for noise reduction. The two-dimensional to three-dimensional transformation is carried out through triangulation principle and feature matching technology, and the rate perception analysis and three-dimensional Gaussian compression algorithm are optimized, and occlusion prediction and completion are finally carried out.

Benefits of technology

It improves the accuracy and speed of bone point detection, significantly improves the accuracy and stability of bone point positioning, solves the problem of bone point occlusion recognition in complex dynamic scenarios, realizes efficient flowable free space perspective reconstruction and multi-view consistency, and improves the accuracy and robustness of 3D bone point reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120411545A_ABST
    Figure CN120411545A_ABST
Patent Text Reader

Abstract

The invention provides a human body skeleton point positioning identification method and system based on OpenPose, and the method comprises the steps: carrying out the preprocessing of an input human body image, and obtaining the preprocessing image data; performing skeleton point detection on the preprocessed image data by adopting a double-domain multi-path self-supervised diffusion model in combination with a convolutional neural network to obtain two-dimensional human skeleton point coordinate data; performing signal-to-noise ratio evaluation and noise reduction processing on the two-dimensional skeleton point coordinate data to obtain two-dimensional skeleton point data after noise reduction; performing two-dimensional to three-dimensional conversion through a triangulation principle and a feature matching technology to obtain preliminary three-dimensional skeleton point coordinate data; performing skeleton point optimization through rate perception analysis and a three-dimensional Gaussian compression algorithm; and carrying out shielding prediction and completion processing on the optimized three-dimensional skeleton points to obtain a complete human skeleton point positioning identification result. According to the invention, the problems of accuracy and stability of skeleton point positioning in a complex dynamic scene are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and particularly relates to a method and system for human skeletal point positioning and recognition based on OpenPose. Background Art

[0002] The human skeletal point positioning and recognition technology is an important research direction in the field of computer vision, and is widely used in fields such as motion capture, pose estimation, and human-computer interaction. This technology realizes the accurate understanding and analysis of human postures by identifying and tracking human key points.

[0003] Traditional human skeletal point recognition technologies mainly rely on methods based on template matching or feature extraction. For example, some methods use HOG features combined with an SVM classifier for human body part detection, and then construct a skeletal model through part connection; other methods use algorithms such as random forests to directly estimate the positions of key points from depth images.

[0004] With the development of deep learning, human pose estimation methods based on convolutional neural networks have made significant progress. Such methods process the input image through a multi-stage convolutional neural network, generate a heat map representing the probability of key point positions, and establish the association between key points through a part affinity field to achieve skeletal point detection in a multi-person scenario. This method can effectively handle complex backgrounds and multi-person overlapping situations, and has high real-time performance and accuracy.

[0005] However, the existing technologies still have obvious deficiencies in dealing with complex postures, severe occlusions, and 2D to 3D conversions. Especially in highly dynamic scenarios such as dancing, due to the complexity and rapid changes of movements, the existing methods are difficult to maintain the stability and accuracy of skeletal point detection. At the same time, traditional 2D to 3D conversion methods are computationally complex, difficult to meet the requirements of real-time applications, and have low accuracy in dealing with 3D reconstruction of occluded parts. Summary of the Invention

[0006] The purpose of the present invention is to provide a method and system for human skeletal point positioning and recognition based on OpenPose to solve the problems existing in the existing technologies in dealing with complex postures, severe occlusions, and 2D to 3D conversions, and improve the accuracy and stability of skeletal point positioning.

[0007] To achieve the above purpose, the present invention provides a method for human skeletal point positioning and recognition based on OpenPose, including:

[0008] Preprocess the input human image to obtain preprocessed image data;

[0009] Use a dual-domain multi-path self-supervised diffusion model combined with a convolutional neural network to perform skeletal point detection on the preprocessed image data to obtain two-dimensional human skeletal point coordinate data;

[0010] Perform signal-to-noise ratio evaluation on the two-dimensional human bone point coordinate data, and perform noise reduction processing on the convolutional neural network based on signal-to-noise ratio unit training and enhanced bone topology map weight mapping to obtain the denoised two-dimensional bone point data;

[0011] Based on the denoised two-dimensional bone point data, perform two-dimensional to three-dimensional conversion through the principle of triangulation and feature matching technology to obtain preliminary three-dimensional bone point coordinate data;

[0012] Based on the preliminary three-dimensional bone point coordinate data, perform bone point optimization through rate perception analysis and three-dimensional Gaussian compression algorithm to obtain the optimized three-dimensional bone point coordinate data;

[0013] Perform occlusion prediction and completion processing on the optimized three-dimensional bone point coordinate data to obtain the complete human bone point positioning and recognition result.

[0014] Optionally, preprocess the input human body image to obtain preprocessed image data, including:

[0015] Adjust the resolution of the input human body image to 1920×1080 pixels and perform cropping processing to obtain image data of standard size;

[0016] Perform adaptive histogram equalization processing on the image data of standard size to obtain image data with enhanced contrast;

[0017] Perform Gaussian filtering and median filtering on the image data with enhanced contrast to obtain the preprocessed image data.

[0018] Optionally, use a dual-domain multi-path self-supervised diffusion model combined with a convolutional neural network to detect bone points in the preprocessed image data to obtain two-dimensional human bone point coordinate data, including:

[0019] Extract features from the preprocessed image data through a multi-scale feature extraction convolutional neural network to obtain a multi-level feature map; enhance and refine the features of the multi-level feature map through the dual-domain multi-path self-supervised diffusion model to obtain an enhanced feature representation;

[0020] Based on the enhanced feature representation, use a multi-stage convolutional neural network to generate a body part heat map and a part affinity field vector to obtain the two-dimensional human bone point coordinate data.

[0021] Optionally, perform signal-to-noise ratio evaluation on the two-dimensional human bone point coordinate data, and perform noise reduction processing on the convolutional neural network based on signal-to-noise ratio unit training and enhanced bone topology map weight mapping to obtain the denoised two-dimensional bone point data, including:

[0022] The quality of the two-dimensional human skeleton point coordinate data is evaluated by a signal-to-noise ratio calculation module to obtain a skeleton point signal-to-noise ratio evaluation result;

[0023] Based on the skeleton point signal-to-noise ratio evaluation result, noise recognition and classification are performed through a convolutional neural network trained by a signal-to-noise ratio unit to obtain the skeleton point noise type and distribution characteristics;

[0024] According to the skeleton point noise type and distribution characteristics, a skeleton point noise suppression model is constructed through a skeleton topology map weight mapping enhancement algorithm to obtain the denoised two-dimensional skeleton point data.

[0025] Optionally, based on the denoised two-dimensional skeleton point data, two-dimensional to three-dimensional conversion is performed through the principle of triangulation and feature matching technology to obtain preliminary three-dimensional skeleton point coordinate data, including:

[0026] The internal and external parameter matrices of the camera are obtained through a camera calibration algorithm to obtain a camera parameter model;

[0027] Based on the camera parameter model and the denoised two-dimensional skeleton point data, multi-view skeleton point depth calculation based on the principle of triangulation is performed to obtain skeleton point depth mapping data;

[0028] The skeleton point depth mapping data is converted into three-dimensional space coordinates through a projective transformation and spatial coordinate reconstruction algorithm to obtain the preliminary three-dimensional skeleton point coordinate data.

[0029] Optionally, based on the camera parameter model and the denoised two-dimensional skeleton point data, multi-view skeleton point depth calculation based on the principle of triangulation is performed to obtain skeleton point depth mapping data, including:

[0030] Multi-view feature matching is performed on the denoised two-dimensional skeleton point data to obtain skeleton point correspondence data;

[0031] Based on the skeleton point correspondence data, parallax information is calculated to obtain the skeleton point depth mapping data.

[0032] Optionally, based on the preliminary three-dimensional skeleton point coordinate data, skeleton point optimization is performed through rate perception analysis and three-dimensional Gaussian compression algorithm to obtain optimized three-dimensional skeleton point coordinate data, including:

[0033] The motion rate and acceleration characteristics of the preliminary three-dimensional skeleton point coordinate data are calculated by a rate perception analysis module to obtain skeleton point dynamic characteristic data;

[0034] Based on the skeleton point dynamic characteristic data, a probability density distribution model of the skeleton point is established to obtain a skeleton point probability distribution representation;

[0035] Optimize and adjust the probability distribution representation of the bone points by minimizing the reprojection error and joint angle constraints to obtain the optimized three-dimensional bone point coordinate data.

[0036] Optionally, establish a probability density distribution model of bone points based on the bone point dynamic characteristic data to obtain a probability distribution representation of bone points, including:

[0037] Analyze the motion speed and acceleration characteristics in the bone point dynamic characteristic data to obtain motion characteristic parameters;

[0038] Perform multivariate Gaussian distribution modeling and probability density calculation based on the motion characteristic parameters to obtain the probability distribution representation of the bone points.

[0039] Optionally, perform occlusion prediction and completion processing on the optimized three-dimensional bone point coordinate data to obtain a complete human bone point positioning and recognition result, including:

[0040] Detect the occluded bone points in the optimized three-dimensional bone point coordinate data through a bone point visibility analysis algorithm to obtain bone point occlusion status marking data;

[0041] Extract the motion history pattern and trend of bone points through a temporal information analysis module to obtain a bone point temporal feature model;

[0042] Predict the positions of the occluded bone points based on the bone point temporal feature model and human biomechanical constraints to obtain the complete human bone point positioning and recognition result.

[0043] Correspondingly, the present application provides a human bone point positioning and recognition device based on OpenPose, including:

[0044] An image preprocessing module for preprocessing the input human image to obtain preprocessed image data;

[0045] A bone point detection module for detecting bone points in the preprocessed image data by using a dual-domain multi-path self-supervised diffusion model combined with a convolutional neural network to obtain two-dimensional human bone point coordinate data;

[0046] A noise reduction processing module for evaluating the signal-to-noise ratio of the two-dimensional human bone point coordinate data and performing noise reduction processing on a convolutional neural network enhanced by signal-to-noise ratio unit training and bone topology map weight mapping to obtain noise-reduced two-dimensional bone point data;

[0047] A coordinate conversion module for performing two-dimensional to three-dimensional conversion based on the noise-reduced two-dimensional bone point data through the principle of triangulation and feature matching technology to obtain preliminary three-dimensional bone point coordinate data;

[0048] A skeletal point optimization module, which is used to optimize the skeletal points based on the preliminary three-dimensional skeletal point coordinate data through rate perception analysis and a three-dimensional Gaussian compression algorithm to obtain optimized three-dimensional skeletal point coordinate data;

[0049] An occlusion processing module, which is used to perform occlusion prediction and completion processing on the optimized three-dimensional skeletal point coordinate data to obtain a complete human skeletal point positioning and recognition result.

[0050] The beneficial effects of the present invention are as follows:

[0051] 1. By combining the dual-domain multi-path self-supervised diffusion model with the traditional convolutional neural network, the accuracy and speed of skeletal point detection are improved;

[0052] 2. Introducing a network noise reduction mechanism with SNR unit training and G-factor mapping enhancement significantly improves the noise reduction efficiency of the convolutional neural network and the accuracy of skeletal point positioning;

[0053] 3. Adopting the rate perception 3D Gaussian compression technology realizes efficient streamable free-space perspective reconstruction and solves the multi-view consistency problem;

[0054] 4. Through an occlusion prediction algorithm based on human biomechanical constraints and temporal feature analysis, the problem of identifying occluded skeletal points in complex dynamic scenes is effectively solved;

[0055] 5. Combining the 2D-3D coordinate conversion method of multi-view triangulation and feature fusion improves the accuracy and robustness of 3D skeletal point reconstruction. Description of the Drawings

[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0057] Figure 1 It is a flowchart of the human skeletal point positioning and recognition method based on OpenPose provided by the embodiment of the present invention;

[0058] Figure 2 It is a detailed flowchart of the 2D skeletal point detection steps based on the improved OpenPose provided by the embodiment of the present invention;

[0059] Figure 3 It is a detailed flowchart of the network noise reduction steps with SNR unit training and G-factor mapping enhancement provided by the embodiment of the present invention;

[0060] Figure 4This is a detailed flowchart of the 2D to 3D bone point coordinate conversion steps provided by the embodiments of the present invention;

[0061] Figure 5 This is a structural block diagram of a human bone point positioning and recognition device based on OpenPose provided by the embodiments of the present invention. Specific embodiments

[0062] To make the above objects, features, and advantages of the present invention more apparent and understandable, the following detailed description of the specific embodiments of the present invention will be given in conjunction with the accompanying drawings.

[0063] Embodiment 1

[0064] Referring to Figure 1 , the embodiments of the present invention provide a human bone point positioning and recognition method based on OpenPose, and the method includes the following steps:

[0065] Step S1: Preprocess the input human image to obtain preprocessed image data. In this step, the system first receives human image data obtained from a monocular or multiocular camera, and these images may contain human postures in various complex scenarios. The system optimizes the original image through a series of image processing operations, including image resolution normalization, cropping, contrast enhancement, and noise suppression, etc. Specifically, the system adjusts the input image to a unified resolution (usually 1920×1080 pixels), then applies adaptive histogram equalization technology to enhance the image contrast, subsequently combines Gaussian filtering and median filtering for noise suppression, and finally corrects problems such as uneven image illumination through brightness and color balance adjustment. This series of processes ensures that the input image has a unified quality standard, reduces the interference of image quality that may be suffered in subsequent bone point detection, and lays a foundation for high-precision bone point detection.

[0066] Step S2: Skeletal point detection is performed on the preprocessed image data using a dual-domain multi-path self-supervised diffusion model combined with a convolutional neural network to obtain two-dimensional human skeletal point coordinate data. In this step, the system inputs the preprocessed, high-quality image data into an innovative neural network framework for processing. This framework first extracts a preliminary feature map using a multi-scale feature extraction convolutional neural network, which is capable of capturing human feature information in the image at different scales and levels of abstraction. The system then inputs the extracted multi-level feature map into the dual-domain multi-path self-supervised diffusion model. This model enhances and refines features by operating in both feature space and image space, incorporating a multi-path information transfer mechanism, significantly improving the accuracy and robustness of feature representation. Based on the enhanced feature representation, the system then generates body part heatmaps and part affinity field vectors using a multi-stage convolutional network. Finally, a non-maximum suppression algorithm is used to extract precise keypoint locations from the heatmaps, and affinity field vectors are used to establish connectivity between keypoints, resulting in highly accurate 2D human skeletal point coordinate data. This innovative network structure design, especially the introduction of the dual-domain multi-path self-supervised diffusion model, effectively solves the detection difficulties of traditional methods under complex backgrounds and posture changes, and improves the accuracy and robustness of skeleton point detection.

[0067] Step S3: The 2D human skeletal point coordinate data is subjected to a signal-to-noise ratio (SNR) assessment. De-noising is then performed using a convolutional neural network (CNN) trained with SNR units and enhanced with skeletal topology weight mapping, yielding the de-noised 2D skeletal point data. This step aims to improve the quality and reliability of the skeletal point coordinates. The system first assesses the quality of each detected skeletal point using a specially designed signal-to-noise ratio (SNR) calculation module. This module analyzes the characteristics of the skeletal point's heat map and its consistency with prior knowledge of human anatomy to generate detailed SNR assessment results. The system then uses these assessment results to conduct in-depth analysis using a convolutional neural network specifically trained with SNR units. This network can accurately identify and classify different types of skeletal point noise. The system then introduces an innovative skeletal topology weight mapping enhancement algorithm. This algorithm first constructs a topological map of the human skeleton. It then assigns specific weights to each skeletal point based on its importance and stability in the overall posture, creating a personalized denoising parameter set for each skeletal point. Finally, the system applies these parameter sets to adaptively denoise the skeletal point coordinates, specifically adjusting the filtering strength and parameters to ensure accurate positioning of important skeletal points while suppressing noise interference. This adaptive denoising method based on signal-to-noise ratio evaluation and bone topology structure significantly improves the quality and stability of bone point coordinates, providing high-quality 2D bone point data for subsequent 3D conversion.

[0068] Step S4: Based on the denoised two-dimensional skeletal point data, perform two-dimensional to three-dimensional conversion through the principle of triangulation and feature matching technology to obtain preliminary three-dimensional skeletal point coordinate data. In this step, the system first obtains the internal parameters (such as focal length, principal point coordinates) and external parameters (such as position, attitude) matrices of all cameras in the system through the camera calibration algorithm, and constructs an accurate camera parameter model. Subsequently, based on these camera parameters and the denoised high-quality 2D skeletal point data, the system applies the principle of triangulation to calculate the depth of skeletal points from multiple perspectives. Specifically, when a skeletal point is recognized in images from multiple perspectives, the system determines the position of this point in 3D space by solving the problem of ray intersection under these perspectives. To improve the matching accuracy, the system uses feature matching technology to analyze the correspondence of skeletal points from different perspectives and solve the consistency problem of skeletal points between perspectives. In addition, the system also combines the least squares method or the RANSAC algorithm to process the noise and errors in multi-perspective data, and introduces temporal information and prior knowledge of human motion for optimization. Finally, through projective transformation and spatial coordinate reconstruction algorithms, the system converts the skeletal point depth mapping data into coordinates in the 3D world coordinate system, and further optimizes the positions of the 3D skeletal points by introducing human skeletal length constraints and joint angle constraints to ensure compliance with human anatomical characteristics. This step lays the foundation for subsequent high-precision optimization through accurate 2D-3D conversion.

[0069] The camera calibration process uses an improved Zhang-Zhengyou calibration method, which is carried out by using a standard 9×6 format checkerboard calibration board. Each corner point on this calibration board has accurate world coordinates. During calibration, the system collects at least 15 calibration board images from different angles to ensure that the calibration board covers different regions of the camera's field of view. For the acquisition of the internal parameter matrix, the system identifies the corner points on the calibration board, establishes the correspondence between the world coordinate system and the image coordinate system, and calculates the internal parameter matrix containing the focal length (fx, fy), principal point coordinates (cx, cy), and skew factor s, in the form of:

[0070]

[0071] For example, for a camera with a resolution of 1920×1080, typical internal parameter matrix values are:

[0072]

[0073] At the same time, the system also calculates the radial distortion parameters (k1, k2, k3) and tangential distortion parameters (p1, p2) to correct the image distortion caused by the lens. For the external parameter matrix, the system solves the Perspective-n-Point (PnP) problem to obtain the rotation matrix R (3×3) and translation vector t (3×1) of each camera relative to the world coordinate system, and forms the external parameter matrix [R|t].

[0074] The construction of the camera parameter model is achieved by combining the internal parameter matrix K, distortion parameters, and external parameter matrix [R|t] obtained above into a unified projection model. For the 3D world coordinate point X = (X, Y, Z, 1) T , its projected pixel coordinates x = (u, v, 1) on the image plane T satisfy the relation: x = K[R|t]X. The system further adopts the Bundle Adjustment global optimization algorithm to optimize the camera parameter model by minimizing the reprojection error of all calibration points, controlling the root mean square value of the reprojection error within 0.3 pixels, and ensuring high-precision parameter estimation.

[0075] Step S5: Based on the preliminary 3D skeleton point coordinate data, perform skeleton point optimization through rate perception analysis and 3D Gaussian compression algorithm to obtain the optimized 3D skeleton point coordinate data. In this step, the system first analyzes the dynamic characteristics of the preliminary 3D skeleton point coordinates through the rate perception analysis module. This module calculates the changes in the positions of skeleton points between adjacent time frames, obtains parameters such as the instantaneous velocity, acceleration, and angular velocity of each skeleton point, pays special attention to the fast motion patterns in highly dynamic scenarios (such as dancing), and performs segmented analysis on the skeleton point trajectories to identify key action segments. Then, based on these dynamic characteristic data, the system applies an innovative 3D Gaussian compression algorithm to construct a probability distribution model of skeleton points. This algorithm represents each skeleton point as a multivariate Gaussian distribution, where the mean vector corresponds to the estimated position of the skeleton point, and the covariance matrix describes the position uncertainty in different directions. The system adaptively adjusts the Gaussian distribution parameters by analyzing the rate characteristics and historical trajectories of skeleton points, making the fast-moving skeleton points have greater uncertainty in the motion direction while maintaining high precision perpendicular to the motion direction. Finally, the system globally optimizes the probability distribution representation of skeleton points through reprojection error minimization and joint angle constraints, adjusts the position parameters of each skeleton point, and ensures requirements such as multi-view observation data consistency, human skeleton length constraints, and joint angle physiological rationality. This optimization method based on rate perception and probability model can effectively handle the skeleton point tracking problem in high-speed motion scenarios and output high-precision and physically reasonable 3D skeleton point coordinate data.

[0076] Step S6: Perform occlusion prediction and completion processing on the optimized 3D bone point coordinate data to obtain a complete human bone point positioning and recognition result. In this step, the system first detects and marks the occluded bone points through a bone point visibility analysis algorithm. This algorithm comprehensively analyzes the detection confidence of bone points from each perspective, multi-perspective consistency, and depth relationships to determine whether a bone point is occluded. The system will pay special attention to those bone points with low confidence in most perspectives or significant inconsistencies in the 3D reconstruction positions from different perspectives. In addition, the system also uses depth information and the human body model for self-occlusion analysis to identify invisible bone points caused by the mutual occlusion of human body parts. Then, the system extracts the motion history patterns of bone points through the temporal information analysis module. This module establishes a time window, collects the historical positions and motion parameters of bone points, and applies a temporal pattern recognition algorithm to learn the motion laws of bone points, especially focusing on the periodic characteristics and transition patterns of actions. Subsequently, based on these temporal feature models, combined with human biomechanical constraints and bone geometric relationships, the system predicts the positions of the occluded bone points. The system first applies bone connection constraints and physiological joint angle limitations to narrow the search range of possible positions, and then further optimizes the prediction results by combining the pose prior knowledge learned from a large-scale action database. Finally, the system evaluates each candidate position through a multi-hypothesis verification and probability fusion algorithm, selects the best position as the final prediction result, and integrates these prediction results with the data of unoccluded bone points to output a complete, coherent, and accurate human bone point positioning and recognition result. This occlusion prediction method based on multi-source information fusion effectively solves the problem of bone point occlusion in complex scenarios and achieves stable and accurate human pose tracking.

[0077] Further, in step S1, the input human body image is preprocessed to obtain preprocessed image data, including:

[0078] Step S11: Adjust the resolution of the input human body image to 1920×1080 pixels and perform cropping processing to obtain image data of a standard size;

[0079] Step S12: Perform adaptive histogram equalization processing on the image data of the standard size to obtain image data with enhanced contrast;

[0080] Step S13: Perform Gaussian filtering and median filtering on the image data with enhanced contrast to obtain the preprocessed image data.

[0081] Further, referring to Figure 2 , in step S2, a dual-domain multi-path self-supervised diffusion model combined with a convolutional neural network is used to detect bone points from the preprocessed image data to obtain two-dimensional human bone point coordinate data, including:

[0082] Step S21: Extract features from the preprocessed image data through a multi-scale feature extraction convolutional neural network to obtain multi-level feature maps. In this step, the system designs an efficient multi-scale feature extraction network structure. This network adopts a hierarchical feature extraction strategy and can capture rich information in the image from different abstraction levels and different spatial scales. Specifically, the front end of the network uses a backbone network (such as ResNet or EfficientNet) for basic feature extraction. These networks are pre-trained and have powerful feature representation capabilities. Subsequently, the system fuses feature maps at different levels through a Feature Pyramid Network (FPN) or a similar structure to generate multi-scale feature representations. This fusion method enables the network to simultaneously focus on the overall structure of the human body at a large scale and local detail features at a small scale. In addition, the system also introduces an attention mechanism, enabling the network to automatically learn to focus on key regions in the image and reduce background interference. Finally, this multi-scale feature extraction network outputs a set of feature maps with different resolutions and abstraction levels, providing a multi-level and multi-angle image representation for subsequent feature enhancement and refinement, greatly improving the robustness and accuracy of skeleton point detection in complex scenarios.

[0083] Step S22: Enhance and refine the multi-level feature maps through the dual-domain multi-path self-supervised diffusion model to obtain enhanced feature representations. In this step, the system introduces an innovative dual-domain multi-path self-supervised diffusion model, which is a major improvement over traditional feature processing methods. This model operates simultaneously in the feature domain and the image domain, improving the accuracy of feature representation through dual-domain information interaction. In the feature domain, the model applies a self-supervised diffusion process, gradually improving the feature quality by iteratively propagating and refining feature information; in the image domain, the model maintains the perception of the original image spatial information to ensure the consistency between features and the actual image structure. The multi-path design enables the model to process features in parallel through different information channels, capturing different types of human pose features, such as edges, textures, and shapes. In addition, the self-supervised learning strategy enables the model to learn useful representations from unlabeled data, significantly reducing the dependence on a large amount of labeled data. The model adopts contrastive learning techniques and consistency regularization during training to further improve the discriminability and stability of the features. After this dual-domain multi-path self-supervised diffusion processing, the original multi-level feature maps are significantly enhanced, and the expression ability and discrimination ability of the features are greatly improved, laying a solid foundation for subsequent skeleton point localization.

[0084] Specifically, the specific construction method of the dual-domain multi-path self-supervised diffusion model adopts a combined design of deep learning and diffusion models. In terms of network architecture, this model uses U-Net as the basic framework, including an encoder-decoder structure, and a total of 12 convolutional layers are designed, with a 3×3 convolutional kernel used in each layer. The encoder part contains 4 downsampling modules, each module contains two convolutional layers and a max pooling layer, and the number of channels is 64, 128, 256, and 512 in sequence; the decoder part correspondingly contains 4 upsampling modules, and transposed convolution is used to restore the size of the feature map. In the dual-domain processing of the feature domain and the image domain, the model designs a feature domain branch and an image domain branch respectively. The feature domain branch focuses on information processing in the high-dimensional feature space, and uses residual connections and attention mechanisms to enhance feature expression; the image domain branch maintains the perception of the original image structure and uses a spatial transformation network to align the image structure and features. These two branches perform information interaction through a specially designed dual-domain fusion module, which contains channel attention and spatial attention units to ensure the consistency of the image structure during the feature enhancement process.

[0085] In terms of multi-path design, the model constructs three parallel information channels, which are responsible for processing edge information, texture information, and shape information respectively. Each channel uses convolutional kernels of different scales (3×3, 5×5, and 7×7) to capture features of different scales. Through a weighted fusion module, the system dynamically adjusts the weights of different channels to optimize the final feature representation. The training of this model adopts a two-stage strategy: first, it is pre-trained on the MS-COCO and MPII human pose datasets (a total of 1.8 million images), and the contrastive learning method is used to construct positive and negative sample pairs; then it is fine-tuned on a specific task dataset, and a combined loss function including reconstruction loss (L1 and L2 norms), perceptual loss, and adversarial loss is adopted. The key hyperparameter settings include: the initial learning rate is 0.001, and the cosine annealing strategy is adopted; the batch size is 32; the number of training epochs is 100 for pre-training and 50 for fine-tuning; the optimizer is Adam, and the momentum parameters are β1 = 0.9 and β2 = 0.999.

[0086] The weighted fusion module is the core component in the dual-domain multi-path self-supervised diffusion model, which is used to integrate the information extracted from multiple parallel feature channels. This module uses the attention mechanism to dynamically adjust the weight contributions of different channels to ensure that the final feature representation can retain the key information required for bone point localization to the greatest extent. Specifically, for the feature maps Fedge, Ftexture, and Fshape extracted from the edge channel, texture channel, and shape channel, the weighted fusion module first extracts channel-level feature vectors vedge, vtexture, and vshape through global average pooling, and then generates channel attention weights through a two-layer fully connected network:

[0087] α = softmax(W2·ReLU(W1·[v edge ; v texture ; v shape + b1) + b2)

[0088] where W1 and W2 are weight matrices, b1 and b2 are bias vectors, and α = (α edge , α texture , α shape ) is the weight vector of three channels. The finally fused feature map F fusion is obtained by weighted summation:

[0089] F fusion = α edge ·F edge + α texture ·F texture + α shape ·F shape

[0090] In implementation, the weight α is not only affected by the input features but also constrained by the context of the current processing task. For example, when dealing with a complex texture background, the system will automatically reduce the weight of the texture channel to reduce interference; while when dealing with a low-contrast scene, it will increase the weight of the edge channel to enhance the ability to detect bone points. In addition, the module also implements a spatial adaptive weight allocation mechanism, allowing different channel weight configurations to be used in different regions of the image to cope with local variations in complex scenes. Through this dynamic weight adjustment mechanism, the weighted fusion module can effectively integrate the advantages of different channels, improve the robustness and discriminative ability of feature representation, and ultimately improve the accuracy of bone point detection, especially the performance under complex backgrounds and lighting conditions.

[0091] In the design of the diffusion process, this model implements a 100-step noise addition and denoising process, using a linear noise scheduling strategy. Self-supervised learning is achieved through feature masking and reconstruction tasks, that is, randomly masking 30% of the input features and then training the model to recover these masked regions. To enhance the generalization ability of the model, consistency regularization technology is also introduced. By applying different data augmentation methods to the same input, the model is required to produce consistent feature representations under different transformations. The model performance is evaluated using the Average Precision (AP) and PCK (Percentage of Correct Keypoints) metrics. The test results on the public dataset show that, compared with traditional methods, the detection accuracy of this model has increased by 15% under complex backgrounds and pose changes, and the processing speed has increased by 23%.

[0092] Step S23: Based on the enhanced feature representation, a multi-stage convolutional neural network is used to generate body part heatmaps and part affinity field vectors, obtaining the two-dimensional human body bone point coordinate data. In this step, the system utilizes the feature-enhanced representation and, through a carefully designed multi-stage convolutional neural network architecture, simultaneously generates two types of key outputs: body part heatmaps and part affinity field vectors. The body part heatmaps are a set of two-dimensional probability maps, with each map corresponding to a human key point type (such as the head, shoulders, knees, etc.), and each pixel value in the map represents the probability that the corresponding key point exists at that position. The part affinity field vectors are a set of two-dimensional vector fields used to encode the association relationships between adjacent bone points, indicating the connection direction and strength between bone points. The multi-stage architecture adopted by the system allows for the gradual refinement of these prediction results, with each stage improving based on the output of the previous stage, forming a progressive optimization process. In the final stage, the system extracts the local maximum points from the heatmaps as the precise positions of the bone points through the non-maximum suppression algorithm and uses the part affinity field vectors to guide the connection between bone points, forming a complete human body bone structure. This dual-path prediction method combining heatmaps and affinity fields can effectively handle the bone point detection problem in multi-person scenarios and maintain accuracy even in complex situations where humans overlap, ultimately outputting high-precision 2D human body bone point coordinate data.

[0093] Further, referring to Figure 3 , in step S3, the signal-to-noise ratio of the two-dimensional human body bone point coordinate data is evaluated, and through convolutional neural network noise reduction processing based on signal-to-noise ratio unit training and bone topology map weight mapping enhancement, the denoised two-dimensional bone point data is obtained, including:

[0094] Step S31: The signal-to-noise ratio calculation module performs quality assessment on the two-dimensional human bone point coordinate data to obtain the bone point signal-to-noise ratio assessment result. In this step, the system introduces a specially designed signal-to-noise ratio (SNR) calculation module to comprehensively evaluate the quality and reliability of each detected bone point. This module first analyzes the peak characteristics of the heat map, including parameters such as the shape, height of the peak, and the contrast with the surrounding area. Sharp and prominent peaks usually indicate high confidence in bone point detection, while a flat or multi-peak heat map may imply positioning uncertainty. Then, the module calculates the statistical distribution characteristics of the pixel values around each bone point, such as variance, skewness, and kurtosis, which can reflect the stability and accuracy of bone point detection. In addition, the system also introduces prior knowledge based on human anatomy to evaluate whether the detected bone points conform to the basic laws of human configuration, such as bone length ratio, joint range of motion, etc. This assessment considers the spatial relationship and biomechanical constraints between bone points and can identify abnormal points that, although having a high response on the heat map, do not conform to human configuration. Finally, the system synthesizes the analysis results from multiple dimensions above to calculate a comprehensive signal-to-noise ratio index for each bone point, which quantifies the quality and reliability of bone point detection and provides an important basis for subsequent precise noise reduction. Through this multi-dimensional quality assessment method, the system can more accurately identify the noise points that need to be processed, improving the pertinence and effectiveness of noise reduction processing.

[0095] The signal-to-noise ratio (SNR) calculation module uses a multi-dimensional feature analysis method to accurately evaluate the bone point quality. This module first performs quantitative analysis on the bone point heat map and extracts peak features, including peak intensity Ipeak, the ratio of peak to background average value Rpb, the sharpness of the peak Speak (calculated by the mean of the second derivative of the peak region), and the shape feature Hshape of the peak region (obtained by calculating the ratio of the major axis to the minor axis of the peak region). Specifically, the module first locates the maximum value point (x max , y max ) and its intensity Ipeak in the heat map, then selects a circular region with a radius of r (usually r is set to 5% of the heat map resolution) centered on this point, calculates the average value I region of the pixel values within this region and the average value I bg of the pixel values outside this region, and then calculates the peak-to-background ratio R pb = I peak / I bg . For the peak sharpness S peak , the module calculates the average absolute value of the Laplacian operator response within the peak region. In addition, the module also analyzes the spatial relationship C spatial between the bone point and its adjacent bone points, and evaluates its structural consistency by calculating the deviation between the current bone point position and the inferred position based on adjacent bone points.

[0096] Step S32: Based on the bone point signal-to-noise ratio (SNR) evaluation results, perform noise identification and classification through a convolutional neural network trained by an SNR unit to obtain the bone point noise type and distribution characteristics. In this step, the system uses a specially trained convolutional neural network to deeply analyze and accurately classify bone point noise. The training process of this network adopts an innovative SNR unit training strategy, that is, different types and degrees of noise are deliberately introduced into the training data, enabling the network to learn to identify various noise patterns. The network architecture adopts a deep residual structure combined with a self-attention mechanism, which can effectively capture the long-range dependencies and local context information between bone points. The data input into this network includes the original bone point coordinates, heatmap features, and the SNR evaluation results generated in step S31. Through multi-layer feature extraction and analysis, the network can accurately classify bone point noise into multiple types, such as random jitter noise (unstable detection caused by image quality problems), systematic offset noise (consistent offset caused by algorithm deviation), heatmap multi-peak interference (multi-candidate point interference caused by complex background or similar human body parts), etc. At the same time, the network can also analyze the spatial distribution characteristics of the noise, such as whether the noise is concentrated in specific body parts, whether it is related to the human body posture, and whether it presents a certain pattern in the time series. These detailed analyses of noise types and distribution characteristics provide accurate guiding information for the next step of personalized noise reduction strategies, enabling the system to adopt the most suitable processing methods for different types of noise and significantly improving the noise reduction effect.

[0097] Specifically, to achieve efficient bone point noise identification and classification, the SNR unit-trained convolutional neural network in this method adopts a specially designed network architecture and a dedicated training strategy. This network structure is improved based on the ResNet-50 backbone and contains a total of 5 stages, with each stage containing 3, 4, 6, 3, and 2 residual blocks respectively. The network input is a multi-channel tensor containing bone point coordinates and corresponding heatmaps, with a size of 224×224×(K + 3), where K is the number of bone points, and the additional 3 channels contain the original RGB image information. The core innovation of the network lies in the specially designed SNR perception units, which are embedded in each residual block and analyze the local and global signal-to-noise characteristics of bone points through a parallel branch structure. Each SNR unit contains three key components: a signal strength evaluation sub-unit that uses 1×1 convolution to extract the signal strength characteristics of each bone point; a noise pattern analysis sub-unit that uses 3×3 dilated convolution to analyze the noise distribution in the surrounding area of the heatmap; and a signal-to-noise ratio calculation sub-unit that combines the outputs of the previous two sub-units to calculate the signal-to-noise ratio characteristics of bone points.

[0098] The signal-to-noise ratio calculation subunit adopts a weighted feature fusion and non-linear transformation mechanism to integrate the outputs of the signal intensity evaluation subunit and the noise pattern analysis subunit into a comprehensive signal-to-noise ratio feature representation. Specifically, let the output of the signal intensity evaluation subunit be S ∈ R C×H×W , and the output of the noise pattern analysis subunit be N ∈ R C'×H×W , where C and C' are the numbers of feature channels respectively, and H and W are the height and width of the feature map. The signal-to-noise ratio calculation subunit first unifies the number of channels of the two feature maps to C” through a 1×1 convolutional layer to obtain the transformed feature maps S' and N', and then calculates the signal-to-noise ratio feature F SNR :

[0099] F SNR = φ(W s ·S t ⊙ W N ·N′)

[0100] where WS and WN are channel attention weight matrices, which are adaptively learned through the Squeeze-and-Excitation network; ⊙ represents the element-wise product operation; φ(·) is a non-linear activation function.

[0101] The training data construction adopts a synthetic noise enhancement strategy. Based on the COCO and MPII human pose datasets (a total of about 250,000 annotated images), training samples containing 5 typical noise types are created: random jitter noise, systematic offset noise, multi-modal interference noise, false detections caused by occlusion, and low-confidence noise. For each noise type, a specific mathematical model is designed for parametric generation. For example, random jitter noise is realized by adding a random offset (mean 0, standard deviation adjustable) that follows a Gaussian distribution; systematic offset noise is simulated by adding a fixed-distance offset in a specific direction to simulate algorithmic bias. The training process adopts a two-stage strategy: in the first stage, the cross-entropy loss function is used to train the network to identify different types of noise, and the learning rate is set to 0.0005; in the second stage, a regularization loss term is introduced to optimize the robustness of the network under different noise intensities, and the learning rate is reduced to 0.0001. The batch size is set to 64, the Adam optimizer parameters are β1 = 0.9, β2 = 0.999, the total number of training epochs is 120, and the learning rate is decayed (multiplied by 0.1) at the 80th and 100th epochs.

[0102] The "specific mathematical model" referred to here refers to a series of mathematical formulas and algorithms that precisely describe different types of noise, used to generate controllable synthetic noise in training data. For random jitter noise, the system adopts a two-dimensional Gaussian perturbation model; for system offset noise, a directional offset model is used; for multi-peak interference noise, the system uses a mixture Gaussian model to generate multiple peaks on the heat map; for low-confidence noise, the system applies a peak suppression and background enhancement model. The above mathematical models are prior art and will not be elaborated in the embodiments of this application.

[0103] To evaluate the network performance, we designed a dedicated test set containing different intensities and combinations of noise types. The evaluation metrics include the accuracy of noise type recognition, the mean absolute error of noise intensity estimation, and the percentage increase in the accuracy of bone points after noise reduction. The experimental results show that the network achieves an accuracy of 92.7% in noise type recognition, the mean absolute error of noise intensity estimation is 0.08, and the average improvement in the localization accuracy of bone points is 17.3%. It performs particularly prominently in complex background and fast-moving scenarios.

[0104] Step S33: According to the bone point noise type and distribution characteristics, construct a bone point noise suppression model through the bone topology graph weight mapping enhancement algorithm to obtain the denoised two-dimensional bone point data. In this step, based on the noise characteristics identified in the previous step, the system introduces an innovative bone topology graph weight mapping enhancement algorithm to construct a personalized noise reduction processing strategy for each bone point. The algorithm first establishes a topological structure graph of the human skeleton, modeling the human body as a graph structure composed of bone points (nodes) and bone connections (edges). Then, the algorithm assigns specific weight coefficients to each node in the graph, and these weights reflect the importance, stability, and sensitivity to noise of the bone points in the overall posture. For example, the bone points in the trunk area are usually more stable than those at the ends of the limbs, so they will be given higher weights. Next, the system maps the noise type and characteristics of the bone points onto this weighted topology graph, and through graph convolution operations or message passing algorithms, propagates and synthesizes information on the graph structure to generate a noise suppression model considering the global bone structure. Based on this model, the system customizes a set of noise reduction parameters for each bone point, including filter type, kernel size, intensity coefficient, and time-domain smoothing factor, etc. Finally, the system applies these personalized parameters to precisely denoise the bone point coordinates, using methods such as adaptive convolution filtering, Kalman filtering, or graph structure-based smoothing algorithms to specifically suppress different types of noise while retaining useful motion details. For bone points with high confidence, the system uses slight filtering to retain the original details; while for bone points with low confidence or high noise, stronger filtering or constraints based on surrounding high-confidence points are applied for correction. In addition, the system also combines temporal information for trajectory smoothing to ensure the coherence and naturalness of the bone point motion. Through this personalized noise reduction processing based on the bone topology structure, the system can effectively remove various types of noise while maximizing the retention of the true details of human motion, output high-quality 2D bone point data, providing a reliable basis for subsequent 3D conversion.

[0105] Specifically, in step S31, based on the 2D human bone point coordinate data output by step S2, the system evaluates the quality of each detected bone point through a specially designed signal-to-noise ratio (SNR) calculation module. Specifically, this module calculates the confidence score and noise interference degree of each bone point by analyzing the distribution characteristics of pixel values around the bone point and the contrast between the peak value of the heat map and the surrounding area. This process takes into account multiple factors such as the shape of the heat map peak, the contrast ratio between the peak and the background, and the consistency between the bone point and the prior knowledge of the human body configuration, so as to generate the SNR evaluation result of the bone point, providing a decision basis for subsequent noise reduction processing.

[0106] The system uses a multi-level feature analysis framework to calculate the confidence score and noise interference degree of the skeleton points. For the confidence score calculation, the system first extracts the skeleton point position (x p ,y p ) of the peak intensity P v =H(x p ,y p ), and then in (x p ,y p ) as the center and the radius r1 (usually set to 7% of the heat map resolution) in the circular area Ω1 to calculate the local statistical features, including the regional average μ1, variance σ1 2 And kurtosis k1:

[0107]

[0108] At the same time, in (x p ,y p ) as the center, inner diameter r1, outer diameter r2 (usually set to 15% of the thermal map resolution) in the annular region Ω2 to calculate the background statistical features, including the regional mean μ2 and variance σ2 2 Based on these statistical features, the system calculates the peak clarity index Cp and the peak significance index Sp:

[0109]

[0110] In step S32, the system leverages the signal-to-noise ratio (SNR) assessment results from step S31 and applies a convolutional neural network, specifically trained with the SNR unit, for further analysis. Pre-trained on a large dataset containing varying noise types and levels, this network accurately identifies and classifies low-confidence skeletal point noise types, such as random jitter, system offset, or multimodal interference in heatmaps. The network architecture utilizes residual connections and an attention mechanism to ensure sensitivity to subtle noise variations. It also outputs detailed skeletal point noise type and spatial distribution characteristics, laying the foundation for precise, targeted noise reduction.

[0111] The so-called "special training" refers to a specific training strategy designed for the task of skeletal point noise recognition, including specialized data augmentation schemes, objective function design, and training process control. This training process first constructs a synthetic noise training set based on the COCO and MPII human pose datasets by adding five typical noises to the standard pose data: random jitter noise (implemented by adding Gaussian noise with σ = 1.0 - 3.0 pixels), systematic offset noise (implemented by adding an offset of 2.0 - 5.0 pixels in a fixed direction), multi-peak interference noise (implemented by adding secondary peaks in the heatmap, with peak intensity being 40% - 70% of the main peak), false detections caused by occlusion (implemented by simulating partial occlusion scenarios and adding error detection points), and low-confidence noise (implemented by reducing the heatmap peak and increasing background noise). For each noise type, the training data contains different intensity levels (mild, moderate, severe), forming diverse training samples.

[0112] The special training process adopts a three-stage progressive learning strategy: In the first stage (pre-training), supervised learning with noisy labels is used, and the cross-entropy loss function is adopted to train the network to recognize different types of noises, with the initial learning rate set to 0.0005; in the second stage (hard sample mining), a specifically designed focal loss function (FocalLoss) part is added, focusing on optimizing the network's ability to recognize difficult-to-distinguish noise samples, and the learning rate is reduced to 0.0003; in the third stage (robustness optimization), adversarial training technology is introduced, and by generating adversarial samples and incorporating them into the training process, the network's generalization ability on unseen noise types is improved, and the learning rate is further reduced to 0.0001. In particular, an SNR-aware sampling strategy is implemented during the training process, that is, the appearance frequency of samples in the training batch is dynamically adjusted according to the signal-to-noise ratio characteristics of the samples, ensuring that the network fully learns the noise characteristics at various SNR levels.

[0113] In step S33, for the noise characteristics identified in step S32, an innovative bone topology map weight mapping enhancement algorithm is introduced. This algorithm first constructs a bone topology map, and then assigns specific weights to each bone point, and these weights reflect the importance and stability of the bone point in the overall pose. By mapping the noise type and weights into a multi-dimensional feature space, the system constructs a personalized noise reduction parameter set for each bone point, including filter kernel size, intensity coefficient, and time-domain smoothing factor, etc., realizing the precise customization of noise reduction processing. Finally, according to these noise reduction parameter sets, the system applies an adaptive convolution filter to each bone point coordinate for precise noise reduction, and outputs high-quality 2D bone point data after noise reduction.

[0114] Among them, the specific implementation of the bone topology map weight mapping enhancement algorithm is designed based on graph theory and human biomechanics principles. First, the algorithm constructs a human bone topology map, where the vertex set represents bone points and the edge set represents bone connection relationships. For each bone point, the algorithm calculates its importance weight, which is determined by the combination of three factors: the centrality factor, the stability factor, and the functional importance factor. The centrality factor is calculated through the centrality measure of the graph, specifically using betweenness centrality, which measures the connection importance of bone points in the overall skeleton structure. The stability factor is calculated based on the coefficient of variation of the positions of bone points in historical frames. The smaller the coefficient of variation, the more stable the bone point, and the higher the corresponding weight. The functional importance factor is preset according to the human biomechanics model. For example, bone points in the trunk area (such as the neck, shoulders, and pelvis) have higher functional importance, while end bone points (such as fingers and toes) have lower importance.

[0115] The coefficient of variation of position is a statistical index that quantifies the position stability of bone points, defined as the ratio of the standard deviation of the position change of bone points in consecutive time frames to the average displacement. Specifically, for bone point v, consider its three-dimensional coordinate sequence in the past N frames (usually N = 30 to 60 frames, about 1 to 2 seconds of data) where t0 represents the current time frame. Where ||·|| represents the Euclidean distance. Then calculate the average value μd and standard deviation σd of these displacements:

[0116]

[0117] The coefficient of variation of position CV_d is defined as the ratio of the standard deviation to the average value, that is:

[0118]

[0119] This coefficient measures the relative variation degree of the displacement of bone points, and the smaller the CV d value, the more stable the movement of the bone point.

[0120] The centrality factor is calculated using the betweenness centrality measure in graph theory, which is used to quantify the connection importance of bone points in the overall skeleton structure. In the human bone topology map G = (V, E), V represents the set of bone points and E represents the bone connection relationship. For bone point v ∈ V, its betweenness centrality BC(v) is defined as:

[0121]

[0122] where σ stDenote the number of shortest paths from skeleton point s to skeleton point t, and σst(v) denote the number of these paths passing through skeleton point v. In the actual calculation process, the system uses the Brandes algorithm to efficiently calculate the betweenness centrality of all skeleton points. The specific steps are as follows:

[0123] For each skeleton point s ∈ V, initialize:

[0124] Distance array dist[t] = ∞ (for all t ≠ s) and dist[s] = 0

[0125] Path array σ[t] = 0 (for all t ≠ s) and σ[s] = 1

[0126] Predecessor node list P[t] = empty list (for all t)

[0127] Node stack S = empty stack

[0128] Use breadth-first search to calculate single-source shortest paths:

[0129] Maintain a queue Q, initially containing s

[0130] When Q is not empty, take out the head node v

[0131] For each neighbor w of v, if dist[w] = ∞, then set dist[w] = dist[v] + 1 and enqueue w

[0132] If dist[w] = dist[v] + 1, then update σ[w] = σ[w] + σ[v] and add v to P[w]

[0133] After searching all neighbors of v, push v onto the stack S

[0134] Accumulate the betweenness centrality values:

[0135] Initialize the dependency value δ[v] = 0 (for all v)

[0136] When the stack S is not empty, pop the top node w

[0137] For each predecessor node v ∈ P[w] of w, calculate δ[v] = δ[v] + (σ[v] / σ[w])·(1 + δ[w])

[0138] If w ≠ s, then accumulate BC[w] = BC[w] + δ[w]

[0139] Normalize the betweenness centrality values:

[0140] For each skeleton point v, calculate the normalized betweenness centrality:

[0141] NBC(v) = BC(v) / ((|V| - 1)·(|V| - 2) / 2)

[0142] The weight calculation adopts the weighted average method, comprehensively considering the above three factors, and the weight coefficient determines the optimal value through cross-validation. After obtaining the weights of the skeleton points, the algorithm constructs the Laplacian matrix of the skeleton topology graph to capture the overall topological features of the skeleton structure. Then, based on the noise type and distribution characteristics of the skeleton points, the algorithm constructs a personalized filtering kernel function: for random jitter noise, a Gaussian filtering kernel is used; for systematic offset noise, a directional correction filtering kernel is used; for multi-modal interference noise, a non-local mean filtering kernel is used.

[0143] The filtering intensity parameter is adaptively adjusted according to the weights of the skeleton points: the higher the weight, the more important or stable the skeleton point is, and the lower the filtering intensity is used to retain more original details; the lower the weight, the higher the filtering intensity is used to more effectively suppress noise. Finally, the noise reduction process for each skeleton point is achieved through an optimization objective, minimizing the noise impact while maintaining the original position characteristics of the skeleton points. The algorithm updates the positions of all skeleton points in each iteration through iterative solution until convergence or the maximum number of iterations is reached. Experimental verification shows that while retaining the human motion characteristics, the algorithm effectively reduces various noise interferences, and the average noise reduction efficiency is increased by 32%, especially having a prominent effect on complex noises in fast-moving scenarios.

[0144] Further, referring to Figure 4 , in step S4, based on the denoised two-dimensional skeleton point data, two-dimensional to three-dimensional conversion is performed through the triangulation principle and feature matching technology to obtain preliminary three-dimensional skeleton point coordinate data, including:

[0145] Step S41: Obtain the camera intrinsic and extrinsic parameter matrices through the camera calibration algorithm to obtain the camera parameter model. In this step, the system first conducts a comprehensive camera calibration to obtain accurate camera geometric parameters, which is the basis for 2D to 3D conversion. The calibration process usually adopts Zhang Zhengyou's calibration method or its optimized variants, combined with a specially designed calibration board for multi-angle shooting. The feature points on the calibration board have a known geometric layout. The system detects the 2D image positions of these feature points from different perspectives and combines their known 3D spatial positions to establish a mapping relationship between the 2D image coordinate system and the 3D world coordinate system. By solving a series of non-linear equations (here, "a series of non-linear equations" refers to the mathematical equations constructed based on the camera projection model, which are used to establish the correspondence between points in the three-dimensional world coordinate system and their projections on the two-dimensional image plane), the system calculates the intrinsic parameter matrix and the extrinsic parameter matrix of each camera. The intrinsic parameter matrix contains the basic optical characteristics of the camera, such as focal length, principal point coordinates (image center), and pixel aspect ratio, etc.; it also includes lens distortion parameters, such as radial distortion and tangential distortion coefficients, which are used to correct the image distortion caused by the lens. The extrinsic parameter matrix describes the position (translation vector) and orientation (rotation matrix) of the camera in the 3D world coordinate system, defining the spatial transformation relationship between the camera coordinate system and the world coordinate system. To improve the accuracy and stability of calibration, the system also adopts a global optimization algorithm, comprehensively considering the observation data from all perspectives, minimizing the reprojection error, and obtaining the optimal parameter estimation. In addition, the system regularly verifies and fine-tunes the calibration parameters to cope with possible slight movements or parameter drifts of the camera. These accurate camera parameter models provide a reliable geometric basis for subsequent triangulation and 3D reconstruction, ensuring the accuracy of coordinate conversion.

[0146] Step S42: Based on the camera parameter model and the denoised two-dimensional skeleton point data, perform multi-view skeleton point depth calculation based on the principle of triangulation to obtain skeleton point depth mapping data. In this step, the system uses the principle of triangulation to infer the position of a point in 3D space from the 2D skeleton point position information under multiple views. Triangulation is a classic 3D reconstruction technique, and its core idea is that when a 3D point is observed by two or more cameras at known positions, the 3D position of the point can be determined by calculating the intersection point of the rays from the camera optical centers to the projection positions of the point. Specifically, when implementing, the system first performs multi-view feature matching on the denoised 2D skeleton point data to determine the corresponding skeleton points under different views. This matching process takes into account factors such as the type of skeleton points, the relative position relationship, and the consistency of the surrounding skeleton structures, and adopts a matching algorithm that combines geometric constraints and appearance features. For successfully matched pairs of skeleton points, the system calculates the equations of the rays from the camera optical centers to the skeleton point positions in the image in 3D space according to their 2D positions in their respective images and the internal and external parameter matrices of the corresponding cameras. Ideally, these rays should intersect at the true 3D position of the skeleton point; however, due to measurement errors and digital precision limitations, the rays usually do not intersect precisely. Therefore, the system uses the least squares method or other optimization techniques to find the 3D point closest to multiple rays as the position estimate of the skeleton point. To further improve the accuracy and robustness of depth calculation, the system also introduces the RANSAC algorithm to process possible matching outliers and optimizes by combining temporal information and prior knowledge of human motion. For example, the system checks whether the calculated 3D skeleton point positions meet the constraint conditions of human skeleton lengths and whether the motion trajectories are smooth and reasonable, etc. Through these comprehensive processes, the system finally generates high-precision skeleton point depth mapping data, providing key information for constructing complete 3D skeleton point coordinates.

[0147] The principle of triangulation is a method of determining the spatial position of a point by observing the same 3D point through multiple cameras at known positions. In actual implementation, the system calculates the skeleton point depth by combining linear triangulation and non-linear optimization. Specifically, assume that the position of a certain skeleton point in the world coordinate system is X = (X, Y, Z, 1) T , and its image coordinates in camera 1 are x1 = (u1, v1, 1) T , and its image coordinates in camera 2 are x2 = (u2, v2, 1) T , then the following relationship is satisfied:

[0148] λ1x1 = K1[R1|t1]X

[0149] λ2x2 = K2[R2|t2]X

[0150] where λ1 and λ2 are scale factors, K1 and K2 are camera intrinsic matrices, and [R1|t1], [R2|t2] are camera extrinsic matrices. The system constructs a linear equation system by eliminating the scale factors:

[0151]

[0152] where represents the i-th row of the projection matrix Kj[Rj|tj] of the j-th camera. For example, for the same skeletal point detected in two cameras, with image coordinates (512, 640) and (498, 629) respectively, by solving the above equation system, the initial position of this skeletal point in the world coordinate system is determined to be (124.5, 176.2, 95.8). When more than two cameras observe the same skeletal point, the system constructs an overdetermined equation system and uses singular value decomposition (SVD) to obtain the least squares solution.

[0153] Step S43: Convert the depth mapped data of the skeletal points into three-dimensional space coordinates through projective transformation and spatial coordinate reconstruction algorithm to obtain the preliminary three-dimensional skeletal point coordinate data. In this step, the system combines the depth mapped data of the skeletal points obtained in the previous step with the 2D skeletal point coordinates, and through projective transformation and spatial coordinate reconstruction algorithm, calculates the complete coordinates of the skeletal points in the 3D world coordinate system. Projective transformation is a mathematical tool that converts 2D image coordinates and depth information into 3D space coordinates. It utilizes the intrinsic and extrinsic matrices of the camera to establish a mapping relationship between the image coordinate system and the world coordinate system. Specifically, the system first uses the intrinsic matrix of the camera to convert the skeletal point coordinates (pixel coordinates) on the image into normalized coordinates in the camera coordinate system, then combines the depth information to calculate the 3D position of the skeletal point in the camera coordinate system, and finally converts it to the world coordinate system through the inverse transformation of the extrinsic matrix. To further optimize the reconstruction result, the system introduces a spatial coordinate reconstruction algorithm, which comprehensively considers multiple constraint conditions to adjust the initial 3D coordinate estimation. First is the multi-view consistency constraint, which ensures that the same skeletal point reconstructed from different views has a consistent 3D position; second is the human skeletal length constraint, which uses the relatively fixed length of human skeletal segments to correct the skeletal deformation caused by depth estimation errors; third is the joint angle physiological constraint, which excludes non-physiological postures that, although satisfying other constraints, exceed the normal range of human joint activities. These constraints are organized into a unified optimization problem, and the system iteratively solves this problem to obtain the global optimal solution that satisfies all constraints. In addition, the system also introduces a temporal smoothing strategy to ensure the coherence and naturalness of the reconstructed skeletal point trajectory in the time dimension. Through this multi-constraint global optimization method, the system can generate high-precision and biomechanically reasonable 3D skeletal point coordinate data, providing a reliable basis for subsequent advanced pose analysis and action recognition.

[0154] Further, in step S42, based on the camera parameter model and the denoised two-dimensional skeleton point data, multi-view skeleton point depth calculation is performed based on the principle of triangulation to obtain skeleton point depth mapping data, including:

[0155] Step S421: Perform multi-view feature matching on the denoised two-dimensional skeleton point data to obtain skeleton point correspondence data. In this step, the core challenge faced by the system is to accurately identify the same skeleton point corresponding to different camera views, which is a prerequisite for triangulation. The system adopts a multi-level and multi-feature matching strategy, comprehensively utilizing the geometric characteristics, semantic characteristics, and context information of the skeleton points. First, the system uses the topological structure characteristics of the human skeleton to perform preliminary matching by comparing the similarity of the skeleton structures under different views. This structure-based matching considers the relative position relationship and connection pattern between the skeleton points and can effectively resist the appearance differences caused by view changes. Secondly, the system introduces the constraint conditions based on epipolar geometry, calculates the epipolar line using the calibrated camera parameters, and restricts the search range of the corresponding points, significantly improving the matching efficiency and accuracy.

[0156] Calculating the epipolar line using the calibrated camera parameters is realized based on the epipolar geometry theory in multi-view geometry. For two calibrated camera views (camera 1 and camera 2), the system first needs to construct a fundamental matrix, which contains the geometric relationship information between the two views. The calculation of the fundamental matrix depends on the internal parameter matrices of the two cameras and their relative position relationship, that is, relative rotation and translation. The system extracts the relative rotation and relative translation information from the external parameter matrices (including rotation matrix and translation vector) of each camera obtained from the calibration process.

[0157] The process of constructing the fundamental matrix first calculates the essential matrix, which is obtained by multiplying the relative rotation matrix and the skew-symmetric matrix of the relative translation vector. Then, the system calculates the fundamental matrix using the internal parameter matrices of the two cameras and the essential matrix. To ensure the stability of numerical calculations, the system usually normalizes the coordinates and performs denormalization after the calculation. In addition, the rank of the fundamental matrix is ensured to be 2 through singular value decomposition technology, which is a theoretical condition that the fundamental matrix must satisfy.

[0158] After obtaining the fundamental matrix, the system can calculate the epipolar line. For any image point in camera 1, the system multiplies it by the fundamental matrix to obtain the epipolar line parameters corresponding to this point in camera 2. This epipolar line represents all possible projection positions in camera 2 of the three-dimensional point that this point in camera 1 may correspond to. According to the epipolar geometry theory, the projection points of the same three-dimensional point under different camera views must be located on the corresponding epipolar lines.

[0159] In the application of skeletal point matching, when the system detects the position of a skeletal point in Camera 1, it calculates the corresponding epipolar line of this point in Camera 2. The system only needs to search for possible matching points within a narrow band area (usually with a width set to 5 to 10 pixels) near this epipolar line, rather than searching throughout the entire image, which greatly reduces the search space and the possibility of false matches.

[0160] For the case of three or more viewpoints, the system extends the application of epipolar line constraints. Through trilinear tensors or multi-view geometric relationships, more stringent matching constraint conditions are established. For fast-moving skeletal points, the system also combines temporal information to predict the possible positions of skeletal points and narrow down the search range on the epipolar line, further improving the matching efficiency.

[0161] By applying such constraint conditions based on epipolar geometry, the system can significantly improve the efficiency and reliability of skeletal point matching under multiple viewpoints while ensuring high accuracy, providing accurate correspondence data for subsequent 3D reconstruction. This technology is particularly important in complex scenarios and fast motion capture, and can effectively address challenges such as partial occlusion and uncertainty in skeletal point recognition.

[0162] For each skeletal point to be matched, the system searches for possible corresponding points in the nearby area along the corresponding epipolar line in the images of other viewpoints, greatly reducing the possibility of false matches. In addition, the system also combines the semantic label information of skeletal points (such as left knee, right elbow, etc.) to ensure that matching is only carried out between skeletal points of the same type, avoiding type confusion. To handle cases of partial occlusion and detection failure, the system has also implemented a robust partial matching strategy. Even if some skeletal points are not detected in a certain viewpoint, the system can still use the visible parts to construct a partial skeletal structure and perform matching based on this. To further improve the robustness of the matching, the system adopts a global optimization algorithm based on graph matching, regarding the matching problem of multiple skeletal points as a whole, comprehensively considering the consistency between all potential matching pairs, and finding the globally optimal matching solution. This method can effectively handle local ambiguity and improve the overall matching accuracy. Through this method of combining multiple strategies, the system finally generates high-precision skeletal point correspondence data, laying a solid foundation for the next step of depth calculation.

[0163] Step S422: Calculate the parallax information based on the skeletal point correspondence data to obtain the skeletal point depth mapping data. In this step, the system uses the established skeletal point correspondence to estimate the depth information of the skeletal points by calculating the parallax. Parallax refers to the difference in the image positions of the same object under different viewpoints, and it has an inverse relationship with the depth of the object - the farther the object, the smaller the parallax, and the closer the object, the larger the parallax. For each pair of matched skeletal points, the system first calculates the position difference between them on the image plane to obtain a parallax vector. Then, the system combines the internal and external parameter matrices of the camera and converts the parallax information into a depth value through the principle of triangulation. In an ideal situation, if there are skeletal point correspondences from two viewpoints, the 3D position of the skeletal point can be determined by solving the intersection point of the rays from the optical centers of the two cameras to the positions of the skeletal points on their respective images. However, due to measurement errors and limitations in digitization accuracy, these two rays usually do not intersect precisely in 3D space. Therefore, the system adopts the least squares method to find the point closest to these two rays as the 3D position estimate of the skeletal point. When there are more than two viewpoints, the problem becomes more complex, and the system then uses multi-view geometry techniques to comprehensively integrate the information from all viewpoints for more accurate depth estimation. Specifically, the system constructs an optimization problem that minimizes the reprojection error, that is, to find a 3D point that minimizes the sum of the distances between the 2D points obtained by projecting it back to each camera and the positions of the actually detected skeletal points. To improve the stability and accuracy of depth estimation, the system also introduces the RANSAC (Random Sample Consensus) algorithm to effectively handle possible abnormal matches. In addition, the system incorporates temporal information and prior knowledge of human motion, and through smoothing constraints and motion consistency checks, excludes those depth estimation results that are physically unreasonable. For example, the displacement of skeletal points between adjacent time frames should be within a reasonable range of human motion speed, and estimates outside this range may be incorrect and need to be corrected. Through this multi-level processing and optimization, the system finally obtains high-precision skeletal point depth mapping data, providing key information for constructing a complete 3D skeletal model.

[0164] Step S41 starts from the denoised high-quality 2D skeletal point data and obtains the internal and external parameter matrices of all cameras in the system through an accurate camera calibration process. This process usually uses Zhang Zhengyou's calibration method or more advanced optimized calibration techniques, combined with a specially designed calibration board for multi-angle shooting. By identifying the calibration points and solving the non-linear equations, the internal parameter matrix including the focal length, principal point coordinates, and radial distortion parameters, as well as the external parameter matrix describing the position and orientation of the camera in the world coordinate system, are calculated. These parameters form the camera parameter model, providing a geometric basis for subsequent 3D reconstruction.

[0165] In step S42, the system calculates the depth information of the bone points based on the obtained camera parameter model and the denoised 2D bone point data. Specifically, when a bone point is recognized in the images of two or more viewpoints, the system can determine the position of this point in the 3D space by solving the problem of ray intersection in these viewpoints. To improve the accuracy, the system uses the least squares method or the RANSAC algorithm to process the noise and errors in the multi-viewpoint data, and at the same time combines the temporal information and the prior knowledge of human motion for optimization, and finally generates the bone point depth mapping data containing depth values.

[0166] Next, in step S43, the bone point depth mapping data output in the previous step is converted into three-dimensional space coordinates. This process uses projective transformation and spatial coordinate reconstruction algorithms to combine the 2D image coordinates of each bone point with the corresponding depth value to calculate its exact position in the 3D world coordinate system. To further improve the accuracy, the system introduces human bone length constraints and joint angle constraints, and adjusts the 3D positions of each bone point through a global optimization algorithm to ensure compliance with human anatomical characteristics. The finally output preliminary 3D bone point coordinate data contains the spatial position information of all key joints of the human body, laying a foundation for subsequent high-precision optimization.

[0167] Furthermore, in step S5, based on the preliminary three-dimensional bone point coordinate data, bone point optimization is performed through rate perception analysis and three-dimensional Gaussian compression algorithm to obtain the optimized three-dimensional bone point coordinate data, including:

[0168] Step S51: The velocity-aware analysis module calculates the motion velocity and acceleration characteristics of the preliminary 3D skeleton point coordinate data to obtain dynamic characteristic data of the skeleton points. In this step, the system introduces an innovative velocity-aware analysis module to perform dynamic characteristic analysis on the preliminary reconstructed 3D skeleton points. This is crucial for accurately capturing the rapid changes in complex movements (such as dance and gymnastics). Based on time series data, this module calculates motion parameters such as displacement, velocity, acceleration, and angular velocity for each skeleton point between consecutive time frames. Specifically, the system first calculates the instantaneous velocity vector by differencing the skeleton point positions between adjacent time frames. The acceleration vector is then calculated by differencing the velocity vectors. Simultaneously, for joint points, the system calculates angular displacement and angular velocity to describe the dynamic characteristics of joint rotation. To capture the multi-scale characteristics of motion, the system employs a multi-time window analysis method, simultaneously considering the dynamic characteristics of both short-term windows (to capture instantaneous changes) and long-term windows (to capture overall trends). This multi-scale analysis effectively distinguishes noise fluctuations from true motion patterns, improving the robustness of feature extraction. For highly dynamic scenes such as dancing, the system pays special attention to the accurate portrayal of rapid movement stages. Through adaptive sampling rate adjustment, the sampling frequency is increased during periods of drastic movement changes to ensure that key movement details are not missed. In addition, the system also performs segmented analysis of the skeletal point trajectory, automatically identifying key action segments such as acceleration, deceleration, and turning, and extracting the characteristic parameters of each segment. Through machine learning algorithms, the system can also learn movement patterns of specific types of movements from historical data, providing a reference for subsequent skeletal point trajectory prediction and optimization. This comprehensive dynamic characteristic analysis enables the system to deeply understand the spatiotemporal characteristics of human movement, providing rich dynamic information for subsequent high-precision optimization, especially when dealing with high-speed movement and complex posture changes. It has significant advantages.

[0169] Step S52: Establish a probability density distribution model for the skeletal key points based on the dynamic characteristic data of the skeletal key points to obtain a representation of the skeletal key point probability distribution. In this step, the system constructs a probability distribution model for the skeletal key points by using an innovative 3D Gaussian compression algorithm based on the dynamic characteristic data extracted in the previous step. Traditional representations of skeletal key points usually use deterministic point coordinates and cannot effectively express the uncertainty of position and the possible range of distribution. The probability model method introduced by this system represents the position of each skeletal key point in 3D space as a multivariate Gaussian distribution, where the mean vector corresponds to the most likely position of the skeletal key point, and the covariance matrix describes the position uncertainty in different directions. This representation method can naturally encode the uncertainty of the skeletal key point position and provides richer information for subsequent optimization and fusion. The system adaptively adjusts the parameters of the Gaussian distribution according to the dynamic characteristics of the skeletal key points, especially the velocity and acceleration characteristics. For example, for a skeletal key point with high-speed movement, the system will increase the variance of the distribution in the movement direction, reflecting the increased position uncertainty along the movement direction; while keeping a smaller variance in the dimension perpendicular to the movement direction to ensure lateral accuracy. This rate-aware parameter adjustment strategy enables the probability model to more accurately reflect the impact of the movement state on the position uncertainty. To further improve the expressive power of the model, the system also considers the historical trajectory and prediction trend of the skeletal key points and incorporates time information into the probability model. For a skeletal key point with an obvious acceleration or deceleration trend, the system will correspondingly adjust the skewness of the distribution to better reflect the asymmetric characteristics of the movement. In addition, the system also implements a complex distribution representation based on a mixture of Gaussian models to handle those skeletal key points that may have multiple possible positions (such as in cases of occlusion or high sensor uncertainty). This probability-based representation method not only reduces the data storage requirements but also retains the uncertainty information of the skeletal key point position, providing a more comprehensive and flexible input for subsequent optimization processing, especially having significant advantages when dealing with challenging situations such as motion blur, occlusion, and sensor noise.

[0170] Specifically, the core idea of the 3D Gaussian compression algorithm is to represent the probability distribution of each skeletal key point in 3D space as a multivariate Gaussian distribution, where the mean vector corresponds to the estimated position of the skeletal key point, and the covariance matrix describes the position uncertainty. The innovation of this algorithm lies in adaptively adjusting the Gaussian distribution parameters according to the movement characteristics of the skeletal key points, especially the shape and direction of the covariance matrix. In specific implementation, the algorithm first represents the covariance matrix as a combination of a rotation matrix and an eigenvalue matrix through eigenvalue decomposition, and then adjusts the eigenvalues according to the rate characteristics of the skeletal key points, increasing the eigenvalues in the main movement direction while keeping the eigenvalues perpendicular to the movement direction small, thus forming an ellipsoidal distribution extending along the movement direction.

[0171] The compression ratio calculation is based on the principles of information theory, comparing the amount of information in the original point cloud representation and the Gaussian distribution representation. The original representation needs to store the complete sequence of 3D point coordinates, and the amount of information is proportional to the number of points; while the Gaussian representation only needs to store the mean vector and the covariance matrix (due to symmetry, actually only a small number of independent parameters need to be stored), and the amount of information is fixed. When the skeleton point trajectory contains a large number of time points, the compression ratio is significantly improved. To ensure the reconstruction accuracy, the algorithm designs an adaptive segmentation strategy: when the skeleton point trajectory deviates from the predicted position of the current Gaussian model by more than a preset threshold, or when the cumulative Mahalanobis distance exceeds the confidence interval threshold, the algorithm will start a new segmentation and construct a new Gaussian model.

[0172] In practical applications, the algorithm realizes a hierarchical representation of the skeleton point trajectory: low-frequency long trajectory segments are represented by a single Gaussian model, and high-frequency complex motion segments are represented by multiple consecutive Gaussian models. Experimental data shows that in typical dance motion capture scenarios, the algorithm achieves an average compression ratio of 15:1 while keeping the root mean square reconstruction error below 1 centimeter. More importantly, this representation method provides a probabilistic basis for subsequent multi-view consistency processing and occlusion prediction by retaining position uncertainty information, enabling the system to quantify the reliability of decisions and reasonably handle edge cases.

[0173] Step S53: Optimize and adjust the probability distribution representation of the skeletal points through reprojection error minimization and joint angle constraints to obtain the optimized 3D skeletal point coordinate data. In this step, the system precisely adjusts and optimizes the probability distribution representation of the skeletal points through a comprehensive global optimization framework. This optimization framework comprehensively considers various constraint conditions, while preserving the probability characteristics of the skeletal points, ensuring the rationality and consistency of the overall skeletal structure. First, the system introduces the reprojection error minimization criterion, that is, when the optimized 3D skeletal points are projected back to each camera view, they should match the positions of the actually detected 2D skeletal points as closely as possible. This criterion directly connects the 3D reconstruction result with the 2D observation data and is a basic requirement to ensure the accuracy of the reconstruction. The system projects the probability distribution of each skeletal point onto the image plane, calculates the difference from the actually detected 2D distribution, and optimizes the 3D distribution parameters by minimizing these differences. Second, the system introduces the skeletal length preservation constraint, that is, the distance between connected skeletal points should conform to human anatomical characteristics and remain relatively fixed. Since the skeletal length remains basically unchanged in a short period of time, this constraint can effectively correct the skeletal deformation problems caused by sensor noise or occlusion. The system calculates the distance between the connected skeletal point distributions through the distance calculation between probability distributions to ensure that the expected value of the distance between the connected skeletal point distributions conforms to the preset skeletal length parameters. Third, the system adds the physiological limit constraint of joint angles to exclude those poses that, although satisfying other constraints, are biomechanically unreasonable. For example, the bending angle of the knee joint has a certain range limit, and poses beyond this range are impossible in reality. The system calculates the distribution characteristics of the joint angles through probability distributions and applies appropriate prior constraints to ensure that the optimization results conform to the range of motion of human joints. In addition, the system also considers the time continuity constraint to ensure the smoothness and coherence of the skeletal point trajectory in the time dimension. All these constraints are organized into a unified optimization problem, and the system uses a global optimization algorithm based on a graph model, such as the weighted least squares method or the variational inference method, to solve this complex constraint satisfaction problem. Through this comprehensive global optimization, the system can, while preserving the uncertainty information of the skeletal point positions, ensure the rationality and consistency of the overall skeletal structure, and finally output high-precision and physically reasonable 3D skeletal point coordinate data, providing a reliable basis for advanced action analysis and understanding.

[0174] Further, in step S52, based on the dynamic characteristic data of the skeletal points, a probability density distribution model of the skeletal points is established to obtain the probability distribution representation of the skeletal points, including:

[0175] Step S521: Analyze the movement speed and acceleration characteristics in the dynamic characteristic data of the bone points to obtain movement characteristic parameters. In this step, the system deeply analyzes the dynamic characteristic data of the bone points obtained in the previous stage, and extracts characteristic parameters that can accurately describe the movement pattern. These parameters will directly affect the construction method of the subsequent probability distribution model and are crucial for capturing the movement characteristics of the bone points. The system first performs principal component analysis on the movement speed vector of each bone point to determine the main movement direction and the secondary movement direction. This decomposition can identify the main axis of the bone point movement, which helps to adjust the distribution parameters in the subsequent orientation. For the direction with a larger principal component, it indicates that the bone point has a larger movement amplitude in this direction, and correspondingly, the position uncertainty is higher; while for the direction with a smaller principal component, the bone point movement is more stable and the position certainty is higher. Then, the system analyzes the magnitude and direction of the acceleration vector to identify the change trend of the movement state. Continuous acceleration or deceleration indicates that the bone point is in a dynamic change stage, and the system needs to adjust the prediction model accordingly; while the acceleration close to zero indicates that the bone point is in a uniform motion or a static state, and at this time, the system can adopt a simpler linear prediction model. In addition, the system also calculates the correlation between the speed vector and the acceleration vector to analyze the curvature characteristics of the movement trajectory. High correlation usually indicates that the bone point moves along a straight line or a smooth curve, while low correlation may mean sharp turns or complex non-linear movements. The system pays special attention to the mutation points of speed and acceleration, which usually correspond to the key turning points of the action, such as the starting point or landing point of a jump in dance. By identifying these critical moments, the system can more accurately segment and analyze complex action sequences. To further improve the accuracy of the analysis, the system also introduces a spectral analysis method. Through Fourier transform or wavelet transform, it extracts the periodicity and frequency characteristics of the bone point movement. This is particularly effective for analyzing actions with obvious rhythm such as dance, and can identify repetitive patterns and synchronization with the music beat. Finally, the system synthesizes the above multi-dimensional analysis results to generate a set of characteristic parameters that comprehensively describe the movement characteristics of the bone points, including the main movement direction, speed amplitude, acceleration characteristics, curvature index, and periodicity characteristics, etc., providing rich and accurate input data for the subsequent probability distribution modeling.

[0176] Step S522: Based on the motion feature parameters, perform multivariate Gaussian distribution modeling and probability density calculation to obtain the probability distribution representation of the bone points. In this step, the system uses the motion feature parameters extracted in the previous step to construct a probability model that accurately describes the uncertainty of the bone point positions. The system selects the multivariate Gaussian distribution as the basic model, which has good mathematical properties and expressive power and can effectively capture the uncertainty and correlation of the bone point positions. For each bone point, the system first defines its mean vector, usually using the currently best estimated 3D position as the central value. Then, based on the motion feature parameters, the system constructs a covariance matrix that describes the position uncertainty and its correlation in different directions. In this process, the system uses the eigenvalue decomposition method to represent the covariance matrix as a combination of eigenvectors and eigenvalues. The eigenvectors define the principal axis directions of the uncertainty ellipsoid, and the eigenvalues define the uncertainty amplitudes along these directions. The system adaptively adjusts the parameters of the covariance matrix according to the motion rate characteristics of the bone points. Specifically, for bone points with high-speed motion, the system increases the corresponding eigenvalues along the main motion direction (i.e., the direction of the velocity vector), reflecting the increased position uncertainty in the motion direction; while keeping smaller eigenvalues in the dimensions perpendicular to the motion direction to ensure the lateral positioning accuracy. At the same time, the system also considers the impact of the acceleration characteristics on the uncertainty. For bone points with large acceleration, the prediction uncertainty of their future positions will increase accordingly, and the system reflects this characteristic by adding an acceleration-related adjustment term to the covariance matrix. To handle more complex position distributions, the system will use a Gaussian mixture model (GMM) when necessary, expanding the single Gaussian distribution into a weighted combination of multiple Gaussian distributions. This mixture model can represent multimodal distributions and is suitable for handling situations where bone points may have multiple candidate positions, such as in severe occlusion or sensor data conflicts. After constructing the probability distribution model, the system calculates the probability density function of each bone point in the 3D space and provides necessary sampling and evaluation interfaces to support the subsequent optimization and decision-making processes. This probability representation method based on the multivariate Gaussian distribution can not only accurately capture the uncertainty characteristics of the bone point positions but also provide rich probability information, providing a solid mathematical foundation for subsequent bone structure optimization and occlusion handling, especially having significant advantages when dealing with high-speed motion and complex pose changes.

[0177] In a possible implementation, in step S51, based on the preliminary 3D bone point coordinate data, the rate perception analysis module deeply analyzes the motion characteristics of the bone points. This module calculates the changes in the positions of bone points between adjacent time frames to obtain dynamic characteristic parameters such as the instantaneous velocity, acceleration, and angular velocity of each bone point. The system pays special attention to the fast motion patterns in highly dynamic scenarios such as dancing. By using the sliding window algorithm to segment and analyze the bone point trajectories, it identifies key action segments such as acceleration, deceleration, and turning, and marks the motion state characteristics of each bone point, thereby outputting the dynamic characteristic data of bone points containing rich spatio-temporal information.

[0178] In a possible implementation, in step S52, based on the dynamic characteristic data of bone points, an innovative 3D Gaussian compression algorithm is applied to construct a probability distribution model of bone points. This algorithm represents the position of each bone point in 3D space as a multivariate Gaussian distribution, where the mean vector corresponds to the estimated position of the bone point, and the covariance matrix describes the position uncertainty in different directions. By analyzing the rate characteristics and historical trajectories of bone points, the system adaptively adjusts the parameters of the Gaussian distribution, increasing the uncertainty of fast-moving bone points along the motion direction while maintaining high precision perpendicular to the motion direction. This probability-based representation method not only reduces the data storage requirements but also retains the uncertainty information of the bone point positions, outputting a more compact and information-rich representation of the bone point probability distribution.

[0179] In a possible implementation, in step S53, the probability distribution representation of bone points is optimized and adjusted through minimizing the reprojection error and joint angle constraints. Specifically, the system scores the distribution representation of each bone point, and the scoring criteria include the consistency with multi-view observation data, the compliance with the human bone length constraints, and the physiological rationality of joint angles, etc. Then, the system adopts a global optimization algorithm, such as the graph structure-based optimization method, to adjust the position parameters of each bone point, minimizing the overall reprojection error while ensuring compliance with the length and angle constraints of the human bones. This optimization strategy ensures that the reconstruction result not only conforms to the visual observation data but also meets the human biomechanical characteristics, finally outputting the optimized 3D bone point coordinate data with high precision and physical rationality.

[0180] Furthermore, in step S6, occlusion prediction and completion processing are performed on the optimized three-dimensional bone point coordinate data to obtain a complete human bone point positioning and recognition result, including:

[0181] Step S61: Detect the occluded bone points in the optimized three-dimensional bone point coordinate data through a bone point visibility analysis algorithm to obtain bone point occlusion status marking data. In this step, the system introduces a comprehensive visibility analysis algorithm to accurately identify and mark the occluded bone points. Occlusion detection is a key challenge in human bone point tracking, especially in multi-person scenarios, complex environments, or special postures, where bone points are often occluded by other objects or the body parts of the person themselves. The system first analyzes the detection confidence of bone points from each perspective. Low confidence is usually the primary indicator of occlusion. Specifically, the system examines the peak intensity, shape features, and consistency with historical data of each bone point to comprehensively evaluate its reliability. For bone points with low confidence in most perspectives, the system will mark them as possibly occluded. Secondly, the system analyzes the consistency of multi-perspective data. For the same bone point, if there are significant differences in the 3D positions reconstructed from different perspectives, this may indicate that the point is occluded in some perspectives. The system quantifies this inconsistency by calculating the variance or dispersion degree of the multi-perspective reconstruction results and sets an appropriate threshold for occlusion judgment. In addition, the system also combines depth relationships for occlusion analysis. By comparing the depth values of different objects or body parts in the scene, the system can infer the front-back relationship and identify the bone points occluded by other objects. Especially for self-occlusion of the human body (such as the arm occluding the torso), the system uses the known human model and pose information to simulate the light propagation path through the ray casting algorithm to determine which bone points are behind the occluding object. The system also implements a supplementary strategy for occlusion detection based on temporal analysis. When a certain bone point suddenly disappears or its quality significantly deteriorates in several consecutive frames while the surrounding bone points are still visible, the system will infer that the point may be temporarily invisible due to occlusion. To improve the reliability of occlusion detection, the system adopts a multi-feature fusion method, comprehensively considering multiple indicators such as confidence, multi-perspective consistency, depth relationship, and temporal coherence, and determines the final occlusion status through a weighted voting or probability inference mechanism. For each identified occluded bone point, the system not only marks its occlusion status but also estimates the severity of occlusion, the possible type of occluding object, and the expected occlusion duration, providing more detailed reference information for subsequent prediction and completion processing. Through this comprehensive and accurate occlusion analysis, the system can reliably identify bone point occlusion phenomena in various complex situations, laying a foundation for subsequent bone point completion.

[0182] Step S62: Extract the motion history patterns and trends of the skeletal points through the temporal information analysis module to obtain the temporal feature model of the skeletal points. In this step, the system utilizes powerful temporal information analysis techniques to mine the historical laws and trend features of the skeletal point motion, providing a key basis for predicting the future positions of occluded skeletal points. The system first establishes a time window for each skeletal point and collects motion parameters such as position, velocity, and acceleration within a certain period in the past (usually including several frames before occlusion occurs). The window size is adaptively adjusted according to the action type, using a smaller window for rapidly changing actions to capture instantaneous changes, and a larger window for slow or periodic actions to identify long-term trends. Then, the system applies a series of temporal pattern recognition algorithms to analyze and extract features from the collected historical data. For simple linear motions, the system uses techniques such as linear regression or Kalman filtering to extract the velocity and acceleration trends of the motion; for complex non-linear motions, the system applies more advanced temporal models, such as recurrent neural networks (RNNs), long short-term memory networks (LSTMs), or temporal convolutional networks (TCNs). These deep learning models can learn and capture complex temporal dependencies and identify non-linear motion patterns. For action scenarios with strong rhythm and repetition, such as dancing, the system particularly focuses on the periodic features of the actions. Through spectral analysis or autocorrelation analysis, the system can identify the repeating patterns in the actions, estimate their period lengths and phase information. This periodic analysis is particularly effective for predicting regular actions in dancing. The system can predict future action trajectories based on the observed periodic parts. In addition, the system also analyzes the co-motion relationships between skeletal points, identifying groups of interrelated skeletal points that usually exhibit similar motion patterns or fixed phase differences. This co-relationship analysis is particularly important when some skeletal points are occluded. The system can use the information of visible skeletal points to assist in predicting the motion of occluded points. The system also introduces prior knowledge learning based on a large-scale action database. By analyzing a large number of pre-collected action sequences, statistical models and transformation rules of various typical actions are constructed. When the current action is identified as similar to a certain type of action in the database, the system can use this prior knowledge to guide the prediction process. Finally, the system synthesizes the above multi-dimensional analysis results to construct a personalized temporal feature model for each skeletal point. This model contains rich information such as linear trends, non-linear patterns, periodic characteristics, and relationships with other skeletal points, providing a comprehensive and accurate temporal basis for the next step of position prediction.

[0183] Step S63: Based on the skeletal point temporal feature model and human biomechanical constraints, predict the positions of occluded skeletal points to obtain the complete human skeletal point localization and recognition results. In this step, the system combines the results of temporal analysis with human biomechanical constraints to achieve accurate position prediction and completion of occluded skeletal points. This process integrates dynamic temporal information and static structural constraints to ensure that the prediction results not only conform to the motion trend but also satisfy human anatomical characteristics. The system first makes a preliminary position prediction for each occluded skeletal point based on the temporal feature model constructed in Step S62. For skeletal points with obvious linear or periodic motion patterns, the system can directly extrapolate their motion trajectories; for complex non-linear motions, the system applies trained deep learning models (such as LSTM or TCN) for sequence prediction. These temporal predictions consider the historical positions, velocity trends, acceleration changes, and possible periodic features of skeletal points to generate a preliminary set of position candidates. Then, the system introduces human biomechanical constraints to refine and optimize the preliminary prediction results. First is the skeletal connection constraint. Utilizing the property of fixed skeletal lengths, the system calculates the possible position range of the occluded skeletal point based on the positions of connected and visible skeletal points. For example, if the upper arm skeletal point is visible while the elbow skeletal point is occluded, the system can determine the spherical region where the elbow may be located based on the length of the upper arm and the shoulder position. Second is the physiological limit constraint of joint angles. The system excludes non-physiological postures that, although satisfying the skeletal length constraint, exceed the normal range of joint motion. For example, the knee joint usually does not bend backward, and there are also certain range limits for the bending angle of the elbow joint. By applying these biomechanical rules, the system significantly reduces the search space of possible positions. In addition, the system also combines pose prior knowledge, learning the statistical characteristics and transformation rules of common postures from a large-scale action database to further optimize the prediction results. When the current action is similar to a certain known action in the database, the system can use these prior distributions to guide the prediction process. Finally, the system evaluates the feasibility and credibility of each candidate position through multi-hypothesis verification and probability fusion algorithms. The system constructs a comprehensive scoring mechanism, considering multiple factors such as the consistency of temporal prediction, the compliance with biomechanical constraints, the similarity with pose priors, and the naturalness of the overall pose, and selects the candidate position with the highest comprehensive score as the final prediction result. For each predicted position, the system also calculates its credibility score and marks its uncertainty degree with visual effects such as color coding or transparency to provide intuitive quality feedback to the user. Finally, the system integrates these prediction results with the data of unoccluded skeletal points to form a complete, coherent, and biomechanically reasonable human skeletal point localization and recognition result, successfully achieving stable skeletal point tracking and recognition in complex occlusion scenarios.

[0184] In a possible implementation, in step S61, based on the optimized 3D bone point coordinate data, occluded bone points are detected and labeled through a bone point visibility analysis algorithm. This algorithm determines whether a bone point is occluded by analyzing the detection confidence of the bone point in each view, multi-view consistency, and depth relationship. For example, when a certain bone point has low confidence in most views, or there are significant inconsistencies in the 3D reconstruction positions under different views, the system will mark it as possibly occluded. In addition, the system also uses depth information and the human body model for self-occlusion analysis to identify invisible bone points caused by the mutual occlusion of the human body's own parts. Through these comprehensive judgments, the system finally outputs labeled data containing the visibility status of each bone point, providing a basis for the prediction of occluded bone points in the subsequent process.

[0185] In a possible implementation, in step S62, based on the bone point occlusion status labeled data, a motion history pattern of the bone points is extracted through a temporal information analysis module. This module first establishes a time window, collects the position, velocity, and acceleration information of each bone point in the past several frames, and then applies a temporal pattern recognition algorithm, such as a recurrent neural network or a temporal convolutional network, to learn and extract the motion rules of the bone points from this historical data. For action scenarios with strong rhythm and repetition, such as dancing, the system pays special attention to the periodic characteristics and transition patterns of the actions, so as to be able to more accurately predict the future motion trends of the bone points. These analysis results form a temporal feature model of the bone points, providing an important reference for the position prediction of occluded bone points.

[0186] In a possible implementation, in step S63, based on the bone point temporal feature model and human biomechanical constraints, the positions of the occluded bone points are predicted and completed. This process first applies bone connection constraints, that is, using the characteristic of fixed bone length, to infer the possible position range of the occluded bone points from the known bone point positions. Then, the system introduces physiological limitations of joint angles to exclude candidate positions that, although meeting the bone length constraints, violate the human joint range of motion. In addition, the system also combines the pose prior knowledge learned from a large-scale action database to further narrow the prediction range and finally determine the best position of the occluded bone points. For each predicted bone point position, the system will also calculate its confidence score and mark its uncertainty with visual effects such as color or transparency, providing more intuitive feedback to the user. Finally, the system integrates these prediction results with the data of unoccluded bone points and outputs a complete, coherent, and accurate human bone point positioning and recognition result, achieving stable bone point tracking and recognition in complex occlusion scenarios.

[0187] Embodiment 2

[0188] An embodiment of the present invention further provides a human skeleton point positioning and recognition device based on OpenPose, and the device includes:

[0189] An image preprocessing module 10, configured to preprocess the input human body image to obtain preprocessed image data;

[0190] A skeleton point detection module 20, configured to perform skeleton point detection on the preprocessed image data by using a dual-domain multi-path self-supervised diffusion model combined with a convolutional neural network to obtain two-dimensional human skeleton point coordinate data;

[0191] A noise reduction processing module 30, configured to evaluate the signal-to-noise ratio of the two-dimensional human skeleton point coordinate data, and perform noise reduction processing on the convolutional neural network based on signal-to-noise ratio unit training and skeleton topology map weight mapping enhancement to obtain denoised two-dimensional skeleton point data;

[0192] A coordinate conversion module 40, configured to perform two-dimensional to three-dimensional conversion on the basis of the denoised two-dimensional skeleton point data through the principle of triangulation and feature matching technology to obtain preliminary three-dimensional skeleton point coordinate data;

[0193] A skeleton point optimization module 50, configured to optimize the skeleton points based on the preliminary three-dimensional skeleton point coordinate data through rate perception analysis and a three-dimensional Gaussian compression algorithm to obtain optimized three-dimensional skeleton point coordinate data;

[0194] An occlusion processing module 60, configured to perform occlusion prediction and completion processing on the optimized three-dimensional skeleton point coordinate data to obtain a complete human skeleton point positioning and recognition result.

[0195] The image preprocessing module 10 is specifically configured to:

[0196] Adjust the resolution of the input human body image to 1920×1080 pixels and perform cropping processing to obtain image data of a standard size;

[0197] Perform adaptive histogram equalization processing on the image data of the standard size to obtain image data with enhanced contrast;

[0198] Perform Gaussian filtering and median filtering on the image data with enhanced contrast to obtain the preprocessed image data.

[0199] The skeleton point detection module 20 is specifically configured to:

[0200] Extract features from the preprocessed image data through a multi-scale feature extraction convolutional neural network to obtain a multi-level feature map;

[0201] Enhance and refine the multi-level feature map through the dual-domain multi-path self-supervised diffusion model to obtain an enhanced feature representation;

[0202] Based on the enhanced feature representation, a multi-stage convolutional neural network is used to generate a body part heat map and a part affinity field vector, and the two-dimensional human bone point coordinate data is obtained.

[0203] The noise reduction processing module 30 is specifically used for:

[0204] The quality of the two-dimensional human bone point coordinate data is evaluated by a signal-to-noise ratio calculation module to obtain a bone point signal-to-noise ratio evaluation result;

[0205] Based on the bone point signal-to-noise ratio evaluation result, a convolutional neural network trained by a signal-to-noise ratio unit is used for noise recognition and classification to obtain the bone point noise type and distribution characteristics;

[0206] According to the bone point noise type and distribution characteristics, a bone point noise suppression model is constructed by a bone topology map weight mapping enhancement algorithm to obtain the denoised two-dimensional bone point data.

[0207] The coordinate conversion module 40 is specifically used for:

[0208] The internal and external parameter matrices of the camera are obtained by a camera calibration algorithm to obtain a camera parameter model;

[0209] Based on the camera parameter model and the denoised two-dimensional bone point data, multi-view bone point depth calculation based on the principle of triangulation is performed to obtain bone point depth mapping data;

[0210] The bone point depth mapping data is converted into three-dimensional space coordinates by a projective transformation and a space coordinate reconstruction algorithm to obtain the preliminary three-dimensional bone point coordinate data.

[0211] The bone point optimization module 50 is specifically used for:

[0212] The motion rate and acceleration characteristics of the preliminary three-dimensional bone point coordinate data are calculated by a rate perception analysis module to obtain bone point dynamic characteristic data;

[0213] Based on the bone point dynamic characteristic data, a probability density distribution model of the bone point is established to obtain a bone point probability distribution representation;

[0214] The bone point probability distribution representation is optimized and adjusted by minimizing the reprojection error and joint angle constraints to obtain the optimized three-dimensional bone point coordinate data.

[0215] The occlusion processing module 60 is specifically used for:

[0216] Detect the occluded skeleton points in the optimized three-dimensional skeleton point coordinate data through the skeleton point visibility analysis algorithm to obtain the skeleton point occlusion status marking data;

[0217] Extract the motion history patterns and trends of the skeleton points through the timing information analysis module to obtain the skeleton point timing feature model;

[0218] Predict the positions of the occluded skeleton points based on the skeleton point timing feature model and human biomechanical constraints to obtain the complete human skeleton point positioning and recognition results.

[0219] The above are only the embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.

Claims

1. A method for human skeletal point positioning and recognition based on OpenPose, characterized in that Including: Preprocess the input human body image to obtain preprocessed image data; Use a dual-domain multi-path self-supervised diffusion model combined with a convolutional neural network to perform skeletal point detection on the preprocessed image data to obtain two-dimensional human skeletal point coordinate data; Evaluate the signal-to-noise ratio of the two-dimensional human skeletal point coordinate data, and perform noise reduction processing on the convolutional neural network based on signal-to-noise ratio unit training and skeletal topology map weight mapping enhancement to obtain denoised two-dimensional skeletal point data; Based on the denoised two-dimensional skeletal point data, perform two-dimensional to three-dimensional conversion through the principle of triangulation and feature matching technology to obtain preliminary three-dimensional skeletal point coordinate data; Based on the preliminary three-dimensional skeletal point coordinate data, perform skeletal point optimization through rate perception analysis and three-dimensional Gaussian compression algorithm to obtain optimized three-dimensional skeletal point coordinate data; Perform occlusion prediction and completion processing on the optimized three-dimensional skeletal point coordinate data to obtain a complete human skeletal point positioning and recognition result.

2. The method according to claim 1, wherein Preprocess the input human body image to obtain preprocessed image data, including: Adjust the resolution of the input human body image to 1920×1080 pixels and perform cropping processing to obtain image data of standard size; Perform adaptive histogram equalization processing on the image data of standard size to obtain image data with enhanced contrast; Perform Gaussian filtering and median filtering on the image data with enhanced contrast to obtain the preprocessed image data.

3. The method according to claim 1, wherein Use a dual-domain multi-path self-supervised diffusion model combined with a convolutional neural network to perform skeletal point detection on the preprocessed image data to obtain two-dimensional human skeletal point coordinate data, including: Extract features from the preprocessed image data through a multi-scale feature extraction convolutional neural network to obtain a multi-level feature map; enhance and refine the multi-level feature map through the dual-domain multi-path self-supervised diffusion model to obtain an enhanced feature representation; Based on the enhanced feature representation, use a multi-stage convolutional neural network to generate a body part heat map and a part affinity field vector to obtain the two-dimensional human skeletal point coordinate data.

4. The method according to claim 1, characterized in that, Evaluate the signal-to-noise ratio of the two-dimensional human skeletal point coordinate data, and perform noise reduction processing on the convolutional neural network based on signal-to-noise ratio unit training and skeletal topology map weight mapping enhancement to obtain denoised two-dimensional skeletal point data, including: Evaluate the quality of the two-dimensional human skeletal point coordinate data through a signal-to-noise ratio calculation module to obtain a skeletal point signal-to-noise ratio evaluation result; Based on the skeletal point signal-to-noise ratio evaluation result, perform noise identification and classification through a convolutional neural network trained by a signal-to-noise ratio unit to obtain skeletal point noise types and distribution characteristics; According to the skeletal point noise types and distribution characteristics, construct a skeletal point noise suppression model through a skeletal topology map weight mapping enhancement algorithm to obtain the denoised two-dimensional skeletal point data.

5. The method according to claim 1, wherein Based on the denoised two-dimensional skeletal point data, perform two-dimensional to three-dimensional conversion through the principle of triangulation and feature matching technology to obtain preliminary three-dimensional skeletal point coordinate data, including: Obtain the camera internal and external parameter matrices through a camera calibration algorithm to obtain a camera parameter model; Based on the camera parameter model and the denoised two-dimensional skeleton point data, perform multi-view skeleton point depth calculation based on the principle of triangulation to obtain skeleton point depth mapping data; Convert the skeleton point depth mapping data into three-dimensional space coordinates through projective transformation and spatial coordinate reconstruction algorithm to obtain the preliminary three-dimensional skeleton point coordinate data.

6. The method according to claim 5, wherein Based on the camera parameter model and the denoised two-dimensional skeleton point data, perform multi-view skeleton point depth calculation based on the principle of triangulation to obtain skeleton point depth mapping data, including: Perform multi-view feature matching on the denoised two-dimensional skeleton point data to obtain skeleton point correspondence data; Calculate the disparity information based on the skeleton point correspondence data to obtain the skeleton point depth mapping data.

7. The method according to claim 1, characterized in that, Based on the preliminary three-dimensional skeleton point coordinate data, perform skeleton point optimization through rate perception analysis and three-dimensional Gaussian compression algorithm to obtain optimized three-dimensional skeleton point coordinate data, including: Calculate the motion rate and acceleration characteristics of the preliminary three-dimensional skeleton point coordinate data through a rate perception analysis module to obtain skeleton point dynamic characteristic data; Establish a probability density distribution model of skeleton points based on the skeleton point dynamic characteristic data to obtain a skeleton point probability distribution representation; Optimize and adjust the skeleton point probability distribution representation through reprojection error minimization and joint angle constraint to obtain the optimized three-dimensional skeleton point coordinate data.

8. The method according to claim 7, wherein Establish a probability density distribution model of skeleton points based on the skeleton point dynamic characteristic data to obtain a skeleton point probability distribution representation, including: Analyze the motion rate and acceleration characteristics in the skeleton point dynamic characteristic data to obtain motion characteristic parameters; Perform multivariate Gaussian distribution modeling and probability density calculation based on the motion characteristic parameters to obtain the skeleton point probability distribution representation.

9. The method according to claim 1, wherein Perform occlusion prediction and completion processing on the optimized three-dimensional skeleton point coordinate data to obtain a complete human skeleton point positioning and recognition result, including: Detect the occluded skeleton points in the optimized three-dimensional skeleton point coordinate data through a skeleton point visibility analysis algorithm to obtain skeleton point occlusion status marker data; Extract the motion history pattern and trend of skeleton points through a time series information analysis module to obtain a skeleton point time series feature model; Predict the positions of occluded skeleton points based on the skeleton point time series feature model and human biomechanical constraints to obtain the complete human skeleton point positioning and recognition result.

10. A human body bone point positioning and recognition device based on OpenPose, characterized in that, Including: An image preprocessing module for preprocessing the input human image to obtain preprocessed image data; A skeleton point detection module for detecting skeleton points of the preprocessed image data by using a dual-domain multi-path self-supervised diffusion model combined with a convolutional neural network to obtain two-dimensional human skeleton point coordinate data; A noise reduction processing module for evaluating the signal-to-noise ratio of the two-dimensional human skeleton point coordinate data, and performing noise reduction processing on the convolutional neural network based on signal-to-noise ratio unit training and skeleton topology map weight mapping enhancement to obtain denoised two-dimensional skeleton point data; A coordinate conversion module, which is used to perform two-dimensional to three-dimensional conversion based on the denoised two-dimensional skeleton point data through the principle of triangulation and feature matching technology to obtain preliminary three-dimensional skeleton point coordinate data; A skeleton point optimization module, which is used to optimize the skeleton points based on the preliminary three-dimensional skeleton point coordinate data through rate perception analysis and three-dimensional Gaussian compression algorithm to obtain optimized three-dimensional skeleton point coordinate data; An occlusion processing module, which is used to perform occlusion prediction and completion processing on the optimized three-dimensional skeleton point coordinate data to obtain a complete human skeleton point positioning and recognition result.

Citation Information

Cited By

  • Dynamic Voronoi skeleton constrained multi-modal door body detection and position correction method and dynamic Voronoi skeleton constrained multi-modal door body detection and position correction system

    CN120580294A

  • Gaussian hand-object interaction rendering denoising method

    CN120931520A

  • Hair style replacement video generation method and device based on diffusion model

    CN121032836A

  • Bone image data model construction method based on CT (Computed Tomography) image

    CN121074239A

  • Video human behavior prediction method based on residual diffusion theory and skeleton points

    CN121768083A