3D human pose estimation and multi-keypoint temporal analysis method

By constructing a 3D pose model through multimodal video access and deep learning models, and performing multi-keypoint temporal analysis, the problems of temporal semantic parsing and stability in human motion analysis are solved, and high-precision analysis of refined motion parsing and individual adaptation is achieved.

CN122024331BActive Publication Date: 2026-07-21CHENGDU UNIV OF INFORMATION TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHENGDU UNIV OF INFORMATION TECH
Filing Date
2026-04-10
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing technologies for human motion analysis suffer from insufficient ability to perform temporal semantic parsing of multiple key points, poor stability of 3D posture temporal sequence, and lack of quantitative analysis of multi-key point deviations. This leads to misjudgment of motion counting, blurred stage boundaries, and a lack of interpretability and targeted guidance in the analysis results.

Method used

A human detection and pose estimation model based on multimodal video access and deep learning is used to construct a 3D pose model and assign a credibility index. Through multi-keypoint temporal feature quantification and temporal stage perception, action temporal analysis is performed, and evaluation results are output to guide action correction.

Benefits of technology

It achieves refined analysis of continuous motion, improves the accuracy and versatility of motion analysis, provides interpretable motion deviation diagnosis and correction guidance, adapts to individual differences and suppresses interference such as lighting and occlusion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122024331B_ABST
    Figure CN122024331B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of human motion analysis, and particularly relates to a 3D human posture estimation and multi-key point time sequence analysis method, comprising the following steps: multi-modal video access, human key point space-time perception and extraction, 3D posture reconstruction and reliability modeling, key point time sequence feature quantization, time sequence stage perception and logical determination, quantitative evaluation and intelligent diagnosis, and result output, that is, outputting the analysis result in the form of score, text, graphics or voice to guide the user to correct the action, solving the problems of insufficient time sequence semantic understanding, poor individual difference adaptability, lack of explainability and targeted guidance of the analysis result in the human posture estimation and action analysis in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of human motion analysis technology, specifically to a 3D human posture estimation and multi-keypoint temporal analysis method. Background Technology

[0002] Human motion analysis technology, with posture estimation at its core, can extract spatial location information of common key points in the human body. It has been applied in fields such as physical education, physical training, and movement evaluation. However, existing technologies have not built a systematic temporal analysis system around multiple key points, resulting in significant deficiencies in the refined and intelligent analysis of continuous human motion. The core problems are as follows:

[0003] 1. Lack of multi-keypoint temporal semantic parsing capability. Existing technologies mostly focus on spatial feature extraction or simple temporal statistics of keypoints in a single frame, without exploring the temporal correlation and collaborative motion patterns between multiple keypoints. They cannot accurately divide the stages of continuous actions or identify the temporal logic of action execution, which can easily lead to problems such as misjudgment and omission of action counting and blurred stage boundaries, making it difficult to achieve refined parsing of the entire cycle of continuous actions.

[0004] 2. The 3D pose temporal stability is poor and there is no individual adaptive mechanism. Affected by factors such as changes in lighting, limb occlusion, and rapid motion jitter, 3D key points are prone to coordinate shifts and jumps. Furthermore, there is a lack of temporal noise reduction optimization schemes for multiple key points, and the estimation error of a single frame is amplified in the temporal dimension. At the same time, the judgment method of using fixed thresholds or rigid comparisons does not take into account the individual physiological characteristics to adapt to the motion range of multiple key points, nor does it make dynamic adjustments to the temporal change rate of key points. It cannot adapt to individual differences and changes in the rhythm of movement, and the objectivity and universality of the analysis results are insufficient.

[0005] 3. The lack of quantitative analysis of deviations at multiple key points results in low interpretability and practicality of the analysis results. Existing technologies can only output simple movement classifications, counts, or comprehensive scores. They cannot accurately quantify the types and degrees of deviations at key points in different time stages, nor can they locate the specific key points and stages of movement errors. They also cannot generate targeted correction suggestions based on the deviation characteristics of multiple key points, making it difficult to meet the actual needs of refined guidance in sports teaching and professional training.

[0006] In summary, existing technologies, due to their lack of integration of 3D pose estimation and multi-keypoint temporal analysis, have significant shortcomings in continuous motion temporal semantic parsing, pose anti-interference and individual adaptation, and refined diagnosis of motion errors. There is an urgent need to propose an innovative technical solution to address these issues and improve the accuracy, versatility, and practicality of human motion analysis. Summary of the Invention

[0007] The purpose of this invention is to provide a 3D human pose estimation and multi-keypoint temporal analysis method to solve the problems of insufficient temporal semantic understanding, poor adaptability to individual differences, and lack of interpretability and targeted guidance in the existing human pose estimation and motion analysis.

[0008] To solve the above-mentioned technical problems, the present invention adopts the following technical solution:

[0009] A method for 3D human pose estimation and multi-keypoint temporal analysis includes the following steps:

[0010] S1, Multimodal video access, which means acquiring video data containing human motion through image acquisition devices;

[0011] S2. Spatiotemporal perception and extraction of human key points, that is, using a human detection model and pose estimation model based on deep learning to extract human key point information from video frames. The key point information includes at least the two-dimensional or three-dimensional coordinates of multiple joints of the human body and their corresponding confidence scores.

[0012] S3, 3D pose reconstruction and credibility modeling, that is, based on the key point information, construct a human 3D pose model and assign credibility indexes to each key point or bone segment to characterize the reliability of the pose estimation result of the key point.

[0013] S4. Key point temporal feature quantification, which involves organizing the 3D key point information in a continuous time series to generate multi-key point temporal features;

[0014] S5. Temporal stage perception and logical judgment, that is, based on the temporal characteristics of multiple key points, perform temporal analysis on human movements, identify different stages of movements, and judge whether the movements meet the preset temporal logical relationship.

[0015] S6. Quantitative assessment and intelligent diagnosis: Based on the time series analysis results, quantitative assessment of human movements is performed, and indicators of movement completion, stability, and standardization are output. At the same time, the types of movement deviations and their occurrence stages are identified.

[0016] S7. Results Output: The analysis results are output in the form of scores, text, graphics, or voice to guide users in correcting their actions.

[0017] A further technical solution is that the credibility modeling in step S3 is constructed from information on kinematic rationality credibility, specifically: constructing a bone length ratio test scoring function; for a bone segment composed of key points p and q, its current length is... The desired length of the preset bone segment and length tolerance parameters The length reasonableness score is: ,in Let L be an exponential function with base e, and L be the actual measured length of the current skeletal segment. This is the expected length of the bone segment. The length tolerance parameter is used; a joint angle limit test scoring function is constructed for the joint angle formed by adjacent points a and b with key point i as the vertex. The normal range of motion of this joint is And define the angle reasonableness score as: ,in This represents the lower limit of the normal range of motion of the joint. This represents the upper limit of the normal range of motion of the joint. For the nearest boundary value, when hour ; hour , This is a preset constant that controls the rate at which the score decreases when the angle exceeds the limit. To scale the squared deviation.

[0018] A further technical solution is that the multi-keypoint temporal features in step S4 include keypoint spatial coordinate change features, joint angles and their rate of change, keypoint trajectory features, and posture stability features; the keypoint spatial coordinate change feature is defined as follows: let the two-dimensional image coordinates of the i-th keypoint in frame t be... The system frame rate is FPS, and the time interval between adjacent frames is... The second, and the corresponding nth order kinematic quantity are uniformly expressed as , Let be the nth-order kinematic vector of the i-th keypoint in frame t, where n is a non-negative integer; n=0 represents position, n=1 represents velocity, and n=2 represents acceleration. For the nth-order forward difference starting from frame t−n+1, The time interval is a power of n; the joint angles and their rates of change are calculated using a confidence-weighted time series fusion strategy, as shown in the following formula. ,in The original joint angle for frame t+k. Let k be the confidence level of the joint center key point in this frame, and k be the time offset for the current frame t, taking the integer value in the range [−k, k]. For time weighting coefficients, To sum all frames within the window, iterate through all frames from t−K to t+K; the keypoint trajectory features quantify the economy and regularity of the keypoint trajectory path, introducing a trajectory efficiency index. This index closely correlates the geometric properties of the trajectory with the energy efficiency of the signal, and is specifically defined as: ,in Let i be the two-dimensional coordinate vector of the i-th key point in frame t. , The Euclidean distance between key points in adjacent frames. This represents the total distance traveled by all keypoints from frame 2 to frame T. This represents the linear displacement of key points between the first and last frames of the window. It is a very small normal number; postural stability features assess the degree of postural stability of the body's core region during movement or static maintenance, and define a multi-keypoint cooperative stability index, defined as... , For set The number of key points Let x be the sample variance of the x-coordinate of the j-th keypoint within the time window. Let be the sample variance of the y-coordinate of the j-th keypoint within the time window. The expected length of the bone segment corresponding to keypoint j. To calculate the values ​​of the expressions corresponding to all key points in set j and sum them.

[0019] A further technical solution is that, in step S5, the action stage based on time-series features is identified as being driven by real-time joint angles and coordinate change rates as inputs, which in turn drive the state machine to migrate according to a preset threshold.

[0020] Compared with the prior art, the beneficial effects of the present invention are:

[0021] 1. Construct a multi-keypoint temporal semantic fusion modeling system to explore the temporal correlation and collaborative motion patterns between key points, accurately divide action stages, identify temporal logic, realize refined analysis of the entire cycle of continuous actions, effectively solve the problems of misjudgment and omission in counting and fuzzy stage boundaries, and improve the accuracy of motion analysis.

[0022] 2. A multi-keypoint temporal denoising scheme is proposed to suppress interference from lighting, occlusion, jitter, etc., and improve the temporal stability of 3D keypoints; an adaptive adjustment mechanism based on individual physiological characteristics is designed to dynamically adapt to the individual's range of motion and rhythm of movement, thereby improving the objectivity and universality of the analysis results. Attached Figure Description

[0023] Figure 1 This is a flowchart illustrating a 3D human pose estimation and multi-keypoint temporal analysis method according to the present invention. Detailed Implementation

[0024] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0025] Example:

[0026] Figure 1 This invention illustrates a preferred embodiment of a 3D human pose estimation and multi-keypoint temporal analysis method. The specific steps of this embodiment include:

[0027] S1, Multimodal Video Access, involves acquiring video data containing human motion through image acquisition devices. Serving as the data source entry point for the entire process, it solves the problems of poor input adaptability and inconsistent raw data in existing technologies, which limit the scope of application. By establishing multi-device video access links, complete and continuous human motion video is acquired, providing standardized raw data for subsequent full-process analysis and ensuring the method's cross-scenario adaptability.

[0028] S2. Spatiotemporal perception and extraction of human key points: This involves using a deep learning-based human detection and pose estimation model to extract human key point information from video frames. The key point information includes at least the two-dimensional or three-dimensional coordinates of multiple human joints and their corresponding confidence scores. As a fundamental feature element of the technical solution, this addresses the problems of insufficient spatiotemporal correlation and lack of reliability verification in existing key point extraction technologies, leading to a lack of solid feature support for subsequent analysis. Relying on a deep learning model, it extracts the 2D / 3D coordinates of multiple human joints and their corresponding confidence scores from continuous video frames, achieving both precise spatial localization within a single frame and preserving temporal correlation, providing core foundational features for subsequent full-process analysis.

[0029] S3, 3D pose reconstruction and reliability modeling, involves constructing a 3D human pose model based on the keypoint information and assigning reliability indices to each keypoint or bone segment to characterize the reliability of the pose estimation results for that keypoint. As a core component ensuring pose data quality, this approach precisely addresses the critical issues of poor temporal stability of 3D poses in the background, temporal amplification of single-frame errors, and the lack of individual adaptive mechanisms. By constructing a 3D human pose model based on keypoint information and employing corresponding bone length and joint angle rationality verification functions to complete reliability modeling, it can identify and filter abnormal pose data, suppress error propagation, and adapt to different individual physiological characteristics, providing high-quality 3D pose data for subsequent temporal analysis.

[0030] S4. Keypoint Temporal Feature Quantization: This involves organizing 3D keypoint information from a continuous time series to generate multi-keypoint temporal features. As a crucial link connecting single-frame pose analysis to continuous temporal analysis, it addresses the shortcomings of existing technologies in uncovering multi-keypoint temporal correlations and lacking a standardized temporal feature quantification system. It systematically organizes discrete single-frame data within a time series, generating multi-dimensional temporal features through corresponding kinematic calculations and confidence-weighted temporal fusion methods. This comprehensively characterizes the collaborative motion patterns of keypoints and simultaneously performs adaptive noise reduction of the temporal data, providing a core quantitative foundation for subsequent motion temporal analysis.

[0031] S5, Temporal Stage Perception and Logical Judgment, analyzes human movements based on multi-keypoint temporal features, identifies different stages of the movement, and determines whether the movement meets preset temporal logical relationships. As a core component of temporal semantic parsing, it addresses industry pain points such as ambiguous boundaries between movement stages, misjudgments and omissions in counting, and inaccurate identification of temporal logic. By mining collaborative motion patterns based on multi-keypoint temporal features, it accurately delineates the execution stages and boundaries of continuous movements, verifies the compliance of the movement's temporal logic, and achieves refined temporal semantic parsing throughout the entire lifecycle of continuous movements, laying the foundation for subsequent phased assessment and diagnosis.

[0032] S6. Quantitative Assessment and Intelligent Diagnosis: Based on time-series analysis results, this function quantitatively assesses human movements and outputs indicators for movement completion, stability, and standardization. It also identifies the types of movement deviations and their stages of occurrence. As a core component of refined movement analysis, it addresses the fundamental problems of unquantifiable movement deviations, lack of interpretability in analysis results, and inability to support refined guidance. Based on time-series analysis results, it performs full-cycle, stage-by-stage quantitative assessments of movement completion, stability, and standardization. Simultaneously, it accurately locates the key points, stages of occurrence, and degrees of movement deviations, ensuring the assessment results are traceable and interpretable, providing precise evidence for refined movement guidance in professional scenarios.

[0033] S7. Results Output: This involves outputting the analysis results in the form of scores, text, graphics, or audio to guide users in correcting their actions. As a closed-loop component of the technical solution, it solves the problem of existing technologies having limited output formats and the inability to translate professional analysis results into actionable guidance. It transforms quantitative assessments and deviation diagnosis results into various easily readable formats such as scores, text, graphics, and audio, intuitively presenting analysis conclusions and correction directions. This completes the closed loop from technical analysis to practical action guidance, significantly enhancing the method's practical value and applicability.

[0034] In step S1, video data containing human motion is acquired using an image acquisition device.

[0035] The system implements a dual video acquisition scheme using a local camera and RTSP streaming media: the former supports device enumeration, start / stop, frame rate control, and anomaly handling to ensure the stability of local real-time acquisition; the latter, through an RTSP streaming module, encodes and pushes the local image to the streaming media server in real time, supporting remote client subscription and playback, meeting the low-latency video sharing requirements in distributed and cloud-edge collaborative scenarios. Utilizing a deep learning-based human detection and pose estimation model, key human point information is extracted from video frames. These key points include at least the two-dimensional or three-dimensional coordinates and corresponding confidence scores of multiple human joints. Step S2 addresses the real-time challenge of high-precision pose estimation under limited computing power on edge devices. An INT8 quantized lightweight detection network reduces computational overhead while maintaining accuracy, enabling rapid localization of human bounding boxes. A CSPNEXt backbone and SimCC coordinate classification scheme replace the traditional heatmap regression method, achieving high-precision real-time output of 2D key points. Through multi-level optimization, including confidence threshold filtering, temporal smoothing, and coordinate normalization, low-quality points, inter-frame jitter, and scale / viewpoint differences are eliminated, outputting stable and reliable structured pose data.

[0036] In the human detection stage, the system employs a lightweight object detection network trained with multi-stage data augmentation and INT8 quantization, enabling efficient inference and accurate localization of human bounding boxes on an edge computing processor. In the pose estimation stage, a top-down model is used, combining the CSPNEXt backbone and the SimCC coordinate classification scheme to output real-time 2D coordinates and confidence scores of general human keypoints. The keypoint processing and optimization stage performs multi-level optimization on the raw output: low-quality points are filtered out using a confidence threshold, temporal smoothing within a sliding window using exponentially weighted moving averages to suppress jitter, and coordinate normalization based on the human bounding box to eliminate scale and viewpoint differences. Finally, each frame outputs structured pose data containing the human bounding box, keypoint coordinates and confidence sequence, and precise timestamps, providing stable and aligned input for subsequent temporal analysis. This step solves the real-time challenge of high-precision pose estimation under the limited computing power of edge devices, transforming raw image data into a structured pose representation, laying the data foundation for subsequent temporal analysis and motion evaluation.

[0037] In step S3, 3D pose estimation: For devices with limited computing power, such as edge computing, a lightweight algorithm based on monocular video is adopted. Using the 2D keypoint sequence output from the top-down model as input, a lightweight Transformer (a lightweight deep learning network based on a self-attention mechanism) is used to learn the mapping from 2D to 3D sequences. This solution solves the high cost and deployment complexity problems of traditional 3D pose estimation relying on depth sensors or multi-view cameras, enabling real-time reconstruction of human 3D pose on edge devices using only a monocular RGB camera. This provides complete spatial information support for subsequent kinematic analysis, motion quality assessment, and time series modeling, while ensuring real-time operation on edge platforms such as the RK3588.

[0038] In step S3, credibility modeling is constructed from information on kinematic plausibility credibility. Specifically, it involves constructing a bone length ratio test scoring function. For a bone segment composed of keypoints p and q, its current length is... , Let p be the three-dimensional coordinate vector of the key point. Let q be the 3D spatial coordinate vector of the keypoint, and let the expected length of the bone segment be preset. and length tolerance parameters The length reasonableness score is: ,in It is an exponential function with the natural constant e as its base. This is the actual measured length of the current bone segment. This is the expected length of the bone segment. The length tolerance parameter is calculated by mapping the deviation between the actual length of the current bone segment and the preset expected length to a score between 0 and 1 using a Gaussian function. This index addresses anomalies in pose estimation that may not conform to human anatomy constraints, such as bone stretching or compression. A score of 1 is awarded when the actual length matches the expected length, and the larger the deviation, the closer the score is to 0. This mechanism effectively filters out geometrically distorted keypoints caused by occlusion and pose estimation errors, providing highly reliable keypoint inputs for subsequent time-series analysis and significantly improving the system's ability to identify and filter abnormal poses. A joint angle limit test scoring function is constructed for joint angles formed by adjacent points a and b with keypoint i as the vertex. The normal range of motion of this joint is And define the angle reasonableness score as: .in This represents the lower limit of the normal range of motion of the joint. This represents the upper limit of the normal range of motion of the joint. For the nearest boundary value, when hour ; hour , This is a preset constant that controls the rate at which the score decreases when the angle exceeds the limit. To normalize the squared deviation, a constant 2 is derived from the canonical form of a Gaussian function, used to assess whether human joint angles conform to physiological range of motion constraints. A score of 1 is awarded when the joint angle is within the normal range; when it exceeds the normal range, the score is smoothly decayed according to the degree of exceedance using a Gaussian function, with lower scores for greater exceedances. This index addresses anomalies in posture estimation, such as reverse bending and hyperextension, which violate human kinematics. Compared to traditional hard thresholding methods, this Gaussian scoring function provides a continuous and differentiable evaluation of reasonableness, enabling the system to distinguish between slight and severe exceedances, providing a refined quantitative basis for motion quality assessment and abnormal posture warnings.

[0039] The multi-keypoint temporal features in step S4 include keypoint spatial coordinate change features, joint angles and their rate of change, keypoint trajectory features, and posture stability features; the keypoint spatial coordinate change feature is defined as follows: let the two-dimensional image coordinates of the i-th keypoint in frame t be... The system frame rate is FPS, and the time interval between adjacent frames is... The second, and the corresponding nth order kinematic quantity are uniformly expressed as , Let be the nth-order kinematic vector of the i-th keypoint in frame t, where n is a non-negative integer; n=0 represents position, n=1 represents velocity, and n=2 represents acceleration. For the nth-order forward difference starting from frame t−n+1, The formula, where the time interval is a power of n, transforms the original coordinate sequence of human keypoints into a multi-order kinematic feature system that is physically meaningful, frame rate independent, and supports causal computation through a combination of discrete difference and time normalization. It effectively solves problems in motion analysis such as inconsistent feature representation, sensitivity to acquisition parameters, and difficulty in standardizing the extraction of high-order motion quantities, providing a reliable basic feature construction method for applications such as video pose analysis and real-time motion recognition. The joint angles and their change rates employ a confidence-weighted temporal fusion strategy, as shown in the following formula. ,in The original joint angle for frame t+k. Let k be the confidence level of the joint center key point in this frame, and k be the time offset for the current frame t, taking the integer value in the range [−k, k]. For time weighting coefficients, To sum over all frames within the window, the formula iterates through all frames from t−K to t+K. Essentially, this elevates joint angle calculation from "static geometric calculation based on a single-frame image" to "dynamic signal reconstruction based on a temporal window." It utilizes the continuity of human motion as prior knowledge and combines the confidence level of keypoint detection as a quality assessment metric. Through a weighted moving average, it addresses the temporal jitter and estimation inaccuracies that traditional methods easily encounter in scenarios involving occlusion, rapid movement, and false keypoint detection, significantly improving system stability, anti-interference capabilities, and temporal consistency of output. Keypoint trajectory features quantify the economy and regularity of keypoint trajectory paths by introducing a trajectory efficiency index. This index closely correlates the geometric properties of the trajectory with the energy efficiency of the signal, and is specifically defined as: ,in Let i be the two-dimensional coordinate vector of the i-th key point in frame t. , The Euclidean distance between key points in adjacent frames. This represents the total distance traveled by all keypoints from frame 2 to frame T. This represents the linear displacement of key points between the first and last frames of the window. As an extremely small positive constant, this formula maps the geometric regularity of the visual trajectory to physical energy efficiency through the "ratio of distance to displacement." It solves the problem that traditional methods cannot uniformly measure "detours" and "redundancy," providing a robust, scale-independent, and physically clear quantitative tool for scenarios such as motion quality assessment, anomaly detection, and behavior segmentation. Postural stability features assess the degree of postural stability of the core body region during motion or static maintenance, defining a multi-keypoint cooperative stability index, defined as... , For set The number of key points Let x be the sample variance of the x-coordinate of the j-th keypoint within the time window. Let be the sample variance of the y-coordinate of the j-th keypoint within the time window. The expected length of the bone segment corresponding to keypoint j. To calculate and sum the values ​​of the expressions corresponding to all key points in set j, this formula constructs a scalar index that is unaffected by body type and reflects the overall coordinated control ability of the core area by normalizing the time series variance of multiple key points. It solves the problems of large interference from individual differences and the disconnect between local and overall aspects in traditional methods, providing a highly reliable and comparable quantitative tool for sports biomechanical analysis, rehabilitation assessment, and sports performance optimization.

[0040] This formula is used to assess the postural stability of the core body during movement or static maintenance. By comprehensively calculating the variance of the coordinates of multiple core key points within a time window and normalizing it by combining the expected length of the corresponding skeletal segments, this indicator solves the problem that single key point stability assessments cannot reflect overall postural control ability. This indicator can effectively quantify the degree of swaying and coordinated stability of the core areas such as the trunk, hips, and shoulders. In the assessment of movements such as holding the lowest point of a squat, plank, and standing balance, a smaller value indicates greater postural stability and stronger core control; a larger value reflects significant postural swaying and insufficient stability. This indicator provides objective statistical basis for movement quality diagnosis, balance ability assessment, and training effect quantification.

[0041] Based on the aforementioned multi-keypoint temporal features, temporal analysis is performed on human movements to identify different stages of the movements and determine whether the movements satisfy preset temporal logic relationships.

[0042] Action phase recognition based on temporal features: taking real-time joint angles and coordinate change rates as inputs, driving the state machine to migrate according to preset thresholds.

[0043] Verification of temporal logic relationships and action determination: The state machine not only identifies the phase but also enforces the verification of the correctness of the temporal logic of action execution. The counting logic requires that the action must completely go through the preset state loop; the timing logic, for hold-type actions, accumulates the duration of the "holding" state, and if the feature deviates, the timing is paused and resumed for the next action.

[0044] Step S5 aims to parse continuous human motion into a sequence of stages with clear semantics. This step takes the multi-dimensional temporal features output from the keypoint temporal feature construction step as direct input and includes the following three core sub-schemes:

[0045] 1. Constructing an action state model based on the temporal features of multiple key points. This step uses the continuous 3D key point sequence generated in the 3D pose estimation and credibility modeling steps. Based on the universal human body key point definition standard, each key point Includes three-dimensional coordinate signals and confidence signal Heteroscedastic Kalman filtering is achieved by dynamically modulating the observation noise covariance using the confidence signal: ,in The baseline noise variance, Zero protection. Recursive filtering of the position signal. is the observed 3D coordinate vector of the i-th keypoint in frame t, and is the original noisy observation value output by the attitude estimation model; It is the true 3D coordinate vector of the i-th key point in frame t, and is the true value that the filtering needs to approximate; Ii is the confidence signal of the i-th keypoint in frame t, representing the reliability of the pose estimation result of the keypoint, which is derived from the output of the 3D pose reconstruction and confidence modeling module mentioned above; I3 is a 3-order identity matrix, representing that the observation noise in the x, y, and z coordinate axes is independent and has the same variance.

[0046] Constructing weighted features based on the filtered position signal:

[0047] Weighted joint angles utilize limb vectors , This is a limb vector pointing from the joint center point b to the key point a. For example, in the knee joint scenario, b is the knee joint key point and a is the hip joint key point. This vector represents the thigh segment vector. It is a limb vector pointing from the joint center point b to the key point c. For example, in the knee joint scenario, c is the ankle joint key point. This vector represents the lower leg segment vector, and is weighted by combining the posterior covariance signal: ,in , It is obtained by normalizing the inverse of the posterior covariance trace of the three key points. For frame t, the optimal estimate of the joint weighted angle formed by abc with b as the vertex is the core output of this formula; The four-quadrant arctangent function, compared to the ordinary arccos vector angle calculation, can avoid the numerical instability problem when the angle is close to 0° / 180°, and has higher calculation accuracy; S is the weighted covariance matrix, which integrates limb vector, reference vector and weight matrix to realize weighted calculation of joint angle. The reference vector for the limb vector is usually taken as the standard limb vector in the neutral position of the joint, which is used as the reference for the coordinate system of the angle calculation. This formula is a confidence-weighted method for calculating joint angles. Unlike the traditional method of directly calculating the angle between vectors, it uses the posterior covariance weighting matrix W to reduce the weight of the coordinates of low-reliability key points, which solves the problems of distortion and large temporal fluctuations in joint angle calculation under occlusion and shaking scenarios, and outputs a smooth and reliable joint angle time series.

[0048] Weighted velocities and accelerations are weighted by Savitzky-Golay filtering using the confidence signal as the weight to obtain smoothed velocities. With acceleration : , Let be the weighted smoothed velocity vector of the i-th keypoint in frame t, and let be the core output feature, representing the speed and direction of the keypoint's motion; Let be the weighted smoothed acceleration vector of the i-th keypoint in frame t, and let be the core output feature that characterizes the velocity change trend and force application state of the keypoint. The keypoint rate scalar (the magnitude of the velocity) is obtained by confidence-weighted Savitzky-Golay filtering. The filtering process uses the confidence of the keypoint as the weight, and the proportion of high-confidence frame data is higher, which ensures the temporal smoothness of the rate. The keypoint acceleration scalar (magnitude of acceleration) is obtained by confidence-weighted Savitzky-Golay filtering. The filtering logic is consistent with the rate, which suppresses high-frequency noise in the acceleration.

[0049] The unit direction vector of the keypoint displacement represents the spatial direction of the keypoint motion and is obtained by normalizing the coordinate difference after filtering of adjacent frames. It is the unit direction vector of the velocity change at the key point, representing the spatial direction of acceleration, and is obtained by normalizing the velocity difference between adjacent frames.

[0050] This formula, based on confidence-weighted Savitzky-Golay filtering, calculates smoothed velocity and acceleration vectors for key points. This solves the problem of traditional difference methods, which suffer from severe temporal fluctuations due to coordinate noise and fail to reflect the true motion trend when calculating velocity / acceleration. The output velocity and acceleration are core temporal features for subsequent action phase division and force rhythm analysis.

[0051] The weighted stability index uses exponentially weighted recursive covariance to analyze the fluctuations of position signals. , , Where λ is the forgetting factor, controlling the decay rate of historical signals; α is the compensation factor, used to correct the deviation between exponential weighting and the expected value of the sliding window. The exponentially weighted mean vector of the i-th keypoint in frame t is recursively updated to represent the central trend of the keypoint's position. The exponentially weighted covariance matrix of the i-th keypoint in frame t is recursively updated to characterize the fluctuation and dispersion of the keypoint position in three-dimensional space, and is the core intermediate quantity for stability assessment. The weighted stability index of the i-th keypoint in frame t is the core output of this formula; the larger the value, the greater the fluctuation of the keypoint position and the worse the attitude stability; the smaller the value, the more stable the attitude. Let be the confidence signal of the i-th keypoint in frame t, consistent with the previous text; the lower the confidence, the lower the weight of the current frame data in the mean and covariance update, to avoid distorted data interfering with the stability assessment; The covariance matrix is ​​the outer product of the current keypoint position with respect to the mean, representing the dispersion of the current frame position. This formula is a confidence-weighted exponentially weighted recursive covariance calculation method, used to recursively calculate the fluctuation of keypoint positions in real time and output a pose stability index. Unlike traditional sliding window variance calculation, this method does not require caching all historical data in the window, has low memory usage, high computational efficiency, and is suitable for real-time deployment on edge devices; simultaneously, it uses confidence... Weighting avoids low-quality data from interfering with stability assessment results and can accurately capture issues such as posture jitter and center of gravity shift during the movement process.

[0052] The multi-domain fusion spatiotemporal feature tensor concatenates the above feature signals into a fusion vector: ,in It is the position feature vector obtained by vectorizing the three-dimensional coordinate matrix of all key points. It is a joint angle feature vector. It is the feature vector of motion velocity. It is an acceleration eigenvector. It is the attitude stability feature vector.

[0053] The statistical fluctuation features are obtained by vectorizing the lower triangular part of the position covariance matrix of each key point within the time window, and these features are then concatenated to obtain the single-frame fused features. ,in This indicates that the feature vector is located at 3D real vector space, The total number of all characteristic components; then the continuous Frames Horizontal stacking forms a spatiotemporal feature tensor. ,in This indicates that the tensor is OK The column of real-valued matrices is used for subsequent temporal analysis. This formula concatenates multi-dimensional features such as keypoint positions, joint angles, motion velocity, acceleration, posture stability, and position covariance into a single-frame fusion vector, which is then stacked along the time dimension to form a spatiotemporal feature tensor. This solves the problems of single features and spatiotemporal fragmentation in traditional methods. This tensor achieves the fusion of geometric, motion, and statistical information, preserving spatial structure and temporal dependencies, making it easy for deep learning models to process directly. Furthermore, the physical meaning of each feature is clear and highly interpretable, providing a structurally clear foundation for subsequent action phase segmentation, state machine modeling, and precise boundary localization.

[0054] 2. Divide the continuous action into multiple action phases, including at least a preparation phase, an execution phase, and a recovery phase: Construct a finite state machine based on F and define the set of states. ,in Indicates the standing phase. Indicates the squatting phase. This indicates the period during which the lowest point is held. This represents the squatting and standing phase. This set of states constitutes the basic state space of a finite state machine, discretizing continuous human motion into a sequence of stages with clear semantics. The transitions between states are precisely triggered by threshold conditions of multi-dimensional temporal features such as joint angles, movement speed, and stability.

[0055] A Bayesian adaptive threshold is generated using the statistical distribution of key point signals. Let's assume the knee joint angle signal during standing... Its prior mean with prior variance Obtained from population statistics; collected online. After each frame, the posterior mean is obtained through Bayesian update. With posterior variance Based on this, the threshold values ​​for each state transition are defined as follows:

[0056] Standing → Squatting Triggering is achieved by statistically analyzing the lower limit using angle signals;

[0057] Squat down → Lowest point: in This is the femur length signal estimated from key points. For reference length, The standard knee flexion angle at the lowest point of the human body is obtained through offline calibration or population statistics.

[0058] Lowest point → Squat up: ,in Estimation of the standard deviation of the hip vertical velocity signal;

[0059] Squat to stand: Updated recursively by index weighting. It is the dynamic mean of the attitude stability characteristics. It is its dynamic standard deviation.

[0060] This Bayesian adaptive thresholding mechanism has the following advantages: First, it provides personalized adaptation by fusing group priors with individual data through Bayesian updates, allowing the threshold to dynamically adjust according to the user's real-time status. Second, it eliminates body size differences by ensuring that depth judgment is not affected by the user's height ratio through femur length normalization. Third, it has strong anti-interference capabilities by using statistical features (such as standard deviation) to set the threshold, effectively resisting transient noise interference. Fourth, it provides real-time adaptation by using exponentially weighted recursive updates to enable the stability threshold to quickly respond to changes in user fatigue and posture control ability. This scheme makes the segmentation of action phases more accurate and robust, providing a reliable temporal framework for subsequent action counting and quality assessment.

[0061] An anti-shake mechanism is introduced: the transfer must meet the condition ≥ 3 times in 5 consecutive frames and the condition must be met in the last frame. The majority voting of the timing signal is used to suppress noise interference.

[0062] 3. Determine the switching boundaries between motion phases by analyzing the trends in joint angle changes and the posture stability indicators of key point velocity changes.

[0063] The state machine provides frame-level coarse judgment of switching moments. Kernelized support vector machines are used to classify multi-domain feature signals, achieving precise sub-frame boundary localization. During offline training, a subset of feature signals from Δ=5 frames before and after the frame switching is extracted. , The knee flexion angle after weighted smoothing is the core discriminative feature for switching between movement phases (such as squatting and standing up), and it originates from the output of the weighted joint angle calculation module mentioned above.

[0064] The smooth motion velocity of the hip joint key points in the vertical direction (y-axis) is the core feature for identifying the transition between the rising and falling phases of a movement, and it originates from the output of the weighted velocity calculation module mentioned above.

[0065] It is a weighted stability index for the trunk region, representing the degree of fluctuation in trunk posture. It is an auxiliary feature for identifying the stable state and dynamic switching of movement, and is derived from the output of the weighted stability index module mentioned above.

[0066] The smooth angular acceleration of the knee joint characterizes the rate of change of velocity in joint flexion and extension. It is a highly sensitive feature for capturing the start and stop of movement and the switching of force. It is obtained by weighted smoothing of the second difference of the joint angle.

[0067] This formula is a multi-domain fusion feature vector for stage boundary recognition, which integrates four core temporal features: joint angle, key point velocity, trunk stability, and joint acceleration. It covers the core kinematic changes during action stage switching, provides highly discriminative input features for SVM classification, and solves the problems of poor anti-interference ability and high misjudgment rate of single feature boundary recognition.

[0068] Training a Gaussian kernel SVM: , The multi-domain fusion feature vector consists of two inputs; during offline training, these are the features of positive samples (stage switching boundary frames) and negative samples (non-boundary frames), respectively; during online inference, they are the features of the current frame and the trained support vector features, respectively.

[0069] The square of the Euclidean distance (L2 norm squared) between two feature vectors represents the distance between the two features in the feature space. The greater the distance, the greater the difference between the features.

[0070] The bandwidth parameter of the Gaussian kernel function is a core hyperparameter of the SVM model, controlling the range of the kernel function: the smaller the value, the more sensitive the model is to feature differences and the stronger its fitting ability; the larger the value, the stronger the model's generalization ability.

[0071] This formula uses the Gaussian kernel function (RBF kernel) of the SVM model. Its core function is to map linearly inseparable temporal features in low-dimensional space to a high-dimensional Hilbert space, making them linearly separable. This enables accurate classification of action phase transition boundaries and non-boundary frames, solving the problems of nonlinearity in action temporal features and insufficient accuracy of traditional linear classifiers. and They represent the first The and the first The feature vector of each sample The squared Euclidean distance between two eigenvectors is given. The kernel function bandwidth parameter controls the rate at which similarity decays. The smaller the value, the more sensitive the kernel function is to feature differences; the larger the value, the smoother the similarity distribution.

[0072] During online inference, in the candidate interval Calculate the discriminant function for the feature signal of each frame: The formula in question is the discriminant function of the SVM model, which is the core formula for identifying stage switching boundaries during online inference. The output value C(t) is the confidence level that the current frame signal is a "stage switching boundary". The larger the value, the higher the probability that the frame is a switching boundary. By selecting the maximum value of C(t) within the candidate interval, the sub-frame level switching boundary can be accurately located. The number of support vectors obtained after the SVM model is trained is automatically determined by the offline training of the model and represents the number of training samples that play a key role in classification decisions.

[0073] is the Lagrange multiplier, which is the optimal parameter obtained from offline training of the SVM model. It corresponds to the weight coefficient of the i-th support vector and determines the contribution of the support vector to the classification decision.

[0074] The label value of the training sample corresponds to the classification label of the i-th support vector: Represents a positive sample (stage switching boundary frame). Represents negative samples (non-boundary frames);

[0075] The value is calculated using the Gaussian kernel function, and the input is the current frame feature. Features of the i-th support vector The similarity between the two is represented by the value of C(t). The point where C(t) reaches its maximum value is selected as the precise switching time. The positioning accuracy reaches 1 / 5 of a frame interval.

[0076] Through the above mechanism, this invention transforms the original key point signals into a structured semantic stage sequence.

[0077] In step S6, based on the time series analysis results, the human movement is quantitatively evaluated, and the movement completion, stability, and standardization indicators are output. At the same time, the types of movement deviations and their occurrence stages are identified.

[0078] In the evaluation dimensions and indicators phase, the system quantifies motion evaluation through three dimensions. Motion amplitude is based on comparing the extreme values ​​of key joint angles within the motion cycle with preset thresholds; posture accuracy is monitored in real-time for specific angles or relative positional relationships; stability is measured by the variance of the motion trajectory of core key points, with smaller variance indicating more stable motion. In the scoring and feedback generation phase, the system uses a weighted summation method, assigning different weights to the three dimensions and synthesizing the sub-scores of each dimension into a comprehensive quality score ranging from 0 to 100. , Let represent the score of the i-th evaluation dimension; Let represent the weight coefficient of the i-th evaluation dimension.

[0079] Motion Deviation Recognition: Based on the geometric relationships of key points and preset rules, the system can automatically identify common types of motion deviations and accurately pinpoint their stages of occurrence. Detection logic includes: knee valgus, excessive forward tilting of the torso, and insufficient squat depth.

[0080] In step S7, the analysis results are output in the form of scores, text, graphics or voice, etc., to guide users in correcting their actions.

[0081] The quantitative scoring module outputs a comprehensive score of 0-100 points in real time, dynamically updating on the data panel. When scores fall below a preset threshold, color or icon warnings are triggered, and historical records are automatically saved to track long-term training trends. It transforms invisible process quality (such as training intensity) into visible digital assets. Its core value lies in solving both the immediate perception of "whether the current state is good" and the long-term decision-making problem of "whether the method is correct," ultimately improving management efficiency and safety margins through digital means. The text feedback module dynamically displays specific and actionable correction prompts during training. After each training session, it generates a concise summary report detailing the type of deviation, frequency of occurrence, and improvement suggestions. It effectively addresses the pain points of traditional training, such as delayed feedback, vague guidance, and inefficient debriefing. It not only ensures the standardization of actions in each training session through "instant correction" but also helps users build clear self-awareness through "data-driven diagnostic reports," achieving a shift from "blind trial and error" to "precise correction." The graphical visualization module overlays general human skeletal lines onto the real-time video stream, highlights incorrect areas with dynamic directional arrows, and synchronously drives a freely rotatable and scalable 3D human model. It also plots the changes in core joint angles in real-time using line graphs and overlays standard movement ranges. This approach visualizes invisible biomechanical data, refines vague experience-based judgments, and makes delayed result analysis real-time. It's not just a display tool, but a motion-level intelligent error correction and quantitative analysis system that effectively shortens the learning cycle, reduces injury risk, and improves movement quality. The voice prompt module supports key status announcements, continuous real-time error reminders, and brief summaries at the end of training, providing users with an immersive and uninterrupted interactive experience.

[0082] This invention integrates 3D pose estimation and multi-keypoint temporal analysis techniques, possessing significant technical advantages and application value, as detailed below:

[0083] This method constructs a multi-keypoint temporal semantic fusion modeling system to uncover the temporal correlations and collaborative motion patterns between key points, accurately divides action stages, identifies temporal logic, and achieves refined analysis of the entire cycle of continuous actions. It effectively solves problems of inaccurate counting and unclear stage boundaries, improving the accuracy of motion analysis. Essentially, this approach represents a paradigm shift from "single-point threshold judgment" to "multi-point spatiotemporal graph analysis." By constructing a "semantic collaboration network" between key points, it effectively addresses the three major problems in traditional visual motion analysis: inaccurate counting, unclear boundaries, and difficulty in logical judgment.

[0084] A multi-keypoint temporal denoising scheme is proposed to suppress interference from lighting, occlusion, and jitter, thereby improving the temporal stability of 3D keypoints. An adaptive adjustment mechanism based on individual physiological characteristics is designed to dynamically adapt to the individual's range of motion and rhythm of movement, improving the objectivity and universality of the analysis results. Essentially, the reliability of data acquisition is solved by "enhancing temporal stability," and the effectiveness of data analysis is solved by "adapting to physiological characteristics." Ultimately, this enables 3D human pose analysis to achieve high precision, high robustness, and high personalization in complex real-world environments.

[0085] Although the invention has been described herein with reference to several illustrative embodiments, it should be understood that many other modifications and implementations can be devised by those skilled in the art, which will fall within the scope and spirit of the principles disclosed herein. More specifically, various variations and modifications can be made to the components and / or layout of the subject matter arrangement within the scope of the disclosure, drawings, and claims. Besides variations and modifications to the components and / or layout, other uses will be apparent to those skilled in the art.

Claims

1. A method for 3D human pose estimation and multi-keypoint temporal analysis, characterized in that, Includes the following steps: S1, Multimodal video access, which means acquiring video data containing human motion through image acquisition devices; S2. Spatiotemporal perception and extraction of human key points, that is, using a human detection model and pose estimation model based on deep learning to extract human key point information from video frames. The key point information includes at least the two-dimensional or three-dimensional coordinates of multiple joints of the human body and their corresponding confidence scores. S3, 3D pose reconstruction and credibility modeling, that is, based on the key point information, construct a human 3D pose model and assign credibility indexes to each key point or bone segment to characterize the reliability of the pose estimation result of the key point. S4. Keypoint Temporal Feature Quantization: This involves organizing 3D keypoint information from a continuous time series to generate multi-keypoint temporal features. These features include keypoint spatial coordinate variation characteristics, joint angles and their rate of change, keypoint trajectory features, and attitude stability features. The keypoint trajectory features quantify the economy and regularity of the keypoint trajectory path, introducing a trajectory efficiency index. This index closely correlates the geometric properties of the trajectory with the energy efficiency of the signal, and is specifically defined as: ,in Let i be the two-dimensional coordinate vector of the i-th key point in frame t. , The Euclidean distance between key points in adjacent frames. This represents the total distance traveled by all keypoints from frame 2 to frame T. This represents the linear displacement of key points between the first and last frames of the window. It is a very small normal number; postural stability features assess the degree of postural stability of the body's core region during movement or static maintenance, and define a multi-keypoint cooperative stability index, defined as... , For set The number of key points Let x be the sample variance of the x-coordinate of the j-th keypoint within the time window. Let be the sample variance of the y-coordinate of the j-th keypoint within the time window. The expected length of the bone segment corresponding to keypoint j. To calculate the values ​​of the expressions corresponding to all key points in set j and sum them; S5. Temporal stage perception and logical judgment, that is, based on the temporal characteristics of multiple key points, perform temporal analysis on human movements, identify different stages of the movements, and judge whether the movements meet the preset temporal logical relationships. S6. Quantitative assessment and intelligent diagnosis: Based on the time series analysis results, the human body's movements are quantitatively assessed, and indicators of movement completion, stability, and standardization are output. At the same time, the types of movement deviations and their occurrence stages are identified. S7. Results Output: The analysis results are output in the form of scores, text, graphics, or voice to guide users in correcting their actions.

2. The method for 3D human pose estimation and multi-keypoint temporal analysis according to claim 1, characterized in that: In step S3, the credibility modeling is constructed from information on kinematic plausibility credibility, specifically as follows: Construct a bone length ratio test scoring function. For a bone segment composed of keypoints p and q, its current length is... The desired length of the preset bone segment and length tolerance parameters The length reasonableness score is: ,in Let L be an exponential function with base e, and L be the actual measured length of the current skeletal segment. This is the expected length of the bone segment; Construct a joint angle limit test scoring function for a joint angle formed by adjacent points a and b with key point i as the vertex. The normal range of motion of this joint is And define the angle reasonableness score as: ,in This represents the lower limit of the normal range of motion of the joint. This represents the upper limit of the normal range of motion of the joint. For the nearest boundary value, when hour ; hour , This is a preset constant that controls the rate at which the score decreases when the angle exceeds the limit. To scale the squared deviation.

3. The method for 3D human pose estimation and multi-keypoint temporal analysis according to claim 1, characterized in that: The key point spatial coordinate change feature in step S4 is as follows: Let the two-dimensional image coordinates of the i-th key point in frame t be... The system frame rate is FPS, and the time interval between adjacent frames is... The second, and the corresponding nth order kinematic quantity are uniformly expressed as , Let be the nth-order kinematic vector of the i-th keypoint in frame t, where n is a non-negative integer; n=0 represents position, n=1 represents velocity, and n=2 represents acceleration. For the nth-order forward difference starting from frame t−n+1, The time interval is the nth power; The joint angle change rate is calculated using a confidence-weighted time-series fusion strategy, as shown in the following formula. ,in The original joint angle for frame t+k. Let k be the confidence level of the joint center key point in this frame, and k be the time offset for the current frame t, taking the integer value in the range [−k, k]. For time weighting coefficients, To sum all frames within the window, iterate through all frames from t−K to t+K.

4. The method for 3D human pose estimation and multi-keypoint temporal analysis according to claim 1, characterized in that: In step S5, the action phase based on time-series features is identified as inputting real-time joint angles and coordinate change rates, driving the state machine to migrate according to a preset threshold.