Adult joint motion range evaluation method and system based on multi-view video

By constructing individualized joint geometry models from multi-view videos and performing incremental optimization using factor graphs, the instability issues caused by occlusion and individual differences in joint mobility assessment were resolved, thereby improving the reliability and usability of joint mobility assessment.

CN121811512APending Publication Date: 2026-04-07THE FIRST AFFILIATED HOSPITAL OF ZHENGZHOU UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-12
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing technologies for assessing joint range of motion in adults suffer from occlusion, observation noise, and individual differences, leading to unstable joint range of motion boundaries. They also lack verifiable and closed-loop re-sampling mechanisms, making it difficult to implement stably in real-world scenarios.

Method used

By acquiring observation datasets from multi-view video, an individualized joint geometric model is constructed. Supplementary acquisition instructions are generated through factor graph incremental optimization and counterfactual consistency testing, enabling traceable, verifiable, and closed-loop updating of joint mobility assessment.

Benefits of technology

It improves the reliability and engineering usability of joint range of motion assessment results. Through multi-view redundant observation support, explicit modeling of individual differences, non-physical posture suppression and segmentation noise robustness enhancement, it achieves continuous and stable output of joint angle sequences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121811512A_ABST
    Figure CN121811512A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of human body kinematics parameter measurement, in particular to an adult joint motion range evaluation method and system based on a multi-view video, and the method comprises the steps: obtaining a multi-view video sequence of a testee, and generating an observation data set containing two-dimensional key point observation, a human body segmentation mask and camera parameter information; constructing an individualized joint geometric model containing a bone segment length, a joint center, a joint principal axis and a joint angle coordinate system based on the observation data set, and generating a three-dimensional symbol distance field; constructing a factor graph and performing incremental optimization solution, and outputting a joint angle sequence and an abnormal frame set; and determining an activity range interval and carrying out anti-fact consistency check, and if the validity is not satisfied, calculating an information matrix by a Jacobian matrix to generate a supplementary collection instruction to update a result, thereby improving robustness and verifiability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of human kinematic parameter measurement technology, and in particular to a method and system for assessing adult joint range of motion based on multi-view video. Background Technology

[0002] Adult joint range of motion assessment is a fundamental quantitative indicator in rehabilitation evaluation, orthopedic follow-up, sports injury management, and occupational health screening, directly impacting triage, treatment efficacy evaluation, and resource allocation. Current clinical practice often relies on manual goniometers or single-view video / wearable sensors, which suffers from high subjectivity, poor repeatability, sensitivity to occlusion and posture shifts, and difficulty in establishing a traceable chain of quality control evidence. While some 3D reconstruction solutions can output joint angles, they often ignore individual bone segment and joint axis differences or lack constraints on the geometrically feasible domain of the human body, leading to non-physical jumps when occluded, segmented jitter, or keypoint drift occurs. Furthermore, the assessment process often focuses solely on "outputting angles," lacking a self-checking mechanism for the effectiveness of range of motion boundaries and a closed-loop process to guide supplementary data collection when information is insufficient, making stable implementation in real-world scenarios difficult. Summary of the Invention

[0003] This invention provides a method and system for assessing adult joint mobility based on multi-view video, which at least solves the problem that joint mobility boundaries are unstable and lack verifiable and closed-loop re-acquisition mechanisms under real acquisition conditions due to occlusion, observation noise and individual differences.

[0004] In a first aspect, the present invention provides a method for assessing the range of motion of adult joints based on multi-view video, comprising the following steps: A multi-view video sequence of the test subject is obtained to generate an observation dataset, which includes two-dimensional human key point observation, human segmentation mask and camera parameter information. An individualized joint geometry model is constructed based on the observation dataset. The individualized joint geometry model includes bone segment length, joint center, joint principal axis and joint angle coordinate system, and a three-dimensional symbolic distance field is generated by camera parameter information and human body segmentation mask. Based on the individualized joint geometry model, observation dataset and three-dimensional symbolic distance field, factor graph incremental optimization solution is constructed to output joint angle sequence and abnormal frame set; The joint range of motion is determined based on the joint angle sequence and the joint range of motion assessment result is obtained through counterfactual consistency test. If the validity condition is not met, an information matrix is ​​calculated based on the Jacobian matrix to generate a supplementary acquisition instruction. The supplementary acquisition instruction is used to re-acquire multi-view video sequences to update the joint range of motion assessment result.

[0005] In one possible implementation, generating the observation dataset includes: temporally aligning the multi-view video sequences to form a unified timeline; performing camera calibration on the multi-view video sequences to generate camera parameter information; determining the human body region of the subject in the multi-view video sequences based on the unified timeline; performing two-dimensional human keypoint detection within the human body region to generate two-dimensional human keypoint observations, and generating keypoint confidence scores for the two-dimensional human keypoint observations; and performing human body segmentation within the human body region to generate a human body segmentation mask corresponding to the unified timeline.

[0006] In one possible implementation, joint geometry calibration action segments are determined by two-dimensional human keypoint observations, and multi-view correspondences are established for the two-dimensional human keypoint observations within the joint geometry calibration action segments to generate a keypoint observation sequence for joint geometry calibration.

[0007] In one possible implementation, weighted triangulation is performed on the two-dimensional human keypoint observations of the same human keypoints from different viewpoints using camera parameter information to generate three-dimensional keypoint trajectories. The weights of the weighted triangulation are determined by the keypoint confidence. Bone segment length consistency constraint optimization is then performed based on the three-dimensional keypoint trajectory to ensure consistency between bone segment length and the three-dimensional keypoint trajectory.

[0008] In one possible implementation, the length of a bone segment is determined by the trajectory of three-dimensional key points; the joint center is determined by the trajectory of three-dimensional key points; the joint principal axis is determined by the relative motion of adjacent bone segments; and a joint angle coordinate system is constructed by the joint principal axis, the joint center, and the bone segment direction vectors of adjacent bone segments.

[0009] In one possible implementation, generating a three-dimensional symbolic distance field based on camera parameter information and a human body segmentation mask includes: performing voxel sculpting on the human body segmentation mask under the constraints of camera parameter information to generate a human body shell pixel model; and generating a three-dimensional symbolic distance field based on the human body shell pixel model, wherein the three-dimensional symbolic distance field is used to characterize the feasible region of the human body shell.

[0010] In one possible implementation, the factor map is constructed by: constructing a reprojection factor from two-dimensional human keypoint observations, camera parameter information, and individualized joint geometry models; constructing an outer shell constraint factor from a three-dimensional symbolic distance field and individualized joint geometry models; constructing a bone segment length factor from bone segment length; constructing a joint principal axis prior factor from joint principal axes; and constructing a temporal smoothing factor from the adjacent frame differences of the joint angle sequence.

[0011] In one possible implementation, the set of output anomalous frames includes: calculating the reprojection factor residual and the shell constraint factor residual during incremental optimization; determining the observation inconsistency frames based on the reprojection factor residual and the shell constraint factor residual; and adding the observation inconsistency frames to the set of anomalous frames.

[0012] In one possible implementation, determining the joint range of motion and performing a counterfactual consistency check includes: mapping the joint angle sequence to a joint orientation angle sequence using a joint angle coordinate system; determining candidate platform segments using the joint orientation angle sequence and a set of abnormal frames, where the candidate platform segment is a time period in which the difference between adjacent frames of the joint orientation angle sequence continuously satisfies a preset stability condition; determining the boundary values ​​of the joint orientation angle sequence within the candidate platform segment to form the joint range of motion; applying perturbation to the boundary values ​​of the joint range of motion to generate a counterfactual joint angle sequence, and calculating the cost difference based on the cost function of the factor graph to determine the validity of the boundary values; and generating supplementary acquisition instructions based on the Jacobian matrix by calculating the information matrix includes: obtaining the Jacobian matrix during the incremental optimization solution of the factor graph; calculating the information matrix based on the Jacobian matrix; and determining supplementary acquisition instructions based on the information matrix.

[0013] Secondly, the present invention provides an adult joint range of motion assessment system based on multi-view video, used to implement the adult joint range of motion assessment method based on multi-view video, the system comprising: The observation generation module is used to acquire multi-view video sequences of the subject and generate an observation dataset, which includes two-dimensional human key point observations, human segmentation masks, and camera parameter information. The geometry construction module is used to build individualized joint geometry models based on the observation dataset. The individualized joint geometry model includes bone segment length, joint center, joint principal axis and joint angle coordinate system, and generates a three-dimensional symbolic distance field from camera parameter information and human segmentation mask. The inversion solution module is used to construct a factor map based on the individualized joint geometry model, observation dataset and three-dimensional symbolic distance field and perform incremental optimization solution, outputting joint angle sequence and abnormal frame set; The evaluation and update module is used to determine the joint range of motion based on the joint angle sequence and obtain the joint range of motion evaluation result through counterfactual consistency test. When the validity condition is not met, the module calculates the information matrix based on the Jacobian matrix to generate supplementary acquisition instructions, and re-acquires multi-view video sequences according to the supplementary acquisition instructions to update the joint range of motion evaluation result.

[0014] Compared with the prior art, the advantages and beneficial effects of the present invention are as follows: By jointly constructing an observation dataset using multi-view keypoints and segmentation masks, redundant observation support for occlusion and viewpoint changes was achieved. Explicit modeling of individual differences was achieved through individualized joint geometry models (bone segment length, joint center, joint principal axis, and joint angle coordinate system), reducing the systematic bias introduced by general models. The introduction of a three-dimensional symbolic distance field as a constraint on the geometrically feasible domain of the human body shape suppressed non-physical postures and improved robustness against segmentation noise. Continuous and stable output of joint angle sequences and synchronous localization of abnormal frames were achieved through factor graph incremental optimization combined with reprojection, external shell, and temporal smoothing constraints. The verifiability of activity boundaries and a closed loop of "evaluation—quality control—supplementary acquisition—update" were realized through counterfactual consistency checks and supplementary acquisition commands driven by an information matrix, improving the reliability and engineering usability of the results. Attached Figure Description

[0015] Figure 1 This is a schematic diagram of the execution flow of the method of the present invention; Figure 2 This is a schematic diagram illustrating the joint angle sequence and range of motion boundary annotations in a specific embodiment of the present invention; Figure 3 This is a schematic diagram of factor graph constraint residual diagnosis in a specific embodiment of the present invention; Figure 4 This is a schematic diagram illustrating the platform segment information strength and supplementary data collection triggering in a specific embodiment of the present invention; Figure 5 This is a top-view schematic diagram of the four-view acquisition arrangement in a specific embodiment of the present invention; Figure 6 This is a structural block diagram of the system of the present invention. Detailed Implementation

[0016] The embodiments of the present disclosure will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the disclosure. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the present disclosure for ease of explanation.

[0017] Multi-view video-based adult joint mobility assessment refers to the simultaneous recording of a subject's joint movements from multiple observation positions within the same time period. The pose cues from each viewpoint are then fused along a unified timeline to reconstruct the joint's trajectory in three-dimensional space and the sequence of joint angle changes over time. This allows for the extraction of assessment elements such as mobility boundaries and stable plateau segments from the joint angle sequence. Compared to single-view video, multi-view video provides complementary observations in spatial geometry. Even when the same joint experiences occlusion, perspective shortening, or similar appearance leading to unstable positioning, constraints can still be provided by other views. Simultaneously, multi-view synchronization creates conditions for the subsequent introduction of camera geometric constraints, shape geometric constraints, and cross-frame continuity constraints, enabling the assessment process to not only "provide numerical values" but also verify the consistency of observations and the validity of results. Based on this, this invention proposes a multi-view video-based adult joint mobility assessment workflow. This workflow involves constructing an observation dataset, individualized joint geometric modeling, incremental optimization of factor graphs, and counterfactual consistency testing. Supplementary acquisition instructions are generated when validity is insufficient, achieving a traceable, verifiable, and closed-loop updated joint mobility assessment.

[0018] like Figure 1 As shown, a method for assessing adult joint range of motion based on multi-view video includes the following steps: A multi-view video sequence of the test subject is obtained to generate an observation dataset, which includes two-dimensional human key point observation, human segmentation mask and camera parameter information. In one embodiment, at least two cameras simultaneously capture images of the subject to obtain a multi-view video sequence. The frame rate and exposure parameters of each camera are kept consistent, and a synchronization trigger signal is output before acquisition begins to align the acquisition times of each camera. The multi-view video sequence is used to cover the main motion plane of the limb containing the target joint of the subject, minimizing mutual occlusion. Temporal alignment processing is performed on the multi-view video sequence to form a multi-view frame sequence with a unified timeline. In each frame corresponding to the unified timeline, two-dimensional human keypoint detection is performed to obtain two-dimensional human keypoint observations, including the two-dimensional coordinates and confidence scores of each keypoint. Human segmentation is performed in the same frame to obtain a human segmentation mask, used to indicate the boundaries between human pixel regions and background pixel regions. The camera is calibrated to obtain camera parameter information, including the camera intrinsic and extrinsic parameter matrices. The two-dimensional human keypoint observations, the human segmentation mask, and the camera parameter information constitute an observation dataset, used for subsequent individualized joint geometry model construction and joint mobility assessment.

[0019] The process of generating the observation dataset includes: temporally aligning multi-view video sequences to form a unified timeline; performing camera calibration on the multi-view video sequences to generate camera parameter information; determining the human body region of the subject in the multi-view video sequences based on the unified timeline; performing two-dimensional human keypoint detection within the human body region to generate two-dimensional human keypoint observations, and generating keypoint confidence scores for the two-dimensional human keypoint observations; and performing human body segmentation within the human body region to generate a human body segmentation mask corresponding to the unified timeline.

[0020] In one embodiment, the multi-view video sequence is acquired simultaneously from at least two cameras, with each camera maintaining a consistent frame rate. At the start of acquisition, synchronization is achieved via a synchronous trigger signal or a unified clock source to align the acquisition times. The multi-view video sequence is then time-aligned to form a unified timeline. This time alignment process includes: extracting the frame timestamps from each camera's video and establishing a frame index mapping relationship; when frame drops or jitter occur, the nearest frame matching method is used to align the frames from each viewpoint to the corresponding frame timestamps on the unified timeline, ensuring that the same timeline index corresponds to the same frame at each viewpoint, facilitating subsequent cross-viewpoint consistency processing.

[0021] Camera calibration is performed on multi-view video sequences to generate camera parameter information. Camera calibration can be completed offline before acquisition or under the same acquisition environment according to preset calibration actions. The camera parameter information includes at least a camera intrinsic parameter matrix and a camera extrinsic parameter matrix. The camera intrinsic parameter matrix is ​​used to characterize imaging parameters such as imaging focal length and principal point position, while the camera extrinsic parameter matrix is ​​used to characterize the pose relationship between the camera coordinate system and the world coordinate system. To ensure the usability of the calibration results, the reprojection error obtained from the calibration can be checked for consistency. When the reprojection error exceeds a preset range, the calibration sequence is reacquired and the camera parameter information is updated.

[0022] Human body regions of the test subject are determined in multi-view video sequences based on a unified timeline. These regions are used to limit the processing range for subsequent keypoint detection and human body segmentation, reducing background interference. Human body regions can be determined using object detection-based bounding box extraction or foreground segmentation-based connected component extraction. When multiple candidate human body regions exist under the same unified timeline index, the corresponding human body region for the test subject is selected based on region area, region positional continuity, and cross-view consistency. The human body region is then smoothly updated between adjacent frames to avoid jitter.

[0023] Two-dimensional human keypoint detection is performed within the human body region to generate two-dimensional human keypoint observations, and keypoint confidence scores are generated for these observations. The output of the two-dimensional human keypoint detection includes the two-dimensional coordinates of a preset set of human keypoints, and the keypoint confidence score characterizes the reliability of the corresponding two-dimensional coordinates. The keypoint confidence score can be obtained from the peak value of the keypoint heatmap, the output score of the keypoint regression branch, or a fusion of both, and is used for subsequent cross-view correspondence establishment and cross-frame stability assessment. To improve the stability of keypoint observations, temporal filtering can be performed on the keypoint coordinates, and when the keypoint confidence score is lower than a preset threshold, the keypoint is marked as a low-confidence keypoint for subsequent processing to reduce its weight or remove it.

[0024] Human segmentation is performed within the human body region to generate a human segmentation mask corresponding to a unified timeline. The human segmentation mask indicates the boundary between the human pixel region and the background pixel region. Human segmentation can employ a semantic segmentation network to output a human category mask, or an instance segmentation network to output a subject instance mask. When multiple human instances exist, the human segmentation mask and the human body region are used together to locate the subject instance. To ensure the consistency of the human segmentation mask's boundaries, morphological closing operations can be performed on the mask to fill holes, and edges can be smoothed to reduce jagged edges.

[0025] The observation dataset consists of the time-aligned multi-view frame sequence, camera parameter information, human body regions, 2D human keypoint observations, keypoint confidence scores, and human body segmentation masks. This observation dataset is used in subsequent steps to construct individualized joint geometry models and generate 3D symbolic distance fields. It also provides observation constraints for incremental optimization of the factor graph, thereby improving the stability of the joint angle sequence and reducing evaluation bias caused by observation inconsistencies under conditions such as occlusion and rapid movement.

[0026] The joint geometry calibration action segment is determined by two-dimensional human key point observation, and a multi-view correspondence relationship is established for the two-dimensional human key point observation within the joint geometry calibration action segment to generate a key point observation sequence for joint geometry calibration.

[0027] In one embodiment, to construct an individualized joint geometry model, joint geometry calibration action segments are first determined by two-dimensional human keypoint observation. The joint geometry calibration action segments are continuous frame intervals used for stable estimation of joint geometry parameters. The selection criteria are: the target joint undergoes sufficient and identifiable rotational changes within the interval, while the overall confidence of the keypoints is high, occlusion is minimal, the subject's posture is relatively stable, and there is no significant translational departure from the field of view. Specifically, under a unified time axis, the confidence of the target limb keypoints in each frame is summarized, and the presence of significant movement is determined based on the inter-frame changes in keypoint coordinates. When a continuous frame satisfies that the keypoint confidence is not lower than a preset threshold and the keypoint coordinate changes show a continuous trend, the continuous frame interval is selected as a candidate segment. To avoid misjudging short-term jitter as calibration actions, the duration of the candidate segments is further required to meet a preset lower limit, and frames containing significant occlusion are removed. The determination of occluded frames can be based on a sudden drop in target keypoint confidence, inconsistency between the human segmentation mask and the keypoint position, or an increase in the proportion of missing keypoints. If multiple candidate segments exist, the candidate segment with the larger change in target joint angle and more complete observation of key points is selected as the joint geometry calibration action segment.

[0028] After determining the joint geometry calibration action segment, a multi-view correspondence relationship is established for the observation of 2D human keypoints within this segment to generate a keypoint observation sequence for joint geometry calibration. The multi-view correspondence establishment is based on a unified time axis, ensuring that frames from different viewpoints under the same time axis index form a set of synchronous observations. The semantic category of the human keypoints is used as the primary correspondence criterion, ensuring that human keypoints of the same semantic category from different viewpoints constitute the same correspondence group. To improve the robustness of the correspondence, for each set of synchronous observations, the coordinates and confidence scores of target keypoints are first extracted from the human body region of each viewpoint. When the confidence score of a target keypoint in a certain viewpoint is lower than a preset threshold, that keypoint in that viewpoint is marked as an invalid observation and does not participate in the correspondence group of this frame. When multiple candidate positions of the same semantic category appear in a certain viewpoint, the selection is made by combining the continuity of the human body region position and the geometric consistency of the camera parameter information. Geometric consistency can be achieved by projecting the corresponding keypoints from other viewpoints onto this viewpoint according to the constraints of the camera parameter information and comparing the distances. For keypoints that are temporarily missing, interpolation can be used to complete them based on the continuity of keypoint trajectories in adjacent frames, and the weighted result of the confidence scores of adjacent frames can be inherited as the completion confidence score. Finally, the keypoint observation sequence records the corresponding groups of multi-view keypoints for each frame in the joint geometry calibration action segment in chronological order. The keypoint observation sequence includes the coordinates of keypoints in each view, the confidence score of keypoints, and invalid observation markers, which are used for subsequent estimation of bone segment length, joint center, and joint principal axis, thereby reducing geometric parameter deviations caused by cross-view mismatch and occlusion.

[0029] The two-dimensional human keypoints observed from different perspectives using camera parameter information are weighted triangulated to generate three-dimensional keypoint trajectories. The weights of the triangulation are determined by the keypoint confidence. Based on the three-dimensional keypoint trajectories, bone segment length consistency constraint optimization is performed to ensure consistency between bone segment length and the three-dimensional keypoint trajectories.

[0030] In one embodiment, weighted triangulation is performed on the 2D human keypoint observations of the same human keypoint from different viewpoints based on camera parameter information to generate 3D keypoint trajectories. Specifically, in each frame on a unified timeline, for each human keypoint, the 2D coordinates and keypoint confidence scores of the human keypoints corresponding to each viewpoint are collected; when the keypoint confidence score of a certain viewpoint is lower than a preset threshold, the 2D coordinates of the human keypoints from that viewpoint are marked as invalid observations and do not participate in triangulation; when the number of effective observation viewpoints is less than two, the human keypoint in that frame is marked as a missing frame and processed later. For human keypoints with at least two effective observation viewpoints, a projection relationship from the 3D spatial point to the imaging plane of each viewpoint is established based on camera parameter information, and the 3D spatial point is solved by minimizing the weighted projection error using the keypoint confidence score as the observation weight, to obtain the 3D coordinates of the human keypoints in that frame. To improve feasibility, the 3D spatial point solution can first obtain initial values ​​from the two viewpoints with the highest confidence through linear triangulation, and then perform a small number of iterative updates using iterative least squares. In each iteration, the reprojection residuals of each viewpoint are calculated. When the residual of a certain viewpoint exceeds a preset upper limit, the weight of that viewpoint is reduced or the observation of that viewpoint is removed to suppress the impact of falsely detected keypoints on the 3D results. Thus, the 3D coordinate sequence of each human body keypoint is obtained frame by frame on the time axis, forming a 3D keypoint trajectory. For missing frames, interpolation can be performed based on the continuity of the 3D keypoint trajectories of adjacent frames to complete the missing frames, and the missing markers are retained for subsequent abnormal frame determination.

[0031] The core calculation of weighted triangulation can be expressed as: ; in, The three-dimensional coordinates of key points on the target human body; The variables to be solved are the three-dimensional coordinates; For perspective indexing; The observation weights are determined by the confidence scores of the key points; The projection function is determined by the camera parameter information; From the perspective Two-dimensional coordinates of the same key points on the human body.

[0032] After obtaining the 3D keypoint trajectory, bone segment length consistency constraint optimization is performed based on the 3D keypoint trajectory to ensure consistency between bone segment length and the 3D keypoint trajectory. Specifically, bone segment connection relationships are predefined, and human keypoint pairs corresponding to adjacent joint centers are taken as bone segment endpoints. Within the joint geometry calibration action segment, frames with high keypoint confidence and small reprojection residuals are selected. The 3D distance between bone segment endpoints is calculated, and the bone segment length is determined as an estimate using the median. Subsequently, the 3D keypoint trajectory obtained by weighted triangulation is used as the initial trajectory. The 3D keypoint coordinates of each frame are slightly corrected to make the distance between each bone segment endpoint close to the corresponding bone segment length estimate, while keeping the deviation between the corrected trajectory and the initial trajectory within a preset range. Furthermore, a smoothing constraint is applied to the 3D coordinate changes of adjacent frames to suppress jitter. The bone segment length consistency constraint can be expressed in the following form: ; in, For the timeline Frame bone segment endpoints Three-dimensional coordinates of key points on the human body; For the timeline Frame bone segment endpoints Three-dimensional coordinates of key points on the human body; Endpoint of bone segment With the end point of the bone segment The corresponding estimated bone segment length. Through the above processing, the 3D keypoint trajectory can still maintain a stable scale and reasonable bone segment geometry even when occluded, falsely detected, or with individual viewpoint failures. This provides consistent input for the subsequent determination of bone segment length, joint center, and joint principal axis, and reduces the risk of amplitude deviation and jumps in joint range of motion assessment.

[0033] An individualized joint geometry model is constructed based on the observation dataset. The individualized joint geometry model includes bone segment length, joint center, joint principal axis and joint angle coordinate system, and a three-dimensional symbolic distance field is generated by camera parameter information and human body segmentation mask. In one embodiment, constructing an individualized joint geometry model based on the observation dataset includes: within the joint geometry calibration action segment, determining the bone segment length between adjacent joint keypoints based on the 3D keypoint trajectory after multi-view correspondence, and taking robust statistical values ​​of the bone segment lengths from multiple frames as the bone segment lengths; fitting the joint center with the 3D keypoint trajectory of the endpoints of adjacent bone segments, so that the joint center satisfies the distance relationship from the endpoints of adjacent bone segments to the joint center; estimating the joint principal axis based on the relative rotation trajectory of adjacent bone segments, so that the joint principal axis is consistent with the direction of the relative rotation; and then constructing a joint angle coordinate system with the joint principal axis as one axis and the direction vector of the joint center and bone segments to unify the positive and negative directions and zero reference of the joint angle. Generating a 3D symbolic distance field from camera parameter information and a human segmentation mask includes: projecting the human segmentation mask back onto a 3D voxel mesh under multi-view conditions to perform voxel sculpting to generate an outer shell pixel model, and calculating the distance from voxel points to the outer shell surface on the outer shell pixel model, with the distance inside the outer shell recorded as a negative value and the distance outside the outer shell recorded as a positive value, to obtain the 3D symbolic distance field.

[0034] The length of a bone segment is determined by the trajectory of three-dimensional key points; the joint center is determined by the trajectory of three-dimensional key points; the principal axis of the joint is determined by the relative motion of adjacent bone segments; and a joint angle coordinate system is constructed by the principal axis of the joint, the joint center, and the bone segment direction vectors of adjacent bone segments.

[0035] In one embodiment, determining bone segment length from 3D keypoint trajectories includes: pre-establishing a human skeleton connection table, defining pairs of human keypoints corresponding to adjacent joints as bone segment endpoint pairs; within the joint geometry calibration action segment, calculating the 3D distance for each frame's bone segment endpoint pairs, and filtering out low-confidence frames based on keypoint confidence and reprojection residuals; using robust statistics to obtain bone segment lengths from the 3D distances of the retained frames, where robust statistics can be the median or truncated mean to reduce bias caused by false detections in individual frames. The bone segment length is used to subsequently constrain the scale consistency of the 3D keypoint trajectories, ensuring that the triangulation results maintain a reasonable geometric proportion even under occlusion or limited effective viewing angles.

[0036] Determining the joint center from 3D keypoint trajectories involves: for the target joint, selecting the 3D keypoint trajectories corresponding to two adjacent bone segments connected to the joint, and fitting the joint center position within the joint geometry calibration motion segment. The fitting method involves finding a set of joint centers across multiple frames, ensuring that the distance relationship from the endpoints of adjacent bone segments to the joint center is consistent with the bone segment length, and maintaining a smooth joint center along the time axis. When some frames contain missing keypoints, interpolated trajectories from adjacent frames are used, or only valid frames are used for fitting. Through this process, the joint center spatially matches the individual dimensions of the test subject, reducing jumps caused by single-frame estimation.

[0037] Determining the joint principal axis based on the relative motion of adjacent bone segments includes: normalizing the bone segment orientation vectors of adjacent bone segments in each frame with the joint center as a reference, and calculating the relative rotation trend of adjacent bone segments on the time axis; within the joint geometry calibration action segment, fitting the direction of relative rotation to obtain the joint principal axis, ensuring consistency of the joint principal axis across multiple frames and robustness to abnormal frames. To avoid inconsistencies in the reverse direction, the direction of the joint principal axis can be oriented based on the bone segment orientation vectors in the subject's initial static posture, ensuring a consistent principal axis direction definition for the same subject across different acquisition batches.

[0038] A joint angle coordinate system is constructed using the joint principal axis, joint center, and bone segment direction vectors of adjacent bone segments. This system comprises: using the joint center as the origin, the joint principal axis as the first coordinate axis, and a reference direction selected in a plane orthogonal to the joint principal axis as the second coordinate axis. The reference direction is obtained by projecting the bone segment direction vector of a preset frame onto this plane. The third coordinate axis is determined by the cross product of the first two coordinate axes, thus forming a right-handed coordinate system. This joint angle coordinate system is used to uniformly map the subsequently solved joint postures into a comparable joint angle sequence, ensuring that the upper and lower boundaries of the joint range of motion have consistent physical meaning and reducing the range of motion estimation bias caused by inconsistencies in the coordinate system.

[0039] The generation of a 3D symbolic distance field based on camera parameter information and human body segmentation mask includes: performing voxel sculpting on the human body segmentation mask under the constraints of camera parameter information to generate a human body shell pixel model; and generating a 3D symbolic distance field based on the human body shell pixel model, which is used to characterize the feasible region of the human body shell.

[0040] In one embodiment, generating a 3D symbolic distance field based on camera parameter information and a human body segmentation mask includes: first, establishing a 3D voxel mesh in the test area, the 3D voxel mesh covering the possible location range of the test subject, and setting a 3D coordinate index for each voxel. For each frame in a unified time axis, acquiring the human body segmentation mask and camera parameter information corresponding to each viewpoint; the human body segmentation mask is used to indicate the human pixel region of the test subject under that viewpoint. The process of performing voxel sculpting on the human body segmentation mask under the constraints of camera parameter information to generate a human body shell model is as follows: for each voxel point in the 3D voxel mesh, the voxel point is projected onto the imaging plane of each viewpoint using the camera parameter information, and it is determined whether the projected point falls into the human pixel region of the human body segmentation mask of the corresponding viewpoint; when the voxel point falls into the human pixel region in all effective viewpoints, the voxel point is retained as a human body shell model; when the voxel point falls into the background pixel region in any effective viewpoint, the voxel point is removed as a non-human body shell model. To improve robustness to occlusion and segmentation errors, boundary smoothing and hole filling can be performed on the human body segmentation mask, and a tolerance band can be introduced into the projection boundary during voxel sculpting. This allows voxel points close to the human body boundary to be retained or removed based on multi-view consistency, thereby reducing the shell gaps caused by single-view missegmentation.

[0041] The process of generating a 3D symbolic distance field based on the human body shell pixel model is as follows: using the boundary voxels of the human body shell pixel model as the shell surface, the shortest distance from each voxel point in the 3D voxel mesh to the shell surface is calculated, and the distance sign is assigned according to the inside-outside relationship of the voxel points relative to the human body shell pixel model; the distance of voxel points located inside the human body shell pixel model is recorded as a negative value, the distance of voxel points located outside the human body shell pixel model is recorded as a positive value, and the distance of voxel points located on the shell surface is zero. To ensure the continuity of the distance field, a fast distance transformation can be used to calculate the shortest distance, and local smoothing is performed on the distance field to suppress the staircase effect caused by voxelization.

[0042] The three-dimensional symbolic distance field is used to characterize the feasible region of the human body outer shell, which is a three-dimensional spatial region that satisfies the constraints of the human body shape. In the subsequent incremental optimization solution of the factor graph, the shape of the bone segments or the position of key points corresponding to the joint geometry model are constrained within the feasible region of the human body outer shell. This provides additional geometric consistency constraints in the event of occlusion, false key point detection, or cross-view mismatch, reduces the anomalies of the three-dimensional reconstruction penetrating the human body outer shell, and reduces the joint angle sequence jumps caused by this.

[0043] Based on the individualized joint geometry model, observation dataset and three-dimensional symbolic distance field, factor graph incremental optimization solution is constructed to output joint angle sequence and abnormal frame set; In one embodiment, the process of constructing a factor map based on an individualized joint geometry model, observation dataset, and 3D symbolic distance field, and performing incremental optimization, is as follows: Each frame's joint angle along a unified time axis is used as a node to be estimated, and bone segment length, joint center, joint principal axis, and joint angle coordinate system are used as geometric constraint nodes. Reprojection constraints are constructed based on 2D human keypoint observations and camera parameter information, ensuring that the 3D skeleton corresponding to the joint angle is projected onto each viewpoint in accordance with the 2D human keypoint observations. Shell constraints are constructed based on the 3D symbolic distance field, ensuring that the 3D skeleton keypoints or bone segment sampling points fall within the feasible region of the human shell. Temporal smoothing constraints are constructed based on the joint angle differences between adjacent frames to suppress jitter in the joint angle sequence. During incremental optimization, new observation constraints and nodes to be estimated are added frame by frame along the time axis. The newly added parts are linearized, and the Jacobian matrix and information matrix are updated. Incremental least squares iteration is used to obtain the joint angle estimate for the current frame, and the estimated values ​​for each frame are output as a joint angle sequence in chronological order. The generation of the abnormal frame set includes: calculating the reprojection constraint residual and the outer shell constraint residual after each incremental optimization; when the residual exceeds the preset threshold or the number of valid observations of key points is insufficient, the corresponding time axis index is added to the abnormal frame set to indicate inconsistent observation frames and to be used for subsequent validity verification and supplementary acquisition command generation.

[0044] The factor map is constructed as follows: a reprojection factor is constructed from two-dimensional human keypoint observations, camera parameter information, and individualized joint geometry models; a shell constraint factor is constructed from three-dimensional symbolic distance fields and individualized joint geometry models; a bone segment length factor is constructed from bone segment lengths; a joint principal axis prior factor is constructed from joint principal axes; and a temporal smoothing factor is constructed from the adjacent frame differences of the joint angle sequence.

[0045] In one embodiment, when constructing the factor graph, the joint angles of each frame on a unified time axis are used as state variables, and the bone segment length, joint center, joint principal axis, and joint angle coordinate system are written into the factor graph as geometric constraint information. When constructing the reprojection factor based on 2D human keypoint observations, camera parameter information, and individualized joint geometry models, the positions of 3D skeleton keypoints are calculated based on the individualized joint geometry model and the joint angles of the current frame. The positions of the 3D skeleton keypoints are then projected onto the imaging planes of each viewpoint using the camera parameter information to obtain predicted 2D keypoints. The difference between the predicted 2D keypoints and the 2D human keypoint observations is used as the reprojection residual, and the residual is weighted by the keypoint confidence to reduce the impact of 2D human keypoint observations with low confidence on the solution. When constructing the outer shell constraint factor based on the 3D symbolic distance field and the individualized joint geometry model, bone segment sampling points are obtained by sampling at fixed intervals on each bone segment. The symbolic distance of the 3D symbolic distance field at the bone segment sampling points is queried. When the symbolic distance indicates that the bone segment sampling point is located outside the feasible domain of the human outer shell, the corresponding residual increases to suppress the 3D skeleton from penetrating the human outer shell. When constructing the bone segment length factor from the bone segment length, a consistency constraint is applied to the distance between adjacent joint centers to keep the bone segment length stable during multi-frame solution. When constructing the joint principal axis prior factor from the joint principal axis, the relative rotation of adjacent bone segments is restricted to the rotation mode defined by the joint principal axis to reduce joint angle drift caused by two-dimensional observation noise. Furthermore, a temporal smoothing factor is constructed from the difference between adjacent frames of the joint angle sequence to ensure that the joint angle changes continuously between adjacent frames, reducing jitter and jumps.

[0046] The objective function corresponding to the above factor graph can be expressed as: ; in, Let be the cost function to be minimized; The residual is formed by the reprojection factor, bone segment length factor, joint principal axis prior factor and time smoothing factor; The weights are determined by the confidence scores of the key points; The symbolic distance of the three-dimensional symbolic distance field at the sampling point of the bone segment; This is a penalty function for converting the symbolic distance into the shell constraint residual.

[0047] During incremental optimization, new 2D human keypoint observations and the corresponding shell constraints of the human segmentation mask are added to the factor graph frame by frame along the time axis. The newly added factors are linearized and the Jacobian matrix and information matrix are updated. The estimated value of the joint angle of the current frame is obtained iteratively and appended as a joint angle sequence. At the same time, the residuals of each factor are recorded. When the reprojection residual or the shell constraint residual continuously exceeds the preset threshold, the corresponding frame is added to the abnormal frame set to indicate inconsistent observation frames and to be used for subsequent validity verification and supplementary acquisition command generation.

[0048] The set of output anomalous frames includes: calculating the reprojection factor residual and the shell constraint factor residual during incremental optimization; determining the observation inconsistency frames based on the reprojection factor residual and the shell constraint factor residual; and adding the observation inconsistency frames to the anomalous frame set.

[0049] In one embodiment, the process of outputting the set of abnormal frames is performed synchronously with the incremental optimization solution of the factor graph. After the incremental optimization of each frame on the unified time axis is completed, the reprojection factor residual and the outer shell constraint factor residual are calculated respectively. The calculation of the reprojection factor residual includes: obtaining the positions of the 3D skeleton key points based on the joint angle estimation and individualized joint geometry model of the current frame, and projecting the predicted 2D key points onto the imaging plane of each viewpoint using camera parameter information; taking the difference between the predicted 2D key points and the observed 2D human key points as the key point level residual, and weighting and summing the key point level residuals by the key point confidence to obtain the reprojection factor residual of the current frame. The calculation of the outer shell constraint factor residual includes: generating bone segment sampling points on each bone segment of the 3D skeleton in the current frame, and querying the symbolic distance of the 3D symbolic distance field at the bone segment sampling points; when the bone segment sampling point is located outside the feasible region of the human outer shell, mapping the corresponding symbolic distance to the constraint penalty and summing them to obtain the outer shell constraint factor residual of the current frame.

[0050] The determination of inconsistent frames based on reprojection factor residuals and shell constraint factor residuals includes: comparing the reprojection factor residuals and shell constraint factor residuals with preset thresholds respectively, and combining this with the number of valid 2D human keypoint observations in the current frame for consistency judgment; when the reprojection factor residual exceeds the preset threshold and the number of valid 2D human keypoint observations is lower than a preset lower limit, the current frame is determined to be an inconsistent frame; when the shell constraint factor residual exceeds the preset threshold and the proportion of bone segment sampling points exceeding the threshold exceeds a preset proportion, the current frame is determined to be an inconsistent frame; when both the reprojection factor residual and the shell constraint factor residual exceed their respective thresholds, the current frame is directly determined to be an inconsistent frame. To avoid misjudgment caused by single-frame jitter, a short-window consistency check is performed on the residual sequences of adjacent frames. When multiple consecutive frames meet the inconsistent judgment conditions, the entire interval of those consecutive frames is determined to be inconsistent frames.

[0051] The process of adding inconsistent observation frames to the anomalous frame set includes: recording the unified time axis index corresponding to the inconsistent observation frame and writing the index into the anomalous frame set; the anomalous frame set is used to mark low-confidence observation segments in the subsequent joint range determination and counterfactual consistency test, thereby reducing the impact of joint angle sequence jumps caused by occlusion, false detection of key points and segmentation errors on the joint range assessment results.

[0052] The joint range of motion is determined based on the joint angle sequence and the joint range of motion assessment result is obtained through counterfactual consistency test. If the validity condition is not met, an information matrix is ​​calculated based on the Jacobian matrix to generate a supplementary acquisition instruction. The supplementary acquisition instruction is used to re-acquire multi-view video sequences to update the joint range of motion assessment result.

[0053] In one embodiment, determining the joint range of motion based on the joint angle sequence includes: removing frames corresponding to the abnormal frame set from the joint angle sequence and performing temporal smoothing on the remaining frames; statistically analyzing the maximum and minimum values ​​of the joint angle sequence using a sliding window on a unified time axis, selecting the maximum and minimum values ​​within a stable window as the upper and lower boundaries of the joint range of motion, and using the window duration and the proportion of abnormal frames as part of the validity condition. Obtaining the joint range of motion assessment result through counterfactual consistency testing includes: applying perturbations to the upper and lower boundaries of the joint range of motion to generate a counterfactual joint angle sequence, and substituting the counterfactual joint angle sequence into the factor graph cost function to calculate the cost difference; when the cost difference meets the preset judgment rule, the joint range of motion assessment result is output as the joint range of motion; otherwise, the validity condition is not met. When the validity condition is not met, a supplementary acquisition instruction is generated based on the Jacobian matrix to calculate the information matrix, including: calculating the information matrix from the Jacobian matrix obtained by incremental optimization. ,in, For information matrix, For Jacobian matrices, The weighted matrix is ​​determined by the confidence of key points; based on the diagonal elements of the information matrix, the time periods and joint degrees of freedom with high uncertainty are determined, and supplementary acquisition instructions are generated to instruct the re-acquisition of multi-view video sequences of the time periods and prompt the subject to complete the corresponding joint movements, and then the joint mobility assessment results are updated.

[0054] Determining the joint range of motion and performing counterfactual consistency checks includes: mapping the joint angle sequence to a joint orientation angle sequence using a joint angle coordinate system; determining candidate platform segments based on the joint orientation angle sequence and the set of abnormal frames, where the difference between adjacent frames of the joint orientation angle sequence continuously satisfies a preset stability condition; determining the boundary values ​​of the joint orientation angle sequence within the candidate platform segments to form the joint range of motion; applying perturbation to the boundary values ​​of the joint range of motion to generate a counterfactual joint angle sequence, and calculating the cost difference based on the cost function of the factor graph to determine the validity of the boundary values; and generating supplementary acquisition instructions based on the Jacobian matrix to calculate the information matrix includes: obtaining the Jacobian matrix during the incremental optimization solution of the factor graph; calculating the information matrix based on the Jacobian matrix; and determining supplementary acquisition instructions based on the information matrix.

[0055] In one embodiment, determining the joint range of motion and performing a counterfactual consistency check includes: mapping the joint angle sequence to a joint direction angle sequence using a joint angle coordinate system. The joint direction angle sequence is used to unify the joint angle representations of different subjects and different acquisition batches to the same angular reference direction. Specifically, the direction of each frame of joint angles is defined according to the joint angle coordinate system to ensure that the positive and negative directions of joint flexion, extension, abduction, adduction, or rotation are consistent, and a mark is retained on the inconsistent frames indicated by the abnormal frame set for subsequent removal or weighting. Determining candidate plateau segments from the joint direction angle sequence and the abnormal frame set includes: calculating the difference between adjacent frames of the joint direction angle sequence on a unified time axis and smoothing the difference between adjacent frames using a short window; when the difference between adjacent frames meets a preset stability condition within consecutive frames and the proportion of abnormal frame set marks within that consecutive frame is lower than a preset proportion, the corresponding consecutive frame interval is determined as a candidate plateau segment. The candidate plateau segment is used to reflect the stable state of the subject during a certain movement phase, thereby reducing extreme value misjudgments caused by rapid swinging and occlusion.

[0056] Determining the boundary values ​​of the joint orientation angle sequence within candidate platform segments to form joint range of motion involves: for each candidate platform segment, extracting the upper and lower quantile values ​​of the joint orientation angle sequence as the boundary values ​​for that segment, and performing consistency screening on the boundary values ​​of multiple candidate platform segments; when the difference between the boundary values ​​of multiple candidate platform segments exceeds a preset threshold, prioritizing the candidate platform segment boundary values ​​with a lower proportion of outlier frame set markers and smaller reprojection factor residuals. Through the above processing, the joint range of motion is obtained, ensuring that the upper and lower boundaries of the joint range of motion simultaneously meet the requirements of stability and observation consistency.

[0057] The boundary values ​​of the joint range of motion are perturbed to generate counterfactual joint angle sequences, and the cost difference is calculated based on the cost function of the factor graph. The validity of the boundary values ​​is determined based on the cost difference. This process includes: applying positive and negative perturbations to the upper and lower boundaries of the joint range of motion to generate multiple sets of counterfactual joint angle sequences; substituting each set of counterfactual joint angle sequences and the original joint angle sequences into the cost function of the factor graph to calculate the corresponding cost, and calculating the cost difference between the counterfactual cost and the original cost; when a small perturbation of the boundary value leads to a significant increase in the cost difference, the boundary value is determined to be identifiable and retained as a valid boundary value; when the cost difference does not change significantly after the boundary value is perturbed, the boundary value is determined to be insufficiently identifiable and the validity condition is not met, thus entering the supplementary acquisition instruction generation process.

[0058] The supplementary data collection instructions based on the Jacobian matrix for calculating the information matrix include: obtaining the Jacobian matrix during the incremental optimization process of the factor graph, and calculating the information matrix based on the Jacobian matrix. The supplementary acquisition instructions, based on the information matrix, include: analyzing the diagonal and coupling elements related to joint angles in the information matrix to determine the joint angle degrees of freedom and time periods with high uncertainty; when uncertainty is concentrated in a specific time period, generating supplementary acquisition instructions to instruct the re-acquisition of multi-view video sequences for that time period; when uncertainty is concentrated in a specific joint angle degree of freedom, generating supplementary acquisition instructions to instruct the subject to perform calibration actions that enhance the observability of that joint angle, and updating the observation dataset and factor graph solution results after re-acquisition to update the joint mobility assessment results. By combining counterfactual consistency testing with information matrix-driven supplementary acquisition, the source of the problem can be located and supplementary acquisition can be guided when the initial acquisition observation is insufficient, thereby reducing the instability of joint mobility interval boundaries caused by occlusion, viewpoint degradation, and false detection.

[0059] In one specific embodiment, a rehabilitation clinic assessed the flexion-extension range of motion of the right elbow joint in an adult subject. The acquisition system used four simultaneous cameras with a resolution of 1920×1080 and a frame rate of 30fps, continuously acquiring data for 12 seconds to obtain a multi-view video sequence of 360 frames. The cameras were positioned around the subject and pointed towards the upper limb's range of motion. Camera calibration was performed on-site to generate camera parameter information, and the calibration results were written into the observation dataset. (See attached image) Figure 5 As shown, a top-view coordinate system with horizontal axes X / meter and Z / meter is established, with the subject located near the origin. Cameras 1, 2, 3, and 4 are arranged around the subject at different circumferential positions, forming an approximately circular distribution. The dashed arc represents the radius of the arrangement. Each camera faces the central area of ​​the subject, ensuring that upper limb movements can be captured from multiple angles. This four-view arrangement geometrically provides more sufficient parallax and redundancy observations, ensuring that weighted triangulation maintains a good condition number for most of the time, thereby reducing the probability of 3D drift caused by single-view occlusion. (See attached diagram.) Figure 3 The fact that the residuals remain at a low level for most of the time period and only show spikes at a few moments indicates that this arrangement can compress systematic degradation into "local short-term anomalies", providing a stable engineering basis for subsequent anomaly frame location and re-sampling triggering.

[0060] When generating the observation dataset, the four video streams are time-aligned based on hardware timestamps to form a unified timeline. For each moment, the human body region is determined in the images from each viewpoint, and a human body segmentation mask is generated. Two-dimensional human keypoint detection is performed within the human body region, outputting the observations of shoulder, elbow, and wrist keypoints, along with their confidence scores. Taking camera 0 as an example, the first two frames of keypoint observation segments are shown below, with two-dimensional coordinates in pixels and keypoint confidence scores ranging from 0 to 1: {"frame":0,"keypoints_px":[[1791.31,1267.70],[1791.08,973.98],[1790.17,1134.52]],"conf":[0.662,0.673,0.749]} {"frame":1,"keypoints_px":[[1785.75,1278.54],[1796.20,983.34],[1788.73,1131.06]],"conf":[0.791,0.599,0.947]} Subsequent weighted triangulation will use the confidence level of key points as the weights, and low-confidence observations caused by occlusion will be naturally downweighted, reducing the spread of false detections from the source.

[0061] In the individualized joint geometry model construction stage, joint geometry calibration action segments are automatically extracted from 2D human keypoint observations. In this embodiment, frames 40 to 120 are located, covering continuous movements from slight flexion to near extension. Multi-view correspondences are established for keypoints at various perspectives within the calibration action segments, forming a keypoint observation sequence. Weighted triangulation is then performed to obtain the 3D keypoint trajectories of the shoulder, elbow, and wrist. Based on the trajectory estimation, bone segment lengths are optimized by applying bone segment length consistency constraints: the upper arm bone segment length converges to 0.312m, and the forearm bone segment length converges to 0.268m, with fluctuations across the entire sequence controlled within ±3mm. The joint center is determined by the robust mean of the elbow keypoint trajectory, the joint principal axis is estimated by the principal direction of relative motion between adjacent bone segments, and the joint angular coordinate system is constructed from the joint principal axis, joint center, and bone segment direction vectors, resulting in the individualized joint geometry model.

[0062] In the 3D symbolic distance field generation stage, voxel sculpting is performed based on human body segmentation masks and camera parameter information from various viewpoints to obtain a human body shell pixel model. The signed distance from the voxel model to the surface of the shell is calculated to generate a 3D symbolic distance field, which is used to characterize the feasible region of the human body shell. This constraint can limit bone segments from "penetrating" the human body shell when there are clothing folds or short-term occlusion at the elbow, thereby suppressing non-physical jumps in 3D reconstruction.

[0063] The incremental optimization solution stage of the factor graph is shown in the attached figure. Figure 3As shown, the horizontal axis represents time per second. The left vertical axis represents the reprojection residual per pixel, represented by a solid curve; the right vertical axis represents the shell constraint residual per millimeter, represented by a dashed curve; and dots mark the locations of anomalous frames. As can be seen in the figure, both types of residuals fluctuate at low levels within most frames, but synchronous spikes appear at certain moments: the reprojection residual can surge to approximately 6-7 pixels, and the shell constraint residual also rises synchronously to approximately 20 millimeters. These spikes correspond to the anomalous frame markers. The "synchronicity" of the residual spikes indicates that the anomalies are not simply sporadic noise from a keypoint at a certain viewpoint, but rather more consistent with the situation where "two-dimensional observation and the geometrically feasible region are simultaneously disturbed," such as occlusion causing keypoint drift, segmentation contour jitter, or short-term tracking failure. Including the corresponding frames in the anomalous frame set based on the residual spikes can promptly block error propagation during incremental optimization, preventing a small number of inconsistent observations from simultaneously increasing the reprojection cost and shell cost, thereby maintaining the continuity and physical rationality of the joint angle sequence. The reprojection factor, shell constraint factor, bone segment length factor, joint principal axis prior factor, and time smoothing factor were used to construct a unified map, and the joint angle sequence and aberration frame set were updated incrementally on a frame-by-frame basis. Residual statistics showed that the median of the reprojection residuals was 6.4px, and the 95th percentile was 18.2px. The shell constraint residuals remained stable in most frames, with spikes appearing in a few frames. Based on the joint determination of observational inconsistencies using the reprojection factor residuals and the shell constraint factor residuals, 14 aberration frames were identified, accounting for 3.9%, mainly concentrated during the rapid flexion-extension transition period of the subjects.

[0064] During the joint range of motion assessment phase, the elbow flexion-extension angles are calculated from three-dimensional key points to form a joint angle sequence. The calculation formula is as follows: ; in, This refers to the elbow flexion and extension angle. The three-dimensional position of the shoulder key point. The three-dimensional position of the elbow key point. The three-dimensional positions of the wrist key points are shown. After removing inconsistent frames from the abnormal frame set, candidate plateau segments are extracted from the joint angle sequence: the extension plateau segment is located at 2.10–3.00s, and the flexion plateau segment is located at 7.20–8.00s. (See attached image.) Figure 2As shown, the horizontal axis represents time / second, and the vertical axis represents joint angle / degree. The solid blue line represents the joint angle sequence; the light blue shaded area represents the candidate plateau segment; the red dashed line represents the lower boundary of activity, and the green dotted line represents the upper boundary of activity; the orange "×" is used to mark the location of abnormal frames. It can be seen that the joint angle fluctuates slightly after entering the plateau segment. The two boundary lines of the plateau segment are located around 58° and 63° respectively. The abnormal frame marking points mainly appear at non-stationary positions such as peaks and sudden drops in the angle curve. From the amplitude of angle fluctuations and the boundary spacing within the plateau segment, it can be seen that the boundary bandwidth of this plateau segment is about 5°, which is within the range of activity boundaries that can be stably extracted. Abnormal frames are not concentrated in the stable region of the plateau segment, but are more distributed at peaks and sudden drops, indicating that the abnormal frame judgment is consistent with "non-physical angle jumps caused by observation inconsistencies." This supports the elimination / reduction of abnormal frames when estimating the plateau segment boundary, avoiding the boundary being skewed by a few peaks, and improving the robustness and reproducibility of the activity boundary. Within the platform segment, robust boundary values ​​were used to obtain the joint range of motion: minimum angle 18.6°, maximum angle 141.8°, and total range of motion 141.8° - 18.6° = 123.2°. A counterfactual consistency test was performed on the boundary values: a ±2° perturbation was applied to the boundary to generate a counterfactual joint angle sequence, and the cost difference was calculated using a factor graph cost function. Within the aforementioned platform segment, the cost difference was significantly positive, indicating that the boundary values ​​are supported by both observations and constraints, and the boundary validity is passed.

[0065] When validating the conditions, as shown in the attached document. Figure 4 As shown, the horizontal axis represents time / second, and the vertical axis represents information intensity / normalization. The solid curve represents the information intensity of the plateau segment, the dashed line represents the trigger threshold, and the dots represent trigger frames. In the figure, the information intensity fluctuates around 1.0 for most moments, but drops significantly in a few moments, reaching a low of approximately 0.55 and below the threshold (approximately 0.75). These moments are marked as trigger frames. Information intensity falling below the threshold means that within that time window, the constraint of observation on joint state "weakens," and the corresponding boundary estimation is more prone to instability or sensitivity to anomalous observations. The trigger frames are concentrated at the low points of information intensity, indicating that using information intensity as the basis for supplementary acquisition triggers can explicitly quantify "insufficient observability" and limit supplementary acquisition actions to necessary time periods. It was found that camera 2 maintained a consistently low confidence level for wrist key points between 6.0 and 7.5 seconds, resulting in insufficient information intensity. An information matrix was calculated based on the Jacobian matrix to generate supplementary acquisition instructions. The information matrix uses: ; in, For information matrix, For Jacobian matrices, This is the weight matrix obtained by mapping keypoint confidence to observation noise. It is calculated in units of the window corresponding to the platform segment. The minimum eigenvalue was used as an information intensity indicator. The original minimum eigenvalue was approximately 0.74, triggering a supplementary acquisition instruction. The instruction was to "move camera 2 to the right rear side of the subject and raise the camera position to ensure that the wrist is not obstructed by the torso, and extend the extension holding time to 1 second." After re-acquiring and recalculating according to the instruction, the number of abnormal frames decreased to 6, accounting for 1.7%, the minimum eigenvalue increased to 1.21, and the activity level results stabilized in the range of 122.6°–123.4°. The final activity level assessment results can be directly used in the clinical records.

[0066] The following is a portion of the core code: import numpyasnp defweighted_triangulation(P_list,uv_list,w_list): A=[] forP,(u,v),winzip(P_list,uv_list,w_list): ifw<=0: continue A.append(w*(u*P[2,:]-P[0,:])) A.append(w*(v*P[2,:]-P[1,:])) A = np.asarray(A) if A.shape[0]<4: returnNone _,_,Vt=np.linalg.svd(A) X = Vt[-1] X = X / (X[3] + 1e-12) returnX[:3] defelbow_angle(p_sh,p_el,p_wr): v1=p_sh-p_el v2=p_wr-p_el cosang=np.dot(v1,v2) / ((np.linalg.norm(v1)+1e-12)*(np.linalg.norm(v2)+1e-12)) cosang=np.clip(cosang,-1.0,1.0) returnnp.degrees(np.arccos(cosang)) definformation_matrix(J,W): returnJ.T@W@J like Figure 6 As shown, an adult joint range of motion assessment system based on multi-view video is used to implement the adult joint range of motion assessment method based on multi-view video. The system includes: The observation generation module is used to acquire multi-view video sequences of the subject and generate an observation dataset. The observation dataset includes 2D human keypoint observations, human segmentation masks, and camera parameter information. The observation generation module includes a multi-path imaging unit, a synchronization triggering unit, and an edge preprocessing unit. The multi-path imaging unit can consist of 3 to 6 cameras (global shutter priority), each fixed on a bracket or tripod to form a surround view coverage. The synchronization triggering unit provides hardware trigger / timestamp references (such as trigger pulses, PTP clocks, or synchronization lines) to each camera to ensure the temporal consistency of the multi-view video sequences. The edge preprocessing unit can be an in-camera ISP / edge computing box (including CPU / GPU / NPU) used to accelerate image distortion correction, resolution and frame rate shaping, keypoint detection, and human segmentation inference, and output 2D human keypoint observations and human segmentation masks. Camera parameter information is generated and stored at the computing end after acquisition by a calibration board / calibration lightbox in conjunction with the camera, or it can be directly read from a camera with factory calibration parameters.

[0067] The geometry construction module is used to build individualized joint geometric models based on the observation dataset. These models include bone segment lengths, joint centers, joint principal axes, and joint angular coordinate systems. A 3D symbolic distance field is generated from camera parameter information and human segmentation masks. The module includes calibration and geometric computation hardware resources, as well as storage and parallel computing units required for 3D voxel / distance field construction. Calibration and geometric computation can run on the CPU and GPU of an industrial PC / workstation: the CPU is responsible for solving camera intrinsic and extrinsic parameters, data organization, and robust statistics; the GPU is responsible for parallel voxel sculpting, distance transformation, and generating a 3D symbolic distance field from the multi-view human segmentation masks. This module typically requires significant video memory / RAM to cache multi-view mask sequences, voxel meshes, and intermediate geometric quantities, and interfaces with the observation generation module via a high-speed bus (such as PCIe).

[0068] The inversion solution module is used to construct a factor graph based on the individualized joint geometry model, observation dataset, and 3D symbolic distance field, and perform incremental optimization to output the joint angle sequence and a set of anomalous frames. The inversion solution module includes a high-performance computing unit and a high-speed cache subsystem for constructing the factor graph and performing incremental optimization. In terms of hardware implementation, a multi-core CPU is typically used for sparse linearization, incremental updates, and scheduling (such as incremental factor addition and variable elimination order maintenance), and a GPU can be optionally used to accelerate some matrix operations (such as parallel residual calculation and Jacobi block assembly). To ensure real-time performance and stability, this module is generally configured with a large amount of memory to store the sparse structure, joint angle state sequence, anomalous frame index, and residual diagnostic data, and records traceable logs via NVMe solid-state drives (for easy review and calibration playback).

[0069] The evaluation and update module is used to determine the joint range of motion based on the joint angle sequence and obtain the joint range of motion evaluation result through counterfactual consistency verification. When the validity condition is not met, it calculates an information matrix based on the Jacobian matrix to generate supplementary acquisition instructions, and re-acquires multi-view video sequences according to the supplementary acquisition instructions to update the joint range of motion evaluation results. The evaluation and update module includes an evaluation calculation unit, a human-computer interaction and supplementary acquisition control interface. The evaluation calculation unit can share the same computing platform with the inversion solution module. It is responsible for identifying platform segments, calculating range of motion boundaries and performing counterfactual consistency verification on the joint angle sequence. When the validity condition is not met, it calculates an information matrix based on the Jacobian matrix to form supplementary acquisition instructions. The human-computer interaction part can be a monitor / touchscreen / keyboard and mouse, used to present the range of motion results, abnormal frame prompts and acquisition suggestions. The supplementary acquisition control interface can be linked with the camera trigger or acquisition box via USB / Ethernet / serial port to send new acquisition duration, frame rate, viewpoint activation strategy or recalibration prompts to the synchronization trigger unit, thereby driving re-acquisition and closed-loop updating of the evaluation results.

[0070] In one possible implementation, the above modules can be integrated into the same workstation / industrial computer and implemented through an external multi-camera array; in another possible implementation, the observation generation module is distributed and deployed in the form of edge boxes, and the geometry construction, inversion solution and evaluation update are concentrated on the back-end server for execution; the corresponding software functions run on the above hardware as firmware, drivers and inference / optimization programs, thus forming different forms of implementation that are entirely hardware, a combination of hardware and software, or software-oriented.

[0071] The above are merely embodiments of the present invention and are not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of the present invention should be included within the scope of the claims of the present invention.

Claims

1. A method for assessing the range of motion of adult joints based on multi-view video, characterized in that, Includes the following steps: A multi-view video sequence of the test subject is obtained to generate an observation dataset, which includes two-dimensional human key point observation, human segmentation mask and camera parameter information; An individualized joint geometry model is constructed based on the observation dataset. The individualized joint geometry model includes bone segment length, joint center, joint principal axis and joint angle coordinate system, and a three-dimensional symbolic distance field is generated by camera parameter information and human body segmentation mask. Based on the individualized joint geometry model, observation dataset and three-dimensional symbolic distance field, factor graph incremental optimization solution is constructed to output joint angle sequence and abnormal frame set; The joint range of motion is determined based on the joint angle sequence and the joint range of motion assessment result is obtained through counterfactual consistency test. If the validity condition is not met, an information matrix is ​​calculated based on the Jacobian matrix to generate a supplementary acquisition instruction. The supplementary acquisition instruction is used to re-acquire multi-view video sequences to update the joint range of motion assessment result.

2. The method according to claim 1, characterized in that, The process of generating the observation dataset includes: temporally aligning multi-view video sequences to form a unified timeline; performing camera calibration on the multi-view video sequences to generate camera parameter information; determining the human body region of the subject in the multi-view video sequences based on the unified timeline; performing two-dimensional human keypoint detection within the human body region to generate two-dimensional human keypoint observations, and generating keypoint confidence scores for the two-dimensional human keypoint observations; and performing human body segmentation within the human body region to generate a human body segmentation mask corresponding to the unified timeline.

3. The method according to claim 1, characterized in that, The joint geometry calibration action segment is determined by two-dimensional human key point observation, and a multi-view correspondence relationship is established for the two-dimensional human key point observation within the joint geometry calibration action segment to generate a key point observation sequence for joint geometry calibration.

4. The method according to claim 2, characterized in that, The two-dimensional human keypoints observed from different perspectives using camera parameter information are weighted triangulated to generate three-dimensional keypoint trajectories. The weights of the triangulation are determined by the keypoint confidence. Based on the three-dimensional keypoint trajectories, bone segment length consistency constraint optimization is performed to ensure consistency between bone segment length and the three-dimensional keypoint trajectories.

5. The method according to claim 4, characterized in that, The length of a bone segment is determined by the trajectory of three-dimensional key points; the joint center is determined by the trajectory of three-dimensional key points; the principal axis of the joint is determined by the relative motion of adjacent bone segments; and a joint angle coordinate system is constructed by the principal axis of the joint, the joint center, and the bone segment direction vectors of adjacent bone segments.

6. The method according to claim 1, characterized in that, The generation of a 3D symbolic distance field based on camera parameter information and human body segmentation mask includes: performing voxel sculpting on the human body segmentation mask under the constraints of camera parameter information to generate a human body shell pixel model; and generating a 3D symbolic distance field based on the human body shell pixel model, which is used to characterize the feasible region of the human body shell.

7. The method according to claim 1, characterized in that, The factor map is constructed as follows: a reprojection factor is constructed from two-dimensional human keypoint observations, camera parameter information, and individualized joint geometry models; a shell constraint factor is constructed from three-dimensional symbolic distance fields and individualized joint geometry models; a bone segment length factor is constructed from bone segment lengths; a joint principal axis prior factor is constructed from joint principal axes; and a temporal smoothing factor is constructed from the adjacent frame differences of the joint angle sequence.

8. The method according to claim 1, characterized in that, The set of output anomalous frames includes: calculating the reprojection factor residual and the shell constraint factor residual during incremental optimization; determining the observation inconsistency frames based on the reprojection factor residual and the shell constraint factor residual; and adding the observation inconsistency frames to the anomalous frame set.

9. The method according to claim 1, characterized in that, Determining the joint range of motion and performing counterfactual consistency checks includes: mapping the joint angle sequence to a joint orientation angle sequence using a joint angle coordinate system; determining candidate platform segments based on the joint orientation angle sequence and the set of abnormal frames, where the difference between adjacent frames of the joint orientation angle sequence continuously satisfies a preset stability condition; determining the boundary values ​​of the joint orientation angle sequence within the candidate platform segments to form the joint range of motion; applying perturbation to the boundary values ​​of the joint range of motion to generate a counterfactual joint angle sequence, and calculating the cost difference based on the cost function of the factor graph to determine the validity of the boundary values; and generating supplementary acquisition instructions based on the Jacobian matrix to calculate the information matrix includes: obtaining the Jacobian matrix during the incremental optimization solution of the factor graph; calculating the information matrix based on the Jacobian matrix; and determining supplementary acquisition instructions based on the information matrix.

10. A multi-view video-based adult joint range of motion assessment system, used to implement the multi-view video-based adult joint range of motion assessment method according to any one of claims 1 to 9, characterized in that, The system includes: The observation generation module is used to acquire multi-view video sequences of the subject and generate an observation dataset, which includes two-dimensional human key point observations, human segmentation masks, and camera parameter information. The geometry construction module is used to build individualized joint geometry models based on the observation dataset. The individualized joint geometry model includes bone segment length, joint center, joint principal axis and joint angle coordinate system, and generates a three-dimensional symbolic distance field from camera parameter information and human segmentation mask. The inversion solution module is used to construct a factor map based on the individualized joint geometry model, observation dataset and three-dimensional symbolic distance field and perform incremental optimization solution, outputting joint angle sequence and abnormal frame set; The evaluation and update module is used to determine the joint range of motion based on the joint angle sequence and obtain the joint range of motion evaluation result through counterfactual consistency test. When the validity condition is not met, the module calculates the information matrix based on the Jacobian matrix to generate supplementary acquisition instructions, and re-acquires multi-view video sequences according to the supplementary acquisition instructions to update the joint range of motion evaluation result.

Citation Information

Cited By

  • Beef cattle body condition scoring system based on computer vision and attitude estimation

    CN122116425A