Sppb fall risk intelligent assessment system and method based on multi-modal perception
The SPPB fall risk intelligent assessment system, which utilizes multimodal perception and image acquisition and deep learning technologies, solves the problems of large human error and poor standardization in existing SPPB tests. It achieves high-precision and automated physical function assessment, generates structured reports, and meets clinical needs.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- PEKING UNIVERSITY FIRST HOSPITAL (PEKING UNIVERSITY FIRST CLINICAL MEDICAL COLLEGE)
- Filing Date
- 2026-04-10
- Publication Date
- 2026-06-12
AI Technical Summary
The existing Simplified Physical Function Test (SPPB) assessment relies on manual operation, which suffers from large human error, cumbersome operation, poor standardization, low recognition accuracy, lack of safety monitoring and insufficient data intelligence, especially in hospital environments where it is difficult to meet the needs of standardized clinical management.
The SPPB fall risk intelligent assessment system, based on multimodal perception, includes an image acquisition module, a sit-stand detection module, a balance detection module, and a gait detection module. Through optical motion capture and deep learning technology, it identifies human posture and gait, and automatically calculates the SPPB score in conjunction with the control module. It can achieve high-precision sit-stand, balance, and gait testing without the need for wearing equipment.
It achieves high-precision, automated sitting-standing, balance, and gait testing, reduces human error, improves the convenience and standardization of testing, comprehensively reflects physical function, generates structured assessment reports, and meets clinical needs.
Smart Images

Figure CN122200747A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent recognition technology, and in particular to an intelligent assessment system and method for SPPB fall risk based on multimodal perception. Background Technology
[0002] The Short Physical Performance Battery (SPPB) assessment relies mainly on manual operation, which has problems such as large human error, cumbersome operation, poor standardization, low recognition accuracy, lack of safety monitoring and insufficient data intelligence. Especially in the hospital environment, manual testing is inefficient and data is difficult to trace, which cannot meet the needs of standardized clinical management. Summary of the Invention
[0003] The purpose of this invention is to provide an intelligent assessment system and method for SPPB fall risk based on multimodal perception, in order to solve one of the technical problems of poor objectivity, insufficient intelligence, and inaccurate posture recognition in SPPB test assessment.
[0004] To achieve the above objectives, the present invention provides the following technical solution: In a first aspect, embodiments of the present invention provide an SPPB fall risk intelligent assessment system based on multimodal perception, including an image acquisition module, a sit-stand detection module, a balance detection module, a gait detection module, and a control module; the image acquisition module is used to acquire images of human posture; The sitting / standing detection module is used to detect the motion transition nodes between sitting and standing based on the first key point captured by optical motion capture. The balance detection module is used to identify the standing posture of a human body based on the relative positional relationship between the marked area and the left and right feet. The standing posture includes one of the following: feet together, half-seam, or full-seam. The gait detection module is used to identify the start and end times of walking based on a second key point or non-contact distance changes. The control module is used to calculate the SPPB score based on the data collected by the sitting / standing detection module, the balance detection module, and the gait detection module, as well as the SPPB standard rules.
[0005] According to at least one embodiment of the present invention, the sitting / standing detection module is further configured to obtain the confidence level of the first key point of each frame image based on a human pose estimation algorithm. When the confidence level of the first key point is greater than or equal to the first threshold, sit-to-stand detection is performed; when the confidence level of the first key point is less than the first threshold, the current frame is discarded and sit-to-stand detection is not performed.
[0006] According to at least one embodiment of the present invention, the first key point includes at least one of the left shoulder key point, the right shoulder key point, the left hip key point, and the right hip key point.
[0007] According to at least one embodiment of the present invention, the first key point further includes a nasal key point; The sitting / standing detection module is also used to acquire the nose key points of each frame image. When the horizontal coordinate of the nose key points is within the first coordinate range of the current frame, sitting / standing detection is performed.
[0008] According to at least one embodiment of the present invention, the sitting / standing detection module is further configured to determine a standing up when the vertical coordinates of the left hip key point and the right hip key point are simultaneously higher than an upper threshold in each frame image; and to determine a sitting down when the vertical coordinates of the left hip key point and the right hip key point are simultaneously lower than a lower threshold; wherein completing one standing up and one sitting down constitutes one valid sitting / standing action.
[0009] According to at least one embodiment of the present invention, the marked area includes four color blocks arranged in an array on the floor mat; The balance detection module is used to convert the four color blocks into computable spatial location points to form four marker points in the current frame for evaluating standing posture.
[0010] According to at least one embodiment of the present invention, the balance detection module is used to convert the four color blocks from RGB color space to HSV color space, and set corresponding HSV threshold ranges for preset target colors respectively, and use the inRange method to perform threshold segmentation to obtain a binarized image of the target color; Morphological denoising and contour detection are performed on the binarized image. The geometric moments of each connected region are calculated. The centroid coordinates of color blocks with an area greater than the area threshold are calculated to obtain the corresponding marker points of the four color blocks.
[0011] According to at least one embodiment of the present invention, the balance detection module is used to determine a standing posture with feet together when two markers are detected in the current frame and the two feet respectively cover the two markers located behind the four markers. When only one marker is detected, and one foot covers two markers on the same side, while the other foot covers the marker located behind the other two markers on the opposite side, it is determined to be a semi-tandem standing posture. When two markers are detected, and one foot covers one marker on one side while the other foot covers the other marker on the same side, the stance is determined to be a fully tandem standing posture.
[0012] According to at least one embodiment of the present invention, of the four marker points, the two marker points located in front are close to the image acquisition module, and the two marker points located behind are far away from the image acquisition module.
[0013] According to at least one embodiment of the present invention, the second key point includes a key point on the left ankle and a key point on the right ankle. The gait detection module is used to extract the coordinates of the second key point in each frame of the image in real time, and output them in a normalized manner. Based on the resolution of the screen, the coordinates of the second key point are mapped to the screen coordinate system. When the left ankle key points and the right ankle key points are detected to cross the starting virtual line to the rear at the same time, and the gait starts when either foot crosses the starting virtual line to the front again; The moment when both the left and right ankle key points cross the virtual endpoint line is determined as the gait termination moment.
[0014] Secondly, embodiments of the present invention provide an intelligent assessment method for SPPB fall risk based on multimodal perception, using the assessment system described in the first aspect for assessment.
[0015] In one or more technical solutions provided in the exemplary embodiments of the present invention, at least one of the following beneficial effects can be achieved.
[0016] In the evaluation system provided by the exemplary embodiment of this invention, the image acquisition module, the sit-stand detection module, the balance detection module, and the gait detection module are all communicatively connected to the control module. Based on optical motion capture, it accurately tracks key points such as the hip, achieving high-precision sit-stand recognition without the need for wearing devices, and can detect individual positional changes, improving the convenience and accuracy of sit-stand testing. Since the shoulders and hips are always visible during the sit-stand process; compared to the lower and upper limbs, they are least likely to be obscured by clothing, handrails, and arms; relative positional changes can stably reflect overall body height changes.
[0017] Furthermore, by using smooth detection based on duration thresholds and occlusion tolerance, the effects of lighting, occlusion, pose shift, and frame jitter on monocular cameras can be reduced, thereby improving the reliability of sitting and standing motion recognition.
[0018] Furthermore, the balance detection module, image acquisition module, and control module work together to allow subjects to automatically complete balance tests without wearing any external devices. This can comprehensively reflect the stability of the whole body posture, is not affected by the wearing position or signal drift, and can automatically identify violations such as foot movement, thereby achieving standardized testing and automated assessment of balance ability.
[0019] Furthermore, the gait detection module, image acquisition module, and control module work together to achieve automatic recognition of the start and end times of walking through multi-state joint determination. This avoids the risk of misjudgment caused by single-frame triggering from the algorithm level and significantly improves the stability and consistency of timing triggering. Attached Figure Description
[0020] The accompanying drawings illustrate exemplary embodiments of the invention and, together with the description thereof, serve to explain the principles of the invention. These drawings are included to provide a further understanding of the invention and are incorporated in and constitute a part of this specification.
[0021] Figure 1 This is a flowchart of an evaluation method according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the evaluation system structure according to an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of a floor mat according to an embodiment of the present invention. Detailed Implementation
[0022] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments.
[0023] Example 1 Figure 2 This is a schematic diagram of the evaluation system structure according to an embodiment of the present invention. (Refer to...) Figure 2 As shown in the exemplary embodiment of the present invention, the SPPB fall risk intelligent assessment system based on multimodal perception includes an image acquisition module, a sit-stand detection module, a balance detection module, a gait detection module, and a control module; the image acquisition module is used to acquire images of human posture. The image acquisition module includes a camera, and the above modules can be integrated into... Figure 2 The terminal shown can integrate medical staff operation control and subject test guidance on the same interface.
[0024] The evaluation system provided by the exemplary embodiment of the present invention achieves full automation and closed-loop process of test evaluation by automatically switching test processes through an embedded program in the terminal.
[0025] Figure 1 This is a flowchart of an evaluation method according to an embodiment of the present invention. Figure 1As shown, the process begins with subject identity verification, followed by a standardized test guidance phase. The system ensures that subjects accurately understand and execute each action through voice prompts, screen animations, and real-time posture feedback. Identity information and biometrics are automatically linked, and data is encrypted throughout the process before being uploaded to the hospital system. After completing the guidance, the terminal sequentially performs sit-stand tests, balance tests, and gait tests according to the functions of each module. After the tests, the control module integrates and analyzes the data obtained from each module according to preset SPPB standard rules, calculates the SPPB score, generates a structured assessment report in real time, and simultaneously pushes it to the electronic medical record system.
[0026] Figure 3 This is a structural schematic diagram of a floor mat according to an embodiment of the present invention. (Combined with...) Figure 2 and Figure 3 As shown, the mat is positioned within the camera's field of view on the terminal. It should be noted that "front and back" of the mat refers to the area in front of the subject when the subject is on the mat and facing the camera on the terminal. Similarly, "left and right" refers to the subject's left and right positions when on the mat and facing the camera on the terminal.
[0027] like Figure 2 As shown, the four color blocks on the mat are labeled 0, 1, 3, and 2 respectively, from the bottom left counterclockwise to the top left. That is, the two color blocks on the left are 2 and 0, and the two color blocks on the right are 1 and 3; color blocks 2 and 3 are located in front of the mat, and color blocks 0 and 1 are located in the back of the mat, forming a four-point positioning layout with double rows in front and back and symmetrical left and right.
[0028] In addition, behind the four colored blocks on the mat is the sitting and standing test area for the subjects to perform sitting and standing actions so that the camera on the terminal can capture the changes in posture during the sitting and standing process.
[0029] It should be noted that during gait detection, the starting line for the subject's walking is located behind the mat, while the ending line can be located either in front of or behind the mat. A 3m or 4m walking test is performed within the area defined by the starting and ending lines; the 4m walking test will be used as an example below. Guided by the terminal system, the subject can sequentially perform the sit-to-stand test, balance test, and 4m walking speed test. That is, after completing the sit-to-stand and balance tests on the mat, the subject must move backward off the mat and behind the starting line to prepare for the walking speed test.
[0030] Specifically, in the sit-stand test, subjects used a standard chair without arm support and completed five consecutive "stand-up-sit" movements, recording the total time for the five movements. In the 4m gait test, subjects started from the starting line and walked 4m to the finish line with a normal gait, recording the time required. The balance test included three standing postures: feet together (standing with feet together), semi-sequential (one heel touching the big toe of the other foot), and full-sequential (one toe touching the heel of the other foot, feet in a straight line), with increasing difficulty. The longest stable time for each posture was recorded. The control module in the terminal could collect key motion data (including sit-stand completion time, balance maintenance time, gait speed, etc.) and automatically calculate the scores for each item and the total score according to the SPPB standard scoring algorithm.
[0031] The image acquisition module, sitting / standing detection module, balance detection module, and gait detection module are all connected to the control module.
[0032] To address the shortcomings of traditional sit-stand tests that rely on seat pressure sensors or wearable gyroscopes—which require chair modifications or device wearing by subjects, resulting in inconvenience, only provide a rough assessment of sitting / standing states, and cannot determine the quality of movement such as arm assistance or whether the subject is fully upright or stable—and are also susceptible to noise and drift, limiting accuracy, this evaluation system utilizes optical motion capture to precisely track key points such as the hip. It achieves high-precision sit-stand recognition without the need for external devices and can detect individual positional changes, thus improving the convenience and accuracy of the sit-stand test.
[0033] The sitting / standing detection module, in conjunction with the image acquisition module and control module, can achieve the following functions: The sitting / standing detection module is used to obtain the confidence score of the first key point in each frame of the image based on a human pose estimation algorithm. When the confidence score of the first key point is greater than or equal to a first threshold, sitting / standing recognition is performed. When the confidence score of the first key point is less than the first threshold, the current frame is discarded and sitting / standing recognition is not performed. The first key point includes at least one of the following: left shoulder key point, right shoulder key point, left hip key point, and right hip key point.
[0034] The sitting / standing detection module is also used to determine a standing up when the vertical coordinates of the left hip key point and the right hip key point are both above the upper threshold in each frame of the image; and to determine a sitting down when the vertical coordinates of the left hip key point and the right hip key point are both below the lower threshold. A standing up and a sitting down constitute a valid sitting / standing action.
[0035] In practical applications, before determining whether a person is sitting or standing, BlazePose (a lightweight real-time human pose estimation algorithm) is used to read the pose recognition results output by BlazePose for each frame of the image. This includes a confidence score for the overall quality of the current human pose recognition. This score comprehensively reflects whether the human body is fully visible in the frame, whether the first key point is clearly visible, and the reliability of the model's judgment on the current pose.
[0036] Only when the pose recognition confidence score of the current frame is ≥ 0.8 is the first key point information of the frame determined to be stable and reliable, and it is allowed to enter the sitting and standing action recognition process; when the confidence score is lower than the threshold, the current frame is skipped directly, and the frame is directly regarded as an invalid frame and ignored.
[0037] When a score < 0.8 is detected, the system will not make any sitting or standing status judgment on the current frame, nor will it update the timing and counting processes related to the action, thereby avoiding false triggering caused by low-confidence frames participating in the calculation.
[0038] It should be noted that the first key points are the left shoulder key point, right shoulder key point, left hip key point, and right hip key point. In selecting the first key points, the system did not rely on a single, easily obscured part (such as the wrist or knee), but instead used the shoulder and hip, which are the most stable key points, as the main detection reference lines. Since the shoulders and hips are always visible during sitting and standing; compared with the lower limbs and upper limbs, they are least likely to be obscured by clothing, armrests, and arms; and their relative position changes can stably reflect changes in overall body height. Based on this, even if the arms obscure the body or the lower limbs are not visible, the sitting and standing detection module can still rely on the trunk key points (shoulder and hip) to complete the height and status determination.
[0039] To address the issues of occlusion and distortion of the first keypoint caused by body shift, sideways movement, or significant forward tilting, the first keypoint also includes the nose keypoint. The sitting / standing recognition module is further used to acquire the nose keypoint in each frame of the image, and performs sitting / standing recognition when the lateral coordinate of the nose keypoint is within the first coordinate range of the current frame.
[0040] By monitoring the lateral position (x-coordinate) of key points on the nose in real time, it can be determined whether the subject has deviated from the camera's effective recognition area: based on the display resolution (1080×1920), if the nose's x-coordinate deviates from the center of the image (360-720) for more than 2 seconds, it indicates that the person is standing too far to the side or not facing the camera directly; the terminal prompts "position error" to guide the subject back to a suitable position. Based on this, the position of the human body in the image can be constrained, reducing the probability of unreliable posture recognition due to positional deviation.
[0041] Understandably, in the sit-stand test, different subjects have significant differences in height, torso proportions, chair height, and depth of sitting posture. Therefore, it is impossible to uniformly determine "sitting down" and "standing up" by using a fixed pixel height or a fixed proportion. To address this, the control module automatically generates a height judgment threshold that matches the subject's body shape based on the subject's body structure information at the beginning of the test, and dynamically uses this threshold to determine the state during the test.
[0042] In practical applications, capture a key frame at the start of the test when the posture is stable: Left shoulder y: la = landmarks
[11] .y; Left hip y: la1 = landmarks
[23] .y; Wherein, la (left shoulder y-coordinate): represents the vertical position of the subject's left shoulder in the current video frame, used to reflect the height of the upper body. la1 (left hip y-coordinate): represents the vertical position of the subject's left hip in the current video frame, used to reflect the height of the pelvis. landmarks 11 (left shoulder keypoint): the keypoint number used in the human pose recognition model to represent the anatomical position of the left shoulder, its coordinates are used to characterize the longitudinal position of the upper body in the image. landmarks 23 (left hip keypoint): the keypoint number used in the human pose recognition model to represent the anatomical position of the left hip, its coordinates are used to characterize the pelvic position and serve as the lower reference point for calculating the trunk length.
[0043] Calculate the trunk length (vertical distance between the left shoulder and left hip), L = la - la1.
[0044] L is used as an individualized scale parameter to eliminate the influence of height differences and camera distance variations on threshold settings.
[0045] After obtaining the torso length L, the system determines the relative position of the sitting / standing threshold within the torso using a scaling factor k. Here, k is the scaling factor, defaulting to 0.5, indicating the threshold is located near the midpoint between the shoulder and hip. The adjustable range of k is set from 0.3 to 0.7 to accommodate different body types or specific sitting posture requirements.
[0046] Upper threshold: laupper = la1 + ((la - la1) * (k + 0.1); Lower threshold: ladown = la1 + ((la - la1) * (k - 0.1); 0.1 indicates an offset of 0.1 times the torso length above and below the scaling factor k, used to construct a decision buffer.
[0047] The interval formed between (k + 0.1) and (k - 0.1) is the margin, with a width of 2 × 0.1 × L. A margin of 0.1 times the torso length is reserved above and below the proportional coefficient k to form two different judgment thresholds, thereby reducing the probability of frequent switching between sitting and standing states when the first key point shakes near the threshold.
[0048] During the test, the sitting-standing detection module simultaneously monitors the y-coordinates of key points on the left and right hips. When both hips are above the upper threshold, it is determined as "standing up" and when they are below the lower threshold, it is determined as "sitting down", thus completing a valid sitting-standing test.
[0049] As shown above, the sit-stand detection module, in conjunction with the image acquisition and control modules, uses optical motion capture and deep learning pose estimation algorithms to identify the positional changes of multiple key points such as the hip and shoulder in real time. The control module analyzes the input video stream in real time, automatically locating and tracking the two-dimensional coordinate changes of multiple key human body points such as the hip and shoulder. By continuously calculating the coordinate change pattern of the hip joint and the vertical displacement trend of the torso, the sit-stand detection module can identify the action transition node between "sitting down" and "standing up"; after identifying five consecutive complete sit-stand cycles, it records the completion time and outputs the test results. Through smooth detection based on duration thresholds and occlusion tolerance, the influence of lighting, occlusion, posture shift, and frame jitter on the monocular camera can be reduced, improving the reliability of sit-stand action recognition.
[0050] Example 2 To address the challenge of achieving objective, precise, and automated assessment of balance ability in balance testing, the balance detection module is further refined based on Example 1, eliminating the need for subjects to wear any devices and remaining unaffected by signal drift.
[0051] Combination Figure 2 and Figure 3 As shown, four 2x2 color blocks are set on the floor mat, and the four color block areas captured by the camera are stably and accurately converted into calculable spatial location points.
[0052] The balance detection module is used to convert the four color blocks from the RGB color space to the HSV color space, and set the corresponding HSV threshold range for each preset target color. The inRange method is used for threshold segmentation to obtain the binarized image of the target color. Morphological denoising processing and contour detection are performed on the binarized image, the geometric moments of each connected region are calculated, and the centroid coordinates of the color blocks with areas greater than the area threshold are calculated to obtain the corresponding marker points of the four color blocks.
[0053] In practical applications, the acquired ground image is converted from the RGB color space to the HSV color space, and corresponding HSV threshold ranges are set for preset target colors such as blue and green (four color blocks). The inRange method is used for threshold segmentation to obtain a binary image of the target color.
[0054] Subsequently, morphological denoising processing is performed on the binary image, specifically including an erosion operation using a 3×3 structuring element to remove discrete noise points, and a dilation operation using an 8×8 structuring element to enhance the connectivity and stability of the target region. Based on this, contour detection is performed on the processed binary image, the geometric moments of each connected region are calculated, and small noise targets are filtered out by area thresholding. The connected region area must be greater than 30×30 pixels to be considered a valid target.
[0055] Finally, the centroid coordinates of the color block regions that meet the area conditions are calculated. After completing the contour detection and filtering out the color block regions that meet the area threshold conditions, the balance detection module calculates the image geometric moments (Moments) of each effective connected region and obtains the centroid coordinates of the region based on the zeroth moment and the first moment.
[0056] Specifically, the zeroth moment (m00) is first calculated, representing the total number of pixels in the connected region (i.e., the region area). Then, the first moments (m10) and (m01) are calculated, representing the weighted sum of the coordinates of all pixels within the region in the (x) and (y) directions, respectively. Based on this, the geometric center of the region, i.e., the centroid coordinates ((cx, cy)), is obtained using the formulas (cx = m10 / m00) and (cy = m01 / m00). This centroid can be considered as the representative position of the corresponding color patch region in the image, stably reflecting the spatial position of the color patch.
[0057] Finally, the calculated centroid coordinates are added to the colorPos list as the set of foot markers in the current frame for subsequent pose recognition and stance determination.
[0058] The subject stands on the corresponding color block area on the mat. When a color block is covered, the balance detection module cannot detect the corresponding foot marker. However, for uncovered color blocks, the corresponding foot markers are detected by the balance detection module. Of the four markers, the two in front are closer to the image acquisition module, and the two in the back are farther away from the image acquisition module.
[0059] The balance detection module determines the following stances in the current frame: when two markers are detected and both feet cover the two rear markers of the four markers respectively, it is a standing posture with feet together; when only one marker is detected and one foot covers two markers on the same side and the other foot covers the rear marker of the two markers on the other side, it is a semi-sequential standing posture; when two markers are detected and one foot covers one marker on one side and the other foot covers the other marker on the same side, it is a fully sequential standing posture.
[0060] (1) The subject stands with his feet together.
[0061] The balance detection module needs to detect two foot markers in a single frame image; if the number of detected markers is insufficient or exceeds this number, the pose determination is invalid. In terms of spatial constraints, the two detected foot markers are matched as color blocks 2 and 3, respectively, meaning that the left foot and the right foot cover color blocks 0 and 1, respectively.
[0062] The control module will officially start timing only after the above conditions are met continuously for more than 1 second. During the timing process, if the number of foot markers is temporarily insufficient due to brief obstruction or recognition jitter, a 0.5-second error tolerance period is allowed. Once the error tolerance threshold is exceeded or the specified upper limit of 10 seconds is reached, the score will be recorded and the current standing posture test will end.
[0063] (2) Subjects in a semi-tandem standing posture.
[0064] The balance detection module can detect only one foot marker in a single frame image to accommodate situations where the front and rear feet overlap or occlude.
[0065] Regarding spatial constraints, the detected foot markers must match either color block 2 or color block 3. That is, the left foot covers color blocks 0 and 2, and the right foot covers color block 1; or the left foot covers color block 0, and the right foot covers color blocks 1 and 3, which is considered a valid standing posture. If two foot markers are detected, the standing posture is considered invalid, and the timing will not begin.
[0066] A 1-second continuous stability is used as the timing start condition, and a short-term fault tolerance mechanism of 0.5 seconds is set during the timing process; when the number of foot markers no longer meets the requirements or the holding time reaches the upper limit of 10 seconds, the standing posture score is recorded and the test ends.
[0067] (3) Subjects in a fully tandem standing posture.
[0068] The balance detection module detects two foot markers in a single frame image to enhance adaptability to occlusion and pose changes.
[0069] In terms of spatial constraints, the two foot markers being detected match color blocks 0 and 2, or color blocks 1 and 3. That is, the left foot covers color block 1 while the right foot covers color block 3; or the left foot covers color block 0 while the right foot covers color block 1.
[0070] A 1-second continuous stability is used as the timing start condition, and a short-term fault tolerance mechanism of 0.5 seconds is set during the timing process; when the number of foot markers no longer meets the requirements or the holding time reaches the upper limit of 10 seconds, the standing posture score is recorded and the test ends.
[0071] As described above, the system uses a camera to detect four color blocks on the mat in real time and employs color gamut recognition technology to extract the foot position area, ensuring that the subject stands at the correct starting point. The system automatically starts a 10-second timer. Once it detects that the foot position has deviated beyond the stable range or that the color block area has been obscured, it determines that the subject has failed to maintain balance and terminates the phase. Compared to existing technologies that rely on manual observation of timing, force platforms, or wearable inertial sensors, which are subject to strong subjectivity and make it difficult to quantify the swaying during the standing process, the balance detection module of this invention allows the subject to automatically complete the balance test without wearing any external devices. It can comprehensively reflect the stability of the whole body posture, is not affected by the wearing position or signal drift, and can automatically identify violations such as foot movement, thereby achieving standardized testing and automated evaluation of balance ability.
[0072] Example 3 In related technologies, gait speed testing often relies on manual timing, ground marking, or wearable devices. Common problems include reliance on operator experience, subjective judgment of start and end times, and insufficient repeatability and consistency. Some electronic solutions require subjects to wear additional sensors or use dedicated ranging devices, increasing the testing burden and making them susceptible to wear position, device drift, and communication stability issues, thus limiting their clinical application. This embodiment, based on Embodiment 1 or Embodiment 2, further optimizes the gait detection module. It automatically identifies the start and end states of walking through the second key point on the human body or non-contact distance changes, achieving full automation and objectivity of the gait speed test process without requiring subjects to wear any external devices. The system can automatically complete start and end triggering, timing, and result recording, effectively reducing human intervention and operational errors, and significantly improving the accuracy, stability, and repeatability of test results.
[0073] The gait detection module's functions include: Second keypoints include left and right ankle keypoints. The gait detection module extracts the coordinates of the second keypoints in real-time from each frame and normalizes the output. Based on the screen's resolution, the coordinates of the second keypoints are mapped to the screen coordinate system. When both the left and right ankle keypoints are detected crossing the starting virtual line to the rear, and either foot crosses the starting virtual line to the front again, the gait start time is determined. When both the left and right ankle keypoints are detected crossing the ending virtual line, the gait end time is determined.
[0074] In practical applications, the gait detection module extracts the second keypoints (landmark 31 for the left ankle and landmark 32 for the right ankle) of the human body in real time from each frame of the image. The coordinates of the second keypoints are output in a 0–1 normalized form. Then, according to the display resolution (1080×1920), the keypoints are mapped to the screen coordinate system, where the ordinate is converted to y = 1920. (pos.y×1920), where pos.y (normalized coordinates) represents the relative vertical position of the key point in the image, so that the foot position can be accurately compared with the preset virtual start and end lines, thus realizing continuous tracking of the lower limb movement trajectory.
[0075] The gait detection module first checks whether both feet have crossed the starting virtual line at the same time (i.e., the ordinates of both ankles are less than the starting threshold up_h). This state is marked as "crossed". Then, the timing is officially triggered and the start time is recorded only when either foot returns to the back of the starting line (the ordinate is greater than up_h again) and the "crossed" condition has been met before.
[0076] Understandably, since the gait test is conducted after the standing or balance test on the mat, after completing the preceding test on the mat, the subject moves to the rear of the mat (the side away from the camera on the terminal) to the preset starting position. The gait detection module first confirms that both feet have crossed the starting virtual line to the rear, confirming the "crossed the line" state. At this time, the system enters the waiting mode, and timing is started immediately as long as one foot steps back. Based on this, a "crossed the line → returned to the line" state transition is introduced, which can effectively filter out false triggers caused by initial position deviation, slight movement of one foot, or shaking of key points, and avoid the timing starting point being locked in advance.
[0077] During the timing phase, the system continuously updates the walking time at a frame-level frequency and simultaneously monitors the key point recognition status. When it detects that the left and right ankle key points simultaneously cross the virtual finish line (both ordinates are greater than the finish line threshold down_h), the control module and gait detection module automatically determine that the test is complete and record the total walking time.
[0078] If the identification of the second key point is lost continuously for more than 3 seconds during the timing process, it is judged as an abnormal test state and automatically terminated to prevent invalid data caused by posture loss, occlusion or leaving the detection area from entering the result statistics.
[0079] To ensure the standardization and security of the testing process, the maximum test duration threshold is 20 seconds. If the test time exceeds this threshold and the effective endpoint is not triggered, the test will be automatically terminated and marked as timed out to avoid abnormal delays from affecting the overall evaluation process.
[0080] Therefore, the gait detection module, image acquisition module and control module work together to achieve automatic recognition of the start and end times of walking through multi-state joint determination. This avoids the risk of misjudgment caused by single-frame triggering from the algorithm level and significantly improves the stability and consistency of timing triggering.
[0081] The control module coordinates the timing and data flow of each module, employing a multi-threaded data acquisition structure to simultaneously receive real-time analysis results from three subsystems: sit / stand detection, balance detection, and gait detection. At different testing stages, it automatically invokes the corresponding algorithm modules to perform keypoint recognition, color gamut analysis, or trajectory detection. The data is automatically fused in the background and ultimately judged uniformly according to SPPB standard rules, generating sub-scores and a total score, achieving a fully automated processing flow from action recognition to result output.
[0082] Example 4 This embodiment provides an intelligent assessment method for SPPB fall risk based on multimodal perception. After completing three tests using the system of Embodiment 1, Embodiment 2 or Embodiment 3, a structured assessment report is automatically generated, covering the scores of each sub-item and the risk level (low / medium / high).
[0083] SPPB's intelligent fall risk assessment methods include: Step S1: Based on the first key point captured by optical motion capture, detect the motion transition node between sitting and standing. Specifically, based on the human pose estimation algorithm, the confidence score of the first key point of each frame image is obtained; When the confidence level of the first key point is greater than or equal to the first threshold, sitting / standing identification is performed; when the confidence level of the first key point is less than the first threshold, the current frame is discarded and sitting / standing identification is not performed.
[0084] The first key point includes at least one of the following: left shoulder key point, right shoulder key point, left hip key point, and right hip key point.
[0085] The first key point also includes the nose key point; the nose key point of each frame image is obtained, and when the horizontal coordinate of the nose key point is within the first coordinate range of the current frame, sitting and standing recognition is performed.
[0086] In each frame of the image, a standing up is defined as when the vertical coordinates of the left hip key points and the right hip key points are both above the upper threshold; a sitting down is defined as when the vertical coordinates of the left hip key points and the right hip key points are both below the lower threshold. A standing up and a sitting down constitute a valid sitting-standing action.
[0087] Step S2: Based on the relative positional relationship between the marked area and the left and right feet, identify the standing posture of the human body. The standing posture includes one of the following: feet together, half-joint, or full-joint. Specifically, the marked area includes four color blocks arranged in an array on the floor mat; The four color blocks are converted into computable spatial location points to form four marker points in the current frame for evaluating standing posture.
[0088] The four color blocks are converted from the RGB color space to the HSV color space, and corresponding HSV threshold ranges are set for the preset target color. The inRange method is used for threshold segmentation to obtain the binarized image of the target color. Morphological denoising and contour detection are performed on the binarized image. The geometric moments of each connected region are calculated. The centroid coordinates of color blocks larger than the area threshold are calculated to obtain the corresponding marker points of the four color blocks.
[0089] In the current frame, when two markers are detected and both feet cover the two markers located behind the other two of the four markers, it is determined to be a standing posture with feet together. When only one marker is detected, and one foot covers two markers on the same side, while the other foot covers the marker located behind the other two markers on the opposite side, it is determined to be a semi-tandem standing posture. When two markers are detected, and one foot covers one marker on one side while the other foot covers the other marker on the same side, the stance is determined to be a fully tandem standing posture.
[0090] Of the four marker points, the two in front are closer to the image acquisition module, while the two in the back are farther away from the image acquisition module.
[0091] Step S3: Identify the start and end times of the walk based on the second key point or non-contact distance changes; Specifically, the second key points include the left ankle key points and the right ankle key points. The coordinates of the second key points are extracted in real time in each frame of the image and normalized. The coordinates of the second key points are mapped to the screen coordinate system based on the screen resolution. When the left ankle key points and the right ankle key points are detected to cross the starting virtual line to the rear at the same time, and the gait starts when either foot crosses the starting virtual line to the front again; The moment when both the left and right ankle key points cross the virtual endpoint line is determined as the gait termination moment.
[0092] Step S4: Calculate the SPPB score based on the data collected by the sit-stand detection module, balance detection module, gait detection module, and SPPB standard rules.
[0093] The technical advantages of the above-mentioned evaluation method compared to the prior art are the same as those of the above-mentioned evaluation system, and will not be repeated here.
[0094] Those skilled in the art should understand that the above embodiments are merely for illustrating the present invention and are not intended to limit the scope of the invention. Those skilled in the art can make other changes or modifications based on the above disclosure, and these changes or modifications still fall within the scope of the present invention.
Claims
1. A multimodal perception-based intelligent assessment system for SPPB fall risk, characterized in that, It includes an image acquisition module, a sitting / standing detection module, a balance detection module, a gait detection module, and a control module; the image acquisition module is used to acquire images of human posture. The sitting / standing detection module is used to detect the motion transition nodes between sitting and standing based on the first key point captured by optical motion capture. The balance detection module is used to identify the standing posture of a human body based on the relative positional relationship between the marked area and the left and right feet. The standing posture includes one of the following: feet together, half-seam, or full-seam. The gait detection module is used to identify the start and end times of walking based on a second key point or non-contact distance changes. The control module is used to calculate the SPPB score based on the data collected by the sitting / standing detection module, the balance detection module, and the gait detection module, as well as the SPPB standard rules.
2. The evaluation system according to claim 1, characterized in that, The sitting / standing detection module is also used to obtain the confidence level of the first key point of each frame image based on the human pose estimation algorithm; When the confidence level of the first key point is greater than or equal to the first threshold, sit-to-stand detection is performed; when the confidence level of the first key point is less than the first threshold, the current frame is discarded and sit-to-stand detection is not performed.
3. The evaluation system according to claim 2, characterized in that, The first key point includes at least one of the left shoulder key point, the right shoulder key point, the left hip key point, and the right hip key point.
4. The evaluation system according to claim 3, characterized in that, The first key point also includes the key point of the nose; The sitting / standing detection module is also used to acquire the nose key points of each frame image. When the horizontal coordinate of the nose key points is within the first coordinate range of the current frame, sitting / standing detection is performed.
5. The evaluation system according to claim 3, characterized in that, The sitting / standing detection module is also used to determine a standing up when the vertical coordinates of the left hip key point and the right hip key point are simultaneously higher than the upper threshold in each frame of the image. A sitting down is defined as when the vertical coordinates of the left hip key points and the right hip key points are both below the lower threshold. A sitting-standing action consists of completing one standing up and one sitting down.
6. The evaluation system according to claim 1, characterized in that, The marked area includes four color blocks arranged in an array on the floor mat; The balance detection module is used to convert the four color blocks into computable spatial location points to form four marker points in the current frame for evaluating standing posture.
7. The evaluation system according to claim 6, characterized in that, The balance detection module is used to convert the four color blocks from the RGB color space to the HSV color space, and set corresponding HSV threshold ranges for the preset target colors respectively. The inRange method is used to perform threshold segmentation to obtain a binarized image of the target color. Morphological denoising and contour detection are performed on the binarized image. The geometric moments of each connected region are calculated. The centroid coordinates of color blocks with an area greater than the area threshold are calculated to obtain the corresponding marker points of the four color blocks.
8. The evaluation system according to claim 7, characterized in that, The balance detection module is used to detect the number of foot markers in the current frame.
9. The evaluation system according to claim 8, characterized in that, Of the four marker points, the two in front are closer to the image acquisition module, while the two in the back are farther away from the image acquisition module.
10. A method for intelligent assessment of SPPB fall risk based on multimodal perception, characterized in that, The evaluation is performed using the evaluation system described in any one of claims 1-9.