Balanced function evaluation method and system based on multi-modal fusion

CN122320489BActive Publication Date: 2026-08-11XIEHE HOSPITAL ATTACHED TO TONGJI MEDICAL COLLEGE HUAZHONG SCI & TECH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-08
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0004]本发明的主要目的在于提供一种多模态融合的平衡功能评估方法及系统,旨在解决现有平衡功能评估技术中评估维度单一、生态效度低以及设备可及性差的技术问题

Benefits of technology

[0015]This invention controls a curved display screen to show visual stimuli tasks via a synchronous controller, and controls a programmable 3D force stage to apply time-synchronized physical perturbations, recording task parameters. A synchronous trigger unit controls a multi-view camera array, a plantar pressure imaging unit, and an EEG acquisition module to simultaneously acquire multimodal data. An edge computing unit extracts posture features, mechanical features, and EEG features based on the multimodal data. These features and task parameters are input into a cloud-based AI large-scale model platform based on a multimodal Transformer architecture, outputting evaluation results. This approach achieves multimodal quantitative evaluation of the entire balance control chain and can generate personalized rehabilitation plans, improving the ecological validity of balance function assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122320489B_ABST
    Figure CN122320489B_ABST
Patent Text Reader

Abstract

This invention belongs to the field of balance function assessment technology for medical rehabilitation equipment, and discloses a multimodal fusion balance function assessment method and system. It includes: controlling a curved display screen to show a visual stimulus task via a synchronous controller, and controlling a programmable three-dimensional force stage to apply time-synchronized physical perturbations, recording task parameters; controlling a multi-view camera array, a plantar pressure imaging unit, and an EEG acquisition module to synchronously acquire multimodal data via a synchronous triggering unit; extracting posture features, mechanical features, and EEG features based on the multimodal data via an edge computing unit; inputting the above features and task parameters into a cloud-based AI large-scale model platform based on a multimodal Transformer architecture, and outputting the assessment results. This method achieves multimodal quantitative assessment of the entire balance control chain and can generate personalized rehabilitation plans, improving the ecological validity of balance function assessment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of balance function assessment technology for medical rehabilitation equipment, and in particular to a multimodal fusion method and system for balance function assessment. Background Technology

[0002] The maintenance of human balance function depends on the coordinated action of the vestibular, visual, and proprioceptive systems. Currently, equipment used for balance function assessment both domestically and internationally mainly falls into three categories: balance assessment devices based on mechanical sensors, balance assessment devices based on wearable sensors, and balance assessment systems based on image analysis. These existing technologies suffer from the following major drawbacks: They offer a single assessment dimension, relying solely on plantar pressure center parameters, failing to capture the dynamic coordinated movements of the head, trunk, and limbs, and easily leading to "false normal" assessment results; they lack an immersive dynamic stimulation environment, resulting in low ecological validity and limited correlation between laboratory results and real-world fall risk; they cannot monitor central nervous system activity, have insufficient ability to differentiate etiologies, and struggle to distinguish between vestibular, central, and functional vertigo; assessment is disconnected from rehabilitation, lacking dynamic early warning capabilities; and the equipment is bulky, expensive, and poorly accessible.

[0003] The above content is only used to help understand the technical solution of the present invention and does not represent an admission that the above content is prior art. Summary of the Invention

[0004] The main objective of this invention is to provide a multimodal fusion method and system for assessing balance function, aiming to solve the technical problems of single assessment dimensions, low ecological validity, and poor equipment accessibility in existing balance function assessment technologies.

[0005] To achieve the above objectives, the present invention provides a multimodal fusion-based method for evaluating the balance function, the multimodal fusion-based method comprising the following steps: The system controls the curved display screen to display the visual stimulus task through a synchronous controller, and controls the programmable three-dimensional force stage to apply physical perturbations that are time-synchronized with the visual stimulus task. The type of the visual stimulus task and the motion parameters of the physical perturbations are recorded as task parameters. During the process of the perturbed environment acting on the subject, the multi-view camera array is controlled by the synchronous triggering unit to acquire multi-view images, the plantar pressure imaging unit is controlled to acquire plantar pressure distribution images, and the EEG acquisition module is controlled to acquire EEG signals. Among them, the brightness change of the preset area of ​​the curved display screen is detected by the photodiode as a time reference for the visual stimulus task display, so as to ensure that the time of each modal data is aligned with the task parameters. The modal data includes the multi-view images, plantar pressure distribution images, and EEG signals. The edge computing unit performs three-dimensional pose reconstruction based on the multi-view images to obtain pose features containing the three-dimensional coordinate sequence of key points of the whole skeleton. The plantar pressure distribution image is processed to obtain mechanical features including the plantar pressure center trajectory, arch index and plantar support area. The electroencephalogram (EEG) signal is analyzed in the frequency domain to obtain EEG features including the power and ratio of each channel. The posture features, mechanical features, EEG features, and task parameters are input into a cloud-based AI large-scale model platform, which outputs evaluation results. The cloud-based AI large-scale model platform is based on the Transformer architecture and uses the task parameters as a query to fuse features from various modalities through a cross-attention mechanism. The evaluation results include a comprehensive balance function score, vertigo subtype classification probability, contribution weight, and dynamic warning of fall risk. The contribution weight includes vestibular contribution weight, visual contribution weight, and proprioceptive contribution weight.

[0006] In one embodiment, the step of reconstructing three-dimensional pose based on the multi-view images to obtain pose features containing a three-dimensional coordinate sequence of key points of the whole skeleton includes: Perform two-dimensional pose estimation on the multi-view images and output the two-dimensional coordinates and confidence scores of several skeletal key points in each view. For each skeletal keypoint, a hierarchical processing is performed based on the confidence level. The hierarchical processing is as follows: if the confidence level of two or more viewpoints is greater than 0.8, the three-dimensional coordinates are solved by direct linear transformation, and weighted least squares optimization is performed with the confidence level as the weight. If the confidence level of one viewpoint is greater than 0.8, the two-dimensional point of that viewpoint is back-projected into a three-dimensional ray, and the optimal three-dimensional point is searched on the ray in combination with the bone length constraints of adjacent reconstructed keypoints. If the confidence level of all viewpoints is less than or equal to 0.8, the corresponding skeletal keypoints are marked as missing and interpolation is performed. Kalman filtering is used to perform temporal smoothing on the reconstructed 3D keypoint coordinate sequence, and the two ends of the skeleton are fine-tuned and corrected using a preset range of human skeleton length to obtain the pose features containing the 3D coordinate sequence of keypoints of the whole skeleton.

[0007] In one embodiment, the image processing of the plantar pressure distribution image to obtain mechanical features including the plantar pressure center trajectory, arch index, and plantar support area includes: The plantar pressure distribution image is converted to HSV color space, and the footprint region image is extracted based on hue, saturation, and brightness thresholds. The footprint area image is subjected to morphological processing to remove noise and fill holes, and then converted into a grayscale image, wherein the grayscale value in the grayscale image represents the pressure level; The image weighted center is calculated based on the grayscale image, and the image weighted center is used as the pressure center coordinate at the current moment. The pressure center coordinates on the time series are recorded to form the trajectory of the plantar pressure center. The footprint is divided into three equal regions: the forefoot region, the midfoot region, and the heel region, based on the minimum bounding rectangle of the footprint. The arch index and the plantar support area are calculated for each region, and the biomechanical characteristics are output. The arch index is calculated from the area of ​​the midfoot region and the total area of ​​the footprint, and the plantar support area is the total area of ​​the footprint.

[0008] In one embodiment, the step of inputting the posture features, the mechanical features, the EEG features, and the task parameters into a cloud-based AI large-scale model platform and outputting evaluation results includes: The pose features are input into a pose encoder consisting of a one-dimensional convolutional neural network (1D-CNN) and a Transformer encoder to generate a pose feature vector. The mechanical features are input into a mechanical encoder consisting of a long short-term memory network (LSTM) and attention pooling to generate a mechanical feature vector. The EEG features are input into an EEG encoder consisting of a 1D-CNN and a Transformer encoder to generate an EEG feature vector. The task parameters are input into a task encoder composed of a multilayer perceptron to generate a conditional feature vector. The posture feature vector, the mechanical feature vector, the EEG feature vector, and the conditional feature vector are input into a multimodal Transformer. Through a cross-attention mechanism, the conditional feature vector is used as the query, and the posture feature vector, the mechanical feature vector, and the EEG feature vector are used as the key and value, respectively, to output fused features. The fusion features are input into the regression head and the classification head respectively, and the balance function comprehensive score and the vertigo subtype classification probability are output. At the same time, the fusion features are input into the attention mechanism layer to dynamically generate vestibular contribution weight, visual contribution weight and proprioceptive contribution weight. The fused features are input into the time series prediction model, and the output fall risk probability is used as a dynamic early warning of fall risk.

[0009] In one embodiment, the method further includes: Based on the comprehensive balance function score, the probability of vertigo subtype classification, the contribution weight, the ratio contained in the EEG features, and the visual dependence index determined based on the type of visual stimulation task, a preset mapping rule base is matched to generate a personalized rehabilitation plan that includes a list of training items, exercise prescriptions, and safety warnings. After completing the periodic assessment, an updated balance function comprehensive score is obtained. The updated balance function comprehensive score is compared with the historical balance function comprehensive score. Based on the comparison results, the training intensity of the current rehabilitation program is maintained or downgraded, the training frequency is reduced, the training intensity is upgraded, or biofeedback assistance is increased or the training intensity is downgraded, and a manual review warning by the therapist is triggered.

[0010] In one embodiment, the mapping rule base includes at least one of the following rules: When the visual dependence index is greater than 0.6, the matching includes adaptive training with VOR×1 / VOR×2; When the vestibular contribution weight is less than 20% and the VOR gain calculated from the pose features is less than 0.7, the matching includes adaptation and alternative training with VOR×1 and saccades or tracking. When the dynamic warning of fall risk is greater than 70%, balance-gait training including weight transfer, standing on a foam mat, and dual-task walking is matched. When the proprioceptive contribution weight is less than 15%, sensory substitution training is matched, including standing with eyes closed, soft support surfaces, and dynamic platforms. When the ratio of the EEG features is greater than 2.5, the matching includes habituation and cognitive-balance dual-task training that includes inducing repetitive actions and cognitive interference.

[0011] In one embodiment, the method further includes: The system records the daily training check-in rate of the subject for the personalized rehabilitation plan through the user terminal. When the daily training check-in rate is less than 70%, a reminder is automatically pushed to the user terminal and the difficulty of the movements in the training plan is simplified. A rehabilitation summary report is automatically generated periodically. The report includes the change curve of the comprehensive balance function score, the gap with the expected goal, the explanation of the adjustment of the plan, and the training suggestions for the next stage.

[0012] Furthermore, to achieve the above objectives, this invention also proposes a multimodal fusion-based balance function assessment system, which is applied to the multimodal fusion-based balance function assessment method described above. The system includes: An immersive visual physical perturbation subsystem is used to control a curved display screen to display a visual stimulus task through a synchronization controller, and to control a programmable three-dimensional force stage to apply physical perturbations that are time-synchronized with the visual stimulus task. The type of the visual stimulus task and the motion parameters of the physical perturbation are recorded as task parameters. A multimodal biosignal acquisition subsystem is used to control a multi-view camera array to acquire multi-view images, control a plantar pressure imaging unit to acquire plantar pressure distribution images, and control an EEG acquisition module to acquire EEG signals during the process of a disturbed environment acting on the subject. A photodiode is used to detect brightness changes in a preset area of ​​the curved display screen as a time reference for the visual stimulus task display, ensuring that the time of each modal data is aligned with the task parameters. The modal data includes the multi-view images, plantar pressure distribution images, and EEG signals. The intelligent fusion analysis subsystem is used to perform three-dimensional pose reconstruction based on the multi-view images through the edge computing unit to obtain pose features including the three-dimensional coordinate sequence of key points of the whole skeleton; to perform image processing on the plantar pressure distribution image to obtain mechanical features including the plantar pressure center trajectory, arch index and plantar support area; and to perform frequency domain analysis on the EEG signal to obtain EEG features including the power and ratio of each channel. The intelligent fusion analysis subsystem is used to input the posture features, mechanical features, EEG features, and task parameters into the cloud-based AI large model platform and output evaluation results. The cloud-based AI large model platform is based on the Transformer architecture and uses the task parameters as a query to fuse features of various modalities through a cross-attention mechanism. The evaluation results include a comprehensive balance function score, vertigo subtype classification probability, sensory contribution weight, and dynamic warning of fall risk.

[0013] In one embodiment, the programmable three-dimensional force stage includes: a base, a balancing force stage, a linear guide module, a servo drive unit, and four three-dimensional force sensors; the four three-dimensional force sensors are respectively installed at the four corners of the bottom of the balancing force stage for real-time acquisition of the force values ​​at each corner; the edge computing unit calculates the plantar pressure center coordinates based on the force values; the synchronization controller receives the plantar pressure center coordinates and dynamically adjusts the parameters of subsequent physical disturbances applied by the programmable three-dimensional force stage according to the plantar pressure center coordinates.

[0014] Furthermore, to achieve the above objectives, the present invention also proposes a multimodal fusion balance function evaluation device, the multimodal fusion balance function evaluation device comprising: a memory, a processor, and a multimodal fusion balance function evaluation program stored in the memory and executable on the processor, the multimodal fusion balance function evaluation program being configured to implement the steps of the multimodal fusion balance function evaluation method as described above.

[0015] This invention controls a curved display screen to show visual stimuli tasks via a synchronous controller, and controls a programmable 3D force stage to apply time-synchronized physical perturbations, recording task parameters. A synchronous trigger unit controls a multi-view camera array, a plantar pressure imaging unit, and an EEG acquisition module to simultaneously acquire multimodal data. An edge computing unit extracts posture features, mechanical features, and EEG features based on the multimodal data. These features and task parameters are input into a cloud-based AI large-scale model platform based on a multimodal Transformer architecture, outputting evaluation results. This approach achieves multimodal quantitative evaluation of the entire balance control chain and can generate personalized rehabilitation plans, improving the ecological validity of balance function assessment. Attached Figure Description

[0016] Figure 1 This is a flowchart illustrating the first embodiment of the multimodal fusion-based balanced function evaluation method of the present invention; Figures 2 to 4 This is a schematic diagram of the overall architecture of the balance function assessment device in the multimodal fusion balance function assessment method of the present invention; Figure 5 This is a structural block diagram of the first embodiment of the multimodal fusion balanced function evaluation system of the present invention.

[0017] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0018] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.

[0019] This invention provides a multimodal fusion-based method for evaluating the balance function, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of a multimodal fusion-based balanced function evaluation method according to the present invention.

[0020] In this embodiment, the multimodal fusion balanced function evaluation method includes the following steps: Step S10: Control the curved display screen to display the visual stimulus task through the synchronous controller, and control the programmable three-dimensional force stage to apply physical perturbations that are synchronized with the timing of the visual stimulus task. Record the type of the visual stimulus task and the motion parameters of the physical perturbations as task parameters.

[0021] In this embodiment, the executing entity is a multimodal fusion balance function evaluation device, which has functions such as data processing, data communication and program execution. The multimodal fusion balance function evaluation device can be a computer terminal device or other network device, or other devices with similar functions. This embodiment does not limit this.

[0022] It should be noted that the maintenance of human balance function depends on the coordinated action of the vestibular, visual, and proprioceptive systems. Currently, devices used for balance function assessment both domestically and internationally are mainly divided into three categories: balance assessment devices based on mechanical sensors, balance assessment devices based on wearable sensors, and balance assessment systems based on image analysis. These existing technologies have the following main drawbacks: They offer a single assessment dimension, relying solely on plantar pressure center parameters, failing to capture the dynamic coordinated movements of the head, trunk, and limbs, and easily leading to "false normal" assessment results; they lack an immersive dynamic stimulation environment, resulting in low ecological validity and limited correlation between laboratory results and real-world fall risk; they cannot monitor central nervous system activity, have insufficient ability to differentiate etiologies, and struggle to distinguish between vestibular, central, and functional vertigo; assessment is disconnected from rehabilitation, lacking dynamic early warning capabilities; and the equipment is bulky, expensive, and poorly accessible.

[0023] To address the aforementioned technical challenges, this embodiment employs a synchronous controller to control a curved display screen to present visual stimulation tasks and a programmable 3D force stage to apply time-synchronized physical disturbances, recording task parameters. A synchronous trigger unit controls a multi-view camera array, a plantar pressure imaging unit, and an EEG acquisition module to simultaneously acquire multimodal data. An edge computing unit extracts posture features, mechanical features, and EEG features based on the multimodal data. These features and task parameters are then input into a cloud-based AI large-scale model platform based on a multimodal Transformer architecture, outputting evaluation results. This approach achieves multimodal quantitative evaluation of the entire balance control chain and can generate personalized rehabilitation plans, improving the ecological validity of balance function assessment.

[0024] In the specific implementation, the multimodal fusion balance function evaluation system in this embodiment consists of several subsystems. I. Immersive visual-physical disturbance subsystem. This subsystem specifically includes (1) Curved display screen: a high refresh rate curved display screen with a curvature radius matching the natural field of vision of the human eye, used to present immersive visual stimulation tasks; (2) Programmable three-dimensional force stage: including a base, a balance force stage, a linear guide rail module and a servo drive unit. The balance force stage is slidably connected to the base through the guide rail and has tilt and translation degrees of freedom, used to apply physical disturbances; (3) Synchronization controller: connects the curved display screen and the three-dimensional force stage to ensure the timing synchronization of visual stimulation and physical disturbances. II. Multimodal biosignal acquisition subsystem, which specifically includes (1) a multi-view camera array: consisting of 4-6 high-definition industrial cameras arranged around the subject to acquire multi-view human images; (2) a plantar pressure imaging unit: including an internal light source and a plantar camera, which uses the principle of total internal reflection to acquire plantar pressure distribution images; (3) an EEG acquisition module: including a dry electrode EEG headband and a signal processing unit, which is used to synchronously acquire the subject's EEG signals; (4) a synchronization triggering unit: connected to each acquisition device, which achieves millisecond-level time synchronization through hardware triggering.

[0025] III. Intelligent Fusion Analysis Subsystem, which specifically includes (1) an edge computing unit: deployed locally for real-time data processing and feature extraction; (2) a cloud-based AI large model platform: based on a multimodal Transformer architecture, fusing posture, mechanical, EEG, and visual task parameters to output balance function assessment results and dynamic warnings; and (3) a user terminal: providing a human-computer interaction interface, real-time feedback, and report generation functions. The overall architecture diagram of the balance function assessment device in this embodiment can be found in the following reference. Figures 2 to 4 As shown.

[0026] It should be noted that the visual stimulus task includes static visual scenes, scenes that move synchronously with body swaying, sinusoidal oscillation scenes, sudden movement scenes, and immersive daily life scenes (such as supermarkets, streets, buses, etc.). The motion parameters of the physical perturbation include tilt angular velocity, translational velocity, etc., which are normalized to 0-1. The above task parameters are recorded for subsequent multimodal fusion analysis.

[0027] Step S20: During the process of the perturbed environment acting on the subject, the multi-view camera array is controlled by the synchronous triggering unit to acquire multi-view images, the plantar pressure imaging unit is controlled to acquire plantar pressure distribution images, and the EEG acquisition module is controlled to acquire EEG signals.

[0028] In its implementation, the multi-view camera array consists of 4-6 industrial cameras arranged around the subject, employing a global shutter CMOS sensor with a resolution of 1920×1080 and a frame rate of 60fps. The plantar pressure imaging unit is based on the principle of total internal reflection: an LED light source is mounted on the side of the platform; when light travels from high-refractive-index glass to low-refractive-index air, total internal reflection occurs, resulting in a bright background on the platform surface. When the sole of the foot contacts the glass, the total internal reflection condition is disrupted, creating a dark area. The greater the pressure, the larger the contact area, and the lower the grayscale value of the dark area. The plantar camera captures images vertically upwards at a frame rate of 60fps. The EEG acquisition module uses an 8-channel dry electrode EEG headband, with electrode positions covering the prefrontal, central, and parietal lobes according to the international 10-20 system, and a sampling rate ≥500Hz. To achieve millisecond-level time synchronization, the synchronization controller in the synchronization trigger unit generates unified TTL trigger pulses based on an FPGA. Meanwhile, to avoid the high latency uncertainty of the HDMI CEC protocol, photodiodes are used to detect brightness changes in a preset area at the corner of the curved display screen (e.g., displaying a white rectangle at the beginning of each frame). The photodiode output signal is connected to the synchronization controller as a precise time reference for the presentation of visual stimuli. All trigger signals are generated uniformly to ensure that the start time deviation of each device is less than 10μs. The synchronization controller also outputs a common reference signal (e.g., a 1Hz square wave) and connects it to an idle channel of each acquisition device. During offline processing, a cross-correlation algorithm is used to calculate the fixed delay between the data stream of each device and the trigger pulse, and the timestamp is calibrated a second time to ensure that the maximum time deviation between multimodal data does not exceed 2ms.

[0029] Furthermore, the multi-view camera array module specifically includes (1) camera configuration, such as the number of cameras: 6 (which can be adjusted from 4 to 6 depending on the size of the test space); sensor type: CMOS global shutter, resolution 1920×1080, frame rate 60fps; lens: 8mm focal length, F1.4 aperture, manual focus. Installation layout: Camera 1: 0° front, height 2.5m, tilt angle 30°; Camera 2: 180° rear, height 2.5m, tilt angle 30°; Camera 3: 45° front left, height 2.5m, tilt angle 30°; Camera 4: 45° front right, height 2.5m, tilt angle 30°. Illumination: LED array supplementary light, color temperature 5500K, illuminance ≥500lux, uniformity ≥0.8. (2) Camera calibration, for example, using the Zhang Zhengyou calibration method, using a 12×9 checkerboard calibration board (grid size 30mm); acquiring 20 sets of calibration images in different poses, calculating the intrinsic parameter matrix, distortion coefficient and extrinsic parameters (rotation matrix, translation vector) of each camera; reprojection error ≤ 0.5 pixels, calibration frequency: once per quarter or recalibrated after equipment movement. (3) Image acquisition and transmission, for example, the camera is connected to the edge computing unit through the GigE interface and is powered by PoE; video data is transmitted using the Real-Time Streaming Protocol (RTSP), and hardware acceleration decoding is performed using the FFmpeg library; all cameras start acquisition uniformly through a synchronous trigger signal to ensure frame-level synchronization.

[0030] The foot pressure imaging unit specifically includes (1) Imaging principle, for example, using internal light source refraction imaging technology: LED light source is installed on the side of the platform. When light travels from a high refractive index medium (glass, refractive index 1.5) to a low refractive index medium (air, refractive index 1.0), total internal reflection occurs, and the platform surface presents a bright background. When the sole of the foot contacts the glass, the conditions for total internal reflection are destroyed, the contact area forms a dark area, and the non-contact area remains bright, thereby accurately capturing the pressure information of the sole of the foot. The greater the pressure, the larger the contact area, and the lower the gray value of the dark area. (2) Hardware configuration, for example, light source: LED light strip is installed around the platform, power 30W, color temperature 6500K, uniformity ≥0.9; foot camera: Basler acA1300-60gm, resolution 1280×1024, frame rate 60fps, global shutter; camera installation: directly below the platform base, shooting vertically upwards, with a polarizing filter added to the lens to eliminate glass reflection; protection: the camera is installed in a sealed protective box, dustproof and waterproof rating IP65. (3) Image processing algorithms, such as color space conversion: RGB→HSV, based on the threshold segmentation of hue H (40-80), saturation S (50-255), and brightness V (50-255) to segment the footprint area. Morphological processing: 3×3 kernel erosion once to remove isolated noise points, 3×3 kernel dilation twice to fill holes. Gray-scale conversion and pseudo-color mapping: the footprint area is converted to a grayscale image (0-255), and pseudo-color mapping is performed using Colormap Jet. The lower the grayscale value (the greater the pressure), the red is mapped, and the higher the grayscale value (the less the pressure), the blue is mapped. Image weighting center: calculate the sum of all pixel values ​​of the grayscale image P_total=ΣI(i,j), and the weighting center coordinates (x_g,y_g)=(Σi·I(i,j) / P_total,Σj·I(i,j) / P_total). Foot Zoning: The footprint is divided into three equal parts based on the minimum bounding rectangle. Perpendicular lines are drawn from the trisection points along the long side of the rectangle to divide the foot into three zones: forefoot (A, 0-33%), midfoot (B, 33-67%), and heelfoot (C, 67-100%). Key Parameter Calculations: Foot Length FL: Length of the long side of the minimum bounding rectangle; Maximum Foot Width MFW: Length of the short side of the minimum bounding rectangle; Arch Index AI = Area of ​​midfoot / (Total area of ​​forefoot + midfoot + heelfoot); Plantar Support Area BoS = Area of ​​zone A + Area of ​​zone B + Area of ​​zone C; Pressure Center Offset = Euclidean distance between the current pressure center and the reference center.

[0031] The EEG acquisition module utilizes commercially available, mature dry electrode EEG acquisition equipment, integrated with the system via a standard interface. Hardware configuration: A commercially available 8-channel dry electrode EEG headband is used, with electrode placement covering the prefrontal, central, and parietal lobes according to the international 10-20 system; sampling rate ≥500Hz, common-mode rejection ratio ≥100dB; data is transmitted wirelessly via Bluetooth or Wi-Fi. Signal interface: The EEG equipment provides a standard data output interface (LSL protocol or TCP / IP), through which the edge computing unit acquires raw EEG data in real time and performs subsequent processing. Integration method: The system sends TTL trigger pulses to the EEG equipment via a synchronization controller, achieving millisecond-level synchronous acquisition with other modalities (time deviation ≤2ms).

[0032] Step S30: The edge computing unit performs three-dimensional pose reconstruction based on the multi-view image to obtain pose features containing the three-dimensional coordinate sequence of key points of the whole skeleton. The plantar pressure distribution image is processed to obtain mechanical features including the plantar pressure center trajectory, arch index and plantar support area. The electroencephalogram (EEG) signal is analyzed in the frequency domain to obtain EEG features including the power and ratio of each channel.

[0033] In its specific implementation, the three-dimensional pose reconstruction process includes performing two-dimensional pose estimation on the multi-view images, outputting the two-dimensional coordinates and confidence scores of several skeletal key points under each viewpoint; for each skeletal key point, a hierarchical processing is performed based on the confidence scores, wherein the hierarchical processing is as follows: if the confidence scores of two or more viewpoints are greater than 0.8, the three-dimensional coordinates are solved by direct linear transformation, and weighted least squares optimization is performed with confidence scores as weights; if the confidence score of one viewpoint is greater than 0.8, the two-dimensional point of that viewpoint is back-projected into a three-dimensional ray, and the optimal three-dimensional point is searched on the ray in combination with the bone length constraints of adjacent reconstructed key points; if the confidence scores of all viewpoints are less than or equal to 0.8, the corresponding skeletal key points are marked as missing and interpolated; the reconstructed three-dimensional key point coordinate sequence is temporally smoothed using Kalman filtering, and the two ends of the bones are fine-tuned and corrected using a preset range of human bone length to obtain pose features containing the three-dimensional coordinate sequence of the whole body skeletal key points.

[0034] It should be noted that, in the specific implementation of 3D pose reconstruction, for example (1) single-view 2D pose estimation, a high-resolution network (HRNet-W32) is used as the backbone network. Based on the pre-training of the COCO dataset, a self-made dataset is used for fine-tuning. The self-made dataset needs to cover different lighting, clothing, background and occlusion conditions, and contain at least 10,000 labeled images. The key points are expanded to 25 (COCO 17 points + heel, toe, etc.) to ensure label consistency. The input image size is uniformly 512×512 pixels, and the output is the 2D coordinates and confidence of 25 key points (normalized by softmax). During inference, soft-argmax is used to extract sub-pixel coordinates from the heatmap to improve accuracy. On the Jetson AGX Orin platform, through TensorRT optimization, the single-frame inference speed can be stabilized at ≤20ms. (2) Multi-view triangulation: The system is equipped with at least 4 industrial cameras, evenly distributed around a 2m×2m×2m test space to ensure a moderate baseline (0.5-2m) and avoid degenerate configurations where the optical center is collinear with the target. All cameras are rigorously calibrated, and their intrinsic parameters, distortion coefficients, and extrinsic parameters (rotation matrix, translation vector) are known, with a reprojection error ≤ 0.5 pixels. For each key point, confidence-based hierarchical processing is performed: If the confidence of ≥ 2 views is > 0.8: Direct linear transformation (DLT) is used to solve for the 3D position, and weighted least squares optimization is performed with confidence as the weight to reduce the impact of low-quality views. If the confidence of only 1 view is > 0.8: The 2D point is back-projected into a 3D ray, and the optimal 3D point is searched on the ray by combining the bone length constraint (using adjacent reconstructed key points). At the same time, the current position is predicted using Kalman filtering as a priori, and an optimization objective function is constructed to solve for the 3D point that minimizes both the projection error and the length constraint. If the confidence of all viewpoints is ≤0.8: mark the missing point, do not rebuild it for the time being, and use subsequent temporal filtering combined with motion model for interpolation. For dual-view cases, if the angle between the rays of the two viewpoints is too small (<15°), the uncertainty of the triangulation depth is large. At this time, switch to single-view processing strategy and introduce skeletal constraints. (3) Temporal filtering and skeletal constraints: use Kalman filtering to smooth the trajectory of the three-dimensional key points. The state vector is position (x,y,z) and velocity (vx,vy,vz). The state transition adopts a uniform velocity model. The observation noise covariance matrix is ​​dynamically adjusted according to the confidence weight and the number of viewpoints during triangulation (small noise is given when the confidence is high and there are multiple viewpoints). For the marked missing keypoints, Kalman filtering is used to predict and fill in the missing points based on the motion model. The skeletal length constraint is based on the prior of human anatomy and establishes a connection relationship matrix (e.g., the average length of the left shoulder to the left elbow is 30cm, allowing ±5% individual difference).After each frame is reconstructed, all bone connections are traversed. If the length of a bone exceeds the preset range, the two ends of the bone are fine-tuned: the two ends are moved towards or away from each other along the bone direction to restore the length to the range, while keeping the overall center of gravity unchanged or minimizing the position adjustment (using gradient descent or direct projection). This correction step is performed after the Kalman filter update as post-processing. (4) Performance indicators and delay analysis, accuracy: In a 2m×2m×2m test space, a high-precision motion capture system (such as OptiTrack) is used to verify that the average error of key point position is ≤15mm and the bone length error is ≤5%. Delay: The single frame processing flow adopts pipeline parallelism: 2D inference (20ms), triangulation and confidence grading (≤5ms), Kalman filtering and bone constraint (≤5ms), total delay ≤30ms. After reserving data transmission and system overhead, the ≤50ms index can be met. Real-time operation is ensured through multi-threading and CUDA acceleration.

[0035] Furthermore, the process of processing the plantar pressure distribution image specifically includes: converting the plantar pressure distribution image to HSV color space; extracting the footprint region image based on hue, saturation, and brightness thresholds; performing morphological processing on the footprint region image to remove noise and fill holes, and converting it to a grayscale image, where the grayscale value in the grayscale image represents the pressure magnitude; calculating the image weighted centroid based on the grayscale image, using the image weighted centroid as the pressure centroid coordinate at the current moment, and recording the pressure centroid coordinates over time to form the plantar pressure center trajectory; dividing the footprint into the forefoot, midfoot, and posterior foot regions based on the minimum bounding rectangle of the footprint, calculating the arch index and plantar support area for each region, and outputting the mechanical features, where the arch index is calculated from the midfoot area and the total footprint area, and the plantar support area is the total footprint area. The HSV color space, also known as the hue, saturation, and value color space, uses hue (H) to represent the color category (e.g., red, green, blue), saturation (S) to represent the vividness of the color, and value (V) to represent the lightness or darkness of the color. For example, when converting a plantar pressure distribution image to HSV color space, the footprint region is extracted based on thresholding hue (H) (40-80), saturation (S) (50-255), and value (V) (50-255). A single 3×3 kernel erosion is used to remove isolated noise, followed by two 3×3 kernel dilation operations to fill holes. The footprint region is then converted to a grayscale image (0-255), and pseudo-color mapping is performed using ColormapJet (lower grayscale values ​​correspond to higher pressure, mapped to red; higher grayscale values ​​correspond to blue). Based on the grayscale values... The weighted centroid of the image is calculated as the pressure centroid coordinate at the current moment. The pressure centroid coordinates are recorded over time to form the plantar pressure center trajectory (x, y coordinates and velocity / acceleration sequence of the CoP trajectory). Based on the minimum bounding rectangle of the footprint, the footprint is divided into three equal areas: the forefoot (0-33%), the midfoot (33-67%), and the heel (67-100%). The arch index is calculated as: midfoot area / (total area of ​​forefoot + midfoot + heel). The plantar support area is calculated as: total footprint area. The mechanical characteristics are output (time window T = 100 points, corresponding to 0.1 seconds, sampled at 1000Hz and downsampled to 100Hz).

[0036] Furthermore, the extraction of EEG features specifically includes: bandpass filtering of the raw EEG signal, extracting the power spectral density of the θ band (4-8Hz), α band (8-13Hz), and β band (13-30Hz) respectively, and calculating the θ, α, and β power and the θ / β ratio of each channel. A sliding window (window length 1 second, extraction every 50ms, 50% overlap) is used to obtain a sequence of 20 time points, with dimensions of 20 (time points) × 8 (channels) × 4 (features).

[0037] Step S40: Input the posture features, mechanical features, EEG features, and task parameters into the cloud-based AI large model platform and output the evaluation results. The cloud-based AI large model platform is based on the Transformer architecture and uses the task parameters as a query to fuse features of various modalities through a cross-attention mechanism. The evaluation results include a comprehensive balance function score, vertigo subtype classification probability, contribution weight, and dynamic warning of fall risk.

[0038] In specific implementation, the process of inputting the posture features, mechanical features, EEG features, and task parameters into a cloud-based AI large-scale model platform and outputting evaluation results specifically includes: inputting the posture features into a posture encoder composed of a one-dimensional convolutional neural network (1D-CNN) and a Transformer encoder to generate a posture feature vector; inputting the mechanical features into a mechanical encoder composed of a long short-term memory network (LSTM) and attention pooling to generate a mechanical feature vector; inputting the EEG features into an EEG encoder composed of a 1D-CNN and a Transformer encoder to generate an EEG feature vector; and inputting the task parameters into a task encoder composed of a multilayer sensor to generate conditional features. Vector; Input the posture feature vector, the mechanical feature vector, the EEG feature vector and the conditional feature vector into the multimodal Transformer, and output the fusion feature through the cross-attention mechanism with the conditional feature vector as the query and the posture feature vector, the mechanical feature vector and the EEG feature vector as the key and value; Input the fusion feature into the regression head and the classification head respectively, and output the balance function comprehensive score and the vertigo subtype classification probability. At the same time, input the fusion feature into the attention mechanism layer to dynamically generate the vestibular contribution weight, visual contribution weight and proprioceptive contribution weight; Input the fusion feature into the temporal prediction model to output the fall risk probability as a dynamic warning of fall risk. Specifically, (1) Modal encoding: Input the posture feature (100×25×3) into the posture encoder composed of 3 layers of 1D-CNN (convolution kernel size 5, stride 2, number of channels 64→128→256) plus position encoding and then the Transformer encoder layer (4 heads, hidden layer 256) to generate the posture feature vector (256 dimensions). The mechanical features (100×5) are input into a mechanical encoder consisting of 2 layers of LSTM (128 hidden layers) plus attention pooling to generate a mechanical feature vector (128 dimensions). The EEG features (20×8×4) are input into an EEG encoder consisting of 2 layers of 1D-CNN (3 kernel size, 1 stride, 32→64 channels) plus Transformer encoder layers (2 heads, 128 hidden layers) to generate an EEG feature vector (128 dimensions). The task parameters (5 types of visual scene one-hot encoding, 2-dimensional platform motion parameters) are input into a task encoder consisting of 2 layers of MLP (256→128) to generate a conditional feature vector (128 dimensions). (2) Cross-modal fusion: The pose, mechanical, and EEG feature vectors are concatenated and input together with the conditional feature vector into a multimodal Transformer (6 layers, 8 heads, 512 hidden layers). Through the cross-attention mechanism, the conditional feature vector is used as the query and other modal feature vectors are used as the key and value to output the fused feature (512 dimensions).(3) Multi-task output: The fused features are input into the regression head (2-layer MLP: 256→128→1) and the classification head (2-layer MLP: 256→128→4) respectively, and the output balance function comprehensive score (0-100) and vertigo subtype classification probability (Meniere's disease / vestibular neuritis / PPPD / central vertigo). At the same time, the fused features are input into the attention mechanism layer to dynamically generate vestibular contribution weight, visual contribution weight and proprioceptive contribution weight (3-dimensional, Softmax normalized). The fused features are input into the temporal prediction model (transformer-based temporal prediction network) to output the fall risk probability in the next 48 hours as a dynamic warning of fall risk.

[0039] In this embodiment, the model employs multi-task joint training. The loss function includes MSE loss (difference between overall score and expert score), cross-entropy loss (subtype classification), and KL divergence loss (consistency between perceived contribution weight and expert annotation). The training data is a self-built dataset (400 patients + 200 healthy controls), using 5-fold cross-validation, with AdamW as the optimizer and a learning rate of 1e-4. Testing showed that the regression MAE ≤ 5 points and the classification accuracy ≥ 85%.

[0040] It should be noted that the multimodal fusion evaluation model specifically includes the following in its implementation: (1) Input features: posture features: a three-dimensional coordinate sequence of 25 key points, time window T=100 frames (corresponding to about 1.67 seconds, 60fps), dimension: 100×25×3; mechanical features: x, y coordinates and velocity and acceleration sequences of CoP trajectory, time window T=100 points (corresponding to 0.1 seconds, downsampled to 100Hz after sampling at 1000Hz), dimension: 100×5; EEG features: θ, α, β power and θ / β ratio sequence of each channel, time window T=10 points (corresponding to 1 second, features are calculated once every 100ms), dimension: 10×8×4; task parameters: visual scene type (5 categories, one-hot encoding), platform motion parameters (tilt angular velocity, translation velocity, normalized to 0-1). (2) Model architecture, ① Modal encoder: Pose encoder: 3-layer 1D-CNN (convolution kernel size 5, stride 2, number of channels 64→128→256) + position encoding + Transformer encoder layer (4 heads, hidden layer 256); Mechanics encoder: 2-layer LSTM (hidden layer 128) + attention pooling; EEG encoder: 2-layer 1D-CNN (convolution kernel size 3, stride 1, number of channels 32→64) + Transformer encoder layer (2 heads, hidden layer 128); Task encoder: 2-layer MLP (256→128).

[0041] ② Cross-modal fusion: The feature vectors output by each modal encoder are concatenated and input into a multimodal Transformer (6 layers, 8 heads, 512 hidden layers); a cross-attention mechanism is adopted, using task features as queries and other modal features as keys and values. ③ Output heads: Regression head: 2-layer MLP (256→128→1), outputting a comprehensive balance function score (0-100); Classification head: 2-layer MLP (256→128→4), outputting the classification probability of vertigo subtypes (Meniere's disease / vestibular neuritis / PPPD / central vertigo); Sensory system contribution head: 2-layer MLP (256→128→3), Softmax normalization, outputting vestibular / visual / proprioceptive contribution weights. (3) Training method, training data: self-built dataset (400 patients + 200 healthy controls), data augmentation including temporal perturbation, adding noise, and random occlusion; loss function: L_total=L_reg+0.5×L_cls+0.3×L_weight; L_reg: MSE loss (difference between comprehensive score and expert score); L_cls: cross-entropy loss (classification of vertigo subtypes); L_weight: KL divergence loss (consistency between the contribution weight of the sensory system and the expert annotation); optimizer: AdamW, learning rate 1e-4, weight decay 1e-5, batch size 32; validation index: regression MAE≤5, classification accuracy≥85%.

[0042] Furthermore, the dynamic early warning model involved in this embodiment includes, for example, (1) input features. To ensure time alignment of each modality, the time window length is uniformly set to 1 second: Posture features: 3D coordinate sequence of 25 key points, sampling rate 60Hz, taking 60 consecutive frames, dimension: 60×25×3. Mechanical features: x, y coordinates and velocity and acceleration sequence of CoP trajectory, downsampled from the original 1000Hz to 100Hz, taking 100 points, dimension: 100×5. EEG features: θ, α, β power and θ / β ratio of each channel, extracted once every 50ms using a sliding window (50% overlap), to obtain 20 time points, dimension: 20×8×4 (8 channels, 4 features). Task parameters: visual scene type (5 categories, one-hot encoding), platform motion parameters (tilt angular velocity, translation velocity, normalized to 0-1), as conditional inputs. All features are synchronously extracted based on the start time of the experiment, and the alignment error is ensured to be <2ms through subsequent timestamp verification. (2) Model Architecture ① Modal Encoder (Unified Output Feature Dimension 128): Pose Encoder: 2-layer Temporal Convolutional Network (TCN, kernel size 5, dilation factor 1, 2, channels 64→128) + Global Average Pooling, outputting a 128-dimensional vector. It can be pre-trained on a large-scale dataset such as Human3.6M and then fine-tuned. Mechanics Encoder: 2-layer LSTM (128 hidden layers), taking the output of the last time step, and obtaining a 128-dimensional vector through attention pooling. EEG Encoder: 2-layer 1D-CNN (3 kernels, stride 1, channels 32→64), followed by global average pooling, outputting a 128-dimensional vector. Task Encoder: 2-layer MLP (256→128), outputting 128-dimensional conditional features. ② Cross-modal Fusion: The four 128-dimensional feature vectors are concatenated to obtain a 512-dimensional joint representation. To avoid overfitting under small sample size, the fusion structure is simplified: a 2-layer MLP (512→256→128) is used for feature interaction, with each layer followed by LayerNorm and Dropout (0.3). At the same time, the cross-attention mechanism is retained as an extension option (which can be replaced if there is sufficient data). ③ Multi-task output head: Regression head: 2-layer MLP (128→64→1), outputting a comprehensive score of balance function (0-100). Classification head: 2-layer MLP (128→64→4), with Softmax outputting the probability of vertigo subtype. Sensory system contribution head: The contribution weights of vestibular / visual / proprioceptive senses are dynamically generated from the fusion features through a self-attention mechanism (3-dimensional, Softmax normalized), without expert annotation, and the implicit learning is jointly optimized by the model. (3) Training method, dataset: self-built dataset (400 patients + 200 healthy controls), with strict evaluation using 5-fold cross-validation. To address the small sample size problem, transfer learning is introduced: the pose encoder is initialized using Human3.6M pre-trained parameters, and the EEG encoder is pre-trained using a publicly available emotion recognition dataset (such as DEAP).Data augmentation: Temporal perturbations (random cropping / scaling), adding Gaussian noise, randomly occluding keypoints (pose), and discarding channels (EEG) are used to generate augmented samples online. Loss function: Multi-task joint optimization, Ltotal = Lreg + 0.5Lcls + 0.2Lattend, where Lreg is the MSE loss (difference between overall score and expert score), Lcls is the cross-entropy loss (subtype classification), and Lattend is the auxiliary attention regularization term (encouraging smooth attention distribution across sensory channels, unsupervised). Expert scores require inter-rater consistency testing (ICC > 0.8). Optimizer: AdamW, learning rate 1e-4, weight decay 1e-5, batch size 32, training epochs 100, early stopping mechanism (stopping if the validation loss does not decrease after 10 epochs). Validation metrics: Regression MAE ≤ 5, classification accuracy ≥ 80% (expectations are adjusted appropriately considering data scale), and a weighted F1-score and confusion matrix are reported.

[0043] In one embodiment, this embodiment generates a personalized rehabilitation plan containing a training program list, exercise prescription, and safety warnings by matching a preset mapping rule base with the balance function comprehensive score, vertigo subtype classification probability, contribution weight, ratio contained in the EEG features, and visual dependence index determined based on the type of visual stimulation task. After completing periodic assessments, an updated balance function comprehensive score is obtained, and the updated balance function comprehensive score is compared with the historical balance function comprehensive score. Based on the comparison results, the training intensity of the current rehabilitation plan is maintained or downgraded, and the training frequency is reduced, the training intensity is upgraded, or biofeedback assistance is increased or the training intensity is downgraded, and a therapist's manual review and warning are triggered.

[0044] It should be noted that the mapping rule base includes at least one of the following rules: when the visual dependence index is greater than 0.6, matching is performed with adaptive training containing VOR×1 / VOR×2; when the vestibular contribution weight is less than 20% and the VOR gain calculated from the posture features is less than 0.7, matching is performed with adaptive and alternative training containing VOR×1 and saccades or tracking; when the dynamic warning of fall risk is greater than 70%, matching is performed with balance-gait training containing center of gravity shift, standing on a foam mat, and dual-task walking; when the proprioceptive contribution weight is less than 15%, matching is performed with sensory alternative training containing standing with eyes closed, soft support surfaces, and dynamic platforms; when the ratio contained in the EEG features is greater than 2.5, matching is performed with habituation and cognitive-balance dual-task training containing induced repetitive movements and cognitive interference. The specific rules mentioned above can be found in Table 1 below, which provides examples of the rules in the rule base. The vestibulo-ocular reflex (VOR) is trained with VOR×1, which is a 1x gain training method. The goal is to increase the gain (output / input ratio) of the vestibulo-ocular reflex, primarily targeting subjects with low VOR gain (gain less than 0.7). VOR×2, or a 2x gain training method, aims not only to train VOR but, more importantly, to force the brain to integrate vestibular and visual signals, strengthening the visual-vestibular interaction. This is used in scenarios with more severe vestibular dysfunction or requiring greater visual-dependent compensation.

[0045] Table 1:

[0046] Based on the multimodal assessment results and matching rule base, a personalized rehabilitation plan conforming to the International Classification of Functioning, Disability and Health (ICF) framework is generated, including the following core elements: ① A list of training items, item name: such as "VOR×1 gaze stability training", action description: following standardized terminology, such as "sitting position, looking straight ahead at a stationary target on the wall, shaking your head horizontally from side to side (frequency about 1Hz), keeping the target clear, for 1 minute", action demonstration: high-definition video demonstration (studies have confirmed that video guidance is superior to text manuals), training mechanism: indicating the mechanism targeted by the training (adaptation / substitution / habituation). ② Exercise prescription; Frequency: X times daily (refer to RCT studies, twice daily adherence is optimal); Duration: Y minutes per session (start with tolerable levels and gradually increase); Intensity levels: Beginner: Seated, slow, single task; Intermediate: Standing, medium speed, simple dual task; Advanced: Dynamic support surface, fast, complex dual task; ③ Safety warnings and precautions; Absolute contraindications (e.g., cervicogenic vertigo, unstable cardiovascular and cerebrovascular diseases); Relative contraindications and coping strategies; Home environment modification suggestions (e.g., adding fall prevention measures when R003 is triggered). ④ Expected goals: Short-term goal (2 weeks): e.g., "20% reduction in visual dependence index"; Medium-term goal (6 weeks): e.g., "≥3 points improvement in Dynamic Gait Index (DGI)" (based on MCID). Long-term goal (12 weeks): such as "a reduction of ≥18 points in the total score of the Vertigo Disorder Scale (DHI)".

[0047] It should be noted that the following criteria apply to the determination of improvement: If the target indicator improves beyond the minimum clinically significant difference (MCID) (e.g., DHI ≥ 18 points, DGI ≥ 3 points), then: maintain or reduce the training intensity (depending on whether the target threshold has been reached), reduce the training frequency (e.g., from twice a day to once a day), or switch to the next stage of training. If there is no change, and the indicator changes within the measurement error range, then: increase the training intensity (e.g., from sitting to standing), and increase auxiliary feedback (e.g., visual / auditory biofeedback). If the indicator deteriorates beyond the MCID or new symptoms appear, then: reduce the training intensity to a safe level and trigger a manual review warning (the system pushes a notification to the therapist).

[0048] In one embodiment, the daily training check-in rate of the subject for the personalized rehabilitation program can also be recorded through the user terminal. When the daily training check-in rate is less than 70%, a reminder is automatically pushed to the user terminal and the difficulty of the movements in the training program is simplified. A rehabilitation summary report is automatically generated periodically. The report includes the change curve of the comprehensive balance function score, the gap with the expected goal, the explanation of the current program adjustment, and the training suggestions for the next stage.

[0049] In this embodiment, a synchronous controller controls a curved display screen to show visual stimulation tasks and controls a programmable 3D force stage to apply time-synchronized physical perturbations, recording task parameters. A synchronous trigger unit controls a multi-view camera array, a plantar pressure imaging unit, and an EEG acquisition module to synchronously acquire multimodal data. An edge computing unit extracts posture features, mechanical features, and EEG features based on the multimodal data. These features and task parameters are input into a cloud-based AI large-scale model platform based on a multimodal Transformer architecture, outputting evaluation results. This approach achieves multimodal quantitative evaluation of the entire balance control chain and can generate personalized rehabilitation plans, improving the ecological validity of balance function assessment.

[0050] Reference Figure 5 , Figure 5 This is a structural block diagram of the first embodiment of the multimodal fusion balanced function evaluation system of the present invention.

[0051] like Figure 5 As shown, the multimodal fusion-based balanced function evaluation system proposed in this embodiment of the invention includes: The immersive visual physical perturbation subsystem 10 is used to control the curved display screen to display the visual stimulus task through a synchronization controller, and to control the programmable three-dimensional force stage to apply physical perturbations that are time-synchronized with the visual stimulus task, and to record the type of the visual stimulus task and the motion parameters of the physical perturbation as task parameters. The multimodal biosignal acquisition subsystem 20 is used to control a multi-view camera array to acquire multi-view images, control a plantar pressure imaging unit to acquire plantar pressure distribution images, and control an EEG acquisition module to acquire EEG signals during the process of a disturbed environment acting on the subject. The brightness change of a preset area on the curved display screen is detected by a photodiode as a time reference for the visual stimulus task display, ensuring that the time of each modal data is aligned with the task parameters. The modal data includes the multi-view images, plantar pressure distribution images, and EEG signals. The intelligent fusion analysis subsystem 30 is used to perform three-dimensional posture reconstruction based on the multi-view images through the edge computing unit to obtain posture features including the three-dimensional coordinate sequence of key points of the whole skeleton, to perform image processing on the plantar pressure distribution image to obtain mechanical features including the plantar pressure center trajectory, arch index and plantar support area, and to perform frequency domain analysis on the EEG signal to obtain EEG features including the power and ratio of each channel. The intelligent fusion analysis subsystem 30 is used to input the posture features, mechanical features, EEG features, and task parameters into the cloud-based AI large model platform and output evaluation results. The cloud-based AI large model platform is based on the Transformer architecture and uses the task parameters as a query to fuse features of various modalities through a cross-attention mechanism. The evaluation results include a comprehensive balance function score, vertigo subtype classification probability, sensory contribution weight, and dynamic warning of fall risk. The contribution weight includes vestibular contribution weight, visual contribution weight, and proprioceptive contribution weight.

[0052] Furthermore, the programmable three-dimensional force stage includes: a base, a balancing force stage, a linear guide module, a servo drive unit, and four three-dimensional force sensors; the four three-dimensional force sensors are respectively installed at the four corners of the bottom of the balancing force stage to collect the force values ​​at each corner in real time; the edge computing unit calculates the coordinates of the plantar pressure center based on the force values; the synchronization controller receives the coordinates of the plantar pressure center and dynamically adjusts the parameters of the subsequent physical disturbances applied by the programmable three-dimensional force stage according to the coordinates of the plantar pressure center.

[0053] In this embodiment, a synchronous controller controls a curved display screen to show visual stimulation tasks and controls a programmable 3D force stage to apply time-synchronized physical perturbations, recording task parameters. A synchronous trigger unit controls a multi-view camera array, a plantar pressure imaging unit, and an EEG acquisition module to synchronously acquire multimodal data. An edge computing unit extracts posture features, mechanical features, and EEG features based on the multimodal data. These features and task parameters are input into a cloud-based AI large-scale model platform based on a multimodal Transformer architecture, outputting evaluation results. This approach achieves multimodal quantitative evaluation of the entire balance control chain and can generate personalized rehabilitation plans, improving the ecological validity of balance function assessment.

[0054] This application embodiment also provides a multimodal fusion balance function evaluation device, including a processor, a communication interface, a memory, and a communication bus. The processor, communication interface, and memory communicate with each other through the communication bus. The memory is used to store the multimodal fusion balance function evaluation program. When the processor executes the program stored in the memory, it implements the above-mentioned multimodal fusion balance function evaluation method.

[0055] The communication bus mentioned in the aforementioned multimodal fusion balanced function evaluation device can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc.

[0056] The communication interface is used for communication between the aforementioned multimodal fusion balance function evaluation device and other devices.

[0057] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0058] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0059] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)).

[0060] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0061] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0062] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

[0063] It should be understood that the above are merely illustrative examples and do not constitute any limitation on the technical solutions of the present invention. In specific applications, those skilled in the art can make settings as needed, and the present invention does not impose any restrictions on this.

[0064] It should be noted that the workflow described above is merely illustrative and does not limit the scope of protection of this invention. In practical applications, those skilled in the art can select some or all of the workflow to achieve the purpose of this embodiment according to actual needs, and no restrictions are imposed here.

[0065] In addition, for technical details not described in detail in this embodiment, please refer to the multimodal fusion balanced function evaluation method provided in any embodiment of the present invention, which will not be repeated here.

[0066] Furthermore, it should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.

[0067] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0068] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as read-only memory (ROM) / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0069] The above are merely preferred embodiments of the present invention and do not limit the scope of the patent. Any equivalent structural or procedural transformations made based on the description and drawings of the present invention, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of the present invention.

[0070] It is understood that the system provided in the embodiments of the present invention corresponds to the method provided in the embodiments of the present invention, and the explanation, examples and beneficial effects of the relevant content can be referred to the corresponding parts of the above methods.

Claims

1. A multimodal fusion method for evaluating balanced function, characterized in that, The multimodal fusion balance function evaluation method includes: The system controls the curved display screen to display the visual stimulus task through a synchronous controller, and controls the programmable three-dimensional force stage to apply physical perturbations that are time-synchronized with the visual stimulus task. The type of the visual stimulus task and the motion parameters of the physical perturbations are recorded as task parameters. During the process of the perturbed environment acting on the subject, the multi-view camera array is controlled by the synchronous triggering unit to acquire multi-view images, the plantar pressure imaging unit is controlled to acquire plantar pressure distribution images, and the EEG acquisition module is controlled to acquire EEG signals. Among them, the brightness change of the preset area of ​​the curved display screen is detected by the photodiode as a time reference for the visual stimulus task display, so as to ensure that the time of each modal data is aligned with the task parameters. The modal data includes the multi-view images, plantar pressure distribution images, and EEG signals. The edge computing unit performs three-dimensional pose reconstruction based on the multi-view images to obtain pose features containing the three-dimensional coordinate sequence of key points of the whole skeleton. The plantar pressure distribution image is processed to obtain mechanical features including the plantar pressure center trajectory, arch index and plantar support area. The electroencephalogram (EEG) signal is analyzed in the frequency domain to obtain EEG features including the power and ratio of each channel. The posture features, mechanical features, EEG features, and task parameters are input into a cloud-based AI large-scale model platform, and the evaluation results are output. The cloud-based AI large-scale model platform is based on the Transformer architecture and uses the task parameters as a query to fuse features of various modalities through a cross-attention mechanism. The evaluation results include a comprehensive balance function score, vertigo subtype classification probability, contribution weight, and dynamic warning of fall risk. The contribution weight includes vestibular contribution weight, visual contribution weight, and proprioceptive contribution weight. The three-dimensional pose reconstruction based on the multi-view images yields pose features containing a sequence of three-dimensional coordinates of key points on the whole skeleton, including: Perform two-dimensional pose estimation on the multi-view images and output the two-dimensional coordinates and confidence scores of several skeletal key points in each view. For each skeletal keypoint, a hierarchical processing is performed based on the confidence level. The hierarchical processing is as follows: if the confidence level of two or more viewpoints is greater than 0.8, the three-dimensional coordinates are solved by direct linear transformation, and weighted least squares optimization is performed with the confidence level as the weight. If the confidence level of one viewpoint is greater than 0.8, the two-dimensional point of that viewpoint is back-projected into a three-dimensional ray, and the optimal three-dimensional point is searched on the ray in combination with the bone length constraints of adjacent reconstructed keypoints. If the confidence level of all viewpoints is less than or equal to 0.8, the corresponding skeletal keypoints are marked as missing and interpolation is performed. Kalman filtering is used to perform temporal smoothing on the reconstructed 3D keypoint coordinate sequence, and the two ends of the skeleton are fine-tuned and corrected using a preset range of human skeleton length to obtain the pose features containing the 3D coordinate sequence of keypoints of the whole skeleton.

2. The multimodal fusion-based balanced function evaluation method as described in claim 1, characterized in that, The image processing of the plantar pressure distribution image yields mechanical features including the plantar pressure center trajectory, arch index, and plantar support area, including: The plantar pressure distribution image is converted to HSV color space, and the footprint region image is extracted based on hue, saturation, and brightness thresholds. The footprint area image is subjected to morphological processing to remove noise and fill holes, and then converted into a grayscale image, wherein the grayscale value in the grayscale image represents the pressure level; The image weighted center is calculated based on the grayscale image, and the image weighted center is used as the pressure center coordinate at the current moment. The pressure center coordinates on the time series are recorded to form the trajectory of the plantar pressure center. The footprint is divided into three equal regions: the forefoot region, the midfoot region, and the heel region, based on the minimum bounding rectangle of the footprint. The arch index and the plantar support area are calculated for each region, and the biomechanical characteristics are output. The arch index is calculated from the area of ​​the midfoot region and the total area of ​​the footprint, and the plantar support area is the total area of ​​the footprint.

3. The multimodal fusion-based balanced function evaluation method as described in claim 1, characterized in that, The process of inputting the posture features, mechanical features, EEG features, and task parameters into a cloud-based AI large-scale model platform and outputting evaluation results includes: The pose features are input into a pose encoder consisting of a one-dimensional convolutional neural network (1D-CNN) and a Transformer encoder to generate a pose feature vector. The mechanical features are input into a mechanical encoder consisting of a long short-term memory network (LSTM) and attention pooling to generate a mechanical feature vector. The EEG features are input into an EEG encoder consisting of a 1D-CNN and a Transformer encoder to generate an EEG feature vector. The task parameters are input into a task encoder composed of a multilayer perceptron to generate a conditional feature vector. The posture feature vector, the mechanical feature vector, the EEG feature vector, and the conditional feature vector are input into a multimodal Transformer. Through a cross-attention mechanism, the conditional feature vector is used as the query, and the posture feature vector, the mechanical feature vector, and the EEG feature vector are used as the key and value, respectively, to output fused features. The fusion features are input into the regression head and the classification head respectively, and the balance function comprehensive score and the vertigo subtype classification probability are output. At the same time, the fusion features are input into the attention mechanism layer to dynamically generate vestibular contribution weight, visual contribution weight and proprioceptive contribution weight. The fused features are input into the time series prediction model, and the output fall risk probability is used as a dynamic early warning of fall risk.

4. The multimodal fusion-based balanced function evaluation method as described in claim 1, characterized in that, The method further includes: Based on the comprehensive balance function score, the probability of vertigo subtype classification, the contribution weight, the ratio contained in the EEG features, and the visual dependence index determined based on the type of visual stimulation task, a preset mapping rule base is matched to generate a personalized rehabilitation plan that includes a list of training items, exercise prescriptions, and safety warnings. After completing the periodic assessment, an updated balance function comprehensive score is obtained. The updated balance function comprehensive score is compared with the historical balance function comprehensive score. Based on the comparison results, the training intensity of the current rehabilitation program is maintained or downgraded, the training frequency is reduced, the training intensity is upgraded, or biofeedback assistance is increased or the training intensity is downgraded, and a manual review warning by the therapist is triggered.

5. The multimodal fusion-based balanced function evaluation method as described in claim 4, characterized in that, The mapping rule base includes at least one of the following rules: When the visual dependence index is greater than 0.6, the matching includes adaptive training with VOR×1 / VOR×2; When the vestibular contribution weight is less than 20% and the VOR gain calculated from the pose features is less than 0.7, the matching includes adaptation and alternative training with VOR×1 and saccades or tracking. When the dynamic warning of fall risk is greater than 70%, balance-gait training including weight transfer, standing on a foam mat, and dual-task walking is matched. When the proprioceptive contribution weight is less than 15%, sensory substitution training is matched, including standing with eyes closed, soft support surfaces, and dynamic platforms. When the ratio of the EEG features is greater than 2.5, the matching includes habituation and cognitive-balance dual-task training that includes inducing repetitive actions and cognitive interference.

6. The multimodal fusion-based balanced function evaluation method as described in claim 4, characterized in that, The method further includes: The system records the daily training check-in rate of the subject for the personalized rehabilitation plan through the user terminal. When the daily training check-in rate is less than 70%, a reminder is automatically pushed to the user terminal and the difficulty of the movements in the training plan is simplified. A rehabilitation summary report is automatically generated periodically. The report includes the change curve of the comprehensive balance function score, the gap with the expected goal, the explanation of the adjustment of the plan, and the training suggestions for the next stage.

7. A multimodal fusion-based balanced function assessment system, characterized in that, The multimodal fusion balance function evaluation system is applied to the multimodal fusion balance function evaluation method as described in any one of claims 1 to 6, the system comprising: An immersive visual physical perturbation subsystem is used to control a curved display screen to display a visual stimulus task through a synchronization controller, and to control a programmable three-dimensional force stage to apply physical perturbations that are time-synchronized with the visual stimulus task. The type of the visual stimulus task and the motion parameters of the physical perturbation are recorded as task parameters. A multimodal biosignal acquisition subsystem is used to control a multi-view camera array to acquire multi-view images, control a plantar pressure imaging unit to acquire plantar pressure distribution images, and control an EEG acquisition module to acquire EEG signals during the process of a disturbed environment acting on the subject. A photodiode is used to detect brightness changes in a preset area of ​​the curved display screen as a time reference for the visual stimulus task display, ensuring that the time of each modal data is aligned with the task parameters. The modal data includes the multi-view images, plantar pressure distribution images, and EEG signals. The intelligent fusion analysis subsystem is used to perform three-dimensional pose reconstruction based on the multi-view images through the edge computing unit to obtain pose features including the three-dimensional coordinate sequence of key points of the whole skeleton; to perform image processing on the plantar pressure distribution image to obtain mechanical features including the plantar pressure center trajectory, arch index and plantar support area; and to perform frequency domain analysis on the EEG signal to obtain EEG features including the power and ratio of each channel. The intelligent fusion analysis subsystem is used to input the posture features, mechanical features, EEG features, and task parameters into the cloud-based AI large model platform and output evaluation results. The cloud-based AI large model platform is based on the Transformer architecture and uses the task parameters as a query to fuse features from various modalities through a cross-attention mechanism. The evaluation results include a comprehensive balance function score, vertigo subtype classification probability, sensory contribution weight, and dynamic fall risk warning. The contribution weight includes vestibular contribution weight, visual contribution weight, and proprioceptive contribution weight.

8. The multimodal fusion-based balanced function evaluation system as described in claim 7, characterized in that, The programmable three-dimensional force stage includes: a base, a balancing force stage, a linear guide module, a servo drive unit, and four three-dimensional force sensors; the four three-dimensional force sensors are respectively installed at the four corners of the bottom of the balancing force stage to collect the force values ​​at each corner in real time; the edge computing unit calculates the coordinates of the plantar pressure center based on the force values; the synchronization controller receives the coordinates of the plantar pressure center and dynamically adjusts the parameters of the subsequent physical disturbances applied by the programmable three-dimensional force stage according to the coordinates of the plantar pressure center.

9. A multimodal fusion-based balanced function assessment device, characterized in that, The multimodal fusion balance function evaluation device includes: a memory, a processor, and a multimodal fusion balance function evaluation program stored in the memory and executable on the processor, wherein the multimodal fusion balance function evaluation program is configured to implement the steps of the multimodal fusion balance function evaluation method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Human body balance evaluation method and system based on differentiated scenes

    CN118121187A

  • Human body balance evaluation equipment, human body balance evaluation method and human body balance evaluation device

    CN120419909A