Old people falling risk dynamic assessment and real-time early warning system and method based on multi-modal deep learning

By using a multimodal deep learning system, combined with the synchronous acquisition of behavioral, physiological and environmental data and long-term prediction using the Transformer architecture, the problems of high false alarm rate and insufficient timeliness of fall warning systems for the elderly have been solved, achieving high accuracy and stable real-time warning.

CN120913834APending Publication Date: 2025-11-07CHANGZHOU UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202511000869.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-21
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing technologies cannot adapt to differences in individuals and environments, resulting in a high false alarm rate and insufficient timeliness of fall warning systems for the elderly, failing to meet the needs for accurate monitoring.

Method used

A fall risk assessment system for the elderly using multimodal deep learning achieves real-time early warning by simultaneously collecting behavioral, physiological, and environmental data, combined with long-term risk prediction and multi-level response mechanisms based on the Transformer architecture, and uses transfer learning for adaptive model updates.

Benefits of technology

It significantly improves the accuracy of fall risk prediction, reduces false alarm rate, provides sufficient intervention time window, adapts to assessment stability in different scenarios, and reduces technology implementation and maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913834A_ABST
    Figure CN120913834A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of artificial intelligence, and particularly relates to an old people falling risk dynamic assessment and real-time early warning system and method based on multi-modal deep learning. The system comprises a multi-modal data acquisition module used for synchronously acquiring behavior data, physiological data and environmental data; the data preprocessing module is used for preprocessing the video data and the sensor data; the multi-modal feature fusion module is used for realizing cross-modal feature interaction of multi-modal data by adopting a hierarchical network architecture and fusing generated comprehensive features; the dynamic risk assessment module is used for realizing long-time-sequence risk prediction based on a prediction model and outputting a risk probability; the real-time early warning module performs real-time early warning based on the predicted risk probability by constructing a multi-level response mechanism; and the model adaptive updating module is used for constructing a closed-loop optimization mechanism and realizing adaptive updating of the model by adopting a transfer learning fine tuning model. The problem of misjudgment caused by one-sided data in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of artificial intelligence, and particularly relates to a dynamic assessment and real-time early warning system and method for fall risk of the elderly based on multi-modal deep learning. BACKGROUND

[0002] With the intensification of global aging, the fall of the elderly has become an important problem affecting their health and safety. The annual incidence of serious injuries such as fractures and intracranial hemorrhage caused by falls in the elderly accounts for more than 10%, which not only increases the medical burden, but also significantly reduces the quality of life.

[0003] In the current smart elderly care scene, the traditional manual monitoring mode has problems such as response lag and high labor cost, and the fall early warning systems constructed by existing technical means generally face three major bottlenecks: single sensor cannot comprehensively capture multi-dimensional risk factors of behavior, physiology and environment, traditional algorithms are difficult to handle the long-time dynamic characteristics of the behavior of the elderly, and fixed architecture systems cannot adapt to the differences of different individuals and environments, resulting in high false alarm rate and insufficient timeliness of early warning in actual application. Specifically:

[0004] 1. Limitations of single-mode data collection: existing systems mostly rely on single camera or wearable accelerometer, which can only obtain single-dimensional information of action behavior. For example, the pure visual solution is easily affected by light obstruction, privacy controversy and other problems, while the solution based on accelerometer cannot perceive environmental risks such as slippery ground and obstacle distribution, resulting in high false alarm rate and high false negative rate in actual application, which cannot meet the demand of accurate monitoring.

[0005] 2. Insufficient time series modeling of algorithm model: traditional machine learning methods (such as SVM, random forest) are difficult to capture the gradual feature changes before the fall of the elderly. The fall of the elderly is usually accompanied by long-time evolution process such as unstable gait and physiological index fluctuation, while existing algorithms can only make instantaneous judgment based on short window data, resulting in too short warning time, for example, the advance warning time for special groups such as Parkinson's patients is less than 30 seconds, which cannot provide an effective time window for intervention measures.

[0006] 3. Poor adaptability of system environment: existing solutions generally lack model adaptation mechanism and cannot cope with the differences in body size and health status of different elderly people and changes in home, nursing home and other scenes. The measured data of a nursing home shows that the assessment accuracy of the same system in different floor environments fluctuates by 25%-40%, and retraining the model consumes a large amount of labeled data, resulting in high cost and long cycle of technology landing. SUMMARY

[0007] The application aims to provide a multi-modal deep learning-based dynamic assessment and real-time early warning system and method for the fall risk of the elderly to solve the technical problems that the prior art cannot adapt to different individuals and environmental differences, resulting in high false alarm rate and insufficient timeliness of early warning in actual application.

[0008] The multi-modal deep learning-based dynamic assessment and real-time early warning system for the fall risk of the elderly comprises:

[0009] A multi-modal data acquisition module is configured to synchronously acquire behavior data, physiological data and environmental data as the basis for constructing multi-dimensional risk assessment;

[0010] A data preprocessing module is configured to preprocess video data and sensor data to obtain 3D coordinate sequences of joints and time-aligned multi-modal data;

[0011] A multi-modal feature fusion module is configured to realize cross-modal feature interaction of multi-modal data by using a hierarchical network architecture to obtain comprehensive features generated by the fusion of action features, physiological features and environmental features;

[0012] A dynamic risk assessment module is configured to realize long-time sequence risk prediction based on a model with a Transformer architecture to output a risk probability;

[0013] A real-time early warning module is configured to perform real-time early warning based on the predicted risk probability by constructing a multi-level response mechanism;

[0014] A model self-adaptive updating module is configured to construct a closed-loop optimization mechanism and realize self-adaptive updating of the model by using transfer learning to fine-tune the model.

[0015] Preferably, in the dynamic risk assessment module, the environmental features include ground flatness and obstacle spatial distribution for calculating obstacle distance weighting values; and the light data in the environmental features are further used to generate an environmental risk index for risk prediction.

[0016] The multi-modal deep learning-based dynamic assessment and real-time early warning system for the fall risk of the elderly further comprises a data management module, which is configured to set a ring buffer, enable local SSD storage when network delay is greater than a threshold value, realize breakpoint resume through MD5 verification after network recovery, and support hot plug function to automatically switch to redundant equipment when a sensor fails.

[0017] The application further provides a multi-modal deep learning-based dynamic assessment and real-time early warning method for the fall risk of the elderly, comprising the following steps:

[0018] S1, multi-modal data acquisition, the multi-modal data comprising behavior data, physiological data and environmental data, the environmental data comprising laser radar point cloud data;

[0019] S2, data preprocessing, for preprocessing video data and sensor data to obtain 3D coordinate sequence of joint nodes and time-aligned multi-modal data;

[0020] S3, multi-modal feature fusion, cross-modal feature interaction is realized by adopting hierarchical network architecture, and comprehensive features generated by fusion of three types of features of action features, physiological features and environmental features are obtained;

[0021] S4, dynamic risk evaluation, long time sequence risk prediction is realized based on the prediction model of the Transformer architecture, and the risk probability is output;

[0022] S5, real-time early warning, through the construction of multi-level response mechanism, real-time early warning is carried out based on the predicted risk probability;

[0023] S6, model adaptive update, a closed-loop optimization mechanism is constructed, and the prediction model is fine-tuned by using the transfer learning strategy.

[0024] Preferably, the step S1 comprises:

[0025] S1.1, collecting behavior data, the behavior data including human posture video stream, action acceleration and plantar pressure distribution data;

[0026] S1.2, collecting physiological data, the physiological data including vital sign data such as heart rate and blood pressure;

[0027] S1.3, collecting environmental data, the environmental data including laser radar point cloud data and environmental light intensity, the laser radar point cloud data being used to obtain ground flatness and obstacle distribution data, and the environmental light intensity being used to perceive potential environmental risks.

[0028] Preferably, the step S2 comprises:

[0029] S21, the 2D coordinates of the joint nodes are detected by the YOLOv8n model, and the 2D coordinates of the joint nodes are converted into 3D coordinates in combination with the camera calibration parameters;

[0030] S22, the acceleration data is corrected for zero drift, and the pressure data is denoised by Kalman filtering;

[0031] S23, a time synchronization algorithm based on cross-correlation is adopted, the video frame timestamp is taken as the reference, the sensor data is resampled by cubic spline interpolation, and the multi-modal data is time-aligned.

[0032] Preferably, the step S3 comprises:

[0033] S31, the 3D coordinate sequence of the joint nodes is input into the I3D network, the action features are extracted by convolution, and the joint relative position change rate is captured;

[0034] S32, the physiological data is segmented by time window, then input into the LSTM network for processing, and the physiological features containing the physiological index fluctuation coefficient are output;

[0035] S33, after the laser radar point cloud data is voxelized and down-sampled, the environment features containing the environmental risk index are extracted by inputting the data into the ResNet18 network, and the environmental risk index is used as the obstacle distance weighting value;

[0036] S34, the action features, physiological features and environmental features are projected to the same dimension by linear projection, and then the comprehensive features are generated by multi-head attention mechanism fusion.

[0037] Preferably, the step S4 comprises:

[0038] S41, the comprehensive features are added with sinusoidal position encoding, and then input into the 6-layer Transformer encoder;

[0039] S42, the action features of the key joint nodes are weighted by the spatial attention layer, and the joint angular velocity of the corresponding key joint nodes is calculated, the key joint nodes being the joint nodes of the lower limbs;

[0040] S43, the risk probability is output by the fully connected layer, and the model is trained by using the cross-entropy loss function combined with the AdamW optimizer;

[0041] During training, the model is optimized by using the cross-entropy loss L, and the loss function is as follows:

[0042]

[0043] Where y t is the true label, and p t is the predicted probability.

[0044] Preferably, the specific algorithm of the obstacle distance weighting value is as follows:

[0045] 1) First, set the obstacle detection distance threshold, when the detected obstacle angle with the walking path is less than the angle threshold, the obstacle is included in the risk calculation, and the distance d0 from the obstacle to the current position of the elderly is obtained at the same time;

[0046] 2) Secondly, the influence degree of the obstacle is quantified by using the distance weighting function to obtain the obstacle weighting value, and the distance reciprocal weighting function or the Gaussian attenuation function is used as the weighting function;

[0047] 3) At the same time, the light data collected by the light sensor is fused, if the fused light intensity is lower than the set threshold, then the risk weight of the environmental features is improved by the light correction factor, and the light correction factor λ l = 1 + β (1- light value / threshold), wherein β is the light influence coefficient.

[0048] 4) Finally, the obstacle weighting value is multiplied by the illumination correction factor to generate a comprehensive environmental risk index as the obstacle distance weighting value.

[0049] Preferably, the step S6 comprises:

[0050] S61, periodically statistics the misjudgment rate, and the updating process is started when the misjudgment rate exceeds a misjudgment threshold;

[0051] S62, a certain amount of labeled data is extracted, the labeled data is labeled through double verification, the double verification needs to be confirmed by a pressure sensor and manual labeling, and the positive and negative sample ratio is 1:3.

[0052] S63, a transfer learning strategy is adopted, the first four layers of the pre-trained model are fixed, the last two layers of the prediction model are fine-tuned, in the training process, a warm-up learning rate strategy is adopted, combined with cosine annealing scheduling and AdamW optimizer, and an early stopping mechanism is used to optimize the model performance.

[0053] The application has the advantages that:

[0054] 1. The evaluation accuracy is significantly improved, and the single mode technology bottleneck is broken: through the multi-modal fusion of behavior video (17 three-dimensional coordinates of key points captured by a 4K camera), physiological indicators (heart rate / blood pressure ± 3mmHg accuracy) and environmental parameters (obstacle detection with 0.25° angle resolution of laser radar), combined with the spatiotemporal feature extraction of I3D network and the attention mechanism of 8-head Transformer, the system can capture the multi-dimensional correlation of "unstable gait-physiological fluctuation-environmental risk". The test data shows that the 10-minute prediction accuracy of the fall risk of the elderly is more than 92%, which is 30% higher than that of the traditional single camera scheme (accuracy of 62%), the false positive rate is reduced from 35% to less than 8%, especially the recognition rate of gait abnormalities of special groups such as Parkinson's patients is increased by 45%, and the misjudgment problem caused by the one-sidedness of single mode data is solved.

[0055] 2. The long-time dynamic early warning capability is prominent, and provides sufficient time window for intervention: the improved Transformer architecture realizes long dependence modeling of behavior time sequence within 12 hours through sine position encoding (maximum position 5000) and knee / hip joint Gaussian weighting (σ=0.2m), and can identify the gradual evolution process from "mild gait abnormality" to "high risk of falling". For example, when the system detects that the gait cycle fluctuation of the elderly is more than 20% for 30 minutes in a row, the heart rate variability SDNN is less than 50ms, and there is an obstacle (included angle less than 60°) within 1.5m in front, a high-risk early warning (risk probability ≥0.7) can be output 10 minutes in advance, which provides more than 20 times intervention window than the traditional algorithm (early warning time <30 seconds), so that the nursing staff has sufficient time to assist or adjust the environment.

[0056] 3. Adaptive closed-loop optimization mechanism to achieve cross-scenario robustness deployment: incremental learning strategy based on misjudgment rate feedback (start updating when misjudgment rate > 10% per week) and data cache hot plug design (SSD local storage ≥ 500MB / s write speed), the system can maintain ≥ 92% evaluation stability in different scenarios such as home, nursing home, hospital, etc. The actual measurement of a certain chain nursing home shows that the same system does not need to be retrained in 30 different floor environments, and the evaluation accuracy fluctuation is only ± 3%, which is significantly improved compared with the traditional scheme (fluctuation 25%-40%); through the transfer learning fine-tuning of 200 groups of new data (positive and negative samples 1:3), the adaptation time of the model to the individual differences of new residents is shortened from 7 days to 1 day, which greatly reduces the technology landing and maintenance cost. BRIEF DESCRIPTION OF DRAWINGS

[0057] Figure 1 The basic flowchart of the elderly fall risk dynamic assessment and real-time warning system and method based on multi-modal deep learning in the application.

[0058] Figure 2 The Transformer architecture model diagram used by the dynamic risk assessment module in the application. DETAILED DESCRIPTION

[0059] The specific embodiments of the application will be further described in detail below with reference to the drawings, and by describing the embodiments, to help those skilled in the art to have a more complete, accurate and in-depth understanding of the inventive concept and technical solutions of the application.

[0060] As shown in Figures 1-2 , the application provides an elderly fall risk dynamic assessment and real-time warning system based on multi-modal deep learning, which comprises:

[0061] A multi-modal data acquisition module is used to synchronously acquire behavior data, physiological data and environmental data, which are used as the basis for constructing multi-dimensional risk assessment.

[0062] A data preprocessing module is used to preprocess video data and sensor data to obtain 3D coordinate sequences of joints and time-aligned multi-modal data.

[0063] A multi-modal feature fusion module is used to realize cross-modal feature interaction of multi-modal data by using a hierarchical network architecture to obtain comprehensive features generated by fusion of action features, physiological features and environmental features.

[0064] A dynamic risk assessment module is used to realize long-time sequence risk prediction based on a prediction model of Transformer architecture, and output risk probability.

[0065] Among them, the environmental features include ground flatness, obstacle spatial distribution, and are used to calculate obstacle distance weighting values; the environmental risk index is further generated by combining the light data in the environmental features, and is used for risk prediction.

[0066] The real-time warning module constructs a multi-level response mechanism and performs real-time warning based on the predicted risk probability.

[0067] The model adaptive updating module constructs a closed-loop optimization mechanism, and adopts transfer learning to fine-tune the prediction model, so as to realize adaptive updating of the model.

[0068] The system also includes a data management module, which sets a 1GB ring buffer, and when the network delay is greater than 500ms, the local SSD storage (write speed is greater than or equal to 500MB / s) is enabled, and after the network recovers, the breakpoint resume transmission is realized through MD5 check, so as to ensure that the data integrity is greater than or equal to 99.9%. The multi-modal data collected by the data acquisition module is sent to the data management module, and the data management module supports the hot plug function, and automatically switches to the redundant device when the sensor fails, and the switching delay is less than or equal to 100ms.

[0069] The application also provides a multi-modal deep learning-based dynamic assessment and real-time warning method for fall risk of the elderly, which comprises the following steps.

[0070] S1, multi-modal data acquisition.

[0071] This step synchronously collects multi-modal data at a frequency of 50Hz through a camera, a smart bracelet, a smart insole and an environment sensor, and stores the data in a ring buffer after 16-bit ADC conversion. The multi-modal data includes behavior data, physiological data and environmental data, wherein the environmental data includes laser radar point cloud data. The specific collection method is as follows:

[0072] S1.1, collect behavior data. This step is specifically as follows.

[0073] 1) Obtain human posture video stream through a 4K resolution camera (frame rate 30fps) arranged in a room.

[0074] 2) Collect motion acceleration by using a wrist-worn smart bracelet (built-in three-axis accelerometer, range ±16g).

[0075] 3) Collect foot pressure distribution data by using an insole sensor (16x16 pressure sensing array, sampling frequency 100Hz).

[0076] The above behavior data is collected to comprehensively capture the motion characteristics of the limbs of the elderly.

[0077] S1.2, collect physiological data.

[0078] This step synchronously collects vital sign data such as heart rate (sampling accuracy ±1 bpm), blood pressure (systolic / diastolic pressure, accuracy ±3 mmHg) through the smart bracelet, providing physiological state basis for risk assessment.

[0079] When the network delay exceeds 500 ms, the local SSD storage module is started, and after the network is restored, the data is synchronized through breakpoint resume.

[0080] S1.3, collect environmental data.

[0081] 1) Collect laser radar point cloud data through an 8-line laser radar (scanning range 0-10 m, angular resolution 0.25°, obstacle detection distance threshold 1.5 m) to obtain ground flatness (error ≤2 mm) and obstacle distribution data.

[0082] 2) Collect ambient light intensity through a photosensitive sensor (sensitivity 0.1-10000 lux) to perceive potential environmental risks.

[0083] In this step, when the laser radar detects that the angle between the obstacle and the walking path is less than the angle threshold (<60°), the system automatically counts it into the environmental risk index.

[0084] S2, data preprocessing.

[0085] This step includes the following sub-steps:

[0086] S21, video data is detected by a YOLOv8n model to detect the 2D coordinates of 17 joints, and the 2D coordinates of the joints are converted into 3D coordinates combined with the camera calibration parameters (focal length f=500 pixels, principal point c=(320, 240)).

[0087] This step is video processing, specifically including: using a YOLOv8n model to detect 17 joints of the human body in the video stream, the joints corresponding to the joints include the neck, shoulders, elbows, hands, hips, knees, ankles, feet, etc., the positioning error is ≤3 pixels, through a 16-frame sliding window (time step 8 frames) to extract the 2D coordinates of the joints, then convert them into 3D coordinates and form a 3D coordinate sequence with these joint data, and finally normalize the data with the center of the torso as the origin, the coordinate range is [-1, 1].

[0088] S22, zero drift correction is performed on the acceleration data, and Kalman filter denoising is performed on the pressure data.

[0089] This step is sensor data processing, specifically including: using a second-order Butterworth low-pass filter (cutoff frequency 10 Hz) for zero drift correction of acceleration data, calculating the mean value of the first 10 seconds of static period as the offset, and subtracting the offset from the data. The pressure data is denoised by Kalman filtering (process noise covariance 0.01, measurement noise covariance 0.05).

[0090] S23, using a cross-correlation-based time synchronization algorithm, taking the video frame timestamp as the reference, performing cubic spline interpolation resampling on the sensor data, and realizing multi-modal data time alignment.

[0091] Based on IEEE1588 clock synchronization protocol (synchronization accuracy ≤1ms), taking the video frame timestamp as the reference, combining the cubic spline interpolation algorithm to synchronize the multi-modal data in time, and ensuring the consistency of the data timing.

[0092] S3, multi-modal feature fusion.

[0093] This step uses a hierarchical network architecture to realize cross-modal feature interaction, specifically including the following sub-steps:

[0094] S31, input the 3D coordinate sequence of the joint node (dimension 17x3x16) into the I3D network, and extract 256-dimensional motion features through 5 layers of convolution.

[0095] This step realizes motion feature extraction: using the I3D network (input size 224x224x16, convolution kernel size 3x3x3, 5 layers of convolution) to extract spatio-temporal features from the 3D coordinate sequence of the joint node, outputting 256-dimensional motion features, and capturing the rate of change of joint relative position (calculation accuracy 1mm / s)

[0096] S32, physiological data is segmented by time window (30 minutes), and then input into a 2-layer LSTM network for processing, outputting 128-dimensional physiological features.

[0097] This step realizes physiological feature extraction: through a 2-layer LSTM network (128 neurons per layer, dropout rate 0.2) to analyze the trend of physiological data within a 30-minute sliding window, outputting 128-dimensional physiological features, including heart rate variability SDNN and other fluctuation coefficients.

[0098] S33, after voxelization (0.1m resolution) and down-sampling of the laser radar point cloud data, input the ResNet18 network to extract 64-dimensional environmental features.

[0099] This step realizes environmental feature extraction: after voxelization (resolution 0.1m) and down-sampling of the laser radar point cloud data, input the ResNet18 network to extract 64-dimensional environmental features, combine the light data to generate an environmental risk index, and the environmental risk index is used as the obstacle distance weighting value.

[0100] S34, the three types of features are linearly projected to 384 dimensions, and then integrated by 8 attention mechanisms to generate comprehensive features.

[0101] This step unifies the three types of features into 384 dimensions by linear projection, and uses 8 attention mechanisms (each with a dimension of 48) for cross-modal interaction. In the Transformer self-attention mechanism, the dimensions of the query matrix, key matrix, and value matrix are all 384. The attention score calculation uses the scaled dot product formula (scaling factor √384), and adds residual connection and layer normalization (ε = 1e-6) to generate comprehensive features.

[0102] S4, dynamic risk assessment.

[0103] The improved network structure is applied in this method, and the main part is to realize long-time sequence risk prediction based on the improved Transformer architecture. The Transformer architecture includes 6 layers of encoder, and each layer has 2 attention heads. After training, the prediction model is obtained. This step includes the following sub-steps:

[0104] S41, add sinusoidal position encoding to the comprehensive features, and then input them into the 6-layer Transformer encoder.

[0105] The time dimension uses sinusoidal position encoding (period 2π, maximum position 5000) to capture the longest 12-hour behavior time sequence dependence.

[0106] S42, weight the action features of the key joints by the spatial attention layer, and calculate the joint angular velocity of the corresponding key joints.

[0107] The spatial dimension performs Gaussian weighting on the coordinates of 8 key joints such as the knee joint and the hip joint. The weight coefficient w is calculated as follows: w = exp(-d 2 / (2σ 2 )), where d is the distance from the joint to the center of gravity, and σ is the spatial standard deviation based on the joint category, with a default value of 0.3m. When the joint is a knee joint or a hip joint, σ = 0.2m. This method focuses on the action features of the lower limbs, so the joint nodes of the lower limbs are selected as the key joints, and the joint angular velocity (Δθ / Δt) is calculated. Δθ is the joint angle change between images, and Δt is the interval time between images.

[0108] Steps S41 and S42 are combined to form the spatio-temporal attention mechanism used in this method.

[0109] S43, output the risk probability through the fully connected layer, and use the cross-entropy loss function combined with the AdamW optimizer for model training.

[0110] The system applying the method calculates the fall risk probability in a certain time (10 minutes) in the future through the full connection layer (output layer activation function sigmoid). When classifying the risk, the specific preset threshold system includes: low risk (risk probability <0.3), medium risk (risk probability between 0.3-0.7), and high risk (risk probability ≥0.7). During training, the cross-entropy loss L is used for model optimization, and the loss function is as follows:

[0111]

[0112] where y t is the true label, and p t is the predicted probability.

[0113] The environmental features include abstract numerical representations of environmental geometric features such as ground flatness and obstacle spatial distribution. In combination with the light data, an environmental risk index is further generated to represent the obstacle distance weighted value, and the specific algorithm is as follows:

[0114] 1) First, set the obstacle detection distance threshold (such as 1.5m for the laser radar as the obstacle detection distance threshold), and when the detected obstacle and the walking path angle is less than the angle threshold, the obstacle is included in the risk calculation, and the distance d0 of the obstacle to the current position of the elderly is obtained.

[0115] 2) Secondly, the distance weighted function is used to quantify the influence degree of the obstacle to obtain the obstacle weighted value. The distance weighted function adopts the distance reciprocal weighted function or the Gaussian decay function. The distance reciprocal weighted function is: α = 1 / d0; and the Gaussian decay function is: α = exp(-d0 2 / (2σ 2 )); where α is the weight coefficient, d0 is the distance of the obstacle to the current position of the elderly, and σ is the decay parameter.

[0116] 3) At the same time, the light data collected by the light sensor is fused. If the fused light intensity is lower than the set threshold (such as 10 lux), the risk weight of the environmental feature is improved through the light correction factor, and the light correction factor λ l = 1 + β(1-illumination value / threshold), where β is the light influence coefficient.

[0117] 4) Finally, the obstacle weighted value is multiplied by the light correction factor to generate a comprehensive environmental risk index as the obstacle distance weighted value, i.e. comprehensive environmental risk index λ = αλ l .

[0118] In this process, the distance threshold, angle constraint, and weight coefficient are optimized through model training to strengthen the relevance of the environmental risk index and the actual fall risk.

[0119] S5, real-time warning.

[0120] This step ensures the timeliness of early warning by building a multi-level response mechanism in specific operations. It includes the following.

[0121] 1) When the risk probability is ≥ 0.7, trigger local audible and visual alarms (buzzer sound intensity 85 dB ± 3 dB, LED flash frequency 2 Hz ± 0.5 Hz).

[0122] 2) Push early warning information to the family end APP through MQTT protocol (QoS level 2), including UTM coordinates (accuracy 1 m) and risk level.

[0123] 3) Send a short message to the nursing center (including the old person's ID, risk level and timestamp), the overall early warning delay time is ≤ 500 ms;

[0124] 4) Use WebSocket protocol (heartbeat interval 5 seconds) to maintain communication connection to ensure real-time transmission of early warning information.

[0125] S6, model adaptive update.

[0126] This step builds a closed-loop optimization mechanism to improve the environmental adaptability of the system. Updates need to be based on historical data, so data management is involved.

[0127] S61, weekly statistics of misjudgment rate, when the misjudgment rate exceeds the misjudgment threshold (> 10%), start the update process.

[0128] S62, extract a certain number of labeled data (200 groups), and label the labeled data through double verification.

[0129] In the labeled data, the positive and negative sample ratio is 1:3, and the label is confirmed by pressure sensor and manual annotation.

[0130] S63, use transfer learning strategy, fix the first 4 layers of the pre-trained model, fine-tune the last two layers of the model, and use early stopping mechanism to optimize model performance during training.

[0131] The training process uses a warm-up learning rate strategy (initial learning rate 1e-5, increased to 1e-4 after 5 epochs), combined with cosine annealing scheduling (initial learning rate 1e-4, minimum learning rate 1e-6), batch size is set to 32, optimizer uses AdamW optimizer (β1 = 0.9, β2 = 0.999, weight decay 0.01), and early stopping mechanism (stop if the validation set loss does not decrease for 5 consecutive rounds).

[0132] Typical implementation cases and effect data of the embodiment are as follows.

[0133] 1. Case: Hypoglycemia fall risk early warning.

[0134] Scenario: 10:22, the old man walks from the bedroom to the dining room.

[0135] Data characteristics: 25% gait cycle fluctuation, heart rate jumps to 120 bpm, and SDNN = 40 ms, there are wet areas on the corridor floor.

[0136] System response: 10:23 trigger warning, 10:25 nursing staff arrive on the scene, timely provide glucose water to avoid falling.

[0137] Traditional solution: single camera solution alarms at 10:28, missing the best intervention time.

[0138] 2. Long-term operation data statistics (3 months statistics) as shown in Table 1.

[0139] Table 1: Long-term operation data statistics table after application of the present application

[0140] Indicator System of the invention Conventional single camera approach Lift magnitude Early warning time 10.2 ± 2.1 minutes 28.5 ± 5.3 seconds 21.5 times Prediction accuracy 92.3% 62.1% +30.2% False alarm rate 7.2% 35.4% -79.7%

[0141] Some technical points in the specific implementation.

[0142] 1) Network bandwidth requirement: multi-modal data transmission requires uplink bandwidth ≥ 12 Mbps, and it is recommended to use gigabit local area network + 5G backup link.

[0143] 2) Edge computing configuration: 1 edge server (CPU ≥ 16 cores, GPU ≥ 16 TOPS computing power) is required for every 100 users.

[0144] 3) Data security guarantee: video data is encrypted and stored, transmission uses TLS1.3 protocol, which meets the privacy protection requirements.

[0145] 4) Maintenance cycle: sensors are calibrated once every quarter, models are automatically updated every week, and hardware failures are replaced quickly through hot plug redundancy design.

[0146] The above has exemplarily described the present application in combination with the drawings, and it is obvious that the specific implementation of the present application is not limited by the above manner, as long as various non-essential improvements are made using the inventive concept and technical solution of the present application, or the inventive concept and technical solution of the present application is directly applied to other occasions without improvement, all within the protection scope of the present application.

Claims

1. A multi-modal deep learning based dynamic assessment and real-time warning system for fall risk of the elderly, characterized in that: Comprise: A multi-modal data acquisition module for synchronously acquiring behavior data, physiological data, and environmental data as the basis for constructing a multi-dimensional risk assessment; A data preprocessing module for preprocessing video data and sensor data to obtain 3D coordinate sequences of key joints and time-aligned multi-modal data; A multi-modal feature fusion module for realizing cross-modal feature interaction of multi-modal data using a hierarchical network architecture to obtain comprehensive features generated by the fusion of action features, physiological features, and environmental features; A dynamic risk assessment module for realizing long-time sequence risk prediction based on a Transformer architecture model to output risk probability; A real-time warning module for realizing real-time warning based on predicted risk probability by constructing a multi-level response mechanism; A model adaptive updating module for constructing a closed-loop optimization mechanism to realize adaptive updating of the model using transfer learning to fine-tune the model.

2. The multi-modal deep learning based dynamic assessment and real-time warning system for fall risk of the elderly according to claim 1, characterized in that: In the dynamic risk assessment module, the environmental features include ground flatness and obstacle spatial distribution for calculating obstacle distance weighting values; and the light data in the environmental features are further used to generate an environmental risk index for risk prediction.

3. The multi-modal deep learning based dynamic assessment and real-time warning system for fall risk of the elderly according to claim 1, characterized in that: Further comprising a data management module, which sets a ring buffer, enables local SSD storage when network delay is greater than a threshold value, realizes breakpoint resume through MD5 verification after network recovery, and supports hot plug function to automatically switch to redundant equipment when a sensor fails.

4. The method for dynamic assessment and real-time early warning of the fall risk of the elderly based on multi-modal deep learning, characterized in that: Comprise the following steps: S1, multi-modal data acquisition, the multi-modal data includes behavior data, physiological data and environmental data, the environmental data includes laser radar point cloud data; S2, data preprocessing, for preprocessing video data and sensor data to obtain 3D coordinate sequences of key joints and time-aligned multi-modal data; S3, multi-modal feature fusion, cross-modal feature interaction is realized by using a hierarchical network architecture to obtain comprehensive features generated by the fusion of action features, physiological features, and environmental features; S4, dynamic risk assessment, a prediction model based on a Transformer architecture is used to realize long-time sequence risk prediction to output risk probability; S5, real-time warning, real-time warning is realized based on predicted risk probability by constructing a multi-level response mechanism; S6, model adaptive updating, a closed-loop optimization mechanism is constructed to fine-tune the prediction model using a transfer learning strategy.

5. The multi-modal deep learning based dynamic assessment and real-time early warning method for fall risk of the elderly according to claim 4, characterized in that: The step S1 comprises: S1.1, acquiring behavior data, including human posture video stream, action acceleration, and plantar pressure distribution data; S1.2, acquiring physiological data, including vital sign data such as heart rate and blood pressure; S1.3, acquiring environmental data, including laser radar point cloud data and environmental light intensity, the laser radar point cloud data is used to obtain ground flatness and obstacle distribution data, and the environmental light intensity is used to perceive potential environmental risks.

6. The multi-modal deep learning based dynamic assessment and real-time warning method for fall risk of the elderly according to claim 4, characterized in that: The step S2 comprises: S21, video data is detected by a YOLOv8n model to obtain 2D coordinates of key joints, which are converted into 3D coordinates combined with camera calibration parameters; S22, zero drift correction is performed on acceleration data, and Kalman filter denoising is performed on pressure data; S23, a time synchronization algorithm based on cross-correlation is adopted to resample the sensor data three times by spline interpolation based on the video frame timestamp, so as to realize time alignment of multi-modal data.

7. The multi-modal deep learning based dynamic assessment and real-time warning method for fall risk of the elderly according to claim 4, characterized in that: The step S3 comprises: S31, inputting the 3D coordinate sequence of the joint node into the I3D network, extracting the action feature through convolution, and capturing the relative position change rate of the joint; S32, the physiological data is segmented according to the time window, and then input into the LSTM network for processing, and the physiological feature is output, including the physiological index fluctuation coefficient; S33, after the laser radar point cloud data is voxelized and down-sampled, the environment feature is extracted by inputting the data into the ResNet18 network, including the environmental risk index, and the environmental risk index is used as the obstacle distance weighting value; S34, the action feature, the physiological feature and the environment feature are projected to the same dimension through linear projection, and then the comprehensive feature is generated through the multi-head attention mechanism.

8. The multi-modal deep learning based dynamic assessment and real-time warning method for fall risk of the elderly according to claim 4, characterized in that: The step S4 comprises: S41, adding the sinusoidal position coding to the comprehensive feature, and then inputting the feature into the 6-layer Transformer encoder; S42, weighting the action feature of the key joint node through the spatial attention layer, and calculating the joint angular velocity of the corresponding key joint node, the key joint node being the joint node of the lower limb; S43, outputting the risk probability through the full connection layer, and training the model by using the cross-entropy loss function combined with the AdamW optimizer; During the training, the model is optimized by using the cross-entropy loss L, and the loss function is as follows: where y t is the true label, p t is the predicted probability.

9. The multi-modal deep learning based dynamic assessment and real-time warning method for fall risk of the elderly according to claim 4, characterized in that: The specific algorithm of the obstacle distance weighting value is as follows: 1) First, set the obstacle detection distance threshold, when the detected obstacle and the walking path included angle is less than the angle threshold, the obstacle is included in the risk calculation, and the distance d0 from the obstacle to the current position of the old person is obtained synchronously; 2) Secondly, the influence degree of the obstacle is quantified to obtain the obstacle weighting value by using the distance weighting function, and the distance reciprocal weighting function or the Gaussian attenuation function is used as the weighting function; 3) At the same time, the light data collected by the fusion photosensitive sensor is fused, and if the light intensity obtained by the fusion is lower than the set threshold, the risk weight of the environmental feature is improved through a light correction factor λ l = 1 + β(1 - light value / threshold), wherein β is a light influence coefficient. 4) Finally, the obstacle weighting value is multiplied by the illumination correction factor to generate a comprehensive environmental risk index as the obstacle distance weighting value.

10. The multi-modal deep learning based dynamic assessment and real-time warning method for fall risk of the elderly according to claim 4, characterized in that: The step S6 comprises: S61, regularly statistics the misjudgment rate, and when the misjudgment rate exceeds the misjudgment threshold, the updating process is started; S62, a certain amount of labeled data is extracted, the labeled data is labeled through double verification, the double verification needs to be confirmed by the pressure sensor and artificial labeling, and the positive and negative sample ratio is 1:

3. S63, a transfer learning strategy is adopted, the first 4 layers of the pre-trained model are fixed, the last two layers of the prediction model are fine-tuned, during the training process, the warm-up learning rate strategy is adopted, combined with the cosine annealing scheduling and the AdamW optimizer, and the early stopping mechanism is used to optimize the model performance.

Citation Information

Cited By

  • Intelligent evaluation method for influence of campus environment on individual health by integrating space-time behaviors

    CN121260480A

  • Intelligent assessment method of campus environment integrating spatio-temporal behavior to affect individual health

    CN121260480B

  • Old people living behavior pattern recognition and risk early warning system based on deep learning

    CN121598172A