Multi-mode AI fusion senile non-inductive assessment method and system and electronic equipment

By deploying non-contact sensors in the daily environment of the elderly to collect multimodal data, and using an attention-based multimodal fusion model, the problems of insufficient cross-scene data linkage and multimodal fusion are solved, realizing full-cycle health management and accurate risk warning for the elderly.

CN121148701APending Publication Date: 2025-12-16ARTIFACT R&D (SUZHOU) CO LTD

Patent Information

Application Number
CN202511321374.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-16
Publication Date
2025-12-16

AI Technical Summary

Technical Problem

Existing elderly health monitoring technologies cannot achieve cross-scenario data linkage between outpatient and home settings, and lack multimodal fusion, making it difficult to capture hidden risk associations.

Method used

By deploying non-contact sensors in the daily life environment of the elderly to collect multimodal data, including visual sensors, audio sensors and millimeter-wave radar sensors, and combining them with an attention-based multimodal fusion model, data preprocessing, feature extraction and weighted fusion are performed to generate comprehensive health assessment results, and early warnings are issued based on health baselines and risk thresholds.

Benefits of technology

It enables full-cycle health management for the elderly, improves the pertinence and accuracy of health status assessment, provides continuous and reliable health risk warnings, and takes into account privacy protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121148701A_ABST
    Figure CN121148701A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent health monitoring, and particularly discloses a multi-modal AI-fused senior non-inductive assessment method and system and electronic device.The method comprises the following steps that in the daily living environment of the senior, original multi-modal data reflecting the physiological and behavior states of the senior are collected in a non-inductive mode through at least two non-contact sensors; and preprocessing the original multi-modal data, and extracting a single-modal health feature vector from the data of each modal by using a preset feature extraction model. According to the invention, non-inductive data acquisition is realized by deploying the non-contact multi-mode sensor in the daily environment of the elderly, the elderly does not need to cooperate deliberately, and the use convenience and acceptability are effectively improved; through preprocessing and feature extraction of multi-modal data, multi-dimensional health features can be mined from visual, audio and radar data, the limitation of single data dimension in the traditional technology is broken through, and comprehensive representation of physiological and behavior states of old people is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent health monitoring technology, and in particular to a multimodal AI-integrated, non-intrusive assessment method, system, and electronic device for the elderly. Background Technology

[0002] There are currently two main types of technical solutions in the field of elderly health monitoring and safety surveillance: the solution with authorization announcement number CN118629659B, which collects multi-dimensional physical data of the elderly, compares the change curves of normal data of the non-disease group, calculates abnormal physical factors and the degree of disease dominance, and screens target elderly patient groups to assess the risk index of chronic diseases; and the solution with authorization announcement number CN118248337B, which relies on AI to identify user activity categories, matches the activity category monitoring frequency library to dynamically adjust the frequency of vital sign collection, and combines a dynamic health assessment model to generate graded early warning and personalized health recommendations.

[0003] The aforementioned existing technologies have significant limitations: First, they are isolated scenarios. CN118629659B focuses on chronic disease monitoring, while CN118248337B emphasizes daily health monitoring. Neither of them achieves cross-scenario data linkage between outpatient and home scenarios, making it impossible to form a closed loop for full-cycle management. Second, they lack multimodal fusion. CN118629659B only focuses on physiological data, while CN118248337B links activity and vital sign data. Neither of them covers deep fusion of physiological, behavioral, and environmental multimodal data, making it difficult to capture hidden risk associations. Summary of the Invention

[0004] This invention aims to at least partially address one of the technical problems in related technologies. Therefore, the objective of this invention is to propose a multimodal AI-based, non-intrusive assessment method, system, and electronic device for elderly individuals, thereby improving the relevance and accuracy of comprehensive health status assessment.

[0005] To achieve the above objectives, a first aspect of the present invention proposes a multimodal fusion AI-based method for sensory-free assessment of the elderly, comprising the following steps: S1. In the daily life environment of the elderly, raw multimodal data reflecting their physiological and behavioral states are collected imperceptibly through at least two non-contact sensors. S2. Preprocess the original multimodal data and use a preset feature extraction model to extract a single modality health feature vector from the data of each modality. S3. Using a multimodal fusion model, the multiple single-modal health feature vectors are weighted, fused, and analyzed over time to generate a quantitative assessment result representing the current comprehensive health status of the elderly. S4. Based on the quantitative assessment results and their dynamic trends, compare them with the preset health baseline and risk threshold. When the warning conditions are met, generate and output the corresponding health risk warning information.

[0006] In some embodiments of the present invention, the raw multimodal data in step S1 includes at least: Video stream data collected by visual sensors is used to analyze the gait, daily living abilities (ADL), and facial expressions of older adults. Audio stream data collected by audio sensors is used to analyze the speech tone, cough frequency, and social interaction activities of older adults. Radar point cloud data collected by millimeter-wave radar sensors is used to monitor the respiratory rate, heart rate, and number of times the elderly turn over during sleep.

[0007] In some embodiments of the present invention, step S2, in which the single-modal health feature vector extracted from the video stream data specifically includes: Scores were given for walking speed, cadence, gait symmetry, duration of opening and closing doors, duration of meals, duration of sitting, and positive facial expressions.

[0008] In some embodiments of the present invention, the multimodal fusion model in step S3 is a time-series fusion model based on an attention mechanism. This model achieves precise focusing on the health status assessment of the elderly by dynamically allocating the weights of different modalities under different time windows.

[0009] In some embodiments of the present invention, the quantitative assessment result in step S3 is used to calculate the comprehensive health score S using the following formula: in: Represents a point in time A comprehensive health score; This represents the total number of data modalities collected. ; For modal indexes; Representing the The modality at time point Health characteristic values; The representation is dynamically calculated by the attention mechanism based on historical data, the first... The modality at time point The fusion weights, and satisfying .

[0010] To achieve the above objectives, a second aspect of the present invention proposes a multimodal fusion AI-based non-intrusive assessment system for the elderly, comprising: The data acquisition module is configured to collect raw multimodal data reflecting the physiological and behavioral states of the elderly without being noticed in their daily environment by using at least two non-contact sensors. The data processing module is connected to the data acquisition module and is configured to preprocess the raw multimodal data and extract single-modal health feature vectors from the data of each modality using a preset feature extraction model. A multimodal fusion assessment module, connected to the data processing module, is configured to use a multimodal fusion model to perform weighted fusion and time-series analysis on the multiple single-modal health feature vectors to generate a quantitative assessment result representing the current comprehensive health status of the elderly. The risk warning module is connected to the multimodal fusion assessment module and is configured to compare the quantitative assessment results and their dynamic change trends with preset health baselines and risk thresholds. When the warning conditions are met, the module generates and outputs corresponding health risk warning information.

[0011] In some embodiments of the present invention, the data acquisition module specifically includes: At least one wide-angle camera, a microphone array, and a millimeter-wave radar sensor are integrated or distributed in the main activity space of the elderly.

[0012] In some embodiments of the present invention, the data processing module includes: A visual analysis unit for identifying key points of the human skeleton and analyzing gait parameters from video stream data; An audio processing unit for separating human voices from audio stream data, performing voiceprint recognition, and emotion calculation; A radar signal processing unit for extracting vital signs signals from radar point cloud data.

[0013] In some embodiments of the present invention, the multimodal fusion evaluation module incorporates a neural network based on the Transformer architecture to implement the above-mentioned attention-based time series fusion module. The risk warning module is connected to a user terminal application or a community health service platform, and the health risk warning information specifically includes: Fall risk level, warning of declining cognitive ability, abnormal sleep quality report, or indication of social isolation tendency.

[0014] To achieve the above objectives, a third aspect of the present invention provides an electronic device including a memory, a processor, and a computer program stored in the memory. When the computer program is executed by the processor, it implements the above-described multimodal fusion AI-based non-invasive elderly assessment method.

[0015] The multimodal fusion AI-based non-contact assessment method, system, and electronic device of this invention achieves non-contact data collection by deploying non-contact multimodal sensors in the daily environment of the elderly, without requiring the elderly to cooperate, effectively improving ease of use and acceptance; through preprocessing and feature extraction of multimodal data, it can mine multi-dimensional health features such as gait, voice, and vital signs from visual, audio, and radar data respectively, breaking through the limitations of traditional single data dimension technology, and realizing a comprehensive representation of the physiological and behavioral state of the elderly; Furthermore, relying on a multimodal fusion model based on an attention mechanism, it can dynamically allocate the weights of different modalities in different time windows, accurately focus on key health information, solve the problem of poor adaptability of traditional static assessments, improve the pertinence and accuracy of comprehensive health status assessments, provide continuous and reliable support for the whole-cycle health management of the elderly, and at the same time, the non-contact collection and data processing process takes into account privacy protection, further ensuring the user experience. Attached Figure Description

[0016] Figure 1 This is a flowchart illustrating the multimodal fusion AI-based non-invasive elderly assessment method provided by the present invention. Figure 2 This is a schematic diagram of the structural framework of the multimodal fusion AI-based non-intrusive assessment system for the elderly provided by the present invention; Figure 3 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0017] Embodiments of the present invention are described in detail below, examples of which are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.

[0018] The following description, with reference to the accompanying drawings, describes an embodiment of the multimodal fusion AI-based sensorless assessment system, method, and electronic device for elderly people.

[0019] Example 1: Figure 1 This is a flowchart illustrating a multimodal fusion AI-based non-invasive elderly assessment method according to an embodiment of the present invention.

[0020] like Figure 1 As shown, the multimodal fusion AI-based non-distracted assessment method for the elderly includes the following steps: S1. In the daily living environment of the elderly, raw multimodal data reflecting their physiological and behavioral states are collected imperceptibly using at least two types of non-contact sensors. The core objective is to comprehensively acquire multi-dimensional raw data reflecting their physiological and behavioral states without interfering with their normal lives or requiring their deliberate cooperation. The raw multimodal data includes at least the following: 1. Video stream data collected by visual sensors is used to analyze the gait, daily activities (ADL), and facial expressions of the elderly. Employing a 1080P resolution wide-angle camera with a frame rate of 30fps to balance acquisition accuracy and storage consumption, the lens covers a 120° wide angle, ensuring a single camera can capture the elderly person's complete activity trajectory within the room. Data acquisition is triggered automatically via infrared human body sensing—when an elderly person enters the acquisition area, the camera automatically switches to working mode; after leaving, it switches to low-power standby mode to avoid invalid data collection. The acquired video stream data is transmitted to the local processing unit in real time. The original video is not stored initially; only keyframes (1 frame every 2 seconds) are cached for subsequent feature extraction, reducing storage pressure and minimizing privacy risks. This video stream data can be used to analyze gait (e.g., calculating stride length and width through skeletal keypoint recognition), daily activity abilities (e.g., recording the time spent opening and closing doors, and the smoothness of grasping utensils during meals), and facial expressions (e.g., judging emotional positivity through facial feature point extraction), providing a direct representation of the elderly person's behavioral state.

[0021] 2. Audio stream data collected by audio sensors is used to analyze the speech tone, cough frequency, and social interaction activities of the elderly. Employing a microphone array with a 16kHz sampling rate, it supports simultaneous acquisition of four audio signals. Deployment avoids high-frequency noise sources such as kitchen range hoods and bathroom exhaust fans to ensure an audio signal-to-noise ratio ≥35dB. The acquisition mode is set to "all-weather low-power monitoring," achieving seamless triggering through a voiceprint wake-up mechanism—automatically increasing sampling accuracy when a human voice signal is detected; maintaining a low sampling rate in standby mode when there is no sound, with power consumption controlled below 50mW to avoid frequent charging and disturbing the elderly. The acquired audio stream data undergoes real-time voice separation, filtering out environmental noise before being temporarily stored as WAV format audio segments (single storage duration ≤30 seconds, continuous voices are automatically spliced). This audio stream data can be used to analyze speech tone (e.g., statistically analyzing speech rate changes and pitch fluctuations to determine cognitive ability trends), cough frequency (identifying cough sounds through voiceprint features, statistically analyzing daily cough frequency and duration to assist in assessing respiratory health), and social interaction activities (identifying dialogue scenarios, statistically analyzing daily social time and number of participants, and warning of social isolation tendencies), supplementing behavioral and social dimensions that physiological data cannot cover.

[0022] 3. Radar point cloud data collected by millimeter-wave radar sensors is used to monitor the respiratory rate, heart rate, and number of times the elderly turn over during sleep. Employing a 77GHz millimeter-wave radar, the output point cloud data boasts a horizontal resolution ≤0.1° and a distance resolution ≤0.5m. The acquisition frequency is dynamically adjusted based on scenario requirements: the radar under the mattress in the bedroom is set to 1Hz to focus on physiological signals during sleep; the radar in the bathroom is set to 5Hz to prioritize activity safety; and the radar in the corridor area is set to 2Hz to assist in gait analysis. The radar utilizes a non-line-of-sight acquisition design, penetrating thin fabrics and wooden furniture without direct contact with the human body, and produces no visible light or sound waves, completely avoiding privacy violations and usage interference. The acquired radar point cloud data is analyzed using micro-Doppler effect and distance-velocity characteristics to extract chest and abdominal undulation signals for calculating respiratory rate (measurement error ≤2 breaths / minute), capture micro-vibrations of heartbeats for monitoring heart rate (measurement error ≤3 breaths / minute), and simultaneously analyze point cloud contour changes to statistically determine the number of times the elderly turn over during sleep (accuracy ≤1 time / hour), achieving precise monitoring of the elderly's nighttime physiological state and sleep quality.

[0023] Furthermore, the aforementioned sensors achieve data time-series alignment through the NTP time synchronization protocol, ensuring that data from different modalities can be correlated and analyzed at the same time. For example, when an elderly person walks, the gait data captured by the camera, the trunk movement data collected by the radar, and the accompanying voice recorded by the microphone can be accurately mapped to the same time point. During data collection, all sensors adopt a low-power design, allowing for continuous operation for more than 30 days on a single charge. The gateway also has a local data caching function, ensuring that data is not lost even during a brief network outage and is automatically retransmitted upon network restoration. In addition, the operating status of all sensors has no obvious audible or visual indication; normal operation is only indicated by a faint green light from the gateway indicator, allowing them to blend seamlessly into the home environment and avoiding any psychological burden for the elderly regarding being monitored.

[0024] S2. Preprocess the original multimodal data and use a preset feature extraction model to extract single-modal health feature vectors from the data of each modality.

[0025] The single-modal health feature vector extracted from video stream data specifically includes: walking speed, cadence, gait symmetry, door opening and closing time, meal time, sedentary time, and facial expression positivity score.

[0026] For example, the preprocessing steps include noise filtering, temporal alignment, and missing value imputation.

[0027] 1. Noise filtering; Video stream: A combined spatiotemporal filtering method is used, first processing the video frames... Gaussian filtering (standard deviation 0.8) followed by 5 frames of temporal median filtering is applied to eliminate pixel noise.

[0028] Audio stream: Audio data acquisition may include background noise, such as electric fan noise, kitchen noise, etc. Noise suppression algorithms (such as Wiener filtering, spectral subtraction) need to be applied to improve the signal-to-noise ratio and ensure that the voice, cough and other features of the elderly can be clearly identified.

[0029] Radar point cloud: Radar signals may contain some irrelevant signals from other objects or the environment. These signals need to be processed using algorithms such as Doppler effect filters and Kalman filters to ensure that accurate vital signs information is extracted.

[0030] 2. Time-series alignment and missing value imputation; Data from different modalities are typically sampled at different frequencies. First, it is necessary to align the data from different modalities using timestamps to ensure that all data are correlated at the same point in time.

[0031] For missing data points, interpolation methods (such as linear interpolation or Lagrange interpolation) can be used to fill them, ensuring the continuity and integrity of the data.

[0032] 3. Feature extraction; (1) Feature extraction from video stream data: Gait analysis: Calculates gait speed, cadence, gait symmetry, etc., by using human skeleton key point recognition algorithms (such as OpenPose) in the video stream; Daily Activity Analysis (ADL): such as the length of time doors are opened and closed, and the duration of meals; Facial expression analysis: By extracting facial feature points, a sentiment analysis algorithm is used to calculate the positivity score of the expression.

[0033] (2) Feature extraction of audio stream data: Speech analysis: Extracting features such as pitch and speech rate from speech using voiceprint recognition technology to determine changes in cognitive ability; Cough frequency analysis: The frequency and duration of coughs are extracted using audio feature recognition technology; Social interaction activities: Analyze daily social interaction time and number of participants to detect social isolation tendencies.

[0034] (3) Feature extraction from radar point cloud data: Physiological data such as respiratory rate and heart rate are extracted using the micro-Doppler effect; Calculate behavioral data such as the number of times a person turns over during sleep.

[0035] For example, Feature vectors extracted from video stream data: Walking speed = 0.75m / s; Step frequency = 80 steps / minute; Gait symmetry = 0.85, indicating high gait symmetry; Door opening / closing time = 4 seconds; Meal duration = 25 minutes; Sedentary time = 2 hours; Positive facial expression score = 7 (out of 10); Feature vectors extracted from audio stream data: A 15% change in tone of voice indicates significant emotional fluctuation. Coughing frequency = 2 times / day; Social interaction activity = 30 minutes; Feature vectors extracted from radar point cloud data: Respiratory rate = 18 breaths / minute; Heart rate = 75 beats / minute; Number of times to turn over = 5 times / hour.

[0036] Assuming that the data for each modality is extracted using a certain algorithm, health feature values ​​are generated. (in (For modal indexing, for time points). For example, gait speed, cadence, etc., obtained from gait analysis are... The features in the vector. For each modality's feature vector. This is then passed as input to the multimodal fusion model for weighted fusion.

[0037] Through step S2, we can extract representative health features from data of different modalities. Preprocessing ensures the accuracy and consistency of the data, laying the foundation for subsequent multimodal fusion and health assessment.

[0038] S3. A multimodal fusion model is employed to perform weighted fusion and time-series analysis on multiple single-modal health feature vectors, generating a quantitative assessment result representing the current comprehensive health status of the elderly. The core of this step is to use a multimodal fusion model to weightedly fuse health features from different modalities and generate a comprehensive health score based on time-series analysis, thereby comprehensively and accurately assessing the health status of the elderly. The multimodal fusion model is a time-series fusion model based on an attention mechanism. This model achieves precise focus in assessing the health status of the elderly by dynamically allocating the weights of different modalities under different time windows.

[0039] Implementation details: 1. Multimodal fusion model: Model type: This invention uses a multimodal fusion model based on an attention mechanism. This model dynamically adjusts the weights of different modalities, enabling it to adaptively focus on the features most important for health assessment of older adults.

[0040] Time series analysis: Each modality of data (such as video, audio, and radar) typically exhibits time-series characteristics; therefore, the impact of the time dimension on health assessment must be considered. Through time series analysis, the model can capture the changing trends of health characteristics within different time windows, improving the accuracy of the assessment.

[0041] 2. Weighted fusion and weight allocation: Weighted fusion: Health feature vector for each modality The importance at time point t is determined by weight. Decision. Weight It is dynamically calculated based on historical data through an attention mechanism, and can reflect the degree of contribution of different modalities to health assessment at different times.

[0042] Weight normalization: To ensure that the contribution of each mode is relatively reasonable, all weights must satisfy the normalization condition: , where n is the total number of modes. It is the weight of the i-th mode at time t.

[0043] 3. Comprehensive Health Score Calculation: Based on health characteristics of different modalities and corresponding weights The overall health score S can be calculated using the following formula: Where: S is the comprehensive health score at time point t; It is the health feature value of the i-th mode at time point t; It is the weighted fusion weight of the i-th modality at time point t.

[0044] 4. Weighting criteria for modalities: Weights of each mode It is trained based on historical data of the model and reflects the impact of modalities on health assessment at different time periods. For example, if an older adult's gait and facial expressions change significantly over a period of time, the weights of the gait and facial expression modalities may be increased, and vice versa.

[0045] For example, suppose there are three modalities (video, audio, and radar data). After preprocessing and feature extraction, health features are obtained from the data of each modality. , ,and .

[0046] Suppose at a certain point in time The health features extracted from each modality are as follows: features of the video modality. (Step speed); characteristics of audio modality (Variation range of speech intonation); Characteristics of radar modes (Respiratory rate); After model training, the weights for each modality are obtained: the weights for the video modality. Weights of audio modalities Radar mode weights ; So, at the point in time Comprehensive health score It can be calculated using the following formula: Step S3, through multimodal fusion and temporal analysis, dynamically adjusts the weights based on the health characteristics of different modalities, generating a comprehensive health score. This method comprehensively considers multiple data sources such as video, audio, and radar, accurately reflecting the changing trends in the health status of the elderly, and possesses high adaptability and accuracy.

[0047] S4. Based on the quantitative assessment results and their dynamic trends, compare them with the preset health baseline and risk thresholds. When the warning conditions are met, generate and output corresponding health risk warning information. The core of this step is to compare the quantitative assessment results with the health baseline and risk thresholds to determine whether there is a health risk and generate corresponding warning information. This process can issue timely warnings based on changes in the health status of the elderly, helping to prevent potential health problems.

[0048] Implementation details: 1. Health baseline and risk threshold: Health baseline: A health baseline is a personalized health status reference value established based on factors such as an older adult's personal health history, age, and gender. It is typically based on extensive data analysis, medical records, and health examination results, reflecting the normal range of health for older adults.

[0049] Risk threshold: A risk threshold refers to a health warning standard set based on the advice of medical experts or health research. When an elderly person's health assessment result exceeds this threshold, the system will automatically trigger an alert. The threshold setting can take into account changes in multiple dimensions of indicators, including physiological, psychological, and behavioral factors.

[0050] 2. Dynamic Trends: Health status is dynamic, therefore it is necessary to track the health scores of the elderly in real time. The trend refers to the rate and direction of change in the assessment results. If an older adult's health score fluctuates significantly or shows an adverse trend (e.g., a sharp decline) in the short term, it indicates that their health status may be problematic.

[0051] To capture this trend, a moving average or exponentially weighted average method can be used to smooth the evaluation results and calculate the rate of change.

[0052] 3. Health risk warning: Warning conditions: When the health score When a patient's health falls below a baseline and the rate of change exceeds a preset threshold, the system will issue a health risk warning. This warning may include information such as fall risk, cognitive decline, and sleep quality problems.

[0053] The formula can be: Calculation of the rate of change in health scores: in, It is the current health score. This is the health score from the previous moment. If If the absolute value of the value exceeds a certain threshold, a health risk warning may be triggered.

[0054] Determining early warning conditions: Setting a lower threshold for health scores and the threshold for rate of change in rating If the following conditions are met: This will generate health risk warning information.

[0055] For example, suppose an older person's health score at a certain point in time... and for: ; And set a lower limit threshold for health scores. Health score change rate threshold .

[0056] Calculate the rate of change in health scores: because Exceeded the set threshold ,and Below the lower limit of health score Therefore, the system will generate and output health risk warning information, indicating that the elderly person may be facing health risks.

[0057] Health risk warnings include: risk of falls, warnings of cognitive decline, and risk of acute respiratory diseases.

[0058] Step S4 compares the health score with a preset health baseline and risk threshold, and considers the score's trend to determine whether a health risk warning has been triggered. When an elderly person's health score declines significantly or fluctuates abnormally, the system will promptly generate a health risk warning and provide timely intervention measures. This warning mechanism based on dynamic data analysis can effectively improve the response speed and accuracy of health management for the elderly, providing comprehensive protection for their health.

[0059] Example 2: Corresponding to the above method embodiments, such as Figure 2 This paper showcases a multimodal AI-based, non-intrusive health assessment system for the elderly. The system collects data using various types of sensors and employs processing, assessment, and early warning modules to monitor and evaluate the health status of senior citizens. The overall system architecture can be divided into the following main modules: The data acquisition module is configured to collect raw multimodal data reflecting the physiological and behavioral states of the elderly without being noticed in their daily environment by using at least two non-contact sensors. The data acquisition module specifically includes at least one wide-angle camera, a microphone array, and a millimeter-wave radar sensor, which are integrated or distributed in the main activity spaces of the elderly.

[0060] The data processing module, connected to the data acquisition module, is configured to preprocess the raw multimodal data and extract single-modal health feature vectors from the data of each modality using a preset feature extraction model. The data processing module includes: A visual analysis unit for identifying key points of the human skeleton and analyzing gait parameters from video stream data; An audio processing unit for separating human voices from audio stream data, performing voiceprint recognition, and emotion calculation; A radar signal processing unit for extracting vital signs signals from radar point cloud data.

[0061] The multimodal fusion assessment module, connected to the data processing module, is configured to use a multimodal fusion model to perform weighted fusion and time-series analysis on multiple single-modal health feature vectors to generate a quantitative assessment result representing the current comprehensive health status of the elderly. The multimodal fusion assessment module has a built-in neural network based on the Transformer architecture to implement the aforementioned attention-based time-series fusion module.

[0062] Corresponding to the above method embodiments, such as Figure 2This paper showcases a multimodal AI-based, non-intrusive health assessment system for the elderly. The system collects data using various types of sensors and employs processing, assessment, and early warning modules to monitor and evaluate the health status of senior citizens. The overall system architecture can be divided into the following main modules: The data acquisition module is configured to collect raw multimodal data reflecting the physiological and behavioral states of the elderly without being noticed in their daily environment by using at least two non-contact sensors. The data acquisition module specifically includes at least one wide-angle camera, a microphone array, and a millimeter-wave radar sensor, which are integrated or distributed in the main activity spaces of the elderly.

[0063] The data processing module, connected to the data acquisition module, is configured to preprocess the raw multimodal data and extract single-modal health feature vectors from the data of each modality using a preset feature extraction model. The data processing module includes: A visual analysis unit for identifying key points of the human skeleton and analyzing gait parameters from video stream data; An audio processing unit for separating human voices from audio stream data, performing voiceprint recognition, and emotion calculation; A radar signal processing unit for extracting vital signs signals from radar point cloud data.

[0064] The multimodal fusion assessment module, connected to the data processing module, is configured to use a multimodal fusion model to perform weighted fusion and time-series analysis on multiple single-modal health feature vectors to generate a quantitative assessment result representing the current comprehensive health status of the elderly. The multimodal fusion assessment module has a built-in neural network based on the Transformer architecture to implement the aforementioned attention-based time-series fusion module.

[0065] I. Based on the monitoring needs of high-frequency activity areas of the elderly (bedroom, living room, bathroom, kitchen), three types of non-contact sensors and local processing gateways are deployed in a differentiated manner using the hardware sensing layer. The specific parameters and deployment logic are as follows: 1. Millimeter-wave radar sensor (Model: HT-MTTR-L1) (1) Deployment locations: under the mattress in the bedroom (1 unit, short-range mode), on the ceiling of the bathroom (1 unit, short-range mode), in the corner of the living room (1 unit, long-range mode). (2) Core parameters include: Operating frequency: 77GHz; Distance measurement range: short distance mode 0.2m~70m, long distance mode 70m~500m; Distance resolution: 0.39m in short-range mode, 1.79m in long-range mode; Angle accuracy: ±1° for short distance mode, ±0.1° for long distance mode; Detection period: 70ms; Target capture rate: ≥98%, supports simultaneous tracking of 256 targets; Power supply and communication: DC12V power supply, ZigBee protocol communication, continuous operation for ≥30 days on a single charge.

[0066] (3) Scene adaptation logic: The bedroom needs to monitor the micro-vibrations of the chest and abdomen (breathing, heart rate) during sleep. The short-range mode can ensure a distance accuracy of 0.1m. The bathroom space is small (4㎡~8㎡). The short-range mode has a 60° horizontal expansion angle to achieve no dead angle coverage. The living room needs to track gait and activity trajectory. The long-range mode has a 9° horizontal expansion angle to cover a range of 15㎡~20㎡, avoiding wall penetration interference.

[0067] 2. Visual sensor (1080P wide-angle camera) (1) Deployment location: living room wall (height 1.8m, downward angle 15°), above bedroom door frame (height 2.2m, downward angle 20°) (2) Core parameters: Resolution: 1920×1080 (1080P); Frame rate: 30fps; Lens wide-angle: 120°; Triggering method: Human infrared sensor (detection distance 0.5m~5m); Data storage: Only keyframes are cached (2-second interval / frame), the original video is not stored, and the keyframe data size is 50KB / frame; Synchronization mechanism: Aligned with millimeter-wave radar via NTP time synchronization protocol, with a time error ≤20ms.

[0068] 3. Audio sensor (4-channel microphone array) (1) Deployment locations: kitchen ceiling (1 set), next to the TV cabinet in the living room (1 set) (2) Core parameters: Sampling rate: 16kHz; Signal-to-noise ratio: ≥35dB; Wake-up method: Voiceprint wake-up, voice recognition threshold 60dB; Power consumption mode: Switch to low power mode when there is no sound, power consumption ≤50mW; Storage format: WAV format, single storage duration ≤ 30 seconds, continuous human voices are automatically spliced; Sound source localization: The direction of the sound source is located using beamforming technology, with a localization error of ≤0.5m.

[0069] 4. Local processing gateway (1) Hardware configuration: Quad-core ARM Cortex-A53 processor (1.5GHz), 16GB local cache; (2) Communication protocol: Supports ZigBee (sensor data reception) and Wi-Fi 6 (cloud data upload); (3) Core functions: receive sensor data and perform preliminary preprocessing, cache key data within 7 days; when the network is interrupted, the basic assessment model can be run locally, such as fall risk judgment, and the data will be automatically retransmitted after the network is restored; support remote configuration of sensor parameters, such as radar acquisition frequency and camera trigger threshold.

[0070] II. The software modules follow a closed-loop logic of data input, processing, fusion, and output, as detailed below: 1. Data Preprocessing Module (1) Noise filtering ① Audio stream noise filtering (improved spectral subtraction) formula: in: The amplitude and phase combination of the processed clean audio signal in the frequency domain; : Discrete frequency point index in the frequency domain (value range: 08000, corresponding to 08kHz frequency at a sampling rate of 16kHz). The amplitude of the noisy signal in the frequency domain; The amplitude of the noise signal in the frequency domain estimated through a silent period (a period without any human voice or activity sound, with a duration of ≥1 second); Noise suppression coefficient (value 1.2, determined through testing with 1000 sets of samples to ensure signal-to-noise ratio is improved to ≥40dB); The phase of the noisy signal in the frequency domain (remains unchanged to avoid audio distortion).

[0071] ② Formula for radar point cloud noise filtering (statistical outlier filtering): like Then the point cloud data points are removed; in, : No. The average distance (in meters) from each radar point cloud data point to its neighboring points. : Index of radar point cloud data points (value range: 0~255, corresponding to a maximum of 256 points per frame); : No. The average distance (in meters) between the 20 neighboring points of a point (selected by the K-nearest neighbor algorithm, K=20). : No. The standard deviation of the average distance (in meters) between 20 neighboring points of a given point.

[0072] ③ Video stream noise filtering (joint spatiotemporal filtering): First use Gaussian filtering (standard deviation) After eliminating spatial domain pixel noise, a 5-frame temporal median filter is applied (to eliminate inter-frame jitter in the temporal domain), resulting in a video frame edge retention rate of ≥95%.

[0073] in, Standard deviation of Gaussian filter (value 0.8, balancing noise reduction and edge preservation).

[0074] (2) Temporal alignment and missing value filling ① Video frame missing padding (weighted average method) The formula can be: in, : The filled pixel matrix of the t-th missing video frame (dimension 1920×1080, corresponding to 1080P resolution). : Timestamp of the next valid frame after the t-th missing frame (in seconds); : Timestamp of the previous valid frame of the t-th missing frame (in seconds); : No. The pixel matrix of valid frames at any given time; : No. The pixel matrix of the valid frames at each time point.

[0075] ② Time alignment: Based on the timestamp of the millimeter-wave radar (denoted as...). ), for video keyframe timestamps ( ), audio segment timestamps ( Alignment is achieved through linear interpolation, resulting in a time error ≤10ms. , .

[0076] 2. Feature Extraction Module (1) Visual feature extraction (based on HRNet model): ①The formula for calculating walking speed is: in, Walking speed of the elderly (unit: m / s, normal range 0.8~1.3m / s). Ankle point Actual displacement over time (unit: m); : The time interval for step speed measurement (valued at 0.5s, for balance accuracy and real-time performance). Ankle point Pixel displacement within a time interval (unit: pixels, calculated by extracting ankle coordinates using the HRNet model, coordinates are...). and , ; Camera calibration coefficient (value 0.02m / pixel, determined using a checkerboard calibration plate to ensure accuracy in converting pixels to actual distance). For example, if... pixels, then m, m / s (within the normal range).

[0077] ②The formula for calculating gait symmetry can be: in, Gait symmetry score (value range [0,1], ≥0.8 is normal); Time delay weight (value 0.6, determined through clinical data; time delay has a greater impact on symmetry). Displacement difference weight (value 0.4, compared with...) The sum is 1 to ensure weight normalization. Time delay of movement at the left and right hip joints (in seconds, determined by cross-correlation function) calculate, This is the time offset. The maximum value corresponding to That is ); Gait cycle duration (unit: seconds, i.e., the time to complete one full gait, determined by the peak interval of the vertical displacement sequence of the hip joint). The maximum difference in horizontal displacement between the left and right ankles (unit: m). , , (These are the horizontal displacement sequences of the left and right ankle points, respectively). Stride length (unit: m, i.e. the actual distance of a single step, determined by the maximum horizontal displacement of the ankle).

[0078] For example, if , , , ,but (Within the normal range).

[0079] (2) Radar feature extraction The following features were extracted using microDoppler effect analysis: respiratory rate Unit: times / minute; measurement error ≤ 2 times / minute; normal range 12~20 times / minute. Heart rate Unit: times / minute; measurement error ≤ 3 times / minute; normal range 60~100 times / minute. Number of times to turn over Unit: times / hour, accuracy ≤ 1 time / hour, normal range for nighttime (22:00~6:00) is 2~5 times / hour.

[0080] (3) Audio feature extraction Extracted based on a ResNet-18 fine-tuned model (training set: 11 classes of emotional speech datasets, 8:2 split between training and validation sets): speech rate change rate Unit: %, Normal range ±15%, i.e. ; Cough frequency Unit: times / day, ≤5 times / day is normal.

[0081] 3. Multimodal fusion evaluation module (1) Fusion model: Transformer model based on attention mechanism, with 13 dimensions of input feature vector, including: visual 6 dimensions: walking speed, gait symmetry, door opening and closing time, meal time, sedentary time, and facial expression positivity score; radar 3 dimensions: breathing frequency, heart rate, and number of times to turn over; audio 4 dimensions: speech rate change rate, cough frequency, speech tone fluctuation, and social communication time.

[0082] (2) Dynamic weight allocation: Let the weight of visual features be... Radar feature weights are The audio feature weights are ,satisfy The weighting varies depending on the risk assessment scenario: Fall risk assessment: , , (Radar displacement characteristics are more sensitive to falls); Cardiovascular risk assessment: , , (Radar heart rate characteristics are more critical).

[0083] (3) The formula for calculating the comprehensive health index is: in, Comprehensive Health Index (range 0-100 points, ≥80 points is excellent, 60-79 points is good, 40-59 points is moderate risk, ≤39 points is severe risk). Visual feature score (0-100 points, obtained by weighted average of 6 visual features, the weights of which are determined according to clinical importance); Radar feature score (0~100 points, obtained by weighted average of 3 radar features); Audio feature score (0~100 points, obtained by weighted average of 4 audio features).

[0084] 4. Early warning output module Warning levels: Mild warning ( ): Send health advice to the elderly person's children's app; Moderate warning ( ): Linking with community doctor terminals to push elderly people's identity information and abnormal characteristics; Severe warning ( ): Triggers emergency resource dispatch and pushes the elderly person's real-time location (based on gateway positioning) to the ambulance dispatch center.

[0085] Response latency: End-to-end latency ms, where data processing time is 1 second. milliseconds (ms) are the time it takes for an alert to be pushed out. ms.

[0086] The complete system implementation process is as follows: 1. Deployment preparation phase (1-2 days) 1.1 Environmental Survey: Record the spatial layout of the elderly person's home, mark 5 core areas including the bedroom and living room, and test the ZigBee communication coverage strength, requiring coverage ≥95%; 1.2 Equipment Selection: Determine the number of sensors (e.g., 3 radars, 2 cameras, and 2 microphones for a two-bedroom, one-living room apartment) and the gateway installation location (center of the living room to ensure optimal signal coverage) based on the survey results.

[0087] 2. Hardware installation phase (1 day) 2.1 Sensor Mounting: Radar: Mounted under the mattress in the bedroom, 1.2m from the head of the bed; installed by drilling a hole in the bathroom ceiling, 2.5m from the ground; wall-mounted in the corner of the living room, 2.8m from the ground; Camera: Mounted on the living room wall, 1.8m from the ground, at a 15° downward angle; and pasted above the bedroom door frame to avoid damaging the wall; Microphone: Recessed in the kitchen ceiling; placed next to the TV cabinet in the living room, 1.2m from the ground. 2.2 Gateway Configuration: Connect to home Wi-Fi and complete sensor pairing (NTP time synchronization, communication protocol settings) through the management platform.

[0088] 3. Software debugging phase (2 days) 3.1 Model Deployment: Load the HRNet, ResNet-18, and Transformer fusion model on the local gateway, import the basic health data of the elderly (age, basic medical history) to generate a personalized health baseline, such as the heart rate baseline of a 65-year-old hypertensive elderly person is 70~85 beats / minute; 3.2 Parameter Calibration: Camera Calibration: Determined using a checkerboard calibration board. Radar accuracy test: Place a simulated target at a known distance and adjust the parameters to make the ranging error ≤0.1m; 3.3 Functional test: Simulate a fall (person falls to the ground) and abnormal cough (play cough audio) to verify that the early warning trigger accuracy rate is ≥95%.

[0089] 4. Trial operation phase (7 days) 4.1 Data Collection: Collect multimodal data continuously for 7 days, and generate a health report daily (including...). Values ​​and anomaly characteristics); 4.2 Error Correction: If the step speed measurement error is >0.1m / s, recalibrate. If the false alarm rate is greater than 5%, adjust the weights of the fusion model. 4.3 User adaptation: Guide family members to use the APP to view reports and set the frequency of alert push notifications (such as pushing daily reports at 9:00 AM).

[0090] 5. Formal Operation and Maintenance Phase 5.1 Daily Operation: Automatically collects data and performs real-time calculations Values ​​that are abnormal trigger an alert; 5.2 Regular maintenance: Remotely calibrate sensor parameters monthly and update the fusion model quarterly (adding new scene data); 5.3 Data Management: Local cache retains data for 7 days, and cloud encrypted backup is performed for 3 months (compliant with GB / T35273-2020). Doctors need authorization to access historical data.

[0091] This system achieves seamless data collection by deploying non-contact multimodal sensors in the daily environment of the elderly, without requiring their deliberate cooperation, thus effectively improving ease of use and acceptance. Through preprocessing and feature extraction of multimodal data, it can extract multi-dimensional health characteristics such as gait, voice, and vital signs from visual, audio, and radar data, breaking through the limitations of traditional single-data-dimensional technologies and achieving a comprehensive representation of the physiological and behavioral state of the elderly.

[0092] Example 3: Corresponding to the above embodiments, the present invention also proposes an electronic device.

[0093] like Figure 3The diagram shows a structural schematic of an electronic device according to the present invention. The electronic device 200 includes a processor 201 and a memory 203. The processor 201 and the memory 203 are connected, for example, via a bus 202. Optionally, the electronic device 200 may further include a transceiver 204. It should be noted that in practical applications, the transceiver 204 is not limited to one unit, and the structure of this electronic device 200 does not constitute a limitation on the embodiments of the present invention.

[0094] Processor 201 may be a CPU, a general-purpose processor, a DSP, an ASIC, an FPGA, or other programmable logic device, transistor logic device, hardware component, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in connection with this disclosure. Processor 201 may also be a combination that implements computational functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0095] Bus 202 may include a path for transmitting information between the aforementioned components. Bus 202 may be a PCI bus or an EISA bus, etc. Bus 202 may be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 3 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0096] The memory 203 stores a computer program corresponding to the multimodal fusion AI-based age-insensitive assessment method of the above embodiments of the present invention. This computer program is executed under the control of the processor 201. The processor 201 executes the computer program stored in the memory 203 to implement the content shown in the aforementioned method embodiments.

[0097] Among them, electronic devices 200 include, but are not limited to: mobile terminals such as laptops and tablets, as well as fixed terminals such as desktop computers. Figure 3 The electronic device 200 shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of the present invention.

[0098] It should be noted that the logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, intelligently monitor, distribute, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, the computer-readable medium can even be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.

[0099] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0100] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0101] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0102] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.

Claims

1. A multimodal fusion AI-based method for sensory-free assessment of the elderly, characterized in that, Includes the following steps: S1. In the daily life environment of the elderly, raw multimodal data reflecting their physiological and behavioral states are collected imperceptibly through at least two non-contact sensors. S2. The original multimodal data is preprocessed, and a single modality health feature vector is extracted from the data of each modality using a preset feature extraction model. S3. Using a multimodal fusion model, the multiple single-modal health feature vectors are weighted, fused, and analyzed over time to generate a quantitative assessment result representing the current comprehensive health status of the elderly. S4. Based on the quantitative assessment results and their dynamic trends, compare them with the preset health baseline and risk threshold. When the warning conditions are met, generate and output the corresponding health risk warning information.

2. The method according to claim 1, characterized in that, The raw multimodal data in step S1 includes at least: Video stream data collected by visual sensors is used to analyze the gait, daily living abilities (ADL), and facial expressions of older adults. Audio stream data collected by audio sensors is used to analyze the speech tone, cough frequency, and social interaction activities of older adults. Radar point cloud data collected by millimeter-wave radar sensors is used to monitor the respiratory rate, heart rate, and number of times the elderly turn over during sleep.

3. The method according to claim 1 or 2, characterized in that, In step S2, the single-modal health feature vector extracted from the video stream data specifically includes: Scores were given for walking speed, cadence, gait symmetry, duration of opening and closing doors, duration of meals, duration of sitting, and positive facial expressions.

4. The method according to claim 1, characterized in that, The multimodal fusion model in step S3 is a time-series fusion model based on an attention mechanism. This model achieves precise focusing on the health status assessment of the elderly by dynamically allocating the weights of different modalities under different time windows.

5. The method according to claim 1, characterized in that, The quantitative assessment result in step S3 is used to calculate the comprehensive health score S using the following formula: in: Represents a point in time A comprehensive health score; This represents the total number of data modalities collected. ; For modal indexes; Representing the The modality at time point Health characteristic values; The representation is dynamically calculated by the attention mechanism based on historical data, the first... The modality at time point The fusion weights, and satisfying .

6. A multimodal AI-based, non-intrusive elderly assessment system, characterized in that, include: The data acquisition module is configured to collect raw multimodal data reflecting the physiological and behavioral states of the elderly without being noticed in their daily environment by using at least two non-contact sensors. The data processing module is connected to the data acquisition module and is configured to preprocess the raw multimodal data and extract single-modal health feature vectors from the data of each modality using a preset feature extraction model. A multimodal fusion assessment module, connected to the data processing module, is configured to use a multimodal fusion model to perform weighted fusion and time-series analysis on the multiple single-modal health feature vectors to generate a quantitative assessment result representing the current comprehensive health status of the elderly. The risk warning module is connected to the multimodal fusion assessment module and is configured to compare the quantitative assessment results and their dynamic change trends with preset health baselines and risk thresholds. When the warning conditions are met, the module generates and outputs corresponding health risk warning information.

7. The system according to claim 6, characterized in that, The data acquisition module specifically includes: At least one wide-angle camera, a microphone array, and a millimeter-wave radar sensor are integrated or distributed in the main activity space of the elderly.

8. The system according to claim 6, characterized in that, The data processing module includes: A visual analysis unit for identifying key points of the human skeleton and analyzing gait parameters from video stream data; An audio processing unit for separating human voices from audio stream data, performing voiceprint recognition, and emotion calculation; A radar signal processing unit for extracting vital signs signals from radar point cloud data.

9. The system according to claim 6, characterized in that, The multimodal fusion evaluation module incorporates a neural network based on the Transformer architecture to implement the attention-based time series fusion module as described in claim 4. The risk warning module is connected to a user terminal application or a community health service platform, and the health risk warning information specifically includes: Fall risk level, warning of declining cognitive ability, abnormal sleep quality report, or indication of social isolation tendency.

10. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory, which, when executed by the processor, implements the multimodal fusion AI-based age-insensitive assessment method as described in any one of claims 1-5.

Citation Information

Patent Citations

  • An AI-based health status monitoring system for the elderly

    CN118248337B

  • A method and system for monitoring and evaluating elderly health based on big data

    CN118629659B

  • AI digital human-based behavior prediction method, system and device, and storage medium

    CN119762930A

  • Pig individual identification and health monitoring system based on biological characteristics

    CN120077966A

  • Health monitoring system and method based on multi-sensor fusion

    CN120319508A

Cited By

  • Multi-modal dynamic evaluation method for intrinsic ability of old people and cloud platform

    CN121768675A