Old people health state analysis and early warning method and system based on multi-modal data, edge computing equipment and medium

By acquiring physiological data, facial videos, and inertial movement data of the elderly, and combining them with edge computing devices to generate health analysis results and early warning information, the problem of low efficiency in traditional health analysis methods has been solved, realizing intelligent monitoring of the health status of the elderly and meeting the needs of the healthy aging model.

CN121938618APending Publication Date: 2026-04-28南京津发健康产业有限公司 +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
南京津发健康产业有限公司
Filing Date
2025-12-16
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Traditional health analysis methods are inefficient and lack intelligence in monitoring the health of the elderly, failing to meet the needs of a healthy aging model.

Method used

By acquiring physiological data, facial videos, and inertial movement data of the elderly, and combining physiological signal characteristic parameters, emotional state, attention level, and body posture, health analysis results and early warning information are generated. Edge computing devices are used to achieve automatic, real-time, long-term, efficient, and intelligent health status monitoring.

Benefits of technology

It enables continuous, accurate, real-time, long-term, objective, efficient, and intelligent monitoring and analysis of the health status of the elderly, improving the efficiency and intelligence of health monitoring and meeting the needs of the healthy aging model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121938618A_ABST
    Figure CN121938618A_ABST
Patent Text Reader

Abstract

The invention provides an old people health state analysis and early warning method and system based on multi-modal data, edge computing equipment and a medium. The method comprises the steps that physiological data, face videos, inertial motion data and basic information of old people are acquired; determining physiological signal characteristic parameters of the old people according to the physiological data; according to the physiological data and the face video, identifying an emotional state and an attention level of the elderly; identifying the body posture of the old person according to the inertial motion data; and according to the physiological signal characteristic parameters, the emotion state and attention level recognition result, the body posture recognition result and the basic information of the elderly, generating a health analysis result of the elderly and / or early warning information for prompting the elderly about abnormal state. According to the invention, the health analysis of the old people is carried out based on the multi-modal data, the health state of the old people can be monitored and analyzed continuously, accurately, in real time, in a long term, integrally, objectively, efficiently and intelligently, and the health monitoring and analysis efficiency and intelligence of the old people are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of intelligent health analysis technology, and in particular to a method and system for analyzing and providing early warning of the health status of the elderly based on multimodal data, as well as edge computing devices and media. Background Technology

[0002] As the elderly population expands, the health risks faced by the elderly are becoming increasingly complex, including sudden illnesses (such as myocardial infarction), deterioration of chronic diseases (such as complications of hypertension), accidental falls, emotional disorders (such as anxiety and grief), and cognitive decline (such as mild cognitive impairment). Society's demand for a "healthy aging model" has become more urgent than ever before.

[0003] Traditional health analysis methods rely on manual observation or periodic testing for data collection, resulting in low efficiency and intelligence in health monitoring and analysis of the elderly, which cannot meet the needs of a healthy aging model. Summary of the Invention

[0004] In view of this, the purpose of this disclosure is to provide a method and system for analyzing and warning the health status of the elderly based on multimodal data, as well as edge computing devices and media, to improve the efficiency and intelligence of health monitoring, analysis and warning for the elderly.

[0005] To achieve the above objectives, this disclosure provides a method for analyzing and warning of the health status of the elderly based on multimodal data, comprising: acquiring physiological data, facial video, inertial movement data, and basic information of the elderly; determining physiological signal characteristic parameters of the elderly based on the physiological data; identifying the emotional state and attention level of the elderly based on the physiological data and the facial video; identifying the body posture of the elderly based on the inertial movement data; and generating health analysis results and / or warning information for indicating abnormal conditions of the elderly based on the physiological signal characteristic parameters, the identification results of emotional state and attention level, the identification results of body posture, and the basic information.

[0006] This disclosure also provides a health status analysis and early warning system for the elderly, comprising: a multimodal data acquisition module for acquiring physiological data, facial video, inertial movement data, and basic information of the elderly; an elderly physiological signal feature parameter recognition module for determining the physiological signal feature parameters of the elderly based on the physiological data; an elderly emotional state and attention level recognition module for recognizing the emotional state and attention level of the elderly based on the physiological data and the facial video; an elderly body posture recognition module for recognizing the body posture of the elderly based on the inertial movement data; and a health analysis and early warning module for generating health analysis results and / or early warning information indicating abnormal conditions of the elderly based on the physiological signal feature parameters, the recognition results of emotional state and attention level, the recognition results of body posture, and the basic information.

[0007] This disclosure also provides an edge computing device, including: a memory for storing a computer program; and a processor for executing the computer program to implement the steps of the method described in any of the foregoing embodiments.

[0008] This disclosure also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the method described in any of the preceding embodiments.

[0009] The technical solution provided in this disclosure acquires physiological data, facial videos, inertial movement data, and basic information of the elderly. Based on the elderly's basic information, physiological signal characteristic parameters determined using physiological data, emotional state and attention level identified using physiological data and facial videos, and body posture identified using inertial movement data, it generates health analysis results and / or early warning information to indicate abnormal conditions in the elderly. This achieves accurate, comprehensive, and objective analysis of the elderly's health status by combining multimodal data, and enables automatic, real-time, long-term, efficient, and intelligent monitoring and analysis of the elderly's health status. In other words, this disclosure enables continuous, accurate, real-time, long-term, holistic, objective, efficient, and intelligent monitoring and analysis of the elderly's health status, improving the efficiency and intelligence of health monitoring and analysis, and meeting the needs of healthy aging models.

[0010] Additional aspects and advantages of this disclosure will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this disclosure. Attached Figure Description

[0011] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 A flowchart illustrating a method for analyzing and providing early warning of the health status of the elderly based on multimodal data, provided for some embodiments of this disclosure; Figure 2 An architecture diagram of an emotion attention recognition model provided in some embodiments of this disclosure; Figure 3 Provided for some embodiments of this disclosure and Figure 2 A training diagram of the corresponding emotion attention recognition model; Figure 4 A schematic diagram illustrating the loss and accuracy changes during dual-task alternating pre-training provided in some embodiments of this disclosure; Figure 5 This is a schematic diagram illustrating the fine-tuning of an emotion state recognition task after alternating training of two tasks, as provided in some embodiments of this disclosure. Figure 6 An architecture diagram of a pose recognition model provided in some embodiments of this disclosure; Figure 7 An architecture diagram of a second model to be trained provided for some embodiments of this disclosure; Figure 8 This is a schematic diagram illustrating the training of a pose recognition model provided in some embodiments of this disclosure; Figure 9 A flowchart illustrating a method for analyzing and providing early warning of the health status of the elderly based on multimodal data, provided for some embodiments of this disclosure; Figure 10 This is a schematic diagram of a health analysis report provided in some embodiments of this disclosure. Detailed Implementation

[0013] The embodiments of this disclosure will be further described in detail below with reference to the accompanying drawings and examples. The detailed description of the embodiments and the accompanying drawings are used to illustrate the principles of this disclosure by way of example, but should not be used to limit the scope of this disclosure. This disclosure can be implemented in many different forms and is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.

[0014] These embodiments are provided in this disclosure to make the disclosure thorough and complete, and to fully express the scope of the disclosure to those skilled in the art. It should be noted that, unless otherwise specifically stated, the relative arrangement of components and steps, the composition of materials, numerical expressions and values ​​set forth in these embodiments should be interpreted as exemplary only and not as limiting.

[0015] All terms used in this disclosure have the same meaning as understood by one of ordinary skill in the art to which this disclosure pertains, unless otherwise specifically defined. It should also be understood that terms defined in general dictionaries should be interpreted as having meanings consistent with their meanings in the context of the relevant art, and not as idealized or highly formalized, unless expressly defined herein.

[0016] Techniques, methods, and equipment known to those skilled in the art may not be discussed in detail, but where appropriate, they should be considered part of the specification.

[0017] It should be noted that, in the optional embodiments of this disclosure, the personnel information and other related data involved require the permission or consent of the personnel involved when the embodiments of this disclosure are applied to specific products or technologies. Furthermore, the collection, use, and processing of the relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. In other words, if the embodiments of this disclosure involve personnel-related data, it must be obtained with the authorization and consent of the personnel, the authorization and consent of the relevant departments, and in accordance with the relevant laws, regulations, and standards of the country and region. If the embodiments involve personal information, the acquisition of all personal information requires the consent of the individual. If sensitive information is involved, the separate consent of the information subject is required, and the embodiments must also be implemented with the authorization and consent of the subject.

[0018] This disclosure provides a method and system for analyzing and issuing early warnings of the health status of the elderly based on multimodal data, as well as edge computing devices and media. This improves the efficiency, intelligence, and user experience of health status analysis and early warning for the elderly. For example, by acquiring physiological data, facial videos, inertial motion data, and basic information of the elderly, and based on this information, physiological signal characteristic parameters determined using physiological data, emotional state and attention level identified using physiological data and facial videos, and body posture identified using inertial motion data, health analysis results and / or early warning information indicating abnormal health status are generated. This achieves accurate, comprehensive, and objective analysis of the health status of the elderly by combining multimodal data, and enables automatic, real-time, long-term, efficient, and intelligent monitoring and analysis of the health status of the elderly. In other words, this disclosure enables continuous, accurate, real-time, comprehensive, long-term, holistic, objective, efficient, and intelligent monitoring and analysis of the health status of the elderly, improving the efficiency and intelligence of health monitoring and analysis, and meeting the needs of healthy aging models.

[0019] See Figure 1 The flowchart below illustrates a method for analyzing and providing early warning of the health status of the elderly based on multimodal data, provided in some embodiments of this disclosure. The method may include the following steps: S11: Obtain physiological data, facial videos, inertial movement data, and basic information of the elderly.

[0020] In some embodiments of this disclosure, physiological data of the elderly can be acquired. For example, wearable devices can be worn on the elderly to acquire their physiological data. Compared to existing invasive tests (such as blood draws and complex instrument examinations) that are prone to causing discomfort, acquiring physiological data from wearable devices allows for the continuous, non-invasive, and low-intrusive collection of multi-dimensional physiological data from the elderly. This reduces discomfort caused by data collection, improves the accuracy of data acquisition, and consequently enhances the accuracy of emotional state and attention level recognition. This, in turn, improves the accuracy of health status analysis and early warning for the elderly, achieving an efficient, intelligent, and humane health and elderly care model. The physiological data involved in this disclosure includes various measurable data signals in the human body, including but not limited to electroencephalogram (EEG) signals, pulse signals, and electrodermal activity (EDA) signals. EEG signals: Macroscopic potential changes generated by the synchronous activity of neuronal groups in the cerebral cortex, recorded through scalp electrodes, reflect the state of neural activity and can be used for the recognition of emotional state, attention level, and sleep state. Pulse signals can be specifically obtained through photoplethysmography (PPG). PPG is a technique that uses light of a specific wavelength to illuminate the skin and detect the local arterial blood volume caused by heartbeats. It can be used to identify emotional states and attention levels, and to calculate cardiovascular-related indicators such as heart rate, heart rate variability (the small fluctuations in the RR interval between consecutive heartbeats, reflecting the autonomic nervous system's regulatory function on the cardiovascular system), pulse pressure, and dicrotic wave characteristics. Of course, pulse signals can also be obtained through other techniques. Electrodermal conductance (EDC) signals are electrophysiological parameters that reflect an individual's autonomic nervous activity and emotional arousal level by measuring the conductivity of the skin surface. They can be used to assess emotional states and attention levels.

[0021] Furthermore, facial videos of elderly individuals can be acquired to assist in the recognition of emotional states and attention levels. Specifically, facial videos of elderly individuals captured by a facial expression capture device can be acquired. For example, facial videos of elderly individuals can be captured using a laptop's built-in camera at a frame rate of 25 fps (frames per second), capturing color video (resolution of 720×1280). Alternatively, facial videos of elderly individuals can be acquired using a high-definition digital camera, action camera, or mobile device (such as a smartphone). This disclosure does not limit the devices used to acquire facial videos or their parameters.

[0022] Furthermore, it is possible to acquire inertial motion data of the elderly in order to identify their body posture based on this data. For example, inertial motion capture data of the elderly can be acquired by an inertial motion capture sensor, which can be 34-channel, with a sampling frequency of 40Hz, and includes a three-axis accelerometer, gyroscope, magnetometer, etc., and can be fixed to multiple parts of the human body (e.g., head, chest, elbow, shoulder, hip, waist, ankle, knee, wrist, etc.). In this way, the movement trajectory of the elderly can be continuously captured in a non-invasive or low-invasive manner.

[0023] In addition, basic information about the elderly can be obtained. This basic information includes, but is not limited to, age, gender, medical history, and lifestyle habits. For example, the basic information of an elderly person could be 75 years old, female, with a 5-year history of hypertension, and a habit of walking for 1 hour daily. Obtaining this basic information allows for analysis of the elderly person's health in conjunction with their individual medical history and lifestyle habits, thereby improving the accuracy of the health analysis.

[0024] It should be noted that physiological data, facial videos, and inertial movement data of the elderly can be continuously acquired to enable continuous, real-time, and long-term monitoring and analysis of their health status. For example, physiological data, facial videos, and inertial movement data can be acquired in 5-second input units. Five-second data sets reflect both the cyclical changes in physiological state and ensure real-time accuracy, thereby improving the accuracy of health status analysis and early warning for the elderly. Furthermore, the elderly health analysis method provided in this disclosure is applicable to scenarios where the aforementioned data is collected by wearing the corresponding acquisition device at fixed times, as well as scenarios where the data is collected by wearing the device for extended periods.

[0025] In some embodiments of this disclosure, after acquiring physiological data, facial videos, and inertial motion data of the elderly, data preprocessing can be performed on the acquired physiological data, facial videos, and inertial motion data to improve the quality of multimodal data. Data preprocessing may include at least one of invalid data removal, outlier removal, standardization processing, and data augmentation processing. Invalid data removal: Data distribution is checked by combining data collection scene records, etc., to remove abnormal segments. Outlier removal: Noise points are removed based on the 3σ criterion. Standardization processing: Z-score standardization is performed on EEG signals, pulse signals, and skin conductance signals, and frame extraction, grayscale processing, and cropping are performed on facial videos. Data augmentation: For time-series data, time axis flipping and adding ±5% Gaussian noise can be performed; for image data, random cropping, brightness adjustment, and horizontal flipping can be performed.

[0026] S12: Determine the physiological signal characteristic parameters of the elderly based on physiological data.

[0027] Based on the physiological data of the elderly, features can be extracted from the physiological data to obtain physiological signal feature parameters. Specifically, physiological signal feature parameters can be calculated using mathematical formulas based on the physiological data of the elderly. For example, for data of different modalities, time domain, frequency domain, and time-frequency domain methods can be used to extract physiological signal feature parameters. Among them, the physiological signal feature parameters mentioned in this disclosure include, but are not limited to: 3 EDA items: average SCR (skin conductance response) peak amplitude, rise time, and number of times per unit time; 6 HRV (heart rate variability) items: SDNN (standard deviation of RR intervals), RMSSD (root mean square of the difference in RR intervals), NN50 (number of times the difference between adjacent RR intervals is >50ms), LF / HF (low-frequency component / high-frequency component, reflecting sympathetic-vagal nerve balance), average rise time, and dicrotic wave ratio; 3 EEG items: peak-to-peak value, standard deviation, and wavelet instantaneous energy; and 3 other items: heart rate, pulse pressure, and sleep duration.

[0028] For example, for EEG signals, time-domain methods can be used to extract peak-to-peak value, standard deviation, and Hjorth parameters; frequency-domain methods can be used to extract relative power (the proportion of energy from a specific channel to the total energy in the brain region); and time-frequency domain methods can be used to extract wavelet instantaneous energy (based on continuous wavelet transform, which can reflect emotional state). Simultaneously, according to international standards, EEG rhythms can be divided into Delta (0.5-4Hz), Theta (4-8Hz), Alpha (8-13Hz), Beta (13-30Hz), and Gamma (>30Hz) frequency bands to create brain topography maps and visualize brain region activity. For pulse signals, heart rate indicators can be calculated, including SDNN, RMSSD, NN50, pNN50 (NN50 percentage), LF / HF ratio, pulse pressure (systolic pressure - diastolic pressure), mean rise time (reflecting left ventricular ejection efficiency), and dicrotic wave characteristics (proportion and amplitude ratio of dicrotic waves). For electrodermal signals, skin conductance level (SCL, the slowly changing baseline component of the ESC, reflecting an individual's basic autonomic nervous activity) can be calculated: SCL slope (reflecting the rate of change of SCL) and SCL standard deviation (reflecting stability), which can be used to assess the stability of emotional state; skin conductance response (the rapidly changing component of the ESC superimposed on the SCL, the transient response caused by stimulation, which can be used to assess emotional arousal and attentional concentration): total number, number of times per unit time, average peak amplitude, and average rise time.

[0029] S13: Identify the emotional state and attention level of the elderly based on physiological data and facial video.

[0030] Based on the acquisition of physiological data and facial videos of elderly individuals, the system can identify their emotional states (e.g., anxiety, calmness, anger, happiness, sadness, irritability, pain, etc.) and attention levels (e.g., highly focused, somewhat focused, moderately focused, somewhat distracted, highly distracted, etc.). For facial videos, the video can be first segmented into frames (e.g., frame extraction, for example, processing the facial video at 1 frame per second) to obtain multiple facial images. Then, the physiological data and multiple facial images are used to identify the elderly person's emotional state and attention level. The facial images used for emotional state and attention level identification can be color images or grayscale images after grayscale processing. The facial images can also be cropped (e.g., 720×1280 → 240×32) to reduce the computational load of the model. Specifically, physiological data and multiple facial images can be input into a pre-trained AI (Artificial Intelligence) model for emotion state and attention recognition to identify the emotional state and attention level of the elderly. This method enables accurate and joint recognition of emotional state and attention level, achieving a lightweight model architecture, reducing model parameters, and making it compatible with low-power wearable devices (such as wristbands and portable sensors), while facilitating real-time monitoring. Alternatively, physiological data and multiple facial images can be input into a pre-trained AI model for emotion state recognition to identify the emotional state of the elderly, and physiological data and multiple facial images can be input into a pre-trained AI model for attention level recognition to identify the attention level of the elderly.

[0031] S14: Identify the body posture of the elderly based on inertial motion data.

[0032] Based on the acquisition of inertial movement data of the elderly, their body postures (such as sitting, falling, running, lying down, etc.) can be identified using this data. This enables real-time and accurate identification of the elderly's body postures, allowing for focused attention on fall risks when falls are detected, thus ensuring the safety of the elderly. Specifically, the inertial movement data of the elderly can be input into a pre-trained AI model for recognizing body postures to identify the elderly's postures.

[0033] S15: Based on the physiological signal characteristic parameters, emotional state and attention level recognition results, body posture recognition results, and basic information of the elderly, generate health analysis results and / or early warning information to indicate abnormal conditions in the elderly.

[0034] Building upon the steps outlined above, health analysis of the elderly can be conducted based on their physiological signal characteristics, emotional state and attention levels, body posture, and basic information. This analysis generates health analysis results and / or early warning information indicating abnormal health conditions, enabling the generation of health analysis results and / or early warning information using multi-dimensional health information. For example, the generated health analysis results may include health status grading (e.g., healthy / borderline / mildly abnormal / moderately severe abnormality, etc.) or health status scores, single-indicator risk assessment results, multi-dimensional risk analysis (identifying ≥2 abnormal indicators) results, trend predictions (e.g., Parkinson's risk, cognitive decline trend, etc.), and actionable intervention recommendations (emergency warnings, exercise / nutrition recommendations, medical follow-ups, etc.), adapting to individual differences (age, medical history, lifestyle habits, etc.) and enhancing the humanistic level of health management.

[0035] For example, the physiological signal characteristic parameters, emotional state and attention level recognition results, body posture recognition results, and basic information of the elderly can be input into a Large Language Model (LLM). The LLM then performs health analysis based on this data to generate health analysis results for the elderly. Therefore, some embodiments of this disclosure integrate multi-dimensional health data of the elderly, utilizing this data to perform health analysis, thereby improving the accuracy of health analysis and enabling real-time, continuous, and long-term health monitoring and analysis. Compared to existing professional assessments that rely on medical personnel and are time-consuming and labor-intensive, some embodiments of this disclosure can automatically analyze and generate health analysis results for the elderly without relying on medical personnel for assessment.

[0036] Another example is the ability to determine whether an elderly person is experiencing an abnormal state based on their physiological signal characteristics, emotional state, attention level, body posture, and basic information. For instance, it can be determined whether the elderly person's physiological signal characteristics are outside the normal range, whether their emotional state has changed abruptly, or whether they are experiencing cognitive decline based on their attention level and physiological signal characteristics. If an abnormal state is identified, early warning information (such as the specific abnormal condition) can be generated to alert the elderly person. This allows for real-time warnings when an abnormal state is detected, thereby reducing disability and mortality rates caused by sudden health risks, achieving "early detection and early intervention," and significantly improving the elderly person's sense of security in independent living. The aforementioned methods enable real-time health status perception, such as immediate fall warnings and alerts for sudden changes in emotional state. Therefore, the health analysis method for the elderly provided in some embodiments of this disclosure has stronger proactive early warning and predictive capabilities. Compared with the existing technologies, which are mostly passive responses (such as ringing a bell after a fall or seeking medical attention after feeling unwell), or statistical risk assessments based on historical data, this disclosure can identify current abnormalities in real time and provide proactive early warnings so as to facilitate timely intervention and slow down the progression of diseases.

[0037] As described above, the elderly health analysis method provided in this disclosure can solve the problems of insufficient real-time performance, single data dimensions, weak predictive ability, and strong device intrusion in existing elderly health analysis methods. This disclosure can achieve accurate and real-time identification of multi-dimensional health indicators: accurately identify the emotional state, attention level, body posture, and physiological signal characteristic parameters of the elderly through multimodal data, and can establish a "real-time monitoring-long-term analysis" system: realize real-time perception of the health status of the elderly through continuous data collection and time series analysis; make trend predictions (such as Parkinson's risk, cognitive decline, etc.) based on long-term accumulated data combined with large language models, proactively warn of potential risks, and provide actionable intervention suggestions, thereby improving the humanization level of health management and filling the gap of "short-term offline analysis" in existing technologies. At the same time, it can automatically analyze and generate health analysis results for the elderly without relying on medical staff for assessment, reducing labor costs and improving the efficiency of health analysis.

[0038] In some embodiments of this disclosure, identifying the emotional state and attention level of an elderly person based on physiological data and facial video may include: inputting various physiological data and multiple facial images obtained by segmenting facial video frames into a pre-trained emotion and attention recognition model to obtain the identification results of the elderly person's emotional state and attention level. The emotion and attention recognition model includes: a first feature extraction layer, a first feature fusion layer, a first emotion state recognition classification layer, and a first attention level recognition classification layer. Inputting various physiological data and multiple facial images obtained by segmenting facial video frames into the pre-trained emotion and attention recognition model to obtain the identification results of the elderly person's emotional state and attention level may include: using the first feature extraction layer to extract physiological features from various physiological data and facial expression features from multiple facial images; using the first feature fusion layer to fuse the physiological features and facial expression features corresponding to various physiological data to obtain fused features; and using the first emotion state recognition classification layer and the first attention level recognition classification layer to identify the emotion state and attention level respectively, to obtain the identification results of the elderly person's emotional state and attention level.

[0039] In some embodiments of this disclosure, various physiological data and multiple facial images obtained by segmenting facial videos can be input into a pre-trained emotion and attention recognition model. The pre-trained emotion and attention recognition model can be used to obtain the recognition results of the emotional state and attention level of the elderly. In this way, a single model can be used to jointly recognize the emotional state and attention level of the elderly, thereby building a lightweight and deployable model architecture. While ensuring recognition accuracy, the model parameters are reduced, meeting the low power consumption and real-time computing requirements of wearable devices.

[0040] The emotion and attention recognition model can include a first feature extraction layer, a first feature fusion layer, a first emotion state recognition classification layer, and a first attention level recognition classification layer. Based on this, the process of inputting multiple physiological data and multiple facial images into the pre-trained emotion and attention recognition model to obtain the recognition results of the elderly person's emotion state and attention level can be as follows: (a) Using the first feature extraction layer to extract corresponding physiological features from each type of physiological data, and using the first feature extraction layer to extract facial expression features from multiple facial images. That is, the first feature extraction layer is a multimodal feature extraction layer, and features of multimodal data are extracted using this first feature extraction layer. (b) Using the first feature fusion layer to fuse the physiological features and facial expression features corresponding to each type of physiological data to obtain fused features. That is, the first feature fusion layer is a multimodal feature fusion layer. (c) Using the first emotion state recognition classification layer to perform emotion state recognition on the fused features to obtain the recognition result of the elderly person's emotion state; using the first attention level recognition classification layer to perform attention level recognition on the fused features to obtain the recognition result of the elderly person's attention level. The structures of the emotion state recognition classification layer and the attention level recognition classification layer can be identical, each containing two Linear layers (fully connected layers) and a Softmax layer (soft maximum value layer or normalized exponential layer). Taking the seven emotion states—anxiety, calmness, anger, joy, sadness, irritability, and pain—as an example, the first emotion state recognition classification layer maps the fused features to the emotion state (7 categories) output space and calculates the classification probability to obtain the recognition result of the elderly person's emotion state. Similarly, taking the five attention levels—very focused, fairly focused, moderately focused, somewhat distracted, and very distracted—as an example, the first attention level recognition classification layer maps the fused features to the attention level (5 categories) output space and calculates the classification probability to obtain the recognition result of the elderly person's attention level. As can be seen from the above, the emotion and attention recognition model has a three-layer architecture of "multimodal feature extraction → multimodal feature fusion → multi-task classification". The emotion state and attention task share the first feature extraction layer and the first feature fusion layer (accounting for more than 98% of the model parameters). It can effectively reduce the model parameters (reduced by nearly 50%) while ensuring recognition accuracy, meet the low power consumption and real-time computing requirements of wearable devices, and reduce computing costs.

[0041] Of course, time-domain / frequency-domain features of various physiological data can also be extracted, concatenated, and input into classifiers such as SVM (Support Vector Machine) and Random Forest to identify emotional states and attention levels, respectively. Alternatively, a single ResNet1D network can be used to process multimodal data. Specifically, various physiological data can be concatenated into time-series data of a unified dimension and directly input into the ResNet1D network for feature extraction and classification to identify emotional states and attention levels, etc.

[0042] In some embodiments of this disclosure, physiological data may include at least one of the following: electroencephalogram (EEG) signals of an elderly person collected by a wearable EEG device, pulse signals of an elderly person collected by a pulse wave sensor, and electrodermal signal of an elderly person collected by a skin conductance sensor. The first feature fusion layer is used to fuse the physiological features and facial expression features corresponding to the physiological data to obtain fused features. This can include: splicing the physiological features and facial expression features corresponding to various physiological data, using a self-attention mechanism to extract features again from the spliced ​​features, and then fusing the results of the second feature extraction to obtain fused features.

[0043] In some embodiments of this disclosure, wearable devices can be worn on elderly individuals to acquire physiological data collected by the wearable devices. Specifically, the wearable device for acquiring the physiological data of the elderly can be at least one of a wearable electroencephalogram (EEG), a pulse wave sensor, and a skin conductance sensor. That is, the acquired physiological data of the elderly can include at least one of the following: brainwave signals acquired by the wearable EEG, pulse signals acquired by the pulse wave sensor, and skin conductance signals acquired by the wireless skin conductance sensor.

[0044] The wearable EEG device can have 32 channels and a sampling frequency of 256Hz. It records the synchronous activity potential of neurons in the cerebral cortex through scalp electrodes, reflecting the state of neural activity. It can capture signals in the Delta (0.5-4Hz), Theta (4-8Hz), and Alpha (8-13Hz) frequency bands. The different brain regions and their locations can be numbered as follows: Fp1, Fpz, Fp2, AF3, AF4, F7, F3, Fz, F4, F8, FC5, FC1, FC2, FC6, T7, C3, Cz, C4, T8, CP5, CP1, CP2, CP6, P7, P3, Pz, P4, P8, Poz, O1, Oz, O2. The pulse sensor can specifically be a PPG wireless blood volume pulse sensor, which can be single-channel (i.e., a PPG wireless blood volume pulse sensor can reflect pulse information in a single channel) with a sampling frequency of 64Hz. It uses green light transmission technology to monitor changes in blood perfusion with the pulse, and can be used to monitor the pulse and calculate cardiovascular indicators such as heart rate, heart rate variability, and pulse pressure. The skin conductance sensor can be an EDA wireless skin conductance sensor, which can be single-channel (i.e., an EDA wireless skin conductance sensor can reflect skin conductivity in a single channel) with a sampling frequency of 64Hz. By measuring the surface conductivity of the skin, it reflects the level of autonomic nervous activity and emotional arousal (e.g., skin conductivity increases during anxiety and anger). Physiological data of the elderly can be collected non-invasively or with minimal impact through various wearable devices. Furthermore, integrated data acquisition using multi-source wearable devices enables spatiotemporal synchronization and low-invasive acquisition of multimodal data, ensuring data quality and reproducibility.

[0045] Based on the above, the method for obtaining fused features using the first feature fusion layer can be as follows: Physiological features and facial expression features corresponding to various physiological data are concatenated; a self-attention mechanism is used to extract features again from the concatenated features; and the results of the second feature extraction are then fused to obtain the fused features. The self-attention mechanism dynamically calculates the correlation weights between elements in the features, optimizing feature interactions to improve the model's ability to capture global information, thereby improving the accuracy of emotional state and attention level recognition. For example, the first feature fusion layer can sequentially include a fully connected layer, a self-attention mechanism layer, and a fully connected layer: the fully connected layer concatenates physiological features and facial expression features corresponding to various physiological data; the self-attention mechanism layer extracts features again from the concatenated features; and the fully connected layer fuses the results of the second feature extraction from the self-attention mechanism layer to obtain the fused features. For example, the first feature fusion layer combines EEG features (256 dimensions), pulse features (128 dimensions), skin conductance features (128 dimensions), and facial expression features (128 dimensions) into 640-dimensional features. It optimizes feature interaction through a self-attention mechanism and outputs 128-dimensional fused features.

[0046] For example, see Figure 2 This is an architectural diagram of the emotion attention recognition model provided in some embodiments of this disclosure. It should be noted that... Figure 2 This explanation uses the PPG pulse signal as an example; however, pulse signals obtained through other methods can also be used. Figure 2 .

[0047] If the physiological data includes EEG signals, then the first feature extraction layer is used to extract EEG features from the EEG signals. The process may include: extracting the first feature (temporal feature) of the EEG signal in parallel using first one-dimensional convolutional kernels of different sizes; extracting the second feature from the first feature using a nonlinear pooling layer (ReLU+Maxpooling) corresponding to the first one-dimensional convolutional kernel; concatenating the second features extracted by each nonlinear pooling layer using a concatenation layer; extracting the third feature from the concatenated second feature using a multi-head self-attention mechanism layer, so as to capture multiple different levels of dependency patterns from the concatenated second feature simultaneously through the multi-head self-attention mechanism, making the model's understanding of the sequence more three-dimensional and comprehensive, and enhancing the model's representation ability; and processing the concatenated third feature using a first fully connected layer (Linear) (e.g., transforming the feature dimension through linear transformation) to obtain the EEG features. For example, the input dimension of the EEG signal can be (B, 32, 1280). Temporal features are extracted in parallel using one-dimensional convolutional kernels of sizes 15, 25, 51, and 65. After dimensionality reduction by max pooling, a multi-head self-attention mechanism is introduced to strengthen temporal dependence, ultimately outputting 256-dimensional features. Therefore, EEG features can be extracted using CNN (Convolutional Neural Network, CNN extracts local features) from EEG signals.

[0048] If the physiological data includes a pulse signal, then the first feature extraction layer is used to extract pulse features from the pulse signal. This process may include: extracting the fourth feature of the pulse signal using a second one-dimensional convolutional kernel; processing the fourth feature using a first rectified linear unit (ReLU) (introducing nonlinearity, feature selection, and sparsification); processing the fifth feature obtained from the first rectified linear unit using a first LSTM (Long Short-Term Memory) network; and processing the first LSTM result obtained from the first LSTM using a second fully connected layer to obtain the pulse features. For example, the input dimension of the pulse features can be (B, 1, 320). Local features are extracted through shallow 1D convolutions, input to an LSTM to model long-term dependencies, and mapped to 128-dimensional features through a fully connected layer. Therefore, a hybrid CNN and LSTM model (CNN extracts local features, LSTM captures temporal dependencies) can be used for pulse signals to improve the accuracy of pulse feature extraction.

[0049] If the physiological data includes electrodermal signal (EDS), the first feature extraction layer extracts EDS features from the EDS signal. This process may include: extracting the sixth feature of the EDS signal using a third one-dimensional convolutional kernel; processing the sixth feature using a second rectified linear unit (RLU); processing the seventh feature obtained from the second RLU using a second LSTM; and processing the second LSTM result obtained from the second LSTM using a third fully connected layer to obtain the EDS features. For example, the input dimension of the EDS signal can be (B, 1, 320). Local features are extracted through shallow 1D convolution, and the input is used to model long-term dependencies using an LSTM, which is then mapped to 128-dimensional features via a fully connected layer. Therefore, a hybrid CNN and LSTM model (CNN extracts local features, LSTM captures temporal dependencies) can be used for EDS signals to improve the accuracy of EDS feature extraction.

[0050] To extract facial expression features from multiple facial images using the first feature extraction layer, the process can include: extracting the eighth feature from multiple facial images using ResNet (residual network); extracting the ninth feature from the eighth feature using a max pooling layer; and processing the ninth feature using a fourth fully connected layer to obtain the facial expression features. The facial images can be color or grayscale images. Furthermore, the ResNet network depth can be 18 or 50, etc. For example, inputting six 240×320 grayscale images, extracting image features using ResNet18, passing through a max pooling layer, and then mapping to 128-dimensional features through a fully connected layer. Therefore, ResNet (for extracting deep image features) can be used for facial images to improve the accuracy of facial expression feature extraction.

[0051] In some embodiments of this disclosure, pre-training an emotion attention recognition model may include: pre-training a first model to be trained using an emotion state attention dataset to obtain a pre-trained first model; freezing the second feature extraction layer and the second feature fusion layer in the pre-trained first model; training the second emotion state recognition classification layer in the pre-trained first model using sample data containing emotion state labels in the emotion state attention dataset; and training the second attention level recognition classification layer in the pre-trained first model using sample data containing attention level labels in the emotion state attention dataset to obtain an emotion attention recognition model.

[0052] In some embodiments of this disclosure, when training the emotion and attention recognition model, various physiological signals and facial videos can be collected from the subject to construct an emotion and attention dataset. Taking physiological signals including electroencephalogram (EEG), pulse, and ductal nerve conductance (TEF) signals as an example, the input to this dataset includes EEG, pulse, TEF signals, and facial videos (or facial images), labeled with multiple emotion states and multiple attention levels, such as 7 emotion states and 5 attention levels. Simultaneously, the subject's inertial motion data can be collected to construct a posture recognition dataset. The input to this dataset is inertial motion data, labeled with multiple body postures (e.g., 4 body postures). It should be noted that a One-Hot (one-bit effective encoding) coding method can be used to assign labels.

[0053] Specifically, taking physiological signals including EEG signals, pulse signals, and skin conductance signals as examples, the equipment selection and parameters can be as follows: Wear-EEG wearable EEG device: 32 channels, acquisition frequency 256Hz, capable of capturing Delta (0.5-4Hz), Theta (4-8Hz), Alpha (8-13Hz) and other frequency band signals; EDA wireless skin conductance sensor: 1 channel, acquisition frequency 64Hz; PPG wireless blood volume pulse wave sensor: 1 channel, acquisition frequency 64Hz; Facial expression acquisition device: using a laptop's built-in camera, frame rate 25fps, capturing color video to assist in emotional state recognition; Inertial motion capture sensor: 34 channels, acquisition frequency 40Hz, including a three-axis accelerometer, gyroscope, and magnetometer, fixed to 14 parts of the human body (head, chest, elbow, shoulder, etc.), capturing human motion trajectory and recognizing sitting posture, falling posture, and other postures.

[0054] Data collection process and scenarios: Data can be collected according to the preset experimental process, covering multiple health scenarios: (1) Baseline resting scenario: sit / stand and relax for 1 minute to collect basic physiological data; (2) Emotional state induction scenario: induce multiple emotional states (e.g., 7 emotional states) through short videos (anger), pictures (pleasant, anxious, sad), mini-games (irritability), etc., and collect data for each emotional state for 2 minutes; (3) Attention level test scenario: have the subjects memorize and write 20 numbers under different noises (e.g., keyboard sounds, cymbal sounds, human voice dialogue, etc.) to simulate the state of attention concentration / distraction, etc.; (4) Posture simulation scenario: collect inertial movement data of various body postures, such as: collect data of four types of body postures: sitting posture, running posture, falling (falling forward, sliding sideways, etc.) and lying flat, as well as sleep and pain stimulation (light / medium / heavy pinching of the arm) scenario data. For details, please refer to Table 1, which is a table of the collection frequency and number of channels of each device: Table 1. Acquisition frequency and number of channels for each device

[0055] Specifically, data can be collected from several elderly subjects for several consecutive days, with several hours of data collection each day, to construct an emotional state attention dataset and a posture recognition dataset. For example, see Table 2, which shows the data collection table for 100 elderly subjects: Table 2 Data Collection Table of 100 Elderly Subjects

[0056] The dataset can be constructed with 5 seconds as an input unit (i.e., a single sample data can be 5 seconds of data). The 5-second data volume can fully reflect a physiological state change cycle of the subject and effectively limit the processing length, thereby enhancing the speed of data analysis and processing and ensuring real-time performance.

[0057] For example, see Table 3, which is a data summary table for the emotion state attention dataset and the pose recognition dataset: Table 3. Data details for the Emotional State Attention Dataset and the Pose Recognition Dataset

[0058] After collecting the raw multimodal data, preprocessing such as cleaning and standardization can be performed to improve data quality and thus enhance the accuracy of model training. Specifically, data preprocessing can include at least one of the following: invalid data removal, valid data extraction, outlier removal, standardization, and data augmentation. Invalid data removal: Based on the data collection scenario (e.g., loose equipment, environmental electromagnetic interference), examine the data distribution and remove abnormal segments. Valid data extraction: Extract corresponding valid data segments according to the experimental procedure (e.g., emotional state induction time period, attention test time period). Outlier removal: Based on the 3σ criterion (data exceeding the mean ± 3 standard deviations is considered outlier), remove noise points from valid segments. Standardization: Perform Z-score standardization on EEG signals, pulse signals, and skin conductance signals (other standardization methods can also be used). Perform frame extraction, grayscale processing, and cropping (e.g., 720×1280 → 240×320) on facial videos to reduce model computation. Data augmentation: For time-series data, time axis flipping and adding ±5% Gaussian noise can be performed; for image data, random cropping, brightness adjustment, and horizontal flipping can be performed.

[0059] Alternatively, a first model to be trained can be constructed, which may include a second feature extraction layer, a second feature fusion layer, a second emotion state recognition and classification layer, and a second attention level recognition and classification layer. For example, see [link to example]. Figure 3 This provides for some embodiments of the present disclosure with Figure 2 A training diagram of the corresponding emotion attention recognition model.

[0060] After constructing the emotion state attention dataset and the first model to be trained, the first model to be trained can be pre-trained using the emotion state attention dataset to obtain the pre-trained first model. The architecture of the first model to be trained and the pre-trained first model are the same as the architecture of the emotion attention recognition model.

[0061] During pre-training, the first model to be trained can be trained alternately using sample data from the emotion state attention dataset with emotion state labels and sample data with attention level labels. For example, in each iteration, forward propagation, loss calculation, backpropagation, and parameter updates can be performed first using sample data with emotion state labels, followed by the same steps. Alternatively, training can be performed first using sample data with attention level labels and then using sample data with emotion state labels. Training can be iteratively continued until a termination condition is met (e.g., until the single-modal validation set performance converges, for example, EEG emotion state recognition accuracy ≥ 75%). By alternating training with emotion state and attention tasks, and sharing the second feature extraction layer and the second feature fusion layer, the task correlation is utilized to improve feature generalization ability and suppress overfitting.

[0062] After obtaining the pre-trained first model, the fine-tuning phase begins. Specifically, the second feature extraction layer and the second feature fusion layer in the pre-trained first model are frozen. The second emotion state recognition classification layer in the pre-trained first model is trained using sample data containing emotion state labels from the emotion state attention dataset, and the second attention level recognition classification layer in the pre-trained first model is trained using sample data containing attention level labels from the emotion state attention dataset, until the corresponding training termination condition is met, thus obtaining the emotion attention recognition model. That is, in the fine-tuning phase, the second feature extraction layer and the second feature fusion layer are frozen, and the two tasks (emotion state recognition task and attention level recognition task) are trained separately. The trained emotion attention recognition model is a pre-trained multi-task neural network that integrates physiological features and multimodal features of facial expressions to achieve accurate and joint recognition of emotion states and attention levels.

[0063] This disclosure employs joint fusion: directly concatenating the pre-trained feature vectors of each modality, inputting the fused feature vector into the fully connected layer + output layer, and simultaneously optimizing the single-modality branch loss and fusion loss. A gradient accumulation strategy can be used to increase the equivalent batch size: training strategies include adding a Dropout layer to the fully connected layer (dropout_rate=0.2); applying L2 regularization (λ=1e-5) to the model weights, etc. A cosine annealing strategy is adopted: base LR (learning rate) = 2e-5, T_max (full annealing cycle) = 50, eta_min (minimum threshold of learning rate) = 1e-6; linear warm-up for the first 1000 steps (increasing from 10% of the base LR to the target value); dynamic gradient clipping (max_grad_norm (maximum norm threshold for gradient clipping) = 1.0), triggering learning rate decay when the global gradient norm > 1.5 × max_grad_norm, with a training cycle of 100.

[0064] The model is pre-trained alternately using emotion state recognition and attention level recognition tasks to obtain efficient shared feature extraction and feature fusion layers. The above embodiments of this disclosure not only fully utilize the correlation between tasks, helping to improve the generalization ability of features, but also effectively suppress overfitting that may occur during single-task training when data is limited. After pre-training, the feature extraction and feature fusion layers are frozen, and single-task training is performed on the classification layers for different tasks, thus balancing the design principles of "feature sharing" and "task specialization." By sharing the feature extraction layer, which has the largest proportion of parameters, the overall number of model parameters can be reduced by nearly half, providing a feasible solution for embedded deployment of multiple physiological tasks. See Table 4 for details, which is a statistical table of model parameters. Table 4 Model Parameter Statistics

[0065] See Figure 4 This is a schematic diagram illustrating the changes in loss and accuracy during dual-task alternating pre-training provided in some embodiments of this disclosure, wherein... Figure 4 (a) in the diagram is a schematic diagram of the loss change during dual-task alternating pre-training (specifically, a schematic diagram of the change in training loss vs. validation loss). Figure 4 (b) in the diagram is a schematic diagram of the accuracy change during dual-task alternating pre-training (specifically, a schematic diagram of the change in training accuracy vs. validation accuracy). Figure 4 In (a), curve 1 represents the training loss, and each point on curve 1 represents the training loss of Task 1 and the training loss of Task 2 (Task 1 and Task 2 are arranged alternately, with the first point being Task 2, the second point being Task 1, and so on). Curve 2 represents the validation loss, and each point on curve 2 represents the validation loss of Task 1 and the validation loss of Task 2 (Task 1 and Task 2 are arranged alternately, with the first point being Task 2, the second point being Task 1, and so on). Figure 4 In (b), curve 3 represents training accuracy, with each point on curve 3 representing the training accuracy of Task 1, Task 2, etc. (the first point represents Task 1, the second point represents Task 2, and so on). Curve 4 represents validation accuracy, with each point on curve 4 representing the validation accuracy of Task 1, Task 2, etc. (Task 1 and Task 2 are alternated, with the first point representing Task 1, the second point representing Task 2, and so on). During the alternating pre-training process, the emotion state recognition task (Task 1) and the attention recognition task (Task 2)... Figure 4As can be seen, during the alternating training process, the overall loss function shows a decreasing trend, while the validation accuracy gradually improves and stabilizes, indicating that the alternating pre-training strategy can effectively promote the model's convergence domain feature learning. The results show that the validation accuracy for the emotion state recognition task reaches a maximum of approximately 87.54%, with a smooth convergence process, and the model's generalization performance on this task is superior to that of the attention recognition task; the validation accuracy for the attention recognition task is approximately 82%, indicating that this task is relatively more difficult to learn. Overall, the alternating pre-training strategy achieved certain performance on both tasks, demonstrating the existence of shared multimodal features for both tasks. The alternating training strategy also helps extract task-related features and improves the model's generalization ability.

[0066] See Figure 5 This is a schematic diagram illustrating the fine-tuning of the emotion state recognition task after alternating training of two tasks, provided by some embodiments of this disclosure. Figure 5 (a) in the figure is a schematic diagram of the loss changes during fine-tuning of the emotion state recognition task (specifically, a schematic diagram of the changes in training loss vs. validation loss). Figure 5 (b) in the diagram illustrates the changes in accuracy during fine-tuning of the emotion state recognition task (specifically, the changes in training accuracy versus validation accuracy). Figure 5 It can be seen that, due to the loading of pre-trained weights, the model's accuracy is above 80% in the initial stage and reaches 96.44% after the final training is completed. This means that fine-tuning based on pre-training can improve the accuracy of emotion state recognition.

[0067] Of course, the emotion state recognition task and the attention recognition task can also be trained independently. In this case, an emotion state recognition model and an attention recognition model can be included to perform emotion state recognition and attention level recognition, respectively. The emotion state recognition model can include a first feature extraction layer, a first feature fusion layer, and a first emotion state recognition classification layer. For example, such as... Figure 2 The architecture consists of a first feature extraction layer, a first feature fusion layer, and a first attention level recognition and classification layer. The attention recognition model may include a first feature extraction layer, a first feature fusion layer, and a first attention level recognition and classification layer, for example, such as... Figure 2 The architecture consists of a first feature extraction layer, a first feature fusion layer, and a first attention level recognition and classification layer.

[0068] See Figure 6This is an architectural diagram of a posture recognition model provided in some embodiments of this disclosure. In some embodiments of this disclosure, recognizing the body posture of an elderly person based on inertial motion data may include: inputting the inertial motion data into a pre-trained posture recognition model to obtain the recognition result of the elderly person's body posture, wherein the posture recognition model may include a first multi-stage feature extraction unit, a multi-branch convolutional refinement unit, and a first fusion classifier; inputting the inertial motion data into the pre-trained posture recognition model to obtain the recognition result of the elderly person's body posture may include: extracting features of the inertial motion data sequentially using a series of third feature extraction layers in the first multi-stage feature extraction unit; wherein the series of third feature extraction layers performs time dimension downsampling and feature dimension expansion on the inertial motion data, and the first multi-branch convolutional refinement unit may include a first convolutional refinement unit corresponding one-to-one with each third feature extraction layer; performing deep feature extraction on the features extracted by the corresponding feature extraction layer using each first convolutional refinement unit; and performing concatenation, feature fusion, and body posture classification processing on the features deeply extracted by each first convolutional refinement unit using a first fusion classifier to obtain the recognition result of the elderly person's body posture.

[0069] In some embodiments of this disclosure, the acquired inertial motion data of the elderly can be input into a pre-trained posture recognition model to identify the elderly person's body posture. This improves the accuracy of body posture recognition for the elderly, thereby enhancing the accuracy of health analysis for them.

[0070] Specifically, in some embodiments provided in this disclosure, the pose recognition model can be a multi-stage, multi-scale cascaded classification network to achieve real-time and accurate recognition of various body poses. The pose recognition model can adopt a structure of "multi-stage feature extraction → multi-branch convolutional refinement → fusion classifier," with the core being "gradual abstraction - fine processing - global fusion." Specifically, the pose recognition model can include a first multi-stage feature extraction unit, a first multi-branch convolutional refinement unit, and a first fusion classifier. The first multi-stage feature extraction unit includes a series of third feature extraction layers. Each third feature extraction layer contains multiple one-dimensional convolutional kernels, and the number of convolutional kernels increases sequentially from the topmost third feature extraction layer to the bottommost third feature extraction layer (i.e., from the shallowest third feature extraction layer to the deepest third feature extraction layer). Each third feature extraction layer can sequentially include a Conv (convolution, specifically a 1D convolution) layer, a BN (batch normalization) layer, a ReLU layer, and a MaxPool layer. The processing flow is Conv layer → BN layer → ReLU layer → MaxPool layer. The first multi-branch convolutional refinement unit includes first convolutional refinement units corresponding one-to-one with the serial third feature extraction layers. Each first convolutional refinement unit may include two L-Conv units (depth-separable convolutional coordinate attention modules) and a MaxPool layer connected in sequence. Each L-Conv unit may include a DWConv (depth-separable convolution) layer, a BN layer, a ReLU layer, and a CA (channel attention mechanism) layer connected in sequence. The CA layer includes an Avg-Pool (global average pooling) layer and a Linear layer. The first fusion classifier is... Figure 6 ClsALL in the model consists of a BN layer, a Linear layer, another BN layer, an ELU (Exponential Linear Unit) layer, and a Linear layer connected in sequence.

[0071] If the input dimension of the obtained inertial motion data of the elderly is in the form of (B, time step, feature dimension), in order to adapt to the third feature extraction layer in the first multi-stage feature extraction unit, the input dimension of the inertial motion data can be transposed to (B, feature dimension, time step).

[0072] Then, the features of the inertial motion data are extracted sequentially using the serial third feature extraction layers in the first multi-stage feature extraction unit. The inertial motion data is downsampled in the time dimension and expanded in the feature dimension through the serial third feature extraction layers. Specifically, the first multi-stage feature extraction unit may include three serial third feature extraction layers: the top third feature extraction layer (i.e., the first stage, corresponding to...) Figure 6 Stage 1 in the middle includes 32 3×3 1D convolutions, BN, ReLU, and MaxPool, with the third feature extraction layer (i.e., the second stage, corresponding to...) in the middle. Figure 6Stage 2 in the code includes 64 3×3 1D convolutions, BN, ReLU, and MaxPool, with the bottom third feature extraction layer (i.e., the third stage, corresponding to...) Figure 6 Stage 3 in the model includes 128 3×3 1D convolutions, BN, ReLU, and MaxPool. Taking inertial motion data with input dimensions (B, 201, 34), where 201 is the time step and 34 is the feature dimension as an example, its transpose becomes (B, 34, 201), and it is input into the first multi-stage feature extraction unit. First, the top-level third feature extraction layer sequentially enters 32 3×3 1D convolutions → batch normalization (BN) → ReLU → 2×2 max pooling, outputting (B, 48, 100) (channels from 34 to 48, time step 201 to 100). Then, the features obtained from the top-level third feature extraction layer enter the middle third feature extraction layer, sequentially entering 64 3×3 1D convolutions → BN → ReLU → 2×2 max pooling, outputting (B, 96, 50). After that, the features obtained from the middle third feature extraction layer enter the bottom-level third feature extraction layer, sequentially entering 128 3×3 1D convolutions → BN → ReLU → 2×2 max pooling, outputting (B, 192, 25). Thus, it can be seen that the first multi-stage feature extraction unit captures features from fine-grained (instantaneous action) to coarse-grained (global pose) through time dimension downsampling (201→100→50→25) and channel expansion (34→48→96→192).

[0073] The first multi-stage feature extraction unit uses deep and shallow layer feature extraction. Each stage uses the classic combination of "convolutional layer + batch normalization + ReLU activation function + max pooling" as the basic unit. Through the mutual learning between low-level texture features and high-level semantic features, high-dimensional information in the feature space is further captured, thereby realizing the gradual abstraction of features and semantic upgrade.

[0074] The top-level third feature extraction layer takes the transposed (B, 34, 201) feature tensor as input. First, it extracts local features using 32 3×3 1D convolutional kernels. The choice of 3×3 kernels is based on their ability to effectively capture local temporal correlations between adjacent time steps while controlling computational complexity to ensure feature extraction capabilities. After convolution, batch normalization normalizes the feature distribution, accelerating network convergence and improving generalization by suppressing overfitting. Then, ReLU activation is used to introduce a non-linear transformation, enhancing the network's ability to fit complex temporal patterns. Finally, 2×2 max pooling compresses the time dimension from 201 to 100, while increasing the number of output channels from 34 to 48, resulting in a feature tensor of shape (B, 48, 100).

[0075] The middle third feature extraction layer inherits the output features of the top third feature extraction layer and further deepens the feature expression through 64 3×3 convolution kernels. The increase in the number of channels (from 48 to 96) aims to increase the dimensional capacity of the features to carry richer semantic information. Similarly, after batch normalization and ReLU activation, 2×2 max pooling is used to reduce the time dimension from 100 to 50, and the shape of the output feature tensor is (B, 96, 50).

[0076] The third feature extraction layer at the bottom layer is the highest level of feature extraction. It uses 128 3×3 convolutional kernels to perform high-order feature extraction and constructs a high-dimensional feature space by expanding the channel size by a larger scale (from 96 to 192). After pooling, the time dimension is reduced to 25, and the final output is a high-order semantic feature of (B, 192, 25).

[0077] It is worth noting that this first multi-stage feature extraction unit achieves co-evolution of feature resolution and semantic hierarchy through progressive downsampling in the time dimension (201→100→50→25) and stepwise expansion in the channel dimension (34→48→96→192): lower stages (such as the top-level third feature extraction layer) retain more temporal details, and their output features are suitable for capturing high-frequency dynamic patterns in time series (such as instantaneous fluctuations and sudden changes); higher stages (such as the bottom-level third feature extraction layer) focus on global trends through large-scale downsampling, and their output features are good at representing low-frequency trend features (such as long-term evolutionary patterns and periodic patterns). This design constructs a complete temporal feature pyramid, enabling the model to adaptively capture key patterns of time series at different time scales, providing comprehensive feature support for subsequent classification tasks.

[0078] After extracting features from the inertial motion data sequentially using the third feature extraction layer in the first multi-stage feature extraction unit, the first convolutional thinning unit in the first multi-branch convolutional thinning unit can be used to perform deep feature extraction on the features extracted by the corresponding feature extraction layers. Specifically, deep processing, nonlinear transformation, and feature interaction can be performed sequentially. Then, the first fusion classifier can be used to concatenate, fuse, and classify the features extracted by the deep processing of each first convolutional thinning unit to obtain the recognition results of the elderly person's body posture.

[0079] In other words, the first multi-stage feature extraction unit, as the underlying architecture of the pose recognition model, achieves hierarchical capture of time series features from fine-grained to coarse-grained through progressive downsampling operations and channel dimension expansion strategies; the first multi-branch convolutional refinement unit performs in-depth processing on the feature tensors output by each third feature extraction layer, and improves the discriminativeness of features by adding nonlinear transformations and feature interactions; the first fusion classifier integrates the high-order features of each third feature extraction layer to achieve the complementarity and enhancement of multi-scale information, and finally outputs the globally optimal decision.

[0080] As can be seen from the above, the first multi-stage feature extraction unit, the first multi-branch convolutional refinement unit, and the first fusion classifier are closely linked through data flow, forming a complete processing link of "feature extraction - fine processing - global fusion", which improves the accuracy of body pose recognition.

[0081] Alternatively, inertial motion data can be input into the ResNet1d model to obtain the body posture recognition results of the elderly using the ResNet1d model.

[0082] See Figure 7This is an architectural diagram of a second training model provided in some embodiments of this disclosure. In some embodiments of this disclosure, pre-training a pose recognition model may include: constructing a second training model; the second training model may include a second multi-stage feature extraction unit, a second multi-branch convolutional refinement unit, a stage classifier layer, and a second fusion classifier. The second multi-stage feature extraction unit may include a serial fourth feature extraction layer, the second multi-branch convolutional refinement unit may include second convolutional refinement units corresponding one-to-one with each fourth feature extraction layer, and the stage classifier layer may include stage classifiers corresponding one-to-one with each second convolutional refinement unit; training the second training model using a pose recognition dataset to obtain a pose recognition model; wherein, each iteration of the training process is as follows: sequentially extracting features from pose recognition sample data using the serial fourth feature extraction layer; performing deep feature extraction on the features extracted by the corresponding fourth feature extraction layer using each second convolutional refinement unit; and performing body pose classification processing on the deep features extracted by the corresponding second convolutional refinement unit using each stage classifier to obtain a pose recognition model. The corresponding stage body pose recognition results are obtained; the features extracted from each second convolutional refinement unit depth by the second fusion classifier are concatenated, fused, and classified to obtain the fused body pose recognition results; a fourth feature extraction layer is selected as the target feature extraction layer from the serial fourth feature extraction layers according to a preset order; a stage loss function is constructed using the stage body pose recognition results and body pose labels of the pose recognition sample data corresponding to the target feature extraction layer; the parameters of the second model to be trained are updated using the stage loss function; the process of extracting features from the pose recognition sample data sequentially using the serial fourth feature extraction layers is repeated until the last fourth feature extraction layer in the serial fourth feature extraction layers is selected as the target feature extraction layer to update the parameters of the second model to be trained according to a preset order; a fusion loss function is constructed using the fused body pose recognition results output by the second fusion classifier and the body pose labels of the pose recognition sample data; the parameters of the second model to be trained are updated using the fusion loss function.

[0083] In some embodiments of this disclosure, a second training model can be constructed to train a pose recognition model using a pose recognition dataset. The constructed second training model may include a second multi-stage feature extraction unit, a second multi-branch convolutional refinement unit, a stage classifier layer, and a second fusion classifier. The second multi-stage feature extraction unit, the second multi-branch convolutional refinement unit, and the second fusion classifier in the second training model have the same structure as the first multi-stage feature extraction unit, the first multi-branch convolutional refinement unit, and the first fusion classifier in the trained pose recognition model, and will not be described again here; the difference lies in the model parameters. Furthermore, compared to the trained pose recognition model, the constructed second training model includes an additional stage classifier layer. This stage classifier layer includes stage classifiers corresponding one-to-one with each second convolutional refinement unit in the second multi-branch convolutional refinement unit. Each stage classifier includes a BN layer, a Linear layer, another BN layer, an ELU layer, and a Linear layer connected sequentially.

[0084] Then, the pose recognition model can be obtained by training the second training model using the pose recognition dataset. The process of each iteration when training the second training model using the pose recognition dataset can be as follows: Step (1): Extract features of pose recognition sample data sequentially using the serial fourth feature extraction layer in the second multi-stage feature extraction unit; Step (2): Perform deep feature extraction on the features extracted by the corresponding fourth feature extraction layer using each second convolutional refinement unit corresponding to the serial fourth feature extraction layer in the second multi-branch convolutional refinement unit; Step (3): Perform body pose classification processing on the features deeply extracted by the corresponding second convolutional refinement unit using the stage classifier corresponding to each second convolutional refinement unit in the stage classifier layer to obtain the corresponding stage body pose recognition result. Furthermore, perform splicing, feature fusion, and body pose classification processing on the features deeply extracted by each second convolutional refinement unit using the second fusion classifier to obtain the fused body pose recognition result; Step (4): Follow a preset order (the preset order can be from the bottom fourth feature extraction layer to the top fourth feature extraction layer, such as...). Figure 7As shown, this is the direction from Stage 3 to Stage 1; of course, it can also be the order from the top fourth feature extraction layer to the bottom fourth feature extraction layer. From the serial fourth feature extraction layers, determine a fourth feature extraction layer as the target feature extraction layer. Use the stage body pose and body pose labels of the pose recognition sample data obtained by the stage classifier corresponding to the target feature extraction layer to construct a stage loss function. Use the stage loss function to backpropagate and update the parameters of the second training model (all parameters that need to be updated in the second training model). Then, return to execute the feature extraction of the pose recognition sample data sequentially using the serial fourth feature extraction layers in the second multi-stage feature extraction unit, that is, return to execute steps (1) to (4) until the last fourth feature extraction layer in the serial fourth feature extraction layer is used as the target feature extraction layer to execute step (4) in the preset order, and update the parameters of the second training model; Step (5): Use the fused body pose recognition result output by the second fusion classifier and the body pose labels of the pose recognition sample data to construct a fusion loss function. Use the fusion loss function to backpropagate and update the parameters of the second training model. As can be seen from the above, in each iteration of the training process, the number of parameter updates of the second model to be trained is equal to the number of the fourth feature extraction layers in sequence + 1. This method can improve data utilization efficiency, accelerate model convergence, and achieve the ability to use a small amount of sample data as if it were a large amount of sample data.

[0085] like Figure 7As shown, if the second multi-stage feature extraction unit includes three sequential fourth feature extraction layers (represented as Stage1, Stage2, and Stage3 from top to bottom respectively), the stage classifiers in the stage classifier layer are represented as Cls1 (whose output corresponds to output1 and Stage1), Cls2 (whose output corresponds to output2 and Stage2), and Cls3 (whose output corresponds to output3 and Stage3). In each iteration of training, steps (1) to (3) above can be performed, and then step (4) can be entered: First, Loss3 is constructed using output3 and body pose labels. Based on Loss3, the parameters of the second model to be trained are updated, and then steps (1) to (3) are returned to be executed. Next, Loss2 is constructed using output2 and body pose labels. Based on Loss2, the parameters of the second model to be trained are updated, and then steps (1) to (3) are returned to be executed. Then, Loss1 is constructed using output1 and body pose labels. Based on Loss1, the parameters of the second model to be trained are updated, and then steps (1) to (3) are returned to be executed. Finally, LossALL is constructed using the fused body pose recognition result output_concat output by the second fusion classifier and body pose labels. Based on LossALL, the parameters of the second model to be trained are updated. Thus, the parameters of the second model to be trained are updated 4 times in each iteration.

[0086] The above approach enables the output layer of the second training model to adopt a multi-branch design strategy, specifically including the output results of each stage classifier and the output results of a second fusion classifier. The advantages of this multi-output structure are reflected in two aspects: First, it preserves the independent discrimination results of features at different scales, which can be used to analyze the classification performance of the model at different feature levels; second, the multi-branch supervision mechanism provides rich gradient information for network training, which not only alleviates the gradient vanishing problem in deep network training, but also promotes the model's balanced learning of features at each level, thereby significantly improving the model's generalization ability.

[0087] See Figure 8 This is a schematic diagram of the pose recognition model training provided in some embodiments of this disclosure, specifically a schematic diagram of training / test accuracy and loss vs. iteration. Figure 8 Curve 5 represents training accuracy, curve 6 represents testing accuracy, curve 7 represents training loss, and curve 8 represents testing loss. Figure 8As can be seen, the model training is stable, the curve is smooth, and the classification accuracy can reach 100%. This fully demonstrates that the multi-stage feature extraction unit, multi-scale classification supervision mechanism and feature fusion strategy of the proposed model work together to form an efficient and stable classification capability, which can accurately capture various features in the data.

[0088] See Figure 9 This document presents a flowchart illustrating a method for analyzing and providing early warning of the health status of the elderly based on multimodal data, as provided in some embodiments of this disclosure. In some embodiments of this disclosure, health analysis results for the elderly are generated based on the identification results of the elderly's physiological signal characteristic parameters, emotional state and attention level, body posture, and basic information. This may include: encapsulating the elderly's physiological signal characteristic parameters, emotional state and attention level, body posture, and basic information into structured data; generating prompt words based on the set role of the large language model as a health management expert, preset input data templates, preset task instructions, preset output requirements, and instructions instructing the large language model to generate a health analysis report using chain reasoning; inputting the prompt words into the large language model, and using the large language model to generate a health analysis report for the elderly based on the structured data.

[0089] In some embodiments of this disclosure, when generating health analysis results for the elderly, the physiological signal characteristic parameters, emotional state and attention level recognition results, body posture recognition results, and basic information of the elderly can be first encapsulated into structured data, such as JSON format, to organically integrate scattered multi-dimensional health information, ensure data accuracy and consistency, and facilitate improved efficiency and accuracy of large language model analysis. That is, the input content of the large language model covers multi-dimensional health information, achieving comprehensive and structured feature expression.

[0090] By encapsulating the identification results of physiological signal characteristic parameters, emotional state and attention level, body posture and basic information of the elderly into structured data, it is possible to organically integrate scattered information through a unified format. For example, "Subject information: Age 75, female, history of hypertension for 5 years, walks for 1 hour daily. Monitoring data: changes in emotional state, changes in attention level, changes in posture (movement), changes in physiological signal characteristic parameters over time." This clearly distinguishes between personal basic information and monitoring data, and makes the logical relationship between various indicators readily apparent. This design enables the large language model to quickly identify key information, improving the ability to perceive and dynamically assess the health status of the elderly in real time.

[0091] Furthermore, in some embodiments of this disclosure, a Prompt can be designed, and a structured Prompt template can be used to generate the Prompt. The structured Prompt template can include four parts: "role definition → input data → task instructions → output requirements." The role definition sets the large language model as a "senior health management expert" with a multidisciplinary background in cardiovascular medicine, neurology, etc. The preset input data template clearly labels the indicator format, such as: EDA indicators 3×50 dimensions, 1 sample every 10 seconds, and coding rules (e.g., emotional state "anger = 1, happy = 2", attention "very focused = 1, very distracted = 5", posture "sitting posture = 1, falling = 3"). The preset task instructions require completion of health status grading (e.g., healthy / borderline / mildly abnormal / moderately severe abnormality), multi-dimensional risk analysis (identifying ≥2 abnormal indicators), tiered intervention plans (emergency warning, exercise / nutrition advice, medical follow-up), and data quality assessment. Preset output requirements: rigorous medical expression, priority of recommendations (e.g., 1-3 ★), and citation of clinical guidelines (e.g., AHA guidelines, WHO reports, etc.).

[0092] Additionally, pre-set instructions can be provided, which instruct the large language model to generate a health analysis report using a chain-like reasoning approach. The chain-like reasoning process is as follows: First, risk assessment for individual indicators (e.g., "SDNN=0ms, indicating arrhythmia" or "EEG peak value=270μV, indicating unstable brain electrical activity"); Second, comprehensive inference (combining individual medical history, such as "high work stress + abnormal HRV → autonomic nervous system disorder" or "alcohol consumption + abnormal EEG → unstable neural excitability"); Third, generating a health analysis report.

[0093] Based on the above, prompt words can be generated according to the set role of the large language model as a health management expert, preset input data templates, preset task instructions, preset output requirements, and instructions to the large language model to generate a health analysis report using chain reasoning. Then, the prompt words are input into the large language model, which uses structured data to generate a health analysis report for the elderly. This achieves the integration of multi-dimensional recognition results through the large language model, generating health assessment reports that conform to individual differences, providing actionable intervention suggestions, and improving the humanization of health management. For example, see [link to example]. Figure 10 This is a schematic diagram of a health analysis report provided in some embodiments of this disclosure. Experimental verification: The health analysis report generated using a large language model can accurately identify the subject's anxiety state (consistent with hospital diagnosis), and the proposed intervention suggestions (such as "30 minutes of aerobic exercise daily" and "limiting alcohol intake") comply with medical standards.

[0094] As can be seen from the above embodiments of this disclosure, this disclosure can achieve: (1) accurate and real-time identification of multi-dimensional health indicators: through standardized data collection and multimodal fusion technology (referring to the technology of fusing information from different modalities (such as EEG time-series signals, skin conductance signals, pulse signals, facial expression images) and achieving more comprehensive cognitive understanding through joint modeling, collaborative learning and other mechanisms, to improve the accuracy of health indicator identification for the elderly), accurately identify the emotional state (such as: 7 types of anger, anxiety, and happiness), attention level (such as: 5 levels of very focused to very scattered), body posture (such as: 4 types of sitting posture, running posture, falling posture, and lying flat posture), and physiological signal characteristic parameters (such as: heart rate, sleep quality, skin conductance response, etc.) of the elderly; (2) constructing a lightweight and deployable model architecture: designing a pre-trained multi-task model. Neural networks and multi-stage, multi-scale cascaded classification networks can reduce model parameters while ensuring recognition accuracy (emotional state recognition ≥96%, posture recognition 100%), and meet the low power consumption and real-time computing requirements of wearable devices; (3) Establish a "real-time monitoring-long-term analysis-dynamic early warning" system: Real-time perception of the health status of the elderly can be achieved through continuous data collection and time-series analysis; Based on long-term accumulated data, trend prediction (such as Parkinson's risk, cognitive decline, etc.) can be carried out in combination with large language models to proactively warn of potential risks; (4) Generate personalized and professional health analysis reports: Relying on large language models and prompt word engineering, multi-dimensional recognition results can be integrated to generate health assessment reports that conform to individual differences (age, medical history, lifestyle habits), provide operable intervention suggestions, and improve the humanization level of health management.

[0095] Compared with existing elderly health management technologies (such as single-modal monitoring devices, offline analysis systems, and general large-scale model applications), some embodiments of this disclosure have the following advantages: 1. Enhanced real-time and continuous performance. Existing technologies rely on discrete data collection methods such as regular physical examinations and outpatient consultations, which cannot capture sudden health events (such as falls) or rapidly changing indicators (such as sudden changes in emotional state). This disclosure achieves continuous multimodal data collection through wearable devices (sampling frequency up to 256Hz), and AI algorithms (multimodal fusion, posture recognition) analyze EEG, skin conductance, and movement signals in real time, providing immediate warnings when falls occur or sudden changes in emotional state occur. This can reduce the disability and mortality rates caused by sudden health risks, achieving "early detection and early intervention," and significantly improving the sense of security for elderly people living independently. 2. More comprehensive multi-dimensional assessment. Existing technologies often focus on single physiological signal parameters (such as blood pressure and blood sugar), relying on subjective questionnaires or observer experience to make judgments, lacking a holistic assessment of health status. This disclosure integrates five types of multimodal data: electroencephalography (EEG), electrical activity of the skin (EDA), pulse (PPG), posture (IMU), and facial expression. Through three recognition paths (calculation of physiological signal parameters, identification of emotional state / attention, and identification of posture), it simultaneously assesses emotional state fluctuations, attention levels (cognitive ability), motor function (fall risk), cardiovascular status, and stress levels. It can reveal potential associated risks (such as anxiety-induced insomnia → increased fall risk), providing a more comprehensive health profile and a scientific basis for personalized intervention. III. Enhanced proactive early warning and predictive capabilities. Existing technologies are mostly reactive (e.g., ringing a bell after a fall, seeking medical attention after feeling unwell) or based on statistical risk assessments using historical data. This disclosure not only identifies current anomalies in real time (e.g., fall warnings worn at fixed times in Scenario 1), but also performs trend analysis and risk prediction (e.g., Parkinson's risk, cognitive decline trend) based on long-term wear data (Scenario 2) combined with a large model, generating predictive insights and preventative recommendations. This shifts the focus from "treating disease" to "preventing disease," buying time for intervention, slowing disease progression, and reducing long-term care needs and medical expenses. IV. Superior Low Intrusion and Personalization. Existing technologies suffer from invasive detection (e.g., blood draws), reliance on professionals (e.g., neuropsychological testing), or privacy violations (e.g., cameras), and health recommendations are based on general standards. This disclosure uses lightweight wearable devices to achieve "non-intrusive / low-invasive" monitoring (especially for long-term wear in Scenario 2), with complex AI analysis running automatically in the background. Combining individual medical history, lifestyle habits, and long-term data, a large model is driven by prompt word engineering to generate reports tailored to individual needs. The recommendations are easy to understand and implement, enhancing the humanistic care aspect of health management.

[0096] This disclosure also provides a health status analysis and early warning system for the elderly, which may include: a multimodal data acquisition module for acquiring physiological data, facial videos, inertial movement data, and basic information of the elderly; an elderly physiological signal feature parameter recognition module for determining the physiological signal feature parameters of the elderly based on the physiological data; an elderly emotional state and attention level recognition module for recognizing the emotional state and attention level of the elderly based on the physiological data and facial videos; an elderly body posture recognition module for recognizing the body posture of the elderly based on the inertial movement data; and a health analysis and early warning module for generating health analysis results and / or early warning information to indicate abnormal conditions of the elderly based on the recognition results of the elderly's physiological signal feature parameters, emotional state and attention level, body posture, and basic information.

[0097] This disclosure also provides an edge computing device, which may include: a memory for storing a computer program; and a processor for executing the computer program stored in the memory to implement the steps of any of the above methods.

[0098] This disclosure also provides a computer-readable storage medium storing a computer program that, when executed by a processor, can implement the steps of any of the above methods.

[0099] For a description of the relevant parts of the elderly health status analysis and early warning system, edge computing device and computer-readable storage medium provided in some embodiments of this disclosure, please refer to the detailed description of the corresponding parts of the elderly health status analysis and early warning method based on multimodal data provided in some embodiments of this disclosure, and will not be repeated here.

[0100] The embodiments of this disclosure have now been described in detail. To avoid obscuring the concept of this disclosure, some details known in the art have not been described. Those skilled in the art can fully understand how to implement the technical solutions disclosed herein based on the above description.

[0101] It should be noted that the logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.

[0102] It should be understood that various parts of this disclosure can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0103] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this disclosure. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0104] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this disclosure, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0105] While specific embodiments of this disclosure have been described in detail by way of examples, those skilled in the art should understand that the examples are for illustrative purposes only and not intended to limit the scope of this disclosure. Those skilled in the art should understand that modifications can be made to the above embodiments or equivalent substitutions can be made to some technical features without departing from the scope and spirit of this disclosure. In particular, as long as there is no structural conflict, the technical features mentioned in the various embodiments can be combined in any manner.

Claims

1. A method for analyzing and predicting the health status of the elderly based on multimodal data, characterized in that, include: Acquire physiological data, facial videos, inertial movement data, and basic information of the elderly; The physiological signal characteristic parameters of the elderly are determined based on the physiological data; The elderly person's emotional state and attention level are identified based on the physiological data and facial video. The elderly person's body posture is identified based on the inertial motion data; Based on the physiological signal characteristic parameters, emotional state and attention level recognition results, body posture recognition results, and basic information of the elderly, health analysis results and / or early warning information for indicating abnormal conditions of the elderly are generated.

2. The method according to claim 1, characterized in that, Identifying the elderly person's emotional state and attention level based on the physiological data and facial video includes: Multiple physiological data and multiple facial images obtained from frame segmentation of the facial video are input into a pre-trained emotion and attention recognition model to obtain the recognition results of the elderly person's emotional state and attention level. The emotion and attention recognition model includes: a first feature extraction layer, a first feature fusion layer, a first emotion state recognition and classification layer, and a first attention level recognition and classification layer; the step of inputting multiple physiological data and multiple facial images obtained by segmenting the facial video into the pre-trained emotion and attention recognition model to obtain the recognition results of the elderly person's emotion state and attention level includes: The first feature extraction layer is used to extract physiological features from various physiological data and facial expression features from multiple facial images; The physiological features corresponding to various physiological data and the facial expression features are fused using the first feature fusion layer to obtain fused features; The first emotion state recognition classification layer and the first attention level recognition classification layer are used to perform emotion state recognition and attention level recognition on the fused features, respectively, to obtain the recognition results of the elderly person's emotion state and attention level.

3. The method according to claim 2, characterized in that, The physiological data includes at least one of the following: the electroencephalogram (EEG) signal of the elderly person collected by a wearable EEG device, the pulse signal of the elderly person collected by a pulse wave sensor, and the skin conductance signal of the elderly person collected by a skin conductance sensor. The physiological features corresponding to the physiological data and the facial expression features are fused using a first feature fusion layer to obtain fused features, including: The physiological features corresponding to the various physiological data and the facial expression features are spliced ​​together. The spliced ​​features are then extracted again using a self-attention mechanism, and the results of the second feature extraction are fused to obtain fused features.

4. The method according to claim 2, characterized in that, Pre-trained emotion attention recognition model, including: The first model to be trained was pre-trained using the emotion state attention dataset to obtain the pre-trained first model. The second feature extraction layer and the second feature fusion layer in the pre-trained first model are frozen. The second emotion state recognition classification layer in the pre-trained first model is trained using sample data containing emotion state labels in the emotion state attention dataset. The second attention level recognition classification layer in the pre-trained first model is trained using sample data containing attention level labels in the emotion state attention dataset, thereby obtaining the emotion attention recognition model.

5. The method according to claim 1, characterized in that, Identifying the elderly person's body posture based on the inertial motion data includes: The inertial motion data is input into a pre-trained posture recognition model to obtain the recognition result of the elderly person's body posture. The posture recognition model includes a first multi-stage feature extraction unit, a first multi-branch convolutional thinning unit, and a first fusion classifier. The process of inputting the inertial motion data into the pre-trained posture recognition model to obtain the recognition result of the elderly person's body posture includes: The features of the inertial motion data are extracted sequentially by the serial third feature extraction layer in the first multi-stage feature extraction unit; wherein, the serial third feature extraction layer performs time dimension downsampling and feature dimension expansion on the inertial motion data, and the first multi-branch convolutional thinning unit includes a first convolutional thinning unit corresponding to each of the third feature extraction layers. The first convolutional thinning unit is used to perform deep feature extraction on the features extracted by the corresponding third feature extraction layer. The first fusion classifier is used to concatenate, fuse, and classify the features extracted from each of the first convolutional thinning units to obtain the recognition result of the elderly person's body posture.

6. The method according to claim 5, characterized in that, Pre-trained pose recognition models include: Construct a second model to be trained; the second model to be trained includes a second multi-stage feature extraction unit, a second multi-branch convolutional refinement unit, a stage classifier layer, and a second fusion classifier. The second multi-stage feature extraction unit includes a serial fourth feature extraction layer. The second multi-branch convolutional refinement unit includes a second convolutional refinement unit corresponding to each of the fourth feature extraction layers. The stage classifier layer includes a stage classifier corresponding to each of the second convolutional refinement units. The second training model is trained using a pose recognition dataset to obtain the pose recognition model; wherein each iteration of the training process is as follows: Features of pose recognition sample data are extracted sequentially using a fourth feature extraction layer. Each of the second convolutional thinning units is used to perform deep feature extraction on the features extracted by the corresponding fourth feature extraction layer; The features extracted from the depth of the corresponding second convolutional refinement unit by each stage classifier are processed for body pose classification to obtain the corresponding stage body pose recognition result; the features extracted from the depth of each second convolutional refinement unit are spliced, fused and processed for body pose classification by the second fusion classifier to obtain the fused body pose recognition result. A fourth feature extraction layer is selected as the target feature extraction layer from the serial fourth feature extraction layers in a preset order. A stage loss function is constructed using the stage body posture recognition result corresponding to the target feature extraction layer and the body posture label of the posture recognition sample data. The parameters of the second model to be trained are updated using the stage loss function. The process of extracting features from the posture recognition sample data sequentially using the serial fourth feature extraction layers is repeated until the last fourth feature extraction layer in the serial fourth feature extraction layers is selected as the target feature extraction layer to update the parameters of the second model to be trained in the preset order. A fusion loss function is constructed using the fused body pose recognition result output by the second fusion classifier and the body pose labels of the pose recognition sample data. The parameters of the second model to be trained are then updated using the fusion loss function.

7. The method according to claim 1, characterized in that, Based on the physiological signal characteristic parameters, emotional state and attention level recognition results, body posture recognition results, and basic information of the elderly person, a health analysis result for the elderly person is generated, including: The physiological signal characteristic parameters, emotional state and attention level recognition results, body posture recognition results, and basic information of the elderly are encapsulated into structured data. Based on the set role of the large language model as a health management expert, the preset input data template, the preset task instructions, the preset output requirements, and the instruction information that instructs the large language model to use chain reasoning to generate a health analysis report, prompt words are generated; The prompt words are input into a large language model, which then generates a health analysis report for the elderly based on the structured data.

8. A health status analysis and early warning system for the elderly, characterized in that, include: The multimodal data acquisition module is used to acquire physiological data, facial videos, inertial motion data, and basic information of the elderly. An elderly physiological signal feature parameter identification module is used to determine the physiological signal feature parameters of the elderly based on the physiological data; An elderly person's emotional state and attention level recognition module is used to recognize the elderly person's emotional state and attention level based on the physiological data and the facial video. An elderly body posture recognition module is used to recognize the body posture of the elderly person based on the inertial motion data; The health analysis and early warning module is used to generate health analysis results and / or early warning information to indicate abnormal conditions of the elderly person based on the physiological signal characteristic parameters, emotional state and attention level recognition results, body posture recognition results, and the basic information.

9. An edge computing device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the steps of the method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the method as described in any one of claims 1 to 7.