A multi-modal human body perception data fusion analysis and risk prediction method and system for smart medical care
By constructing a multi-source acquisition terminal network, timestamp calibration, and cross-modal attention fusion network, and dynamically updating the personalized analysis model, the problems of spatiotemporal alignment and coupling of multimodal data are solved, enabling real-time monitoring and personalized intervention of the health status of the elderly, and improving the accuracy and practicality of risk prediction in the smart medical and elderly care system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGZHOU INST OF RAILWAY TECH
- Filing Date
- 2026-02-11
- Publication Date
- 2026-05-29
AI Technical Summary
In existing smart healthcare systems, multimodal data lacks effective spatiotemporal alignment and deep correlation mining, making it impossible to accurately capture the coupling relationship between physiological changes, behavioral patterns and environmental factors. The analysis models are static and rigid, making it difficult to adapt to the dynamic evolution of the individual health status of the elderly. Risk judgment relies on ex-post threshold triggering, resulting in delayed early warnings and a high false alarm rate, failing to achieve real-time monitoring, early warning and precise intervention.
A multi-source acquisition terminal network is constructed, and spatiotemporal registration of multimodal data is achieved through a timestamp calibration algorithm. Uncertainty quantification processing and cross-modal attention fusion network are introduced to dynamically update the personalized analysis model and generate personalized intervention suggestions by combining a scenario-based rule base.
It enables dynamic perception and intelligent analysis of the health status of the elderly, improves the accuracy of health risk prediction and the real-time nature of personalized intervention, breaks through the problems of model rigidity and early warning lag in traditional methods, realizes the transformation from passive response to proactive early warning, and improves the overall efficiency of the smart medical and elderly care system.
Smart Images

Figure CN122117381A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data analysis and smart healthcare technology, specifically to a method and system for multimodal human perception data fusion analysis and risk prediction for smart healthcare. Background Technology
[0002] With the rapid development of smart healthcare and elderly care services, real-time and accurate health risk monitoring and early warning have become crucial for improving the quality of life and safety of the elderly. Currently, the main approach relies on single or simple multi-source data acquisition technologies such as wearable devices and environmental sensors to independently analyze or superficially integrate the collected physiological, behavioral, and environmental information. Common analysis methods include anomaly alarms based on fixed thresholds, parallel display of multi-source data, and direct application of general machine learning models for health status assessment. While these methods have achieved monitoring of user status to some extent, they still have significant limitations: there is a lack of effective spatiotemporal alignment and deep correlation mining among multimodal data, making it impossible to accurately capture the coupling relationship between physiological changes, behavioral patterns, and environmental factors; analysis models are generally static and rigid, making it difficult to adapt to the dynamic evolution and personalized differences in the health status of individual elderly people; risk judgment often relies on ex-post threshold triggering, resulting in delayed early warnings and a high false alarm rate, making it difficult to achieve true risk prediction.
[0003] Existing technologies often overlook noise interference, missing data, and uncertainties during the data acquisition process, and fail to optimize analysis strategies for high-frequency activity scenarios of the elderly (such as bathrooms and bedrooms). As a result, the accuracy, timeliness, and practicality of the overall monitoring system are insufficient to meet the core requirements of "real-time monitoring, early warning, and precise intervention" in smart healthcare scenarios. Therefore, there is an urgent need for an analysis method and system that can deeply integrate multimodal human sensory data, dynamically adapt to individual characteristics, achieve early risk prediction, and possess high reliability and scenario adaptability for health monitoring of the elderly. Summary of the Invention
[0004] To address the issues regarding the effectiveness of health monitoring for the elderly mentioned in the background section, this invention provides a method and system for multimodal human perception data fusion analysis and risk prediction in smart healthcare and elderly care.
[0005] The above-mentioned objective of this application is achieved through the following technical solution:
[0006] A method for multimodal human perception data fusion analysis and risk prediction for smart healthcare and elderly care includes the following steps:
[0007] A multi-source acquisition terminal network is constructed to simultaneously collect users' physiological data, behavioral data, and environmental data, and spatiotemporal registration is performed on the multimodal data;
[0008] Data cleaning and uncertainty quantification are performed on the spatiotemporally registered multimodal data to generate preprocessed data with confidence labels;
[0009] Feature extraction is performed on the preprocessed data with confidence labels, and the extracted multimodal features are fused based on an attention mechanism to generate a fused feature vector;
[0010] The fused feature vector is input into a dynamically updated personalized analysis model to predict health risks and output the risk level.
[0011] Based on the risk level and the user's scenario, personalized intervention suggestions are generated and output.
[0012] By adopting the above-mentioned approach and constructing a multi-source sensing network and a multi-modal fusion analysis system, dynamic perception and intelligent analysis of the full-dimensional status of medical and elderly care users were achieved. Dispersed physiological, behavioral, and environmental data were transformed into structured feature vectors with spatiotemporal consistency, effectively solving the problem of spatiotemporal alignment between multi-source heterogeneous data. By introducing an uncertainty quantification mechanism and a cross-modal attention fusion network, the dynamic coupling relationship between fluctuations in physiological indicators, changes in behavioral patterns, and environmental disturbances was deeply explored, improving the accuracy of health risk prediction. Combined with dynamically updated personalized analysis models and scenario-based rule bases, real-time tracking and precise intervention of the individual health status of the elderly were achieved, effectively overcoming the shortcomings of traditional methods such as model rigidity and delayed early warning. This realized a shift from passive response to proactive early warning, providing precise and forward-looking health protection support for smart medical and elderly care.
[0013] In a preferred embodiment, this application can be further configured as follows: the construction of a multi-source acquisition terminal network to simultaneously acquire users' physiological data, behavioral data, and environmental data, and to perform spatiotemporal registration on the multimodal data, includes the following steps:
[0014] The multi-source acquisition terminal network is constructed by integrating wearable physiological monitoring devices, UWB positioning base stations, foot pressure sensing pads, and environmental monitoring nodes.
[0015] A timestamp calibration algorithm is used to unify the data collected by the multi-source acquisition terminal network to the same time axis;
[0016] Based on the spatial coordinate data obtained from the UWB positioning base station, a spatial correlation mapping is established between physiological data, behavioral data, and environmental data.
[0017] By adopting the above technical solutions and constructing a heterogeneous acquisition network by integrating multiple types of sensors, synchronous acquisition and spatiotemporal alignment of physiological, behavioral, and environmental data are achieved. A high-precision timestamp calibration algorithm eliminates clock deviations between multiple devices, enabling comprehensive synchronous capture of human physiological signals, behavioral trajectories, and environmental parameters. This overcomes the fragmented nature of traditional data acquisition in terms of time and space, aligning all data streams to a unified time reference through a precise timestamp calibration algorithm, eliminating timing misalignment caused by device heterogeneity. Furthermore, a high-precision spatial coordinate system established using UWB positioning technology maps and correlates previously isolated multi-source data in physical space, forming a multidimensional state field with spatiotemporal consistency. This not only provides a strictly aligned underlying data foundation for subsequent data fusion and analysis but also makes it possible to accurately reconstruct and correlate the overall state of a user at a specific time and location, fundamentally ensuring the integrity and interpretability of multimodal sensing data.
[0018] In a preferred embodiment, this application can be further configured such that: the method of using a timestamp calibration algorithm to unify the data collected by the multi-source acquisition terminal network to the same time axis includes:
[0019] Provide a unified time reference for the multi-source acquisition terminal network;
[0020] Compensation is provided for time deviations caused by data transmission delays, and the time synchronization error of multi-source data is controlled within a preset threshold.
[0021] By adopting the above technical solution, a unified time reference source, such as a high-precision atomic clock or a network time protocol server, is set for the multi-source acquisition terminal network, ensuring that all sensor devices have a consistent time starting point at the start of data acquisition. To address time deviations caused by network latency and device processing delays during data transmission, a dynamic compensation algorithm is used to calculate the time offset of each data stream in real time, and abnormal timestamps are corrected through interpolation or translation operations. Ultimately, the time synchronization error of multimodal data is controlled within a preset threshold at the millisecond level. This effectively solves the timing misalignment problem caused by clock drift and transmission delay in heterogeneous sensor networks, providing a precise time alignment basis for subsequent spatiotemporal registration and feature fusion, and ensuring the consistency of multimodal data in the time dimension.
[0022] In a preferred embodiment, this application can be further configured as follows: the step of performing data cleaning and uncertainty quantification on the spatiotemporally registered multimodal data to generate preprocessed data with confidence labels includes the following steps:
[0023] Multi-level data cleaning is performed on the spatiotemporally registered data, including outlier removal, noise filtering, and missing value completion.
[0024] A probability model is introduced to assess the uncertainty of the cleaned data and calculate the confidence level of each data point.
[0025] Data with confidence levels below a preset threshold are labeled and their weight in subsequent fusion analysis is reduced.
[0026] By adopting the above technical solution, multi-level data cleaning is performed. Outliers in physiological signals are identified and removed using statistical distribution-based anomaly detection algorithms (such as the 3σ principle). Simultaneously, wavelet transform algorithms are used to denoise behavioral trajectory data to eliminate high-frequency noise interference. For missing values in environmental monitoring data, a forward imputation method combined with LSTM prediction algorithm is used to complete the data, taking into account the continuity characteristics of time series. Subsequently, a Bayesian probability model is introduced to quantify the uncertainty of the cleaned multimodal data. By calculating the local density and neighborhood consistency index of each data point in the feature space, a confidence score reflecting the reliability of the data is generated. Finally, data points with confidence scores below a threshold (such as 0.7) are marked with reduced weights. In the subsequent feature fusion process, their contribution weights are dynamically adjusted through an attention mechanism, effectively suppressing the interference of low-quality data on the analysis results and improving the robustness of the preprocessed data.
[0027] In a preferred embodiment, this application can be further configured as follows: The step of extracting features from the preprocessed multimodal data and fusing the extracted multimodal features based on an attention mechanism to generate a fused feature vector includes the following steps:
[0028] Convolutional neural networks are used to extract temporal features from physiological data, graph neural networks are used to extract spatial features from behavioral data, and fully connected layers are used to extract scene features from environmental data.
[0029] Construct a cross-modal attention fusion network and calculate the attention weights for different modal features;
[0030] The temporal features, spatial features, and scene features are weighted and fused according to the attention weights to generate the fused feature vector.
[0031] By employing the above technical solutions, when using convolutional neural networks (CNNs) to extract temporal features from physiological data, the sliding window scanning of multiple convolutional kernels captures the periodic fluctuation patterns of physiological signals such as heart rate and blood pressure. Simultaneously, pooling layers compress the feature dimensions to retain key temporal features. When using graph neural networks (GNNs) to process behavioral data, spatial coordinates obtained from UWB positioning base stations are used as nodes to construct a behavioral trajectory map. Graph convolution operations capture the user's movement patterns and behavioral habits in space. Fully connected layers are used to extract scene features from environmental data. Multi-layer nonlinear transformations transform environmental parameters such as temperature, humidity, and light intensity into scene features reflecting the environmental state. The cross-modal attention fusion network calculates the correlation strength between features of different modalities through a self-attention mechanism, generating dynamic attention weights. It can automatically identify potential coupling relationships between physiological temporal features, behavioral spatial features, and environmental scene features. For example, when an abnormal increase in a user's heart rate is detected, the system will enhance the fusion weight of behavioral trajectory data and environmental temperature data within the corresponding time period. Finally, the three types of features are weighted and summed according to the attention weights to generate a fusion feature vector containing multi-dimensional correlation information. This vector not only retains the original features of each modal data, but also strengthens the interaction between key features through the attention mechanism, providing a richer feature expression basis for subsequent health risk prediction.
[0032] In a preferred example, this application can be further configured as follows: the construction of the cross-modal attention fusion network and the calculation of attention weights for different modal features include the following steps:
[0033] The mutual information entropy algorithm is used to measure the correlation strength among physiological characteristics, behavioral characteristics, and environmental characteristics.
[0034] Attention weights are dynamically assigned based on the correlation strength, with higher correlation strength modal features being assigned higher weight coefficients.
[0035] By adopting the above technical solution, the mutual information entropy algorithm quantifies the weight coefficients between different modal features. Taking physiological and behavioral features as an example, it calculates the relative entropy values of their joint probability distribution and marginal probability distribution in the feature space. The larger the value, the stronger the correlation between the two modal features. Based on the calculated mutual information entropy value, feature pairs with a correlation strength higher than a preset threshold are assigned higher attention weight coefficients. By dynamically adjusting the attention weight coefficients, the cross-modal attention fusion network can prioritize focusing on feature combinations with strong correlations. For example, when heart rate fluctuations in physiological features and temperature changes in environmental features show mutual information entropy, the system will automatically enhance the weight allocation of these two types of features in the fusion process. This data-driven dynamic weight allocation mechanism breaks through the limitations of traditional fixed weight fusion methods, enabling the fused feature vector to more accurately reflect the complex coupling relationship between multimodal data.
[0036] In a preferred embodiment, this application can be further configured as follows: the step of inputting the fused feature vector into a dynamically updated personalized analysis model to predict health risks and output risk levels includes the following steps:
[0037] The initial personalized analysis model was trained based on multi-sample population data;
[0038] The parameters of the personalized analysis model are periodically updated based on the user's latest health status feedback data.
[0039] The fused feature vector is input into the updated personalized analysis model, and a time-series prediction algorithm is used to predict the probability of health risks and output the corresponding risk level.
[0040] By adopting the above technical solution, the initial personalized analysis model is pre-trained by integrating a large-scale multi-center medical and elderly care dataset. This dataset covers elderly samples of different ages, genders, and health conditions, including long-term monitoring records of physiological indicators, behavioral patterns, and environmental parameters. During the training process of the initial personalized analysis model, a transfer learning algorithm is introduced to transfer the ability to extract general health features to the personalized adaptation stage, thereby solving the overfitting problem in small sample scenarios. After the initial personalized analysis model is deployed, the system achieves dynamic updates through user feedback loops: it automatically collects the user's physiological fluctuation data, behavioral trajectory records, and environmental exposure logs for the past 7 days every week, and combines them with health event tags (such as falls and abnormal heart rates) marked by medical experts, and adjusts the model parameters through a gradient descent algorithm. The time series prediction algorithm adopts a bidirectional LSTM algorithm to capture long-term dependencies in the fused feature vectors, and combines an attention mechanism to strengthen the weight allocation of key time windows, ultimately outputting a health risk probability value. The risk level classification sets a four-level warning threshold based on clinical guidelines. When the predicted probability exceeds 85%, a red highest-level warning is triggered and simultaneously pushed to the user terminal, family members' mobile phones, and community medical platforms to achieve multi-level risk response.
[0041] In a preferred embodiment, this application can be further configured as follows: the periodic iterative update of the parameters of the personalized analysis model based on the user's latest health status feedback data includes:
[0042] Set a fixed model update cycle;
[0043] At the end of each update cycle, collect user multimodal data and corresponding health status tags for that cycle;
[0044] The training sample set is constructed using the multimodal data and health status labels, and the parameters of the personalized analysis model are optimized using the gradient descent algorithm.
[0045] By adopting the above technical solution, the model update cycle is set to once a week, which can capture short-term fluctuations in physiological indicators in a timely manner while avoiding model oscillations caused by excessive updates. On the last day of each update cycle, the system automatically triggers the data acquisition module to acquire the user's physiological data (such as heart rate and blood pressure), behavioral data (such as gait and activity range), and environmental data (such as temperature, humidity, and light intensity) for seven consecutive days through a multi-source acquisition terminal network. At the same time, it simultaneously extracts health event tags (such as falls and arrhythmias) confirmed by medical experts within that cycle from the electronic health record. The collected multimodal data first undergoes spatiotemporal registration and data cleaning, and then forms a training sample set together with the health status tags. The weight parameters of the personalized analysis model are iteratively optimized using the gradient descent algorithm. Finally, the updated model parameters are saved to the cloud database to ensure that the personalized analysis model can continuously adapt to the dynamic changes in the user's health status and improve the long-term accuracy of health risk prediction.
[0046] In a preferred embodiment, this application can be further configured as follows: generating and outputting personalized intervention suggestions based on the risk level and the user's scenario includes the following steps:
[0047] Call the contextual rule library corresponding to the user's scenario;
[0048] The risk level is optimized based on the aforementioned scenario-based rule base;
[0049] Based on the optimized risk level and individual user characteristics, personalized intervention suggestions are generated and output.
[0050] By adopting the above technical solution, the system's built-in scenario-based rule library covers high-frequency activity scenarios such as bedrooms, bathrooms, living rooms, and balconies. Each scenario library contains specific risk response strategies for that scenario. For example, in the bedroom scenario, there are early warning intervention measures for the risk of falling at night, and in the bathroom scenario, there are prevention suggestions for the risk of slipping due to excessive humidity. When the system outputs a health risk level, it first obtains the spatial coordinates of the user's current scenario through a UWB positioning base station and matches them with the corresponding scenario-based rule library. For example, when the user is in the bathroom and the risk level is yellow (medium risk), the system automatically calls the humidity threshold parameter in the bathroom scenario rule library. Based on the conditional constraints in the scenario-based rule library, the original risk level is dynamically adjusted. If the bathroom humidity exceeds the safety threshold, the medium risk is upgraded to orange (high risk), and the anti-slip intervention strategy is triggered. Finally, personalized intervention suggestions are generated by combining the user's individual characteristics (such as age and history of underlying diseases). For example, a command to "immediately turn on the ventilation equipment and lay down the anti-slip mat" is pushed to a hypertensive user over 70 years old, and the intervention plan is simultaneously pushed to the family member's mobile phone, realizing precise linkage intervention between scenario, risk, and individual.
[0051] The second objective of this invention is achieved through the following technical solution:
[0052] A multimodal human perception data fusion analysis and risk prediction system for smart healthcare and elderly care includes:
[0053] The data acquisition and registration module integrates wearable physiological monitoring devices, UWB positioning base stations, plantar pressure sensing pads, and environmental monitoring nodes to build a multi-source acquisition terminal network, simultaneously collecting users' physiological data, behavioral data, and environmental data, and using a timestamp calibration algorithm to complete the spatiotemporal registration of multimodal data;
[0054] The data preprocessing and quantization module performs multi-level data cleaning on the spatiotemporally registered multimodal data, including outlier removal, noise filtering, and missing value completion. It also introduces a probability model to assess the uncertainty of the cleaned data and generates preprocessed data with confidence labels.
[0055] The feature fusion and analysis module uses convolutional neural networks to extract temporal features of physiological data, graph neural networks to extract spatial features of behavioral data, and fully connected layers to extract scene features of environmental data. It also uses a cross-modal attention fusion network to perform weighted fusion of multimodal features to generate a fused feature vector.
[0056] The risk prediction and decision-making module inputs the fused feature vector into a dynamically updated personalized analysis model, predicts the probability of health risks through a time-series prediction algorithm, and outputs the risk level in combination with a scenario-based rule base.
[0057] The intervention suggestion output module generates and outputs personalized intervention suggestions based on the risk level and the user's scenario, by calling the scenario-based rule base.
[0058] By adopting the above technical solutions, the data acquisition and registration module integrates various types of sensor devices to construct a multi-source acquisition terminal network that comprehensively covers user physiological, behavioral, and environmental information, ensuring data diversity and integrity. An advanced timestamp calibration algorithm effectively solves the inconsistency problem of multimodal data in the time dimension, laying a solid foundation for subsequent data processing and analysis. The data preprocessing and quantization module improves data quality and reduces the interference of noise and outliers on the analysis results through multi-level data cleaning operations. By introducing a probabilistic model for uncertainty assessment and assigning confidence labels to the data, subsequent analysis can more accurately identify and process low-quality data, improving the system's robustness. The feature fusion and analysis module extracts features from different modalities and achieves weighted fusion of features through a cross-modal attention fusion network. This process generates a fusion feature vector containing rich correlation information. This not only preserves the original features of each modality's data but also strengthens the interaction between key features through an attention mechanism, providing a more comprehensive feature representation for health risk prediction. The risk prediction and decision-making module, through a dynamically updated personalized analysis model combined with a time-series prediction algorithm, achieves accurate prediction of user health risks. Simultaneously, by combining the risk levels output from the scenario-based rule base, the system can provide more realistic early warning information based on the risk characteristics of different scenarios. The intervention suggestion output module intelligently calls the scenario-based rule base based on the risk level and the user's scenario to generate and output personalized intervention suggestions. This not only improves the targeting and effectiveness of intervention measures but also ensures that users can receive necessary help and support in a timely manner through a multi-level risk response mechanism, effectively enhancing the overall efficiency of the smart healthcare system.
[0059] In summary, this application includes at least one of the following beneficial technical effects:
[0060] 1. By constructing a heterogeneous integrated multi-source acquisition network and adopting high-precision timestamp calibration and UWB spatial positioning technology, the problem of the separation of physiological, behavioral and environmental data in the time and space dimensions is solved; enabling the system to accurately restore the user's overall state at a specific time and location, providing a reliable foundation for mining the dynamic coupling relationship between multimodal data;
[0061] 2. By introducing a probabilistic model to assess the uncertainty and label the confidence level of the data, and constructing a cross-modal attention fusion network to dynamically allocate feature weights, the interference of noise, anomalies and low-quality information during the data acquisition process is effectively suppressed; it can automatically focus on key and reliable modal features and their correlations, thereby improving the discriminative power of feature expression and the overall anti-interference ability of the system.
[0062] 3. By continuously adapting to the evolution of users' health status using dynamically updated personalized analysis models and combining them with a scenario-based rule base to optimize risk levels in real time, highly personalized intervention suggestions are generated and output. This achieves a leap from generalized analysis to individualized and accurate prediction, and ensures that early warning information and intervention measures are closely integrated with the user's actual scenario and individual characteristics. This elevates smart healthcare from a passive response stage to a new stage of proactive early warning and precise intervention, comprehensively improving the effectiveness of health and safety protection for the elderly population. Attached Figure Description
[0063] Figure 1 This is a flowchart illustrating an embodiment of a method for multimodal human perception data fusion analysis and risk prediction for smart healthcare in this application.
[0064] Figure 2 This is a flowchart illustrating the implementation of step S10 in an embodiment of a multimodal human perception data fusion analysis and risk prediction method for smart healthcare in this application.
[0065] Figure 3 This is a flowchart illustrating the implementation of step S20 in an embodiment of a multimodal human perception data fusion analysis and risk prediction method for smart healthcare in this application.
[0066] Figure 4 This is a flowchart illustrating the implementation of step S30 in an embodiment of a multimodal human perception data fusion analysis and risk prediction method for smart healthcare in this application.
[0067] Figure 5 This is a flowchart illustrating the implementation of step S40 in an embodiment of a multimodal human perception data fusion analysis and risk prediction method for smart healthcare in this application.
[0068] Figure 6 This is a flowchart illustrating the implementation of step S50 in an embodiment of a multimodal human perception data fusion analysis and risk prediction method for smart healthcare in this application.
[0069] Figure 7 This is a schematic diagram of an embodiment of a multimodal human perception data fusion analysis and risk prediction system for smart healthcare and elderly care, as described in this application. Detailed Implementation
[0070] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0071] In one embodiment, such as Figure 1As shown, this application discloses a method for multimodal human perception data fusion analysis and risk prediction for smart healthcare, including the following steps:
[0072] S10: Construct a multi-source acquisition terminal network to simultaneously collect users' physiological data, behavioral data, and environmental data, and perform spatiotemporal registration on the multimodal data.
[0073] In this embodiment, the multi-source acquisition terminal network is a data acquisition system composed of various devices such as wearable devices, indoor positioning base stations, foot sensors, and environmental monitoring nodes. Physiological data refers to parameters reflecting the body's internal state, such as heart rate, blood pressure, and blood oxygen. Behavioral data refers to the user's external activities and location information, such as gait, trajectory, and posture. Environmental data refers to the physical parameters of the user's space, such as temperature, humidity, light, and ground conditions. Multimodal data is a collection of the above three types of heterogeneous data. Spatiotemporal registration is the process of unifying data collected from different devices at different times and locations to the same spatiotemporal reference through an algorithm.
[0074] Specifically, a multi-source acquisition terminal network is constructed using various devices to simultaneously collect three types of data; a timestamp calibration algorithm is used to align the time of each device with a unified time base and compensate for transmission delays; based on the spatial coordinates obtained from the UWB positioning base station, the collected data are associated and mapped with the specific location of the user to form a data sequence with spatiotemporal consistency.
[0075] S20: Perform data cleaning and uncertainty quantification on the spatiotemporally registered multimodal data to generate preprocessed data with confidence level labels.
[0076] In this embodiment, data cleaning refers to identifying and processing outliers, noise, and missing values in the data through algorithms; uncertainty quantification is the process of evaluating the reliability of data using a probability model, calculating a confidence score between 0 and 1 for each data point; preprocessed data with confidence labels refers to structured data with a reliability score attached to each data point after cleaning and quantization.
[0077] Specifically, the spatiotemporally registered multimodal data is first cleaned in multiple stages: statistical methods (such as the 3σ principle) are used to remove outliers caused by sensor malfunctions; signal processing algorithms (such as wavelet transform algorithms) are used to filter physiological data noise caused by motion interference; for short-term missing data in the environmental data, time series-based prediction methods (such as LSTM prediction algorithms) are used to complete the data; a Bayesian probability model is introduced to assess the uncertainty of the data points based on the consistency of local features and to calculate the confidence level; finally, data points with confidence levels below a preset threshold are marked, and a preprocessed dataset with confidence scores for all data points is generated for subsequent analysis.
[0078] S30: Extract features from the preprocessed data with confidence labels, and fuse the extracted multimodal features based on an attention mechanism to generate a fused feature vector.
[0079] In this embodiment, feature extraction refers to identifying and extracting representative feature information from preprocessed data; attention mechanism is a deep learning technology that simulates human visual attention, which can automatically assign weights to different features and enhance the expressive power of key features; fusion feature vector is a comprehensive feature representation formed by weighting multimodal features through attention mechanism.
[0080] Specifically, a specific neural network model is first used to extract features from three types of data: a convolutional neural network (CNN) is used to capture the temporal fluctuation patterns of signals such as heart rate and blood pressure from physiological data; a graph neural network (GNN) is used to process the behavioral trajectory map constructed based on UWB coordinates to extract the user's spatial movement patterns; a fully connected layer is used to transform environmental data to form features representing the scene state; a cross-modal attention fusion network is constructed, which dynamically assigns different attention weights to various features by calculating the correlation strength between different modal features (e.g., using mutual information entropy); based on these weights, the extracted physiological temporal features, behavioral spatial features, and environmental scene features are weighted and summed to generate a unified, compact fusion feature vector rich in multi-dimensional correlation information, which serves as the input to the subsequent risk prediction model.
[0081] S40: Input the fused feature vector into the dynamically updated personalized analysis model to predict health risks and output the risk level.
[0082] In this embodiment, the personalized analysis model refers to a machine learning model that is continuously optimized based on the user's personal historical data to assess their specific health status; health risk prediction refers to using the model to analyze current and historical characteristics to predict the probability of a specific health event (such as a fall or sudden illness) occurring in the future; and risk level is a classification label used to indicate the severity of risk, based on the predicted risk probability and according to preset clinical standards.
[0083] Specifically, the generated fused feature vector is first input into the user's personalized analysis model. This model is pre-trained and initialized based on large-scale group data and a fixed period (e.g., weekly). Using the latest multimodal data collected within this period and its corresponding health status labels confirmed by medical experts, the model parameters are iteratively updated using a gradient descent algorithm to achieve dynamic personalized adaptation. Internally, the model employs a temporal prediction algorithm (e.g., bidirectional LSTM) to process the input feature sequence, combined with an attention mechanism to capture long-term dependencies between features, predicting the probability of a user experiencing a specific health event (e.g., fall, cardiovascular accident) in the future. Based on the predicted... The system determines the probability of risk and combines it with preset clinical risk grading standards (such as dividing risk into low, medium, and high levels, each corresponding to different intervention strategy thresholds) to output the risk level label for the current moment. For example, when the model predicts that an elderly user has a fall risk probability of over 60% in the next 24 hours, the system automatically marks the user's risk level as "high risk" and triggers the corresponding warning and intervention process. At the same time, the personalized analysis model can dynamically adjust its internal parameters by continuously learning new user data, ensuring that the prediction results are always highly consistent with the evolution of the user's actual health status, effectively solving the problem of decreased prediction accuracy caused by changes in the user's health status in traditional models.
[0084] S50: Based on the risk level and the user's scenario, generate and output personalized intervention suggestions.
[0085] In this embodiment, the user's scenario refers to the specific physical space area where the user is currently located, such as a bedroom, bathroom, or living room, as determined by UWB positioning technology; personalized intervention suggestions refer to targeted guidance or measures generated based on the user's personal health status, real-time risk level, and specific risk factors of the scenario.
[0086] Specifically, the system first invokes a pre-built scenario-based rule base, which defines differentiated risk assessment thresholds and response strategies for different high-frequency activity scenarios (such as bathrooms and bedrooms). The system matches the user's current scenario based on UWB location information and invokes the corresponding rule base. Combining the user's individual characteristics (such as age and medical history), the system fine-tunes and optimizes the previously determined risk level in a scenario-based manner. Based on the optimized final risk level, the system matches and combines rules from the rule base to generate specific and actionable intervention suggestions. These suggestions include immediate reminders for users, automatic control commands for environmental devices, or notifications to family members and medical staff, and are output through multiple channels such as voice, screen, or mobile terminals.
[0087] In one embodiment, such as Figure 2As shown, in step S10, namely, constructing a multi-source acquisition terminal network to synchronously collect users' physiological data, behavioral data, and environmental data, and performing spatiotemporal registration on the multimodal data, the steps include:
[0088] S11: The multi-source acquisition terminal network is constructed by integrating wearable physiological monitoring devices, UWB positioning base stations, foot pressure sensing pads, and environmental monitoring nodes.
[0089] In this embodiment, wearable physiological monitoring devices refer to smart devices worn by users for continuous monitoring of physiological indicators such as heart rate, blood pressure, and blood oxygen saturation; UWB positioning base stations are devices deployed indoors that use ultra-wideband technology to achieve centimeter-level high-precision real-time positioning; foot pressure sensing pads are sensing devices laid in specific areas to sense the distribution of foot pressure and gait characteristics of users; and environmental monitoring nodes are sensors deployed in living spaces to collect environmental parameters such as temperature, humidity, light intensity, and air quality.
[0090] Specifically, a multi-source data acquisition terminal network is constructed by installing the above four types of hardware devices in the user's living environment (such as at home or in a retirement home); wearable devices are worn on the user's wrist or chest; UWB positioning base stations are installed in room corners or on the ceiling to form a positioning field covering the user's activity range; foot pressure sensing pads are mainly laid in key areas prone to falls, such as bathrooms and doorways; environmental monitoring nodes are distributed throughout the room; all devices are connected to a unified IoT gateway via wired or wireless means, and device registration and network synchronization are completed during system initialization, forming a multimodal data acquisition terminal network that can work collaboratively, providing a hardware foundation for subsequent synchronous data acquisition and fusion analysis.
[0091] S12: Use a timestamp calibration algorithm to unify the data collected by the multi-source acquisition terminal network to the same time axis.
[0092] In this embodiment, the timestamp calibration algorithm refers to a technical method for aligning and correcting timestamps from different acquisition devices; the data acquired by the multi-source acquisition terminal network refers to the original data generated asynchronously by heterogeneous devices such as wearable devices and positioning base stations, each with its own local timestamp; the same time axis refers to a unified, high-precision time reference system, under which all data are converted and mapped to ensure that they are comparable and correlated in the time dimension.
[0093] Specifically, the entire multi-source acquisition terminal network provides a unified high-precision time reference source, such as through the Precise Time Protocol (PTP) or a built-in high-stability clock. When data is generated, each acquisition device records its own local timestamp and simultaneously records this unified reference time. To address the delays introduced during data transmission and processing, the timestamp calibration algorithm calculates the time offset of each data stream in real time and uses interpolation or extrapolation methods to dynamically compensate and correct the timestamps. The timestamp errors of all physiological, behavioral, and environmental data from different devices are controlled within a preset threshold at the millisecond level and are strictly aligned to the same time axis, laying a precise time foundation for subsequent spatiotemporal correlation and fusion analysis.
[0094] S13: Based on the spatial coordinate data obtained by the UWB positioning base station, establish a spatial correlation mapping between physiological data, behavioral data and environmental data.
[0095] In this embodiment, spatial coordinate data refers to the precise location information of the user in three-dimensional space, which is output in real time by the UWB positioning system and is usually represented in the form of (X, Y, Z) coordinates; spatial association mapping refers to logically binding the physiological data, behavioral data and environmental data collected at a certain moment with a specific physical location or scene area through the spatial coordinates at that moment, and establishing a one-to-one correspondence between data and space.
[0096] Specifically, the system receives and processes user location coordinates sent by UWB positioning base stations in real time. For each moment, the system binds various types of data acquired at that moment: physiological indicators read from wearable devices, gait data obtained from foot pressure pads, and parameters collected from various environmental sensors are all associated with the corresponding UWB spatial coordinates. Through this association, the system can clearly know "at a certain time, at the bedside location in the bedroom, the user's heart rate underwent a specific change, and the ambient temperature at that location is also known." This constructs a spatiotemporal data point that integrates multi-dimensional information with spatial coordinates as the link, providing a foundation for subsequent analysis of the user's comprehensive state in different spatial locations.
[0097] In one embodiment, in step S12, the step of using a timestamp calibration algorithm to unify the data collected by the multi-source acquisition terminal network to the same time axis includes the following steps:
[0098] S121: Provide a unified time reference for the multi-source acquisition terminal network;
[0099] S122: Compensate for time deviations caused by data transmission delays and control the time synchronization error of multi-source data within a preset threshold.
[0100] In this embodiment, a unified time reference refers to a high-precision and stable clock source, which serves as a common reference for all devices in the network to record time. The time deviation caused by data transmission delay refers to the difference between the actual timestamp and the ideal time caused by network congestion, different device processing speeds, etc., during the data generation, transmission, and reception processing process. The preset threshold refers to the maximum allowable time error range set to ensure the accuracy of subsequent fusion analysis, which is usually in the millisecond range.
[0101] Specifically, the entire multi-source acquisition terminal network is configured with a unified time reference, such as by connecting to the Precise Time Protocol (PTP) or configuring a high-precision synchronization clock module, to ensure that all devices have a consistent time starting point at the start of data acquisition. The system monitors the transmission status of each data in real time. For the unavoidable delays in data transmission and processing, the timestamp calibration algorithm adopts a dynamic compensation mechanism. By calculating the transmission delay of data from the acquisition device to the IoT gateway in real time and combining it with the device's own clock drift parameters, the timestamp of each data packet is interpolated and corrected. The time synchronization error of all calibrated data is strictly controlled within the preset millisecond threshold, achieving accurate alignment of multimodal data in the time dimension.
[0102] In one embodiment, such as Figure 3 As shown, in step S20, which involves data cleaning and uncertainty quantification of the spatiotemporally registered multimodal data to generate preprocessed data with confidence labels, the steps include:
[0103] S21: Perform multi-level data cleaning on the spatiotemporally registered data, including outlier removal, noise filtering, and missing value completion.
[0104] In this embodiment, outlier removal refers to identifying and removing erroneous or invalid data points that significantly deviate from the normal data distribution range; noise filtering refers to using signal processing techniques to weaken or eliminate high-frequency random fluctuations in the data caused by equipment interference, motion artifacts, etc.; missing value completion refers to using reasonable methods to fill or reconstruct gaps in the data sequence caused by brief signal interruptions or equipment non-response; multi-level data cleaning refers to performing the above three operations sequentially and layer by layer to systematically improve data quality.
[0105] Specifically, targeted cleaning operations are performed on various data sequences after spatiotemporal registration. For physiological data sequences, statistical methods (such as the 3σ principle) are used to identify and remove abnormal abrupt values caused by poor sensor contact. Wavelet transform algorithms are used to filter out physiological data noise caused by motion interference. For behavioral trajectory data, smoothing algorithms are used to filter out high-frequency jitter in positioning coordinates. For environmental data, time series trend prediction methods (such as LSTM prediction algorithms) are used to interpolate and complete short-term gaps caused by transmission packet loss. After all cleaning operations are completed, a cleaner, more complete, and reliable multimodal data sequence is output for subsequent uncertainty quantification.
[0106] S22: Introduce a probability model to assess the uncertainty of the cleaned data and calculate the confidence level of each data point.
[0107] In this embodiment, the probabilistic model is a mathematical framework for quantifying the reliability of data. It assesses the uncertainty of data based on the statistical characteristics of the data itself or its distribution in the feature space. Uncertainty assessment refers to the systematic analysis of the credibility of each data point due to factors such as sensor noise, environmental interference, or model error. The confidence level is a value between 0 and 1, used to quantify the probability of the reliability of the data point. The higher the value, the more reliable the data.
[0108] Specifically, the cleaned multimodal data is input into a pre-defined Bayesian probability model. This model calculates a posterior probability as a confidence level label for each data point based on historical data distribution and current observation characteristics. For physiological data, the model considers sensor accuracy, individual user differences, and historical fluctuation range to assess the reliability of the current reading. For behavioral trajectory data, the model quantifies the positioning error probability of coordinate points by combining the spatial distribution density and signal strength of positioning base stations. For environmental data, the model assesses the stability of parameter readings based on sensor calibration history and dynamic environmental changes. All data points are assigned a confidence level value between 0 and 1, where 1 represents complete reliability and 0 represents complete unreliability, generating a pre-processed multimodal dataset with confidence level labels to provide a reliable basis for subsequent fusion analysis. For example, when physiological data at a certain moment experiences abnormal readings due to a brief period of sensor malfunction, the model assigns a lower confidence level value, indicating that subsequent analysis should handle this data point with caution.
[0109] S23: Mark data with confidence levels below a preset threshold and reduce their weight in subsequent fusion analysis.
[0110] In this embodiment, the preset threshold refers to a confidence threshold value (e.g., 0.7) set in advance based on the tolerance of the actual application for data quality; marking refers to adding a specific identifier to the metadata of the data points to intuitively distinguish their low confidence status; reducing their weight in subsequent fusion analysis refers to systematically reducing the impact of these low confidence data points on the final analysis results in the multimodal feature fusion stage.
[0111] Specifically, the confidence level of each data point calculated in step S22 is first compared with a preset threshold. For all data points with a confidence level lower than the threshold, the system adds a specific "low confidence" label to their data structure. In subsequent feature extraction and fusion stages, such as when using an attention mechanism to calculate feature weights, the system automatically identifies these labels. For the labeled data points, their corresponding original feature values or representations in the feature space are assigned a significantly reduced weight coefficient, or even partially masked in some cases. In this way, the negative impact of low-quality, high-uncertainty data on the overall health risk prediction results is effectively suppressed, thereby improving the robustness of the system and the reliability of the analysis conclusions.
[0112] In one embodiment, such as Figure 4 As shown, in step S30, which involves extracting features from the preprocessed multimodal data and fusing the extracted multimodal features based on an attention mechanism to generate a fused feature vector, the steps include:
[0113] S31: Use convolutional neural networks to extract temporal features of physiological data, use graph neural networks to extract spatial features of behavioral data, and use fully connected layers to extract scene features of environmental data.
[0114] In this embodiment, a convolutional neural network is a deep learning model that excels at processing data with a grid-like topological structure (such as time series and images). It is used to automatically learn the regular patterns of physiological data over time, i.e., time-series features. A graph neural network is a model specifically designed to process graph-structured data. Here, it is used to learn behavioral trajectory maps composed of UWB positioning points to extract the spatial correlation and patterns of user movement patterns, i.e., spatial features. A fully connected layer is a basic neural network layer that abstracts environmental data into feature vectors that can characterize the overall environmental conditions, i.e., scene features, by performing nonlinear transformations and combinations on the environmental data.
[0115] Specifically, for physiological data, a multi-layer convolutional neural network model is constructed. The input is preprocessed time-series physiological data. The model slides the convolutional kernel along the time dimension to automatically capture the periodic patterns and abrupt changes in indicators such as heart rate and blood pressure over time, and outputs a fixed-dimensional temporal feature vector. For behavioral data, a temporal graph (nodes are location points, and edges represent movement relationships) is constructed based on UWB coordinate sequences and input into a graph neural network. Through operations such as graph convolution, it learns and outputs spatial feature vectors representing the user's activity range, movement patterns, and gait stability. For environmental data, a fully connected layer receives multi-dimensional environmental parameters such as temperature and humidity, performs feature combination and abstraction, and generates feature vectors that reflect the overall state of the current environment. The three feature vectors are unified to the same dimensional space through independent feature transformation layers, providing structurally consistent multimodal feature inputs for subsequent attention fusion.
[0116] S32: Construct a cross-modal attention fusion network and calculate the attention weights for different modal features.
[0117] In this embodiment, the cross-modal attention fusion network is a specialized neural network architecture whose core function is to evaluate and integrate features from different data sources (modalities). Calculating the attention weights of different modal features means that the network automatically analyzes the intrinsic correlation and importance differences between physiological temporal features, behavioral spatial features and environmental scene features through internal mechanisms, and dynamically assigns a coefficient representing the relative importance of each type of feature. The sum of these coefficients is usually 1.
[0118] Specifically, the three types of feature vectors extracted in step S31 are used as inputs and fed into the constructed attention fusion network. This network first uses an attention mechanism to calculate the correlation between features within the same modality, capturing local dependencies within each modality. The network further analyzes the global correlation between features from different modalities, such as the potential link between heart rate changes in physiological features and gait stability in behavioral features, as well as the possible impact of ambient temperature on physiological indicators. Through this multi-level correlation analysis, the network calculates an attention weight for each feature vector, reflecting the relative importance of that feature in the overall health assessment. For example, when a user is in motion, gait data in behavioral features is given a higher weight, while when a user is at rest, heart rate and blood pressure data in physiological features dominate. These dynamically adjusted attention weights ensure that the fused feature vectors accurately reflect the user's comprehensive health status in different scenarios, providing a more accurate basis for subsequent risk prediction.
[0119] S33: The temporal features, spatial features, and scene features are weighted and fused according to the attention weights to generate the fused feature vector.
[0120] In this embodiment, weighted fusion refers to linearly combining feature vectors that respectively represent physiological, behavioral, and environmental states according to their respective importance weights calculated under the attention mechanism; the fused feature vector refers to a new, unified feature representation generated through this weighted combination.
[0121] Specifically, the system receives three sets of attention weights (e.g., denoted as α, β, and γ, respectively, with α+β+γ=1) calculated by a cross-modal attention fusion network, corresponding to physiological temporal features, behavioral spatial features, and environmental scene features. These three sets of weights are then multiplied element-wise (weighted) with their corresponding feature vectors. All weighted feature vectors are summed to generate a single, fixed-dimensional fusion feature vector. This fusion feature vector not only contains information from all original modalities but also highlights the feature components most relevant to the current health status through attention weights, providing a more discriminative and interpretable input for subsequent personalized risk prediction models.
[0122] In one embodiment, step S32, which involves constructing a cross-modal attention fusion network and calculating attention weights for different modal features, includes the following steps:
[0123] S321: Use the mutual information entropy algorithm to measure the correlation strength between physiological characteristics, behavioral characteristics and environmental characteristics.
[0124] In this embodiment, the mutual information entropy algorithm refers to an information theory-based method used to quantify the correlation between two random variables, that is, the degree to which the information content of one variable can reduce the uncertainty of another variable; the correlation strength refers to the value calculated by mutual information entropy, which reflects the amount of information shared between different modal features, and the larger the value, the stronger the correlation.
[0125] Specifically, for the physiological temporal feature vector, behavioral space feature vector, and environmental scene feature vector extracted in step S31, the mutual information entropy values between each pair are calculated. For example, the mutual information entropy between physiological features and behavioral features is calculated to assess the correlation between the user's heart rate changes and gait stability. Similarly, the mutual information entropy between physiological features and environmental features, and between behavioral features and environmental features, is calculated to obtain three sets of correlation strength values. These values provide an objective basis for the subsequent allocation of attention weights, ensuring that the weights can accurately reflect the intrinsic connections between different modal features. For example, when the mutual information entropy between physiological features and environmental features is high, it indicates that environmental factors have a significant impact on physiological state, and should be given higher weights during fusion.
[0126] S322: Dynamically allocate attention weights based on the correlation strength, with higher correlation strength modal features being assigned higher weight coefficients.
[0127] In this embodiment, dynamic attention weight allocation refers to assigning an appropriate importance ratio to each modality feature in real time and adaptively based on the actual correlation strength between the current input features; the weight coefficient is a specific value assigned to each modality feature, which determines the proportion of the feature in the final fusion result.
[0128] Specifically, based on the three sets of mutual information entropy values calculated in step S321 (i.e., the correlation strength between physiology-behavior, physiology-environment, and behavior-environment), a weight allocation function based on correlation strength is set (e.g., F_fuse = α×F_phys + β×F_behavior + ...). The function γ×F_env, where α+β+γ=1 (α, β, and γ are attention weights), maps the mutual information entropy value to a weight coefficient between 0 and 1, and satisfies the condition that the sum of all modal weights is 1. It maps the association strength value of each modality to the interval [0,1], and ensures that the sum of the weight coefficients of the three modalities is 1. For example, when the mutual information entropy between physiological features and environmental features is significantly higher than other combinations, the environmental modality will be assigned a higher weight coefficient, while the modality with lower association strength will have its weight reduced accordingly. The dynamic weight allocation mechanism ensures that the fusion process can make full use of the intrinsic connection between different modal features, so that the final generated fusion feature vector more accurately reflects the user's comprehensive health status. For example, in a high-temperature environment, the user's physiological indicators (such as heart rate and blood pressure) may change due to heat stress. At this time, the association strength between environmental features and physiological features will be enhanced, and the system will automatically increase the weight of the environmental modality, thereby more accurately capturing the impact of environmental factors on the user's health. In this way, the cross-modal attention fusion network realizes the intelligent integration of multimodal features, providing more reliable and comprehensive feature inputs for subsequent risk prediction.
[0129] In one embodiment, such as Figure 5 As shown, in step S40, the step of inputting the fused feature vector into the dynamically updated personalized analysis model to predict health risks and output risk levels includes the following steps:
[0130] S41: Train the initial personalized analysis model based on multi-sample population data.
[0131] In this embodiment, multi-sample population data refers to historical multimodal data and corresponding labeled health event records collected from a large number of different individuals (covering different ages, genders, and health conditions); the initial personalized analysis model refers to a pre-trained machine learning model.
[0132] Specifically, historical multimodal data from multiple individuals of different ages, genders, and health conditions are collected. This data includes physiological data (such as heart rate and blood pressure), behavioral data (such as gait and activity range), and environmental data (such as temperature and humidity). Simultaneously, records of corresponding health events annotated by medical experts, such as disease onset and falls, are acquired. Next, the collected data undergoes preprocessing, including cleaning, normalization, and annotation, to ensure data quality and consistency. A transfer learning algorithm is then used to train an initial personalized analysis model. This involves dividing the preprocessed multi-sample population data into training and test sets. The training set is used to train the model, adjusting model parameters to minimize prediction error. The test set is used to evaluate the model's performance, ensuring good generalization ability. This results in an initial personalized analysis model trained on multi-sample population data. This model can initially identify the correlation patterns between multimodal data and health risks, providing a foundation for subsequent personalized updates and risk prediction.
[0133] S42: Periodically update the parameters of the personalized analysis model based on the user's latest health status feedback data.
[0134] In this embodiment, "regular" refers to triggering the update process according to a preset fixed cycle (e.g., weekly or monthly); the user's latest health status feedback data refers to the user's multimodal data collected by the system in the most recent cycle, as well as the real health outcome labels corresponding to that cycle confirmed through user reports, medical records, etc.; iterative update refers to using new data to adjust and optimize the internal parameters of the existing model so that its predictions are more adapted to the user's latest health status.
[0135] Specifically, the system automatically triggers a model update process according to a preset cycle (e.g., once a week). First, it collects the user's multimodal sensory data from the past cycle, including physiological indicators (such as continuous monitoring values of heart rate and blood pressure), behavioral data (such as daily activity trajectories and gait analysis results), and environmental parameters (such as records of indoor and outdoor temperature and humidity changes). Simultaneously, through user-initiated reporting (such as health diaries and symptom descriptions) or by connecting with data interfaces from medical institutions, it obtains the user's actual health outcome labels for the same cycle, such as whether a fall occurred or whether there was an acute attack of an illness. These newly collected data are then integrated with historical data in a time-series fashion to form an incremental dataset with actual health labels. Gradient descent is then applied to this dataset. The algorithm iteratively updates the model parameters on the incremental dataset; it calculates the gradient of the loss function with respect to the model parameters through backpropagation, and uses the gradient descent algorithm to gradually adjust the model parameters, so that the prediction error of the model on the new data is continuously reduced; after iterative updates, the parameters of the personalized analysis model are optimized, which can better adapt to the user's latest health status changes and improve the accuracy and personalization of health risk prediction; for example, if a user frequently experiences abnormally high heart rate in the most recent period and eventually develops cardiovascular-related diseases, the model will pay more attention to changes in heart rate indicators after the update, and assign higher weights to heart rate features in subsequent risk predictions, thereby discovering potential health risks more timely and accurately.
[0136] S43: Input the fused feature vector into the updated personalized analysis model, use the time series prediction algorithm to predict the probability of health risks, and output the corresponding risk level.
[0137] In this embodiment, the time series prediction algorithm is an algorithm that can process sequence data and predict future trends (such as LSTM); the health risk probability refers to the possibility calculated by the model of a specific health risk event (such as a fall or sudden illness) occurring within a specific time period in the future, usually expressed as a value between 0% and 100%; the risk level is based on preset clinical guidelines or management strategies, mapping the predicted probability value to several discrete levels with clear operational meanings (such as "low risk", "medium risk", "high risk", "urgent").
[0138] Specifically, the fused feature vector sequence representing the user's recent state, generated in step S30, is input into the currently updated personalized analysis model. The model typically includes a time-series prediction module (such as a bidirectional LSTM network), which analyzes the long-term dependencies and evolution trends in the feature sequence. The model ultimately outputs predicted probability values for one or more health risk events. The system then compares these probability values with preset grading thresholds (e.g., a probability > 85% corresponds to an "emergency" level) to determine the final risk level. The system then standardizes, encapsulates, and outputs this level result for use by the subsequent intervention suggestion module.
[0139] In one embodiment, in step S42, the periodic iterative update of the parameters of the personalized analysis model based on the user's latest health status feedback data includes:
[0140] S421: Set a fixed model update cycle;
[0141] S422: At the end of each update cycle, collect the user's multimodal data and corresponding health status labels for that cycle.
[0142] In this embodiment, the model update cycle refers to the time interval that the system pre-configures and regularly triggers model retraining and parameter adjustment; user multimodal data refers to physiological, behavioral and environmental data that are continuously collected and preprocessed by various sensors within a single cycle; and health status labels refer to labeled information that reflects the user's actual health status and is confirmed by medical personnel or obtained through reliable clinical records within the same time period.
[0143] Specifically, during the initialization phase, the system sets a fixed model update cycle, such as weekly or monthly, based on user needs and the sensitivity of health monitoring, to ensure that the model can adapt to changes in the user's health status in a timely manner. At the end of each preset update cycle, the system automatically triggers a data collection process, collecting multimodal data within that cycle through various sensors deployed around the user (such as wearable devices and environmental monitors), including but not limited to physiological indicators such as heart rate and blood pressure, behavioral data such as gait and activity range, and environmental parameters such as temperature and humidity. At the same time, the system obtains the user's real health status labels within the same cycle through user-initiated reports (such as health diaries and symptom descriptions) or by connecting with data interfaces of medical institutions, such as whether a fall has occurred or whether there has been an acute onset of disease. These labels provide key supervisory information for the iterative updates of the model.
[0144] S423: Using the multimodal data and health status labels to construct a training sample set, the parameters of the personalized analysis model are optimized using the gradient descent algorithm.
[0145] In this embodiment, the training sample set is a dataset used for model training, consisting of data collected in step S422 and label pairings; the gradient descent algorithm is an optimization method that minimizes the loss function by calculating the gradient of the loss function with respect to the model parameters and iteratively updating the parameters in the reverse direction of the gradient.
[0146] Specifically, the multimodal data collected in step S422 is paired and organized with corresponding health status labels to construct a training sample set. This sample set contains both the user's latest health status information and retains the time-series characteristics of the data. Next, the training sample set is input into the current personalized analysis model. The model calculates the predicted output based on the input data and compares it with the actual health status labels to determine the prediction error. The gradient descent algorithm is used to calculate the gradient of the loss function with respect to the model parameters through backpropagation, i.e., the degree of influence of each parameter on the prediction error. Based on the calculated gradient, the model parameters are iteratively updated along the gradient's reverse direction, with each update adjusting the model parameters in a direction that reduces the prediction error. After multiple iterations, the model parameters are optimized, enabling a better fit to the user's latest health status data, thereby improving the accuracy and personalization of health risk prediction. For example, during the update process, if it is found that the model is not sensitive enough to recent abnormal increases in the user's heart rate, the gradient descent algorithm will automatically adjust the weights of parameters related to heart rate features, making the model pay more attention to changes in heart rate indicators in subsequent predictions.
[0147] In one embodiment, such as Figure 6 As shown, in step S50, which involves generating and outputting personalized intervention suggestions based on the risk level and the user's scenario, the steps include:
[0148] S51: Call the contextual rule library corresponding to the user's current scenario;
[0149] S52: Optimize the risk level based on the scenario-based rule base;
[0150] S53: Based on the optimized risk level and individual user characteristics, generate and output personalized intervention suggestions.
[0151] In this embodiment, the scenario-based rule base is a predefined knowledge base that stores specific risk assessment rules, intervention thresholds, and coping strategies for different physical space areas (such as bedrooms, bathrooms, and living rooms). Optimizing the risk level refers to fine-tuning the initial risk level output by the general model by combining the inherent risk factors of the scenario (such as slippery bathroom floors) to make it more consistent with the urgency level in the actual environment. User individual characteristics include static health record information such as the user's age, medical history, and medication. Personalized intervention suggestions are executable guidance or operation instructions formed by integrating the optimized risk level, specific scenario, and user's personal circumstances.
[0152] Specifically, the system determines the user's current scenario (e.g., "master bedroom bathroom") based on real-time UWB positioning data and calls the corresponding rule base for that scenario. The rule base defines risk correction factors and special thresholds for that scenario (e.g., in the bathroom, the fall risk level corresponding to the same gait instability should be automatically increased by one level). The system uses these rules to reassess and optimize the initial risk level output in step S40, obtaining a final risk level adapted to the scenario. The system combines this final risk level with the user's individual characteristics (e.g., the user has a history of osteoporosis) to match and combine intervention strategies from the rule base to generate specific suggestions. These suggestions include voice reminders for the user, automatic control commands for environmental devices (e.g., turning on the bathroom exhaust fan, brightening the lights), and warning information pushed to family members or caregivers. The generated suggestions are output in real time through preset communication channels (e.g., indoor voice terminal, mobile APP, nursing platform), completing the closed loop from risk perception to intervention execution.
[0153] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0154] In one embodiment, a multimodal human perception data fusion analysis and risk prediction system for smart healthcare is provided. This system corresponds one-to-one with the multimodal human perception data fusion analysis and risk prediction method for smart healthcare described in the previous embodiment. Figure 7 As shown, this multimodal human perception data fusion analysis and risk prediction system for smart healthcare includes:
[0155] The data acquisition and registration module integrates wearable physiological monitoring devices, UWB positioning base stations, plantar pressure sensing pads, and environmental monitoring nodes to build a multi-source acquisition terminal network, simultaneously collecting users' physiological data, behavioral data, and environmental data, and using a timestamp calibration algorithm to complete the spatiotemporal registration of multimodal data;
[0156] The data preprocessing and quantization module performs multi-level data cleaning on the spatiotemporally registered multimodal data, including outlier removal, noise filtering, and missing value completion. It also introduces a probability model to assess the uncertainty of the cleaned data and generates preprocessed data with confidence labels.
[0157] The feature fusion and analysis module uses convolutional neural networks to extract temporal features of physiological data, graph neural networks to extract spatial features of behavioral data, and fully connected layers to extract scene features of environmental data. It also uses a cross-modal attention fusion network to perform weighted fusion of multimodal features to generate a fused feature vector.
[0158] The risk prediction and decision-making module inputs the fused feature vector into a dynamically updated personalized analysis model, predicts the probability of health risks through a time-series prediction algorithm, and outputs the risk level in combination with a scenario-based rule base.
[0159] The intervention suggestion output module generates and outputs personalized intervention suggestions based on the risk level and the user's scenario, by calling the scenario-based rule base.
[0160] As described above, it is understood that each component of the multimodal human perception data fusion analysis and risk prediction system for smart healthcare proposed in this application can realize the function of any one of the multimodal human perception data fusion analysis and risk prediction methods for smart healthcare as described above, and the specific structure will not be repeated.
[0161] The present invention and its embodiments have been described above. This description is not restrictive. The accompanying drawings are only one embodiment of the present invention. The actual structure is not limited to this. In short, if a person skilled in the art is inspired by this description and designs a similar structure and embodiment without departing from the spirit of the present invention, such design should fall within the protection scope of the present invention.
Claims
1. A method for multimodal human perception data fusion analysis and risk prediction for smart healthcare, characterized in that, Including the following steps: A multi-source acquisition terminal network is constructed to simultaneously collect users' physiological data, behavioral data, and environmental data, and spatiotemporal registration is performed on the multimodal data; Data cleaning and uncertainty quantification are performed on the spatiotemporally registered multimodal data to generate preprocessed data with confidence labels; Feature extraction is performed on the preprocessed data with confidence labels, and the extracted multimodal features are fused based on an attention mechanism to generate a fused feature vector; The fused feature vector is input into a dynamically updated personalized analysis model to predict health risks and output the risk level. Based on the risk level and the user's scenario, personalized intervention suggestions are generated and output.
2. The method for multimodal human perception data fusion analysis and risk prediction for smart healthcare as described in claim 1, characterized in that, The construction of a multi-source acquisition terminal network to simultaneously collect users' physiological data, behavioral data, and environmental data, and to perform spatiotemporal registration of the multimodal data, includes the following steps: The multi-source acquisition terminal network is constructed by integrating wearable physiological monitoring devices, UWB positioning base stations, foot pressure sensing pads, and environmental monitoring nodes. A timestamp calibration algorithm is used to unify the data collected by the multi-source acquisition terminal network to the same time axis; Based on the spatial coordinate data obtained from the UWB positioning base station, a spatial correlation mapping is established between physiological data, behavioral data, and environmental data.
3. The method for multimodal human perception data fusion analysis and risk prediction for smart healthcare as described in claim 2, characterized in that, The process of using a timestamp calibration algorithm to unify the data collected by the multi-source acquisition terminal network to the same time axis includes: Provide a unified time reference for the multi-source acquisition terminal network; Compensation is provided for time deviations caused by data transmission delays, and the time synchronization error of multi-source data is controlled within a preset threshold.
4. The method for multimodal human perception data fusion analysis and risk prediction for smart healthcare as described in claim 1, characterized in that, The process of cleaning and uncertainty quantification of the spatiotemporally registered multimodal data to generate preprocessed data with confidence labels includes the following steps: Multi-level data cleaning is performed on the spatiotemporally registered data, including outlier removal, noise filtering, and missing value completion. A probability model is introduced to assess the uncertainty of the cleaned data and calculate the confidence level of each data point. Data with confidence levels below a preset threshold are labeled and their weight in subsequent fusion analysis is reduced.
5. The method for multimodal human perception data fusion analysis and risk prediction for smart healthcare as described in claim 1, characterized in that, The step of extracting features from the preprocessed multimodal data and fusing the extracted multimodal features based on an attention mechanism to generate a fused feature vector includes the following steps: Convolutional neural networks are used to extract temporal features from physiological data, graph neural networks are used to extract spatial features from behavioral data, and fully connected layers are used to extract scene features from environmental data. Construct a cross-modal attention fusion network and calculate the attention weights for different modal features; The temporal features, spatial features, and scene features are weighted and fused according to the attention weights to generate the fused feature vector.
6. The method for multimodal human perception data fusion analysis and risk prediction for smart healthcare as described in claim 5, characterized in that, The construction of the cross-modal attention fusion network and the calculation of attention weights for different modal features include the following steps: The mutual information entropy algorithm is used to measure the correlation strength among physiological characteristics, behavioral characteristics, and environmental characteristics. Attention weights are dynamically assigned based on the correlation strength, with higher correlation strength modal features being assigned higher weight coefficients.
7. The method for multimodal human perception data fusion analysis and risk prediction for smart healthcare as described in claim 1, characterized in that, The step of inputting the fused feature vector into a dynamically updated personalized analysis model to predict health risks and output risk levels includes the following steps: The initial personalized analysis model was trained based on multi-sample population data; The parameters of the personalized analysis model are periodically updated based on the user's latest health status feedback data. The fused feature vector is input into the updated personalized analysis model, and a time-series prediction algorithm is used to predict the probability of health risks and output the corresponding risk level.
8. The method for multimodal human perception data fusion analysis and risk prediction for smart healthcare as described in claim 7, characterized in that, The periodic iterative update of the parameters of the personalized analysis model based on the user's latest health status feedback data includes: Set a fixed model update cycle; At the end of each update cycle, collect user multimodal data and corresponding health status tags for that cycle; The training sample set is constructed using the multimodal data and health status labels, and the parameters of the personalized analysis model are optimized using the gradient descent algorithm.
9. The method for multimodal human perception data fusion analysis and risk prediction for smart healthcare and elderly care according to claim 1, characterized in that, The process of generating and outputting personalized intervention suggestions based on the risk level and the user's scenario includes the following steps: Call the contextual rule library corresponding to the user's scenario; The risk level is optimized based on the aforementioned scenario-based rule base; Based on the optimized risk level and individual user characteristics, personalized intervention suggestions are generated and output.
10. A multimodal human perception data fusion analysis and risk prediction system for smart healthcare, characterized in that, include: The data acquisition and registration module integrates wearable physiological monitoring devices, UWB positioning base stations, plantar pressure sensing pads, and environmental monitoring nodes to build a multi-source acquisition terminal network, simultaneously collecting users' physiological data, behavioral data, and environmental data, and using a timestamp calibration algorithm to complete the spatiotemporal registration of multimodal data; The data preprocessing and quantization module performs multi-level data cleaning on the spatiotemporally registered multimodal data, including outlier removal, noise filtering, and missing value completion. It also introduces a probability model to assess the uncertainty of the cleaned data and generates preprocessed data with confidence labels. The feature fusion and analysis module uses convolutional neural networks to extract temporal features of physiological data, graph neural networks to extract spatial features of behavioral data, and fully connected layers to extract scene features of environmental data. It also uses a cross-modal attention fusion network to perform weighted fusion of multimodal features to generate a fused feature vector. The risk prediction and decision-making module inputs the fused feature vector into a dynamically updated personalized analysis model, predicts the probability of health risks through a time-series prediction algorithm, and outputs the risk level in combination with a scenario-based rule base. The intervention suggestion output module generates and outputs personalized intervention suggestions based on the risk level and the user's scenario, by calling the scenario-based rule base.