Method and system for monitoring dietary nutrition status of the elderly based on multi-source data fusion
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-11
- Publication Date
- 2026-08-11
AI Technical Summary
[0003]本发明提供了一种基于多源数据融合的老年人膳食营养状况监测方法及系统,解决了现有技术中老年人膳食营养系统存在监测精度、个性化程度及实用性均无法满足健康老龄化背景下养老服务需求,特征提取不充分、数据融合无因果约束及个体适配性差的技术问题
[0020] According to a fourth aspect of the present invention, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the method according to a first aspect of the present invention.
Smart Images

Figure CN122552124A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of health monitoring and artificial intelligence technology applications, and in particular to a method and system for monitoring the dietary nutrition status of the elderly based on multi-source data fusion. It is applicable to various scenarios such as elderly care institutions, community-based elderly care, and home-based elderly care, and can realize the collection, fusion analysis, and risk warning of multi-dimensional data such as dietary nutrition intake, physiological state, and behavioral characteristics, belonging to the field of intelligent health management technology. Background Technology
[0002] Current methods for monitoring the diet and nutrition of the elderly mainly rely on traditional approaches such as manual recording and regular physical examinations. These methods suffer from problems such as untimely data collection, limited data dimensions, and delayed analysis, making it difficult to accurately match individual differences among the elderly (e.g., chronic diseases, medication history). Existing technologies, while some dietary and nutrition monitoring systems attempt to integrate multi-source data, exhibit significant shortcomings: First, feature extraction lacks adaptation to the scale differences of multimodal data; the feature representation of heterogeneous data such as images, time series, and structured data is insufficient, leading to incomplete data value mining. Second, the fusion of multimodal data does not consider causal logic, easily introducing future information leakage issues, violating the objective laws of nutritional status evolution in the elderly, and affecting the reliability of analysis results. Third, they cannot dynamically adapt to individual health differences among the elderly, lacking sufficient adaptability to the nutritional needs of special groups such as patients with chronic diseases and those on medication, making it difficult to provide personalized monitoring and early warning services. Consequently, the monitoring accuracy, personalization, and practicality of existing dietary and nutrition monitoring systems for the elderly fail to meet the needs of elderly care services in the context of healthy aging; existing systems suffer from insufficient feature extraction, lack of causal constraints in data fusion, and poor individual adaptability. Summary of the Invention
[0003] This invention provides a method and system for monitoring the dietary nutrition status of the elderly based on multi-source data fusion. It solves the technical problems of existing dietary nutrition systems for the elderly, such as insufficient monitoring accuracy, personalization and practicality, which cannot meet the needs of elderly care services in the context of healthy aging, insufficient feature extraction, lack of causal constraints in data fusion and poor individual adaptability.
[0004] According to a first aspect of the present invention, a method for monitoring the dietary nutritional status of the elderly based on multi-source data fusion is provided, comprising:
[0005] By using computer vision, wearable devices, smart sensors, and standardized medical interfaces, dietary images, multi-parameter physiological indicators, behavioral and environmental data, and structured health records are collected and aggregated to form a raw multi-source dataset.
[0006] The original multi-source data is cleaned, standardized, and formatted, and aligned according to a unified time window to generate a standardized dataset for model training and inference.
[0007] Image, temporal, and structured features are extracted from a standardized dataset using a feature pyramid network, a temporal Transformer, and a pre-trained language model, respectively. Multimodal feature fusion is performed through a causal cross-attention mechanism. Finally, the model is processed by an LSTM+Transformer hybrid model to output the prediction results and risk probabilities of future nutritional status.
[0008] It adopts a distributed database architecture that combines local and cloud computing to classify and persistently store real-time high-frequency multi-source data, historical archived multi-source data, future nutritional status prediction results and risk probabilities, and encrypted medical records.
[0009] Based on comprehensive dietary guidelines, individual health records, and historical data, combined with future nutritional status predictions and risk probabilities, a multi-factor nonlinear model is used to calculate and dynamically update individualized nutrient intake and physiological indicator baselines.
[0010] The system compares future nutritional status predictions with risk probabilities, real-time multi-source data with dynamically updated individualized physiological indicator baselines, identifies abnormal risk levels based on preset multi-level collaborative judgment rules, and triggers real-time early warning information pushes to multiple roles.
[0011] According to a second aspect of the present invention, a monitoring system for the dietary nutrition status of the elderly based on multi-source data fusion is provided for implementing the aforementioned monitoring method for the dietary nutrition status of the elderly based on multi-source data fusion, comprising: a data acquisition module, a data preprocessing module, an AI intelligent analysis module, a data storage module, a risk assessment module, a system interface module, and a security and privacy protection module;
[0012] Among them, the data acquisition module is used to collect multi-dimensional heterogeneous data, including dietary image recognition and nutritional composition calculation through computer vision algorithms and ResNet152 models carried by mobile terminal APP, physiological indicator data through wearable devices and smart sensors, behavioral and environmental data through smart home sensors, and electronic medical record structured data through HL7 and FHIR standard interfaces.
[0013] The data preprocessing module is used to collect data and upload it through the transport layer, then synchronize it to the data layer via HTTP / HTTPS protocol to perform preprocessing operations such as data cleaning, data standardization, and model training data partitioning.
[0014] The AI intelligent analysis module includes a multimodal multi-scale feature extraction unit, a causal cross-attention fusion unit, and a time-series prediction unit. The multimodal multi-scale feature extraction unit addresses the heterogeneity of dietary images, physiological time-series data, and structured health record data by extracting multi-scale visual features through a feature pyramid network, extracting time-dependent features through a temporal Transformer, and encoding textual semantic features through a DistilBERT model. It then extracts image, temporal, and structured features from a standardized dataset using the feature pyramid network, temporal Transformer, and a pre-trained language model, respectively. Multimodal feature fusion is performed through a causal cross-attention mechanism. Finally, the data is processed by an LSTM+Transformer hybrid model to output future nutritional status predictions and risk probabilities.
[0015] The data storage module is used to store real-time, high-frequency multi-source data, historical archived multi-source data, future nutritional status prediction results and risk probabilities, and encrypted medical records through local and cloud databases; it adopts off-site multi-copy backup, and medical sensitive data is additionally encrypted and stored.
[0016] The risk assessment module includes a dynamic baseline establishment unit and an anomaly collaborative identification unit. The dynamic baseline establishment unit constructs a dynamic baseline based on individual health records and historical data using a multi-factor weighted model. It integrates dietary guidelines, individual health records, and historical data, and combines future nutritional status predictions and risk probabilities to calculate and dynamically update individualized nutrient intake and physiological indicator baselines using a multi-factor nonlinear model.
[0017] The system interface module includes internal and external interface design, enabling the setting of interfaces between the perception layer and transmission layer and the medical system.
[0018] The security and privacy protection module is used to set up security and privacy protection measures.
[0019] According to a third aspect of the present invention, an electronic device is provided. The electronic device includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the program to implement the method according to the first aspect of the present invention.
[0020] According to a fourth aspect of the present invention, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the method according to a first aspect of the present invention.
[0021] Compared with existing technologies, the advantages and positive effects of this invention are:
[0022] This invention utilizes computer vision technology, wearable devices, smart sensors, and standardized medical interfaces to achieve comprehensive collection and aggregation of dietary images, physiological indicators, behavioral environmental data, and health records, forming a rich multi-source dataset. Data cleaning, standardization, and formatting ensure data consistency and alignment within a unified time window, providing a standardized dataset for model training and inference. Feature pyramid networks, temporal Transformers, and pre-trained language models are used to extract image, temporal, and structured features respectively, and a causal cross-attention mechanism is employed to fuse multimodal features, improving the model's interpretability and prediction accuracy. The LSTM+Transformer hybrid model processing comprehensively enhances the model's... Taking into account both time series and complex structured data, the system outputs more accurate predictions of future nutritional status and risk probabilities. Employing a distributed database architecture combining local and cloud-based systems, it achieves the classification and persistent storage of real-time high-frequency data, historical archived data, model outputs, and encrypted medical records, improving data security and access efficiency. By integrating dietary guidelines, individual health records, and historical data, it calculates and dynamically updates individualized nutrient intake and physiological indicator baselines through a multi-factor nonlinear model, making predictions more personalized and accurate. By comparing AI predictions, real-time data, and dynamic baselines, and based on preset multi-level collaborative judgment rules, it identifies abnormal risk levels and triggers real-time early warning information pushes, improving the timeliness and accuracy of risk management.
[0023] It should be understood that the description in the Summary of the Invention is not intended to limit the key or essential features of the embodiments of the present invention, nor is it intended to restrict the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0024] The above and other features, advantages, and aspects of the various embodiments of the present invention will become more apparent from the accompanying drawings and the following detailed description. The drawings are provided for a better understanding of the invention and are not intended to limit the invention. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:
[0025] Figure 1 A flowchart of a method for monitoring the dietary nutrition status of the elderly based on multi-source data fusion according to an embodiment of the present invention is shown;
[0026] Figure 2 A block diagram of a dietary nutrition status monitoring system for the elderly based on multi-source data fusion according to an embodiment of the present invention is shown.
[0027] Figure 3 A block diagram of an exemplary electronic device capable of implementing embodiments of the present invention is shown;
[0028] Figure 4 This illustrates the effect of an application layer terminal APP according to an embodiment of the present invention. Figure 1 . Detailed Implementation
[0029] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0030] Furthermore, the term "and / or" in this invention is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this invention generally indicates that the preceding and following related objects have an "or" relationship.
[0031] Figure 1 This diagram illustrates a flowchart of a method 100 for monitoring the dietary nutrition status of the elderly based on multi-source data fusion, as shown in an embodiment of the present invention. Figure 1 As shown, method 100 includes:
[0032] S110: Through computer vision, wearable devices, smart sensors and standardized medical interfaces, it collects and aggregates dietary images, multi-parameter physiological indicators, behavioral and environmental data, and structured health records to form a raw multi-source dataset.
[0033] Optionally, in some embodiments, the process of forming the original multi-source dataset specifically includes the following steps:
[0034] S111: Dietary images are preprocessed by a mobile terminal APP equipped with a ResNet152 model, uniformly sized to 224×224 pixels, randomly flipped / rotated, normalized, feature extracted, and matched with a recipe database to obtain preliminary dietary data including calorie, protein, fat, and carbohydrate intake, and data confidence labels are added; the data is transmitted via Bluetooth BLE5.0 or Wi-Fi 802.11b / g / n protocol, and a breakpoint resume mechanism is enabled.
[0035] S112: Physiological signals are collected by wearable devices and smart sensors that support IP67 waterproof rating, and combined with data from a smart body fat scale to obtain physiological indicators such as blood pressure, dynamic blood glucose, blood oxygen, heart rate, sleep quality, BMI, and muscle mass; the data is transmitted via Bluetooth BLE to a smartphone or indoor gateway, and then via Wi-Fi / 5G / 4G networks with AES-256 encryption.
[0036] S113: Behavioral and environmental signals are monitored by smart home sensors deployed in key areas to obtain auxiliary data such as daily activity levels, gait characteristics, and ambient temperature and humidity; all data is stored with end-to-end encryption and hierarchical access control.
[0037] S114: External medical records are transmitted via HL7 and FHIR standard interfaces and dedicated lines. The system automatically synchronizes and extracts chronic disease history, medication history, and ADL assessment scores from electronic medical records to obtain structured medical record data. Field-level encryption is used for storage, and a full access log is recorded. The update frequency is synchronized with the medical system.
[0038] It should be noted that, in this embodiment, image-based dietary intake recognition employs a mobile terminal APP equipped with computer vision (CV) algorithms. The food images are preprocessed using a ResNet152 model (uniform size 224×224 pixels, random flipping / rotation, normalization), visual features are extracted, and matched against a menu library (supporting ≥200 common elderly meal dishes). The system automatically calculates calorie, protein, fat, and carbohydrate intake. The APP interfaces with the transmission layer via Bluetooth BLE 5.0 and Wi-Fi 802.11b / g / n protocols, with a transmission latency of ≤3 seconds under standard network conditions. A breakpoint resume mechanism is enabled during network fluctuations to ensure data integrity. (Note: The system encourages users to photograph their main daily meals. For missing or poor-quality photos, estimations can be made using weighing data from smart tableware or dietary patterns from previous days, and the data confidence level (high / medium / low) will be labeled in the system. Data with medium or low confidence levels will have reduced weight in AI analysis.)
[0039] Computer vision (CV) algorithms: Image recognition algorithms deployed on mobile terminal apps, used to identify food types, portion sizes, and calculate nutritional components from food photos.
[0040] In this process, the food images are adjusted to a fixed-size standard grid by the geometric correction unit built into the mobile terminal. The standard grid images then enter the pixel normalization unit, where pixel values are redistributed according to a pre-stored statistical model to meet the numerical stability requirements of the processing. The normalized images are input into a multi-layer feature abstraction structure. The first level of this structure, the spatial perception layer, performs dense scanning of the image to generate a set of primary feature maps that preserve the original object edges and texture relationships. This set of primary feature maps is then fed into the second level, the semantic abstraction layer. The semantic abstraction layer performs multiple iterative nonlinear transformations and information compression on the primary feature maps, gradually filtering out spatial details, and finally outputting a fixed-dimensional feature vector that represents the global semantic content of the image. The feature vector is compared with a locally pre-stored food feature library. Each food template in the feature library stores a baseline feature vector generated during the training phase. The core of the comparison process is to calculate the cosine angle distance between the input feature vector and each template vector. This distance value is obtained through a... The coprocessor calculates the ratio of the dot product of two vectors to the product of their respective magnitudes. The result ranges from -1 to 1; the closer the value is to 1, the more consistent the directions of the two vectors, meaning the more similar the visual content. The system selects three templates with a cosine distance closest to 1 as candidates. The system presets a high-confidence threshold. If the cosine distance value of the best candidate is higher than this threshold, the recognition is directly determined to be successful, and the dish identifier is output. The distance value is then mapped to a "high" confidence level using a linear piecewise function. If the distance value of the best candidate is lower than the high threshold but higher than a medium threshold, the dish identifier is also output, but mapped to a "medium" confidence level. If the distance values of all candidates are lower than the medium threshold, the system determines that the recognition is unreliable and outputs a "low" confidence identifier. The recognized dish identifier and its confidence level are input into the nutrition component mapping unit. This unit first extracts the corresponding calories, protein, fat, and carbohydrate content per 100 grams from the standard nutrition component database based on the dish identifier, as basic nutritional data. Simultaneously, the system embeds a confidence-weight conversion table, which maps the three levels of "high, medium, and low" to three different weight coefficients. Finally, each value in the basic nutrition data is multiplied by this weight coefficient to generate final dietary nutrition data with calibration properties and confidence indicators, which is used for subsequent transmission and multi-source data fusion analysis.
[0041] ResNet152 model: A deep convolutional neural network model used to extract visual features from dietary images, possessing strong feature representation capabilities. 224×224 pixels: The standardized size of dietary photos, providing a uniform input specification for the ResNet152 model.
[0042] Random flipping / rotation: An enhancement operation in dietary image preprocessing used to improve the model's generalization ability.
[0043] Normalization: A standardized operation in dietary image preprocessing that makes the image pixel values conform to the distribution requirements during model training.
[0044] Calories / Protein / Fat / Carbohydrates: Core nutrient indicators automatically calculated by the system, reflecting dietary nutrient intake.
[0045] Bluetooth BLE 5.0: A short-range communication protocol between the mobile app and the transmission layer, ensuring short-range transmission of dietary recognition data.
[0046] Wi-Fi 802.11b / g / n: A wireless network communication protocol between mobile apps and the transport layer, supporting long-distance data transmission.
[0047] Resume interrupted transmission mechanism: A transmission guarantee mechanism to cope with network fluctuations, which can resume transmission after interruption and avoid data loss.
[0048] Smart tableware: Tableware with weighing function, used to supplement or calibrate dietary intake data.
[0049] Data confidence level (high / medium / low): The system's rating of the reliability of dietary data. Data with medium and low confidence levels will have reduced weight in AI analysis.
[0050] Multi-parameter physiological indicator sensing: Real-time data collection of blood pressure, dynamic blood glucose, blood oxygen, heart rate, and sleep quality is achieved through wearable devices and smart sensors. Combined with a smart body fat scale, BMI and muscle mass changes are obtained. The wearable device features a low-power design, supports IP67 waterproof rating, and has a battery life of ≥3 days. It connects to a smartphone or indoor gateway via Bluetooth BLE, and the gateway then uploads the data to the cloud via Wi-Fi / 5G / 4G networks. In areas with poor wireless network coverage, a LoRa gateway can be used for relay transmission. The device samples at a frequency of ≥1 time / hour and supports offline storage (capacity ≥7 days). Data is transmitted after AES-256 encryption.
[0051] Wearable devices: Smart devices worn by the elderly to collect physiological data in real time.
[0052] Smart sensors: Sensing elements integrated into wearable devices or home environments to collect physiological, behavioral, and environmental data.
[0053] Blood pressure / dynamic blood glucose / blood oxygen / heart rate / sleep quality: core physiological indicators collected by the system, reflecting the health status of the elderly.
[0054] BMI (Body Mass Index): A health assessment indicator calculated based on height and weight, used for correlation analysis of nutritional status.
[0055] Muscle mass: an indicator reflecting muscle reserves in older adults, providing a reference for nutritional interventions.
[0056] IP67 waterproof rating: The waterproof and dustproof standard for wearable devices, ensuring reliability for daily use.
[0057] Smartphone / Indoor Gateway: A data relay device for wearable devices, enabling data transmission from the device to the cloud.
[0058] 5G / 4G network: A wide area network protocol for communication between the gateway and the cloud, supporting remote data transmission.
[0059] LoRa gateway: A relay device in areas with poor wireless network coverage, extending the data transmission distance.
[0060] AES-256 encryption: An encryption algorithm for data transmission and storage, ensuring the security of sensitive data such as physiological and dietary information.
[0061] Behavioral and Assistive Data Monitoring: Smart home sensors are used to monitor the daily activity levels (energy consumption data), gait characteristics, and environmental temperature and humidity of the elderly. Sensors are deployed in key areas and interface with the transmission layer via RS485 serial ports. Transmission latency is ≤5 seconds in a standard network environment, and sampling frequency is once every 2 hours. For monitoring related to excretion and metabolism, written informed consent from the user is required, and ethical review requirements are strictly followed, employing end-to-end encrypted storage and tiered access control.
[0062] Smart home sensors: Environmental and behavioral sensing devices deployed in homes or elderly care facilities.
[0063] Daily activity level (energy consumption data): A quantitative indicator reflecting the amount of exercise of older adults, used for nutritional needs matching analysis.
[0064] Gait characteristics: Data related to the walking posture of the elderly to assist in the assessment of health status.
[0065] Ambient temperature / humidity: key parameters of the living environment for the elderly, providing a reference for the analysis of factors affecting nutritional status.
[0066] RS485 serial port: A wired communication protocol between smart home sensors and the transmission layer, ensuring stable data transmission.
[0067] Excretion and metabolic monitoring: Health monitoring related to excretion in the elderly requires the user's written informed consent.
[0068] Ethical review: A compliance review process for privacy-sensitive monitoring projects to protect user rights.
[0069] End-to-end encrypted storage: This method protects data by encrypting it throughout the entire process from collection to storage, preventing privacy leaks.
[0070] Tiered access control: A tiered access control mechanism for sensitive monitoring data to restrict the scope of data access.
[0071] Multidimensional Health Record Integration: Through HL7 and FHIR standard interfaces and dedicated line transmission, electronic medical records (EMR) are automatically synchronized. This extracts the elderly person's chronic disease history, medication history (such as drugs affecting nutrient absorption), and ADL (Activities of Daily Living) assessment scores, enabling data interoperability with elderly care facility clinics and community hospital systems. Medical data is stored with field-level encryption, access logs are recorded throughout the process, and record updates are synchronized with the medical system.
[0072] HL7 / FHIR Standard Interface: A standardized interface for data exchange in medical systems, enabling cross-system synchronization of health records.
[0073] Dedicated line transmission: A dedicated transmission channel for medical data, enhancing the security of electronic medical record (EMR) transmission.
[0074] Electronic Medical Records (EMR): Health records of the elderly stored in the healthcare system, including medical history, medication history, etc.
[0075] Chronic disease history: Records of previously diagnosed chronic diseases in the elderly provide a basis for nutritional matching.
[0076] Medication history: Current and past medication records of the elderly, with a focus on drugs that affect nutrient absorption.
[0077] ADL (Activities of Daily Living) assessment score: an indicator that quantifies the elderly’s ability to live independently and assists in the development of nutritional care plans.
[0078] Field-level encrypted storage: A refined encryption method for medical data, with separate encryption for key and sensitive fields.
[0079] Access Log: A log file that records medical data access behavior to ensure that data operations are traceable.
[0080] Synchronized with the medical system: Health records are updated at the same frequency as the medical system to ensure data accuracy.
[0081] S120: Clean, standardize, and format the raw multi-source data, and align it according to a unified time window to generate a standardized dataset for model training and inference.
[0082] Optionally, in some embodiments, the process of generating a standardized dataset for model training and inference specifically includes the following steps:
[0083] S121: After the original multi-source data is uploaded through the transport layer, it is synchronized to the data layer via HTTP / HTTPS protocol for data cleaning. The 3σ criterion is used to identify and remove outliers in time-series physiological and behavioral data. For data points with a missing rate of no more than 10%, linear interpolation is used to fill in the missing data based on the values of the adjacent time points to obtain cleaned and regularized data.
[0084] S122: The cleaned and standardized data enters the data standardization process. Based on a predefined data dictionary, the units of data from different sources are converted into the system's internal standard units, and text and category code data are formatted and encoded to generate standardized format data.
[0085] S123: During the system development phase, the standardized format data is divided into two independent subsets in a 9:1 ratio according to time sequence: the first 90% of the data forms the training set for subsequent model parameter learning; the last 10% of the data forms the test set for evaluating model performance; during the actual system operation phase, the newly generated real-time standardized format data is used as the inference analysis dataset, thereby ultimately generating a standardized dataset for model training or inference.
[0086] It should be noted that, in this embodiment, after the collected data is uploaded via the transport layer, it is synchronized to the data layer via HTTP / HTTPS protocol, where the following preprocessing operations are performed:
[0087] Data cleaning: Abnormal data such as sensor malfunctions and manual input errors are removed using the 3σ criterion. When the missing rate is ≤10%, it is supplemented by interpolation. For transmission interruptions caused by network fluctuations and device offline, breakpoint resume and data retransmission mechanisms are enabled.
[0088] For time-series data missing points identified during the data cleaning step with a missing rate of no more than 10%, the following sequential processing is performed to generate imputation values:
[0089] The system first scans the cleaned and organized data in time-stamp order to locate all data points marked as missing.
[0090] For each missing data point, the system searches forward and backward in its time series to locate the nearest valid data point before the missing point (called the predecessor point) and the nearest valid data point after the missing point (called the successor point). The values and timestamps of the two valid data points are extracted to form a temporary data pair for calculation.
[0091] The system establishes a linear numerical transition relationship based on the differences between the values of the predecessor and successor points and their corresponding timestamps. Specifically, this transition relationship stipulates that the ratio of the difference between the estimated value of the missing point and the value of the predecessor point to the difference between the values of the missing point and the successor point should be equal to the ratio of the difference between the timestamp of the missing point and the timestamp of the predecessor point to the difference between the timestamp of the successor point and the timestamp of the missing point. Based on this relationship, a unique estimated value for the missing point can be derived.
[0092] The estimated values derived from the established transition relationships are directly assigned to the corresponding missing data point locations, generating imputation values for those data points. After all missing data points have been processed, the generated imputation values are written back into the original regularized data, replacing the original "missing" markers, thus obtaining a complete, cleaned, and regularized dataset without data gaps for use in standardized processes.
[0093] Data standardization: unify data formats (such as units for physiological indicators and units for dietary intake) and build a data dictionary to ensure consistency.
[0094] Model training data partitioning (development phase): In the early stages of model building, the historically accumulated labeled data is divided into training set and test set in a 9:1 ratio. The training set is used for model training, and the test set is used for performance verification. When the system is actually running, real-time data is directly used for inference analysis without the need for dynamic partitioning.
[0095] HTTP / HTTPS protocol: A communication protocol that synchronizes data from the transport layer to the data layer, ensuring the standardization and security of data transmission.
[0096] Data layer: The functional layer in the system responsible for data storage and preprocessing.
[0097] 3σ Criterion: A standard for identifying outliers in data cleaning, removing outliers that deviate from the mean by more than three standard deviations.
[0098] Sensor malfunction / manual data entry error: These are the main causes of data anomalies and are eliminated through data cleaning.
[0099] Missing rate ≤10%: This is the acceptable range for missing data. Data exceeding this range must be supplemented using other methods.
[0100] Interpolation: A method for supplementing missing data by extrapolating missing values from existing data to ensure data integrity.
[0101] Data retransmission mechanism: A mechanism to supplement transmitted data after network recovery in response to transmission loss caused by network interruption.
[0102] Data standardization: The process of unifying the format and units of data from different sources.
[0103] Physiological indicator units / dietary intake units: the core object of data standardization, ensuring that data can be compared and integrated.
[0104] Data dictionary: A reference document that records the data format, units, and meanings to ensure data consistency.
[0105] 9:1 ratio: The ratio of training set to test set during the model development phase, balancing the model training effect and generalization ability verification.
[0106] Training set: A dataset used for learning model parameters, containing labeled nutrient status tags.
[0107] Test set: An independent dataset used to verify the model's generalization ability; it was not used in model training.
[0108] Model training: The process of optimizing model parameters through data learning to improve the accuracy of model prediction and analysis.
[0109] Performance validation: The process of evaluating the model's prediction accuracy, early warning accuracy, and other metrics using a test set.
[0110] Real-time data: New data collected during system runtime that is not used in model training.
[0111] Inference analysis: The process by which the model uses real-time data to predict nutritional status and assess risks.
[0112] S130: The feature pyramid network, temporal Transformer, and pre-trained language model are used to extract image, temporal, and structured features from the standardized dataset, respectively; multimodal feature fusion is performed through causal cross-attention mechanism; and finally, the LSTM+Transformer hybrid model is used to process the data and output the prediction results and risk probability of future nutritional status.
[0113] Optionally, in some embodiments, the process of outputting future nutritional status predictions and risk probabilities specifically includes the following steps:
[0114] S131: 30 days of standardized dietary image data, temporal physiological behavior data, and expanded structured archive data are processed separately by parallel multi-scale visual feature extraction networks, temporal feature extraction networks, and text feature encoding networks to obtain corresponding image feature sequences, temporal feature sequences, and structured feature sequences.
[0115] S132: Three feature sequences are input into a fusion computation unit based on temporal causal constraints. This unit uses each time step feature in the temporal feature sequence as a query benchmark, and performs correlation calculation and weighted fusion only with image feature fragments and structured features at the same or earlier time steps in chronological order, generating an intermediate fusion feature set containing all historical causal correlation pairs.
[0116] S133: The intermediate fusion feature set is aggregated and compressed in the time dimension and regularized into a daily fusion feature sequence that strictly corresponds to the original 30-day time window.
[0117] S134: The daily fusion feature sequence is input into a stacked prediction structure containing temporal memory units and feature interaction layers. After multi-level nonlinear transformations within it, the final output is the quantified value of nutrient intake matching and the probability value of nutritional risk for a specified number of days in the future.
[0118] It should be noted that, in this embodiment, AI intelligent analysis serves as the system's logical hub, integrating multimodal data fusion and dynamic evaluation functions. It achieves accurate analysis based on deep learning algorithms. The specific technical solution is as follows:
[0119] 3.1 Multimodal and multiscale feature extraction: Feature pyramid network (FPN) is used to process multi-source heterogeneous data, adapt to the feature expression needs of different types of data, and ensure alignment with the 30-day time window.
[0120] 3.1.1 Extract image features (dietary photos): Construct an image sequence of length 30, strictly aligned with a 30-day time window, based on the principle of one representative dietary photo per day. For each individual photo... The pixel-level food photos are first processed by a Feature Pyramid Network (FPN) to extract multi-scale visual features, generating... Feature maps at three resolutions are then processed using the core feature encoding method of the Vision Transformer (ViT).
[0121] Patch splitting: Meal photos The pixels are cut into uniform, non-overlapping blocks of a fixed size, resulting in a total of An image patch.
[0122] Linear embedding: Each The (RGB three-channel) image patch is flattened into a 768-dimensional vector, and then mapped to a 128-dimensional feature vector (i.e., image token) through a learnable linear transformation matrix, finally resulting in 196 image tokens with a dimension of 128.
[0123] Global Token Concatenation: An additional learnable 128-dimensional global token (Class Token) is introduced. This token is used to aggregate the feature information of all image patches, ultimately generating 197 token features (dimensions) for a single image. ).
[0124] Sequence Construction: The dietary photos from the past 30 days are arranged chronologically to form a sequence of 30 individual image features, which are then used to generate an image feature tensor. (where the first dimension) The first dimension represents a 30-day time step, the second dimension (197) represents the total number of tokens in a single image, and the third dimension (128) represents the feature dimension of each token.
[0125] The image feature tensor of the daily meal photo sequence, where Represents the real number space; : The value of the first dimension of the image feature tensor, representing a 30-day time step; FPN (Feature Pyramid Network): A neural network model used to extract multi-scale visual features from food photos, generating feature maps at different resolutions; Linear transformation matrix: A learnable parameter matrix used to map 768-dimensional patch vectors to 128-dimensional feature vectors; Global Token (Class Token): An additional learnable 128-dimensional token used to aggregate feature information from all image patches.
[0126] 3.1.2 Extracting Temporal Features (Physiological / Behavioral Data): Input a 30-day continuous time-series data sequence and extract time-dependent features using a temporal Transformer. This Transformer contains four layers of self-attention and feedforward neural networks, with each layer having a multi-head attention mechanism. Hidden layer dimensions Output time series feature vector .
[0127] : Feature vector of 30-day time series data, where Represents the real number space; Temporal Transformer: A transformer model used to extract time-dependent features from physiological / behavioral temporal data.
[0128] 3.1.3 Extracting structured features (health records / assessment indicators): Text-based indicators are encoded using the DistilBERT model, and standardized feature vectors are obtained after mean pooling and linear transformation. Expanded through broadcast mechanism Aligned with the time dimension of temporal features and image features.
[0129] DistilBERT model: A lightweight pre-trained language model for encoding text data such as health records / assessment metrics; Mean pooling: A pooling operation performed on the output features of the DistilBERT model to reduce dimensionality and aggregate global text features; : The initial feature vector after encoding structured text data, where Represents the real number space; broadcast mechanism: will Expanding the feature vector of dimension to This operation enables alignment with the temporal dimension of time-series / image features; : The expanded structured feature vector is adapted to the dimensional specifications of a 30-day time window.
[0130] 3.2 Causal Cross-Attention (CC Attention) Mechanism: Based on the causal logic that "future states are only influenced by historical factors," CC Attention limits the time frame of attention computation to avoid future information leakage while dynamically adapting to individual differences. For nutritional status analysis of the elderly, their nutritional status at a given moment is only related to previous dietary intake, physiological changes, and health baseline. By constraining the temporal matching relationship between attention queries and key values, it ensures that feature fusion conforms to real-world causal laws.
[0131] CC Attention uses temporal feature vectors ( For historical time steps, ) is the query (Q), which corresponds to the spatial feature tokens of historical and earlier time steps. and structured features Let q be the key (K) and value (V), where q is a non-negative integer representing the maximum number of time steps to look back from the current time step i, i.e., the maximum range of historical information for querying Q; p is a variable representing the time step of key K; the fused features are calculated using the standard cross-attention formula:
[0132]
[0133] in, The learnable parameter matrix (all dimensions are...) ), Given a key dimension of (256), output the fused feature vector. (465 represents the number of causal pair combinations for 30-day time series data:) To adapt to the daily granularity of subsequent prediction models, [the following was implemented]. Daily average pooling is performed to aggregate the 30-step sequence features. This reduces computational complexity while preserving core daily information. For specific disease populations, the weights of key features are increased through an attention mechanism (e.g., the association weight between sugar intake and blood glucose fluctuations in diabetic patients, and the association weight between sodium intake and blood pressure changes in hypertensive patients); for dietary data with low to medium confidence levels, the feature weights are automatically reduced to minimize their impact on the analysis results.
[0134] CC Attention (Causal Cross-Attention): A causal cross-attention fusion mechanism designed based on the causal logic that future states are only affected by historical factors. It is used to fuse multimodal features and avoid the leakage of future information.
[0135] : as the temporal feature vector of attention query (Q), where For historical time steps; : Historical time step marker, representing a specific historical moment within a 30-day time window; : The range of values for the historical time step, where 0 represents the current time step and 29 represents the earliest historical time step within a 30-day window; Q (query): The query vector in the attention mechanism, composed of the time-series feature vector. act as; : Image feature tokens serving as attention keys (K), corresponding to dietary photo features from historical and earlier time steps; Constrain the time step range of image feature tokens to ensure that only image features from the historical and current time steps are used, which conforms to causal logic; : As the structured feature vector of the attention value (V), it corresponds to Simultaneous time-step health record / assessment indicator features; K (key): the key vector in the attention mechanism, composed of image feature tokens. V (value): The value vector in the attention mechanism, composed of structured feature vectors. Act as; Attention Standard cross-attention calculation function, used for fusion. Vectors yield fused features; softmax: a normalized exponential function used to convert attention scores into weight values in the 0-1 range, thus achieving attention weight allocation; Query vector ( The learnable parameter matrix of Q is used to perform a linear transformation on Q; : The learnable parameter matrix of the key vector (K), used to perform linear transformations on K; : The learnable parameter matrix of the value vector (V) used to perform linear transformations on V; The parameter matrix has a dimension of 256 for both input and output. Key vector The transpose operation after transformation is used to perform an inner product calculation with the transformed result of the query vector; The square root of the key dimension is used as a scaling factor for the attention score to prevent the inner product from becoming too large and causing the softmax gradient to vanish. The feature dimension of the key vector (K) has a value of 256; : The initial fused feature vector output by CC Attention, where Represents the real number space; 465: The value of the first dimension of the initial fused feature vector, which is the number of causal pairing combinations of 30-day time series data, calculated as follows: ; ( First dimension): The feature dimension of the initial fused feature vector is consistent with the input feature dimension of the attention mechanism; Average pooling: ... The aggregation operation performed on a daily basis compresses the 465-step fusion feature into 30 steps, which is suitable for daily prediction granularity. : The 30-step sequence feature vector obtained after average pooling is the input of the subsequent prediction model; 30 (first dimension of H): the time step dimension of the sequence feature vector, corresponding to a 30-day time window, which is consistent with the daily prediction granularity; 256 (second dimension of H): the feature dimension of the sequence feature vector, which is consistent with the initial fused feature dimension.
[0136] 3.3 Time-series prediction model (LSTM + Transformer): A hybrid LSTM and Transformer model is constructed for short-term nutritional status prediction. The model contains two LSTM layers (hidden layer dimension...). dropout rate ) and 2-layer Transformer encoder (multi-head attention head count) Feedforward network dimension LSTM is responsible for capturing long-term temporal dependencies, and Transformer enhances key feature interactions; the input is the pooled 30-step sequence features. The LSTM layer uses the tanh activation function, the Transformer feedforward network uses the ReLU activation function, and the output layer uses the sigmoid activation function. The loss function is Mean Absolute Error (MAE), as shown in the following formula:
[0137]
[0138] The input time series prediction model's sequence feature vector is obtained by average pooling after causal cross-attention fusion. Mean absolute error loss function, used to calculate the deviation between predicted and true values, guiding model training and optimization; : True nutritional status values, which are the labeled data used for model training (such as the matching degree of actual nutrient intake, nutritional risk values, etc.). The predicted nutritional status output by the model; : Number of samples involved in MAE calculation; : A summation function that accumulates the absolute errors of n samples; The absolute error between the true value and the predicted value of a single sample is used to avoid the cancellation of positive and negative errors. The mean calculation coefficient converts the accumulated absolute error into an average error, reflecting the overall prediction accuracy of the model.
[0139] 3.4 Model Training:
[0140] Training data source: Multi-source labeled data (including daily dietary records, physiological indicators, health records, and nutritional assessment results) from 1000+ elderly people over three consecutive months were integrated and used for model training after being desensitized.
[0141] Training environment: Model training is performed using a 2×NVIDIA A100 GPU cluster to meet the needs of large-scale data processing.
[0142] Training parameters: Based on the PyTorch framework, using the AdamW optimizer, initial learning rate = 1e-4, batch size = 32, training epochs = 50, L1 regularization (coefficient = 1e-5) and exponential moving average (EMA) are introduced to suppress overfitting; the model performance is verified every 10 epochs using the test set, and an early stopping strategy is adopted (training is stopped if there is no performance improvement for 5 consecutive epochs).
[0143] Deployment optimization: After training, the model is optimized through model quantization (INT8) and pruning (retaining 70% of the core parameters) to adapt to edge computing devices or ordinary servers.
[0144] Input / output logic:
[0145] Input: Multimodal fusion features from 30 consecutive days (30×197×128 image features, 30×1×256 temporal features, and 30×1×256 structured features), which are fused and pooled to obtain a 30×256 input sequence.
[0146] Output: Nutritional status prediction results for the next 1-3 days (including the matching degree of nutrient intake such as protein, carbohydrates, fat, and vitamins) and nutritional risk warning probability (0-1 range).
[0147] 1000+: The number of elderly samples participating in model training, ensuring the diversity and representativeness of the training data; 3 months: The time span covered by the training data, including multi-source data of the elderly for 90 consecutive days.
[0148] Multi-source labeled data: The training data type collection covers daily dietary records, physiological indicators, health records, and nutritional assessment results, all labeled with corresponding nutritional status tags; De-identification processing: Privacy protection operations are performed on the training data to remove sensitive information such as elderly identification marks, in compliance with data compliance requirements.
[0149] 2×NVIDIA A100 GPU Cluster: The hardware environment used for model training consists of two NVIDIA A100 graphics cards forming a cluster, meeting the needs of large-scale data parallel training. PyTorch Framework: The deep learning framework on which model training is based, providing complete interfaces for network construction, training, and optimization.AdamW Optimizer: The optimizer used for model training. It adds weight decay to Adam to improve optimization stability and generalization ability; Initial Learning Rate = 1e-4: The initial learning rate for model training (0.0001), controlling the parameter update step size; Batch Size = 32: The number of samples used for each model parameter update, balancing training efficiency and memory usage; Training Epochs = 50: The maximum number of times the model fully traverses the training set, ensuring the model fully learns data patterns; L1 Regularization: A regularization method used to suppress overfitting. It reduces model complexity by adding an L1 norm penalty term to the parameters; 1e-5: The coefficient of L1 regularization (0.00001), controlling the strength of the regularization penalty; Exponential Moving Average (EMA): By applying an exponentially weighted average to the model parameters, it smooths the parameter update process, suppresses overfitting, and improves model stability; Every 10 Epochs: The interval between model performance validation epochs. This means evaluating the model's performance on a test set every 10 training epochs. The test set is an independent subset of data used to validate model performance; it is not involved in model training and objectively reflects generalization ability. Early stopping strategy: This is the termination strategy for model training. Training stops when there is no performance improvement on the test set for 5 consecutive epochs to avoid overfitting. No performance improvement for 5 consecutive epochs: This is the trigger condition for the early stopping strategy, preventing the model from overfitting the training data in the later stages of training. Model quantization (INT8): This is a model deployment optimization method that converts model parameters from floating-point to 8-bit integers, reducing model storage and inference computation. Pruning (retaining 70% of core parameters): This is a model deployment optimization method that removes 30% of redundant parameters, retaining only 70% of the core parameters crucial to the prediction results, adapting to the computing power of edge devices. 70%: The percentage of core parameters retained after pruning, balancing model lightweighting and prediction accuracy. Edge computing devices: One of the target hardware devices for model deployment (such as NVIDIA). Jetson Xavier NX), adapted for local low-latency inference scenarios in elderly care institutions; 30×197×128 (image features): the image feature dimension of the model input, corresponding to the token feature sequence of 30 days of dietary photos; 30×1×256 (temporal features): the temporal feature dimension of the model input, corresponding to the temporal feature sequence of 30 days of physiological / behavioral data; 30×1×256 (structured features): the structured feature dimension of the model input, corresponding to the encoded feature sequence of 30 days of health records / assessment indicators; 30×256 (input sequence): the final input dimension after fusion and pooling of multimodal features, adapted to the input specifications of temporal prediction models; Nutrient intake matching degree: one of the core prediction indicators output by the model, covering categories such as protein, carbohydrates, fat, and vitamins, reflecting the degree of matching between actual intake and individual baseline; Nutritional risk warning probability: one of the core prediction indicators output by the model, with a value range of 0-1, where 0 represents no risk and 1 represents extremely high risk, quantifying the risk of nutritional imbalance in the elderly.
[0150] S140: It adopts a distributed database architecture that combines local and cloud computing to classify and persistently store real-time high-frequency multi-source data, historical archived multi-source data, future nutritional status prediction results and risk probabilities, and encrypted medical records.
[0151] It should be noted that, in this embodiment, data storage adopts a distributed storage method using a local database and a cloud database.
[0152] Local database: Deploy MySQL 8.0 or above to store high-frequency data collected in real time (such as physiological monitoring and dietary intake data), and automatically back up daily.
[0153] Cloud database: Deploy MongoDB 4.0 or above to store historical data, model training results and health records. Use off-site multi-copy backup, and additionally encrypt and store sensitive medical data.
[0154] S150: Combining dietary guidelines, individual health records, and historical data, and integrating future nutritional status predictions and risk probabilities, it uses a multi-factor nonlinear model to calculate and dynamically update individualized baselines for nutrient intake and physiological indicators.
[0155] It should be noted that, in this embodiment, real-time risk assessment and automated threshold warning are achieved based on AI intelligent analysis results, and the specific scheme is as follows.
[0156] Based on the "Dietary Guidelines for the Elderly in China," individual health records, and historical data, a dynamic baseline is constructed, and a multi-factor nonlinear model is used to optimize the calculation logic.
[0157]
[0158]
[0159] AI automatically adjusts the warning threshold every 7 days based on the basic vital signs of the elderly.
[0160] B: Dynamic baseline, a personalized reference threshold for assessing the nutritional status of the elderly, dynamically calculated by integrating multiple dimensions of factors; α: Weighting coefficient of the guideline recommendation standard (S), used to quantify the influence of the recommended values in the "Dietary Guidelines for the Elderly in China"; S: Guideline recommendation standard, based on the basic reference values of nutrient intake, physiological indicators, etc. determined in the "Dietary Guidelines for the Elderly in China"; β: Weighting coefficient of the historical health data mean (H), used to quantify the influence of the elderly's own historical health data; γ: Mean of historical health data for the elderly, referring to the average level of nutritional and physiological indicators of the elderly over a past period; γ: Weighting coefficient of chronic disease history parameter (D), used to quantify the impact of chronic diseases on nutritional needs; D: Chronic disease history parameter, a quantitative parameter set according to whether the elderly have chronic diseases such as diabetes and hypertension (e.g., 1 for those with the disease, 0 for those without, or graded according to the severity of the disease); δ: Weighting coefficient of medication impact coefficient (M), used to quantify the impact of drugs on nutrient absorption / metabolism; M: Medication impact coefficient, set according to the type and dosage of drugs that affect nutrient absorption taken by the elderly. The quantification coefficients; α+β+γ+δ=1: the constraint condition for the weight coefficients, ensuring that the sum of the weights of each dimension factor is 1, guaranteeing the rationality of the baseline calculation; HTTP / HTTPS interface: the communication interface protocol for pushing warning information from the system's bottom layer to the application layer, ensuring data transmission security and standardization; multi-role interaction module: a functional module responsible for pushing warning information to different terminals, covering roles such as medical staff, family members, and caregivers; local caching and resending mechanism: an abnormal handling mechanism to deal with network interruptions, first caching the warning information locally, and automatically resending it after the network is restored, avoiding the loss of warning information.
[0161] S160: It compares the predicted results of future nutritional status with risk probabilities, real-time multi-source data and dynamically updated individualized physiological indicator baselines, identifies abnormal risk levels according to preset multi-level collaborative judgment rules, and triggers real-time early warning information push to multiple roles.
[0162] It should be noted that, in this embodiment, the abnormal collaborative identification and early warning system makes the following collaborative judgment based on the predicted value of nutrient intake matching degree for the next 1-3 days and the probability of nutritional risk warning output by the AI intelligent analysis module, combined with real-time physiological indicators.
[0163] General warning (yellow): The single nutrient intake matching degree deviates from the baseline by ±20%, there are no abnormal physiological indicators, and the risk probability is <0.3.
[0164] Important warning (orange): The intake matching degree of ≥2 nutrients deviates from the baseline by ±30%, or one physiological indicator is abnormal, and the risk probability is 0.3≤0.7.
[0165] Emergency Warning (Red): Nutrient intake matching degree deviates from the baseline by ±50%, or ≥2 physiological indicators are continuously abnormal, and the risk probability is ≥0.7.
[0166] Warning information is pushed to the application layer via an internal HTTP / HTTPS interface, and then pushed to the terminals of medical staff, family members or caregivers in real time by the multi-role interaction module, with a push delay of ≤10 seconds; in case of abnormal situations such as network interruption, a local caching and resending mechanism is enabled.
[0167] This embodiment utilizes computer vision technology, wearable devices, smart sensors, and standardized medical interfaces to comprehensively collect and aggregate dietary images, physiological indicators, behavioral environmental data, and health records, forming a rich multi-source dataset. Data cleaning, standardization, and formatting ensure data consistency and alignment within a unified time window, providing a standardized dataset for model training and inference. Feature pyramid networks, temporal Transformers, and pre-trained language models are used to extract image, temporal, and structured features respectively, and a causal cross-attention mechanism is employed to fuse multimodal features, improving the model's interpretability and prediction accuracy. The LSTM+Transformer hybrid model processing enables... Taking into account both time series and complex structured data, the system outputs more accurate predictions of future nutritional status and risk probabilities. Employing a distributed database architecture combining local and cloud-based systems, it achieves the classification and persistent storage of real-time high-frequency data, historical archived data, model outputs, and encrypted medical records, improving data security and access efficiency. By integrating dietary guidelines, individual health records, and historical data, it calculates and dynamically updates individualized nutrient intake and physiological indicator baselines through a multi-factor nonlinear model, making predictions more personalized and accurate. By comparing AI predictions, real-time data, and dynamic baselines, and based on preset multi-level collaborative judgment rules, it identifies abnormal risk levels and triggers real-time early warning information pushes, improving the timeliness and accuracy of risk management.
[0168] This embodiment proposes a multimodal, multi-scale feature extraction technique. Addressing the heterogeneity of dietary images, physiological time-series data, and structured health record data, it fully mines the core information of different data types through multi-scale feature extraction and standardized coding, providing high-quality feature support for fusion analysis. Simultaneously, it proposes a causal cross-attention (CCAttention) fusion mechanism. Based on the causal logic that future states are only influenced by historical factors, it limits the time range of attention calculation to avoid future information leakage. For the nutritional status analysis scenario of the elderly, it clarifies that the nutritional status at a certain moment is only related to previous dietary intake, physiological changes, and health foundation. By constraining the temporal matching relationship between attention queries and key values, it ensures that feature fusion conforms to real-world causal laws. This embodiment achieves dynamic adaptation of the fusion mechanism to individual differences. Through dynamic adjustment of attention weights, it strengthens the influence of individual characteristics such as chronic diseases and medication history on nutritional analysis, improving the system's adaptability to special groups. It constructs a real-time, accurate monitoring and early warning system, enabling short-term prediction and risk assessment of the nutritional status of the elderly based on the above technologies. This provides timely early warning information for caregivers and families, assisting in personalized nutritional intervention decisions.
[0169] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, because according to the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to the present invention.
[0170] The above is an introduction to the method embodiments. The following system embodiments will further illustrate the solution of the present invention.
[0171] Figure 2 A block diagram of a multi-source data fusion-based dietary nutrition monitoring system 200 for the elderly, according to an embodiment of the present invention, is shown. A multi-dimensional health profile data pool for the elderly is constructed through intelligent terminal devices, and efficient data transfer is achieved through standardized interfaces. This system adapts to the needs of deploying applications in various scenarios such as elderly care institutions, communities, and homes. Figure 2 As shown, system 200 includes: a data acquisition module 210, a data preprocessing module 220, an AI intelligent analysis module 230, a data storage module 240, a risk assessment module 250, a system interface module 260, and a security and privacy protection module 270. These modules work together to achieve real-time monitoring and risk warning of the dietary nutrition status of the elderly.
[0172] The data acquisition module 210 is used to collect multi-dimensional heterogeneous data, including dietary image recognition and nutritional composition calculation through computer vision algorithms and ResNet152 models mounted on mobile terminal APP, physiological indicator data through wearable devices and smart sensors, behavioral and environmental data through smart home sensors, and structured data of electronic medical records through HL7 and FHIR standard interfaces. The data preprocessing module 220 collects data and uploads it through the transmission layer, then synchronizes it to the data layer through the HTTP / HTTPS protocol to perform preprocessing operations such as data cleaning, data standardization, and model training data partitioning.
[0173] The AI intelligent analysis module 230 includes a multimodal multi-scale feature extraction unit, a causal cross-attention fusion unit, and a time-series prediction unit. The multimodal multi-scale feature extraction unit, targeting the heterogeneous characteristics of dietary images, physiological time-series data, and structured health record data, extracts multi-scale visual features through a Feature Pyramid Network (FPN), extracts time-dependent features through a temporal Transformer, and encodes textual semantic features through a DistilBERT model. The causal cross-attention fusion unit (CC Attention) constrains the temporal matching relationship between attention queries and key values based on the causal logic that future states are only affected by historical factors, fusing only historical and current time-step data to avoid future information leakage. The time-series prediction unit employs a hybrid LSTM and Transformer model to predict short-term nutritional status based on the fused features.
[0174] Causal Cross-Attention Fusion Unit (CC Attention) uses temporal feature vectors For historical time steps, For query (Q), use image feature tokens corresponding to historical and earlier time steps. ( and structured features For the key (K) and value (V), use the standard cross-attention formula Calculate the fusion features and output the initial fusion feature vector. Then, by daily average pooling, we obtain... The sequence features of dimension 465, where 465 represents the number of causal pair combinations in the 30-day time series data, calculated as follows: .
[0175] The Causal Cross-Attention Fusion Unit (CC Attention) can dynamically adapt to individual differences. For patients with chronic diseases such as diabetes and hypertension, it automatically increases the association weight of key features. For dietary data with medium to low confidence, it automatically reduces the feature weight. The confidence level is labeled by the system as high, medium, or low based on image quality and data integrity.
[0176] The time series prediction model contains two LSTM layers (hidden layer dimension). dropout rate ) and 2-layer Transformer encoder (multi-head attention head count) Feedforward network dimension The LSTM layer uses the tanh activation function, the Transformer feedforward network uses the ReLU activation function, the output layer uses the sigmoid activation function, and the loss function is the mean absolute error (MAE), the formula of which is: The input is The sequence features of the dimension output the nutrient intake matching degree for the next 1-3 days and the probability of nutritional risk warning in the 0-1 range.
[0177] The data storage module 240 is used to store high-frequency data (such as physiological monitoring and dietary intake data) collected in real time through local databases and cloud databases; historical data, model training results and health records are backed up in multiple locations, and medical sensitive data is additionally encrypted and stored.
[0178] Risk assessment module 250 includes a dynamic baseline establishment unit and an anomaly collaborative identification unit. The dynamic baseline establishment unit is based on the "Dietary Guidelines for the Elderly in China," individual health records, and historical data, and adopts a multi-factor weighted model. A dynamic baseline is established, and the AI automatically corrects the warning threshold every 7 days; the anomaly collaborative identification unit triggers three levels of warnings—general warning (yellow), important warning (orange), and emergency warning (red)—based on the predicted value, risk probability, and real-time physiological indicators output by the AI module.
[0179] System interface module 260 includes internal design interfaces and external design interfaces.
[0180] The internal interfaces are designed as follows: Sensing layer and transmission layer interfaces: Utilizing RS485 serial port, Bluetooth BLE5.0, and Wi-Fi 802.11b / g / n protocols, seamless integration between the data acquisition terminal and the transmission module is achieved; a LoRa gateway serves as a supplementary relay interface, adapting to data transmission in remote areas. Transmission layer and data layer interfaces: Employing HTTP / HTTPS protocols, supporting encrypted data transmission, ensuring secure synchronization of multi-source data, and a transmission latency of ≤5 seconds. Data layer and application layer interfaces: Providing RESTful API interfaces, supporting application layer calls to data storage, fusion, and preprocessing functions, ensuring real-time feedback of assessment results and early warning information.
[0181] The external interfaces include: Medical system interface: supporting HL7, FHIR standards and dedicated line transmission, enabling data interoperability with physical examination systems and electronic medical record systems, sharing health records and nutritional status data, with full field-level encryption and access auditing. Elderly care service management system interface: providing a JSON format data exchange interface, supporting synchronization with personnel management systems and care record systems, facilitating the tracking of intervention implementation. Third-party service interface: reserving interfaces with fresh food delivery platforms and catering delivery institutions, supporting the integration of food procurement and delivery based on dietary recommendations.
[0182] The security and privacy protection module 270 is used to set up security and privacy protection; the compliance principle strictly follows the "Data Security Law of the People's Republic of China", "Personal Information Protection Law of the People's Republic of China" and "Administrative Measures for Network Security of Medical and Health Institutions" and other laws and regulations. All data collection, transmission, storage and use comply with compliance requirements, and regular compliance audits are conducted.
[0183] Key safeguards: Data transmission security: All data is transmitted using AES-256 encryption. Sensitive medical data is transmitted via dedicated lines to prevent leakage during transmission. Data storage security: Both local and cloud databases use encrypted storage, and medical data is encrypted at the field level; off-site multi-copy backups are implemented to prevent data loss. Access control: A tiered access control system is established. Access to medical data and private data requires multi-level approval, and access logs are recorded throughout the process to ensure traceability. User rights protection: Before collecting personal health data, users are clearly informed of the data's purpose, storage method, and usage period, and written informed consent is obtained; users have the right to query, correct, and delete their personal data, and the system provides convenient operation channels.
[0184] The system deployment and operating environment of this embodiment includes the following hardware deployment: the intelligent dietary weighing device is deployed on a flat surface with a power supply voltage of 220V±10%; the wearable device is adapted to the size of elderly people; the smart home sensor is deployed in key areas to ensure signal stability; and the interactive terminal screen size is ≥7 inches and supports touch operation.
[0185] Server and edge device configuration:
[0186] On-site deployment of elderly care facilities: NVIDIA Jetson Xavier NX edge computing devices are deployed to run lightweight AI models, enabling low-latency real-time analysis and preliminary early warning; paired with local servers (CPU ≥ Intel Xeon E5, memory ≥ 32GB, hard disk capacity ≥ 2TB) to store high-frequency real-time data.
[0187] Cloud-based: Utilizes elastic cloud servers, supporting on-demand scaling for in-depth analysis, model iteration, and long-term storage of historical data.
[0188] Gigabit Wi-Fi coverage is deployed in elderly care facilities, 5G / 4G networks are supported in home and community settings, and LoRa gateways are used to supplement coverage in remote communities to ensure data transmission stability.
[0189] Software runtime environment deployment:
[0190] Operating systems: Local servers support Windows Server 2016 and above, and Linux CentOS 7.0 and above; cloud servers support mainstream cloud operating systems; mobile apps support Android 8.0 and above, and iOS 12.0 and above, with an elderly mode (larger font, simplified operation); edge computing devices support Linux Ubuntu 18.04 and above.
[0191] Database: MySQL 8.0 and above (local), MongoDB 4.0 and above (cloud), supporting distributed storage and off-site backup, with encrypted storage module enabled for medical sensitive data.
[0192] Development environment: The backend uses Java (SpringBoot framework); the frontend uses Vue.js framework; the AI algorithm module supports Python 3.7 and above, and depends on TensorFlow and PyTorch frameworks. During the deployment phase, it is adapted to the inference environment of CPU / GPU / edge computing devices.
[0193] As shown in Tables 1 and 2, following strict inclusion and exclusion criteria, the five institutions selected for the experiment were: a municipal industrial and commercial health and wellness center, a nursing home in a certain city, a nursing home in a certain district, and a nursing home in a certain city. A total of 100 subjects were included in the experiment, with 29 males and 71 females, and a mean age of 78.90 ± 7.09 years. Figure 4 The example demonstrates the effect of the application layer terminal APP.
[0194] Table 1. Experimental Subject Data
[0195]
[0196] Table 2. Experimental Subject Data
[0197]
[0198] This embodiment offers more comprehensive feature extraction and significantly improves data value utilization: Multimodal, multi-scale feature extraction technology designs differentiated extraction schemes for different types of data, enabling multi-scale detail capture of image data, time-dependent mining of time-series data, and semantic encoding of structured data. Compared to traditional single-scale feature extraction, feature expression is more comprehensive, laying a solid foundation for fusion analysis. The fusion logic is more rigorous, and the analysis results are more reliable: The causal cross-attention fusion mechanism avoids future information leakage through causal constraints, ensuring that data fusion conforms to the objective laws of nutritional status evolution in the elderly. This solves the problem of existing technologies where the fusion process violates real-world logic, making the analysis results more consistent with actual conditions and significantly improving monitoring accuracy. It also offers stronger individual adaptability and outstanding personalized service capabilities: The fusion mechanism can dynamically adjust feature weights, automatically strengthening key features for special groups such as patients with chronic diseases and those using medication. The system, which tracks the impact of factors such as the correlation between sugar intake and blood glucose fluctuations in diabetic patients, can more accurately adapt to individual health differences among the elderly compared to existing general monitoring systems, providing personalized monitoring and early warning services. It offers more timely monitoring responses and higher practical value: the system achieves real-time collection, rapid fusion analysis, and automated early warning of multi-source data, predicting nutritional status for the next 1-3 days and issuing tiered warnings with a push delay of ≤10 seconds. Compared to traditional periodic monitoring methods, it can promptly identify nutritional risks, allowing time for intervention measures and effectively reducing health risks caused by malnutrition or nutritional imbalance. It also boasts good adaptability to multiple scenarios and flexible deployment: the system supports deployment in various scenarios such as elderly care institutions, communities, and homes. Through an edge computing and cloud-based collaborative architecture, it balances low-latency real-time analysis with large-scale data storage, while possessing a robust data security and privacy protection mechanism. This meets the actual application needs of elderly care services and has significant promotional value.
[0199] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the described module can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0200] According to embodiments of the present invention, the present invention also provides an electronic device and a readable storage medium.
[0201] Figure 3A schematic block diagram of an electronic device 300 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown in this invention, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0202] Electronic device 300 includes a computing unit 301, which can perform various appropriate actions and processes according to a computer program stored in ROM 302 or a computer program loaded into RAM 303 from storage unit 308. RAM 303 can also store various programs and data required for the operation of electronic device 300. The computing unit 301, ROM 302, and RAM 303 are interconnected via bus 304. I / O interface 305 is also connected to bus 304.
[0203] Multiple components in electronic device 300 are connected to I / O interface 305, including: input unit 306, such as keyboard, mouse, etc.; output unit 307, such as various types of displays, speakers, etc.; storage unit 308, such as disk, optical disk, etc.; and communication unit 309, such as network card, modem, wireless transceiver, etc. Communication unit 309 allows electronic device 300 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0204] The computing unit 301 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 301 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 301 performs the various methods and processes described above, such as the method for monitoring the dietary nutrition status of the elderly based on multi-source data fusion. For example, in some embodiments, the method for monitoring the dietary nutrition status of the elderly based on multi-source data fusion can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as storage unit 308. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 300 via ROM 302 and / or communication unit 309. When the computer program is loaded into RAM 303 and executed by the computing unit 301, one or more steps of the method for monitoring the dietary nutrition status of the elderly based on multi-source data fusion described above can be performed. Alternatively, in other embodiments, the computing unit 301 may be configured to perform the method for monitoring the dietary nutrition status of the elderly based on multi-source data fusion by any other suitable means (e.g., by means of firmware).
[0205] Various embodiments of the systems and techniques described above in this invention can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0206] The program code used to implement the methods of the present invention can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code can be executed entirely on the machine, partially on the machine, as a standalone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0207] In the context of this invention, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0208] It should be understood that the various forms of processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this invention can be achieved, and this invention does not impose any limitations on them.
[0209] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method for monitoring the nutritional status of the elderly based on multi-source data fusion, characterized in that, Includes the following steps: By using computer vision, wearable devices, smart sensors, and standardized medical interfaces, dietary images, multi-parameter physiological indicators, behavioral and environmental data, and structured health records are collected and aggregated to form a raw multi-source dataset. The original multi-source data is cleaned, standardized, and formatted, and aligned according to a unified time window to generate a standardized dataset for model training and inference. Image, temporal, and structured features are extracted from a standardized dataset using a feature pyramid network, a temporal Transformer, and a pre-trained language model, respectively. Multimodal feature fusion is performed through a causal cross-attention mechanism. Finally, the model is processed by an LSTM+Transformer hybrid model to output the prediction results and risk probabilities of future nutritional status. The LSTM layer uses the tanh activation function, the Transformer feedforward network uses the ReLU activation function, and the output layer uses the Sigmoid activation function. It adopts a distributed database architecture that combines local and cloud computing to classify and persistently store real-time high-frequency multi-source data, historical archived multi-source data, future nutritional status prediction results and risk probabilities, and encrypted medical records. Based on comprehensive dietary guidelines, individual health records, and historical data, combined with future nutritional status predictions and risk probabilities, a multi-factor nonlinear model is used to calculate and dynamically update individualized nutrient intake and physiological indicator baselines. The system compares future nutritional status predictions with risk probabilities, real-time multi-source data with dynamically updated individualized physiological indicator baselines, identifies abnormal risk levels based on preset multi-level collaborative judgment rules, and triggers real-time early warning information pushes to multiple roles. The multi-scale visual feature extraction network employs a causal cross-attention fusion mechanism based on the causal logic that future states are only affected by historical factors, thus limiting the time range of attention computation; at the same time, it dynamically adapts to individual differences. Using time series feature vectors , For historical time steps, To query Q, use spatial feature tokens corresponding to historical and earlier time steps. and structured features Given keys K and values V, calculate the fused features using the standard cross-attention formula; q is a non-negative integer representing the maximum number of time steps to trace back from the current time step i; p is a variable representing the time step of key K.
2. The method for monitoring the dietary nutrition status of the elderly based on multi-source data fusion according to claim 1, characterized in that, The process of forming the original multi-source dataset includes the following steps: Dietary images are preprocessed using a mobile app equipped with a ResNet152 model. The images are uniformly sized to 224×224 pixels, randomly flipped / rotated, normalized, and feature extracted and matched with a recipe database to obtain preliminary dietary data containing calorie, protein, fat, and carbohydrate intake. Data confidence labels are also added. The data is then transmitted via Bluetooth or other protocols with a resume mechanism enabled. Physiological signals are collected by wearable devices and smart sensors, and combined with data from a smart body fat scale to obtain physiological index data; the data is then transmitted encrypted over the network via Bluetooth to a smartphone or indoor gateway. Behavioral and environmental signals are monitored by smart home sensors deployed in key areas to obtain auxiliary data on daily activity levels, gait characteristics, and ambient temperature and humidity; and end-to-end encrypted storage and hierarchical access control are adopted. External medical records are transmitted via standard interfaces and dedicated lines, automatically synchronizing and extracting chronic disease history, medication history, and ADL assessment scores from electronic medical records to obtain structured medical record data; field-level encrypted storage is used, and access logs are recorded throughout the process, with update frequency synchronized with the medical system.
3. The method for monitoring the dietary nutrition status of the elderly based on multi-source data fusion according to claim 1, characterized in that, The process of generating a standardized dataset for model training and inference includes the following steps: After the original multi-source data is uploaded through the transport layer, it is synchronized to the data layer via HTTP / HTTPS protocol for data cleaning. The 3σ criterion is used to identify and remove outliers in time-series physiological and behavioral data. For data points with a missing rate of no more than 10%, linear interpolation is used to fill in the missing data based on the values of the adjacent time points to obtain cleaned and regularized data. After cleaning, the standardized data enters the data standardization process. Based on a predefined data dictionary, the units of data from different sources are converted into the system's internal standard units, and text and classification code data are formatted and encoded to generate standardized format data. During the system development phase, the standardized format data is divided into two independent subsets in a 9:1 ratio according to time sequence: the first 90% of the data forms the training set for subsequent model parameter learning; the last 10% of the data forms the test set for evaluating model performance; during the actual system operation phase, the newly generated real-time standardized format data is used as the inference analysis dataset to generate a standardized dataset for model training or inference.
4. The method for monitoring the dietary nutrition status of the elderly based on multi-source data fusion according to claim 1, characterized in that, The process of outputting future nutritional status predictions and risk probabilities includes the following steps: The standardized dietary image data, temporal physiological behavior data, and expanded structured archive data of 30 days were processed by parallel multi-scale visual feature extraction network, temporal feature extraction network, and text feature encoding network to obtain the corresponding image feature sequence, temporal feature sequence, and structured feature sequence. Three feature sequences are input into a fusion computing unit based on time causality constraints. Each time step feature in the time-series feature sequence is used as the query benchmark. According to the time sequence, the correlation degree is calculated and weighted fusion is performed only with the image feature fragments and structured features of the same or earlier time steps to generate an intermediate fusion feature set containing all historical causal correlation pairs. The intermediate fusion feature set is aggregated and compressed in the time dimension and then regularized into a daily fusion feature sequence corresponding to the original 30-day time window. The daily fusion feature sequence is input into a stacked prediction structure containing temporal memory units and feature interaction layers. After undergoing multi-level nonlinear transformations within the structure, the final output is a quantified value of nutrient intake matching for a specified number of future days and a nutrient risk probability value.
5. The method for monitoring the dietary nutrition status of the elderly based on multi-source data fusion according to claim 4, characterized in that, The temporal prediction model of the temporal feature extraction network adopts a hybrid model of LSTM and Transformer for short-term nutritional status prediction. The temporal prediction model consists of a two-layer LSTM and a two-layer Transformer encoder. The LSTM is responsible for capturing long-term temporal dependencies, while the Transformer enhances key feature interactions. The input is a pooled 30-step sequence of features. .
6. The method for monitoring the dietary nutrition status of the elderly based on multi-source data fusion according to claim 1, characterized in that, The calculation and dynamic updating of individualized nutrient intake and physiological indicator baselines employs a multi-factor nonlinear model to optimize the calculation logic. Based on the predicted nutrient intake matching degree and nutritional risk warning probability output by the AI intelligent analysis module for the next 1-3 days, and combined with real-time physiological indicators, a collaborative judgment is made. The warning information is pushed to the application layer via an interface, and then pushed to the terminal in real time.
7. A system for monitoring the dietary nutritional status of the elderly based on multi-source data fusion, used to implement the method for monitoring the dietary nutritional status of the elderly based on multi-source data fusion as described in any one of claims 1-6, characterized in that, include: The system includes a data acquisition module, a data preprocessing module, an AI intelligent analysis module, a data storage module, a risk assessment module, a system interface module, and a security and privacy protection module. Among them, the data acquisition module is used to collect multi-dimensional heterogeneous data, including dietary image recognition and nutritional composition calculation through computer vision algorithms and ResNet152 models carried by mobile terminal APP, physiological indicator data through wearable devices and smart sensors, behavioral and environmental data through smart home sensors, and electronic medical record structured data through HL7 and FHIR standard interfaces. The data preprocessing module is used to collect data and upload it through the transport layer, then synchronize it to the data layer via HTTP / HTTPS protocol to perform preprocessing operations such as data cleaning, data standardization, and model training data partitioning. The AI intelligent analysis module includes a multimodal multi-scale feature extraction unit, a causal cross-attention fusion unit, and a time-series prediction unit. The multimodal multi-scale feature extraction unit addresses the heterogeneity of dietary images, physiological time-series data, and structured health record data by extracting multi-scale visual features through a feature pyramid network, extracting time-dependent features through a temporal Transformer, and encoding textual semantic features through a DistilBERT model. It then extracts image, temporal, and structured features from a standardized dataset using the feature pyramid network, temporal Transformer, and a pre-trained language model, respectively. Multimodal feature fusion is performed through a causal cross-attention mechanism. Finally, the data is processed by an LSTM+Transformer hybrid model to output future nutritional status predictions and risk probabilities. The data storage module is used to store real-time, high-frequency multi-source data, historical archived multi-source data, future nutritional status prediction results and risk probabilities, and encrypted medical records through local and cloud databases; it adopts off-site multi-copy backup, and medical sensitive data is additionally encrypted and stored. The risk assessment module includes a dynamic baseline establishment unit and an anomaly collaborative identification unit. The dynamic baseline establishment unit constructs a dynamic baseline based on individual health records and historical data using a multi-factor weighted model. It integrates dietary guidelines, individual health records, and historical data, and combines future nutritional status predictions and risk probabilities to calculate and dynamically update individualized nutrient intake and physiological indicator baselines using a multi-factor nonlinear model. The system interface module includes internal and external interface design, enabling the setting of interfaces between the perception layer and transmission layer and the medical system. The security and privacy protection module is used to set up security and privacy protection measures.
8. The dietary nutrition status monitoring system for the elderly based on multi-source data fusion according to claim 7, characterized in that, The causal cross-attention fusion unit is based on the causal logic that the future state is only affected by historical factors, constrains the time matching relationship between attention query and key value, and only uses historical and current time step data for fusion; the time series prediction unit adopts a hybrid model of LSTM and Transformer, and realizes short-term nutritional status prediction based on the fused features.
9. The dietary nutrition status monitoring system for the elderly based on multi-source data fusion according to claim 7, characterized in that, The anomaly collaborative identification unit triggers a three-level early warning based on the predicted value, risk probability, and real-time physiological indicators output by the AI module, according to preset conditions.