Multi-modal health monitoring method and system based on large model driving, terminal and storage medium
By collecting and fusing multimodal data and using large-model-driven intelligent decision-making to generate personalized health recommendations, the problems of multimodal data fusion and insufficient autonomous decision-making in existing health monitoring technologies are solved, and high-precision, safe, contactless health management is achieved.
Patent Information
- Application Number
- CN202510781449.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-09-26
AI Technical Summary
Existing health monitoring technologies have problems such as insufficient multimodal data fusion capabilities, limited autonomous decision-making and personalization capabilities, poor interactive experience, insufficient support for home scenarios, immature non-contact monitoring technology, and data privacy and security issues.
A multimodal health monitoring method driven by a large model is adopted. By collecting video stream data, audio data and physiological monitoring data, preprocessing and feature extraction are performed, and a multimodal health large model is used for data fusion. Dynamic intervention strategies are generated based on the intelligent agent decision logic, and health monitoring data are output in combination with a multimodal interactive interface.
It has achieved the deep integration of non-contact health monitoring and personalized health management, improved the autonomous decision-making and personalization capabilities of health monitoring technology, enhanced monitoring accuracy and user experience, and ensured data security.
Smart Images

Figure CN120708891A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of smart home technology, and in particular to a multimodal health monitoring method, system, terminal and storage medium driven by a large model. Background Art
[0002] Currently, health monitoring technologies are mainly concentrated in wearable devices, smartphone applications, and some smart home devices. The following are existing technical solutions: 1) Wearable device health monitoring: Monitor heart rate, blood oxygen, sleep and other data through smart bracelets, watches and other devices, and synchronize them to mobile phone apps to generate health reports.
[0003] 2) Smartphone health applications: Use the phone’s camera and microphone to detect heart rate and respiratory rate, and combine AI algorithms to provide health advice.
[0004] 3) Smart home health devices: Monitor health data through devices such as smart mirrors and smart scales, and upload them to the cloud to generate reports.
[0005] 4) TV health applications: Display health data on TV (such as connecting wearable devices), and provide fitness videos and health advice.
[0006] Although existing technologies have made progress in some aspects, they still have obvious shortcomings in the following aspects: Insufficient multimodal data fusion capabilities: Existing technologies mostly rely on a single data source (such as wearable devices only monitoring heart rate), lack the deep integration of multimodal data (images, voice, sensors), and have serious data silos, making it difficult to fully reflect the user's health status.
[0007] Limited autonomous decision-making and personalization capabilities: Existing systems mostly rely on preset rules (such as triggering an alert when the heart rate exceeds a threshold), lack autonomous decision-making capabilities, and health advice is mostly general content and lacks personalized adaptation.
[0008] Therefore, the existing technology needs to be improved. Summary of the Invention
[0009] The technical problem to be solved by the present invention is that, in response to the defects of the existing technology, the present invention provides a multimodal health monitoring method, system, terminal and storage medium driven by a large model to solve the problems of insufficient multimodal data fusion capability and low autonomous decision-making and personalization capabilities in the existing health monitoring technology.
[0010] The technical solutions adopted by the present invention to solve the technical problems are as follows: In a first aspect, the present invention provides a multimodal health monitoring method based on a large model drive, comprising: Collecting video stream data, audio data, and physiological monitoring data of the target object, and preprocessing the collected data to obtain multimodal health data; Extracting features from the multimodal health data, and performing multimodal health data fusion processing on the extracted features based on a multimodal health macro model to obtain fused health features; Based on the fused health characteristics, a dynamic intervention strategy is generated using the intelligent agent decision logic driven by the large model; Based on the dynamic intervention strategy, voice reminders and health video playback are performed, and multimodal health monitoring data are dynamically output through a multimodal interactive interface.
[0011] In one implementation, collecting the video stream data, audio data, and physiological monitoring data of the target object includes: Capturing video stream data of the target object through a camera; The microphone array collects ambient sound and uses beamforming technology to focus on the mouth and nose area of the target object to collect audio data of the target object; The physiological monitoring data is obtained by collecting non-contact body temperature through an infrared sensor and synchronizing blood pressure and blood oxygen data through a wearable device.
[0012] In one implementation, preprocessing the collected data to obtain multimodal health data includes: Performing face detection on the video stream data based on a multi-task convolutional neural network, cropping a region of interest, and obtaining video stream data of the facial region of interest; Based on the posture estimation algorithm, the human body key points in the video stream data are detected in real time to obtain the video stream data of body posture and movement; Based on the independent component analysis method, the audio data is subjected to respiratory sound separation and frame and window processing to obtain preprocessed audio data.
[0013] In one implementation, the extracting features from the multimodal health data may include: Use public health datasets to learn multimodal associations, and combine user personalized data to optimize model parameters to obtain a trained multimodal health model.
[0014] In one implementation, extracting features from the multimodal health data includes: Performing independent component analysis on the RGB channels of the video stream data of the facial region of interest, selecting independent components corresponding to the heartbeat features, removing noise using a Butterworth bandpass filter, and calculating the real-time heart rate based on a dynamic threshold algorithm to obtain the heart rate features; Extracting the fundamental frequency of the respiratory sound signal in the preprocessed audio data, and calculating the respiratory cycle according to the extracted fundamental frequency to obtain a respiratory frequency feature; The posture and action are recognized based on the video stream data of the posture and action to obtain posture and action features.
[0015] In one implementation, the multimodal health data fusion processing is performed on the extracted features based on the multimodal health macro model to obtain fused health features, including: The heart rate features, the respiratory rate features, and the posture and movement features are input into a multimodal deep learning model, the correlation between different modalities is learned through a cross-attention mechanism, and weighted voting is performed on the multimodal analysis results to obtain the fused health features.
[0016] In one implementation, generating a dynamic intervention strategy based on the fused health features using a large model-driven intelligent agent decision logic includes: Based on the fused health characteristics, the agent decision logic driven by the large model is used to generate health recommendations and intervention measures for exercise plans and dietary adjustments; When the heart rate variability of the target subject is improved, a positive reward signal is given to the intelligent agent to encourage the intelligent agent to continuously explore and generate strategies that can continuously improve the health indicators of the target subject; Acquire feedback information from the target object, integrate usage information of the target object into the intervention measure, and optimize the intervention measure according to the usage information of the target object.
[0017] In a second aspect, the present invention provides a multimodal health monitoring system driven by a large model, comprising: A multimodal data acquisition module is used to collect video stream data, audio data, and physiological monitoring data of the target object, and pre-process the collected data to obtain multimodal health data; A multimodal data fusion module is used to extract features from the multimodal health data and perform multimodal health data fusion processing on the extracted features based on the multimodal health macro model to obtain fused health features; An agent decision module, configured to generate a dynamic intervention strategy based on the fused health features using agent decision logic driven by a large model; The multimodal interaction module is used to provide voice reminders and health video playback based on the dynamic intervention strategy, and dynamically output multimodal health monitoring data through a multimodal interaction interface.
[0018] In a third aspect, the present invention provides a terminal comprising: a processor and a memory, wherein the memory stores a multimodal health monitoring program driven by a large model, and when the multimodal health monitoring program driven by a large model is executed by the processor, it is used to implement the operation of the multimodal health monitoring method driven by a large model as described in the first aspect.
[0019] In a fourth aspect, the present invention also provides a computer-readable storage medium, which stores a large-model-driven multimodal health monitoring program. When the large-model-driven multimodal health monitoring program is executed by a processor, it is used to implement the operation of the large-model-driven multimodal health monitoring method as described in the first aspect.
[0020] The present invention adopts the above technical solution to achieve the following effects: The present invention collects video stream data, audio data, and physiological monitoring data of the target object and pre-processes the collected data. It can extract features from multimodal health data and perform multimodal health data fusion processing on the extracted features based on a multimodal health big model. It can also generate dynamic intervention strategies based on the fused health features using the big model-driven intelligent agent decision logic. It can then perform voice reminders and health video playback based on the dynamic intervention strategy, and dynamically output multimodal health monitoring data through a multimodal interactive interface. The present invention achieves a deep integration of non-contact health monitoring and personalized health management through technologies such as multimodal sensor fusion, big model-driven intelligent agent decision-making, and natural interaction design, thereby improving the autonomous decision-making and personalization capabilities of health monitoring technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the structures shown in these drawings without paying any creative work.
[0022] Figure 1 This is a flow chart of the multimodal health monitoring method driven by a large model in the present invention.
[0023] Figure 2 It is a functional principle diagram of a terminal in one implementation of the present invention.
[0024] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION
[0025] In order to make the purpose, technical solutions and advantages of the present invention more clear and distinct, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0026] Exemplary Methods Existing technologies have made progress in some aspects, but have obvious shortcomings in the following aspects: 1) Insufficient multimodal data fusion capabilities: Existing technologies mostly rely on a single data source (e.g., wearable devices only monitor heart rate) and lack the deep integration of multimodal data (images, voice, sensors). Data silos are serious, making it difficult to fully reflect the user's health status.
[0027] 2) Limited autonomous decision-making and personalization capabilities: Existing systems mostly rely on preset rules (such as triggering an alert when the heart rate exceeds a threshold), lack autonomous decision-making capabilities, and health advice is mostly general content and lacks personalized adaptation.
[0028] 3) Poor interactive experience: Existing technologies have a single interactive mode (such as mobile phone apps relying on touch), which is not suitable for multi-user interaction in home scenarios and lacks support for natural interaction (such as voice and gestures), resulting in a poor user experience.
[0029] 4) Insufficient support for home scenarios: Existing technologies are mostly targeted at individual users and lack multi-user collaborative support in home scenarios, making it impossible to provide collective health management services for family members.
[0030] 5) Immature non-contact monitoring technology: Existing non-contact monitoring technologies (such as camera heart rate detection) have low accuracy, are easily affected by environmental interference, lack multimodal data verification mechanisms, and are insufficiently reliable.
[0031] 6) Data privacy and security issues: Existing technologies mostly rely on cloud storage, which poses a risk of data leakage and lacks localized data processing and privacy protection mechanisms.
[0032] In response to the above technical problems, an embodiment of the present invention provides a multimodal health monitoring method based on large model driving, which includes: collecting video stream data, audio data and physiological monitoring data of the target object, and preprocessing the collected data to obtain multimodal health data; extracting features from the multimodal health data, and performing multimodal health data fusion processing on the extracted features based on the multimodal health large model; generating dynamic intervention strategies based on the fused health features using the intelligent agent decision logic driven by the large model; performing voice reminders and health video playback based on the dynamic intervention strategy, and dynamically outputting multimodal health monitoring data through a multimodal interactive interface. The present invention realizes the deep integration of non-contact health monitoring and personalized health management through technologies such as multimodal sensor fusion, large model-driven intelligent agent decision-making, and natural interaction design, thereby improving the autonomous decision-making and personalization capabilities of health monitoring technology.
[0033] like Figure 1 As shown, an embodiment of the present invention provides a multimodal health monitoring method based on large model driving, comprising the following steps: Step S100 , collecting video stream data, audio data, and physiological monitoring data of the target object, and preprocessing the collected data to obtain multimodal health data.
[0034] In this embodiment, the multimodal health monitoring method driven by a large model is implemented through a multimodal health monitoring system driven by a large model, wherein the system adopts a layered design and is divided into a perception layer, an edge computing layer, a cloud collaboration layer, and an interaction layer; specifically, the perception layer is used for multimodal health data collection and preprocessing; the edge computing layer is used to deploy lightweight AI models (such as TensorFlow Lite models) to complete real-time data processing; the cloud collaboration layer is provided with a multimodal health large model and an intelligent agent decision engine to realize multimodal health data fusion processing and dynamic intervention strategy generation; the interaction layer is used to realize voice interaction and gesture interaction functions, and dynamically output multimodal health monitoring data, charts, etc. through a visual interface.
[0035] Specifically, in one implementation of this embodiment, step S100 includes the following steps: Step S101, capturing video stream data of the target object through a camera; Step S102, collecting ambient sound through a microphone array, and focusing on the mouth and nose area of the target object through beamforming technology to collect audio data of the target object; Step S103: collecting non-contact body temperature through an infrared sensor, and synchronizing blood pressure and blood oxygen data through a wearable device to obtain the physiological monitoring data.
[0036] In this embodiment, the perception layer includes a hardware module and a data preprocessing module; wherein the hardware module of the perception layer includes a built-in sensor of the television and an external device.
[0037] Specifically, the television built-in sensor includes: Camera: It uses a high-frame rate (≥60fps) RGB + infrared camera solution to support the capture of facial microvascular changes in low-light environments.
[0038] Microphone array: has directional sound pickup and noise suppression functions, used to separate breathing sounds from ambient sounds.
[0039] Infrared sensor: non-contact body temperature monitoring (accuracy ±0.2℃).
[0040] The external device includes: Wearable devices: Synchronize heart rate, blood pressure, and blood oxygen data via Bluetooth / WiFi.
[0041] Environmental sensors: temperature, humidity, and air quality detection.
[0042] In this embodiment, the hardware module based on the perception layer collects video stream data, audio data and physiological monitoring data of the target object; wherein the video stream data includes: face video stream data and video stream data containing posture and movement of key points of the human body; physiological monitoring data includes: non-contact body temperature data, blood pressure data, blood oxygen data, heart rate data, etc.
[0043] As an example, for video stream data, this embodiment uses a camera with high frame rate RGB image and infrared sensor functions to collect facial video stream data and human body key point video stream data respectively; and, in the process of collecting facial video stream data, the camera captures the RGB video stream of the facial ROI area (region of interest, such as forehead, cheek), so as to facilitate the subsequent implementation of non-contact heart rate detection function.
[0044] For audio data, this embodiment uses a microphone array with directional sound pickup and noise suppression capabilities for acquisition. During the acquisition process, the microphone array captures ambient sound and uses beamforming technology to focus on the user's mouth and nose to separate breathing sounds from ambient sound. For physiological monitoring data, non-contact body temperature is collected using an infrared sensor, and blood pressure and blood oxygen data are synchronized with the wearable device to obtain the physiological monitoring data.
[0045] Specifically, in one implementation of this embodiment, step S100 further includes the following steps: Step S104, performing face detection on the video stream data based on a multi-task convolutional neural network, cropping the region of interest, and obtaining video stream data of the facial region of interest; Step S105, detecting key points of the human body in the video stream data in real time based on a posture estimation algorithm to obtain video stream data of body posture and movement; Step S106 , performing respiratory sound separation and frame-by-frame windowing processing on the audio data based on an independent component analysis method to obtain pre-processed audio data.
[0046] In this embodiment, after the hardware module of the perception layer collects the video stream data, audio data and physiological monitoring data of the target object, the data preprocessing module of the perception layer preprocesses the collected data. The specific processing method is as follows: For camera data (video stream data), face detection is performed using a real-time face detection algorithm (for example, the MTCNN algorithm, a face detection algorithm based on deep learning), and the ROI (Region of Interest) area is cropped to obtain video stream data of the facial region of interest.
[0047] For the microphone data (audio data), respiratory sound separation is performed using a respiratory sound separation method (for example, based on ICA independent component analysis), and pre-processed audio data is obtained through frame division and windowing processing.
[0048] like Figure 1 As shown, an embodiment of the present invention provides a multimodal health monitoring method based on large model driving, comprising the following steps: Step S200 , extracting features from the multimodal health data, and performing multimodal health data fusion processing on the extracted features based on a multimodal health big model to obtain fused health features.
[0049] In this embodiment, after collecting and preprocessing multimodal health data, real-time feature extraction is performed through the edge computing layer; wherein, the edge computing layer is deployed with a lightweight AI model (for example, a TensorFlow Lite model) to complete real-time data processing.
[0050] In this embodiment, the physiological parameters extracted in real time by these lightweight AI models include heart rate, respiratory rate, and body posture recognition. Heart rate data can be extracted using PPG (photoplethysmography) signal processing; respiratory rate data can be obtained using LSTM time series analysis; and body posture recognition data can be obtained using OpenPose skeleton detection. After real-time data processing is completed at the edge computing layer, AES-256 encryption is performed before local storage, and only desensitized data is uploaded to the cloud to achieve data encryption.
[0051] In this embodiment, the cloud collaboration layer of the system deploys a multimodal health model. The model architecture of the multimodal health model is a Transformer-based multimodal fusion network (similar to the CLIP structure), and the input includes: Visual modality: facial microvascular change sequence (temporal CNN features).
[0052] Speech modality: Breathing audio spectrogram (Mel spectrum + ResNet features).
[0053] Sensor modality: time series data such as heart rate and body temperature (LSTM encoding).
[0054] In this embodiment, before performing feature extraction and fusion processing on the multimodal health data, it is necessary to train it.
[0055] Specifically, in one implementation of this embodiment, the following steps are included before step S200: Step S201a: Use the public health data set to learn multimodal associations, and optimize model parameters in combination with user personalized data to obtain a trained multimodal health model.
[0056] In this embodiment, during the pre-training phase of the multimodal health model, a public health dataset (such as the MIMIC-III dataset) is used to learn multimodal associations; during the fine-tuning phase, the model parameters are optimized in combination with user personalized data (user authorization is required).
[0057] Specifically, in one implementation of this embodiment, step S200 includes the following steps: Step S201, performing independent component analysis on the RGB channels of the video stream data of the facial region of interest, selecting independent components corresponding to the heartbeat features, removing noise using a Butterworth bandpass filter, and calculating the real-time heart rate based on a dynamic threshold algorithm to obtain a heart rate feature; Step S202, extracting the fundamental frequency of the respiratory sound signal in the preprocessed audio data, and calculating the respiratory cycle according to the extracted fundamental frequency to obtain a respiratory frequency feature; Step S203 : performing posture and action recognition based on the video stream data of the posture and action to obtain posture and action features.
[0058] In this embodiment, PPG signal extraction involves performing ICA separation on each of the RGB channels to select independent components related to the heartbeat. A Butterworth bandpass filter (filter parameters: 0.7Hz - 4Hz, corresponding to 42-240bpm) is then used to remove noise. Finally, a dynamic threshold algorithm is used to calculate the real-time heart rate and obtain heart rate features. To improve heart rate detection accuracy, forehead and cheek signals are analyzed simultaneously, and motion artifacts are reduced through weighted averaging. In low-light environments, the infrared camera is switched to avoid visible light interference.
[0059] The method for extracting respiratory frequency features is to use NMF (non-negative matrix factorization) to separate respiratory sounds from background noise, extract the fundamental frequency of the respiratory sound signal (0.1Hz - 0.5Hz), calculate the respiratory cycle, and obtain the respiratory frequency features.
[0060] The method for extracting posture and motion features is to detect human key points (for example, 17 skeleton nodes) in real time based on the OpenPose algorithm. Then, posture and motion recognition are performed based on defined health-related motion patterns. The defined health-related motion patterns include: Sedentary detection: If the angles of the hip and knee joints remain >120° for more than 30 minutes, it is identified as a sedentary posture.
[0061] Fall detection: sudden change in speed at key points (such as head height drop rate > 2m / s 2 ), which is recognized as a falling posture.
[0062] In this embodiment, after feature extraction is performed on the multimodal health data, the extracted features are input into a multimodal health big model, and multimodal health data fusion processing is performed on the extracted features based on the multimodal health big model to obtain fused health features; wherein, the multimodal health big model is set in the cloud collaboration layer.
[0063] Specifically, in one implementation of this embodiment, step S200 further includes the following steps: In step S204, the heart rate features, the respiratory rate features, and the posture and movement features are input into a multimodal deep learning model, the correlation between different modalities is learned through a cross-attention mechanism, and weighted voting is performed on the multimodal analysis results to obtain the fused health features.
[0064] In this embodiment, during the multimodal data fusion process, features such as heart rate (time series signal), respiratory rate (time series signal), and body posture (skeleton coordinates) are input into a multimodal Transformer model (multimodal deep learning model). The model learns the correlation between different modalities through a cross-attention mechanism, for example, whether an increase in heart rate during rapid breathing is related to exercise (combined with body posture judgment). In addition, during the multimodal data fusion process, decision-level fusion is also required. The specific method is as follows: Perform weighted voting on multimodal analysis results: If the camera heart rate detection conflicts with the wearable device data, the wearable device data (with higher confidence) is used first.
[0065] like Figure 1 As shown, an embodiment of the present invention provides a multimodal health monitoring method based on large model driving, comprising the following steps: Step S300: Generate a dynamic intervention strategy based on the fused health characteristics using the agent decision logic driven by the large model.
[0066] In this embodiment, after the extracted features are subjected to multimodal health data fusion processing based on the multimodal health big model, a dynamic intervention strategy is generated through the big model-driven intelligent agent decision logic; wherein, the big model-driven intelligent agent decision logic is set by the intelligent agent decision engine set in the cloud collaboration layer.
[0067] Specifically, in one implementation of this embodiment, step S300 includes the following steps: Step S301, generating exercise plans, dietary adjustment health recommendations and intervention measures based on the fused health features using the agent decision logic driven by the large model; Step S302: When the heart rate variability of the target subject is improved, a positive reward signal is given to the intelligent agent to motivate the intelligent agent to continuously explore and generate strategies that can continuously improve the health indicators of the target subject; Step S303: Acquire feedback information from the target object, integrate the usage information of the target object into the intervention measure, and optimize the intervention measure according to the usage information of the target object.
[0068] In this embodiment, the agent decision engine is based on dynamic strategy generation using reinforcement learning (PPO algorithm), where the reward function plays a key role and mainly includes the following two core dimensions: Improved user health indicators: for example, increased heart rate variability. Heart rate variability refers to the variation in the differences between heartbeats, which reflects the activity of the heart's autonomic nervous system. Generally speaking, higher heart rate variability is closely related to better cardiovascular health. When the system provides health advice and interventions such as exercise plans and dietary adjustments, and these measures successfully improve the user's heart rate variability, the agent will receive a positive reward signal. This positive feedback motivates the agent to continuously explore and generate strategies that can continuously improve the user's health indicators, making the health management solutions provided by the system more effective.
[0069] User interaction satisfaction: For example, feedback ratings. While using the system, users can provide feedback on the services provided, such as recommended health plans and interactive experiences. High user ratings indicate satisfaction with the system interaction, and the agent will be rewarded for this. This reward mechanism encourages the agent to not only focus on improving health indicators when making decisions, but also to fully consider the user experience, thereby continuously improving the humanization and friendliness of the service and providing a better user experience.
[0070] In this embodiment, the large model-driven intelligent agent decision logic includes two parts: health status assessment and dynamic intervention strategy generation. Among them, the health status assessment is evaluated by the health risk stratification model, and the dynamic intervention strategy is generated by the personalized recommendation engine.
[0071] Specifically, the health risk stratification model: Input: multimodal feature vector (heart rate, respiration, posture, etc.).
[0072] Output: Health risk level (normal, warning, high risk).
[0073] Implementation: Use the XGBoost classifier combined with SHAP values to explain risk factors (e.g., “low heart rate variability”, with a contribution of 60%).
[0074] Personalized recommendation engine: Recommended content: Exercise plans, diet advice, meditation classes, and more.
[0075] Recommendation logic: Build a knowledge graph based on historical user data (e.g., "User A's yoga course completion rate is 80%"). Use graph neural networks (GNNs) to mine potential preferences (e.g., recommending Pilates courses selected by similar users).
[0076] Real-time adjustment: If the user fails to complete the plan for three consecutive days, the agent will automatically reduce the difficulty or change the content type.
[0077] As an example, a health intervention case is: Scenario: The user is detected to have been sedentary for more than 1 hour.
[0078] Decision-making process: The agent uses posture analysis to identify sedentary status. It then queries user history data: If the user prefers light exercise, it recommends a "5-minute stretching video." It then issues a voice reminder: "We've detected you've been sedentary for an hour. Would you like to see a stretching tutorial?" If the user accepts, the video automatically plays the tutorial in full screen. If the user declines, the feedback is recorded and the frequency of similar reminders is reduced.
[0079] like Figure 1 As shown, an embodiment of the present invention provides a multimodal health monitoring method based on large model driving, comprising the following steps: Step S400: Perform voice reminders and health video playback based on the dynamic intervention strategy, and dynamically output multimodal health monitoring data through a multimodal interactive interface.
[0080] In this embodiment, multimodal health monitoring data is dynamically output based on a multimodal interaction approach. Specifically, the multimodal interaction implementation details are as follows: 1) Voice interaction: Wake-up word detection: custom wake-up word (such as "health assistant") + localized voice fingerprint verification.
[0081] Intent recognition: A semantic understanding model based on BERT supports complex queries (such as "show me my average heart rate over the past week").
[0082] 2) Gesture interaction: Gesture library definition: 6 core gestures (such as "open palm" to switch pages, "make fist" to confirm selection).
[0083] Dynamic gesture recognition: Uses optical flow (Lucas-Kanade algorithm) to track hand motions and combines it with a temporal convolutional network (TCN) to determine gesture type.
[0084] 3) Visual interface: Generate interactive charts using D3.js: Heart rate variability (HRV) trend graph: color-mapped stress level (green → red).
[0085] Health score radar chart: displays scores in multiple dimensions such as sleep, exercise, and diet.
[0086] The system in this embodiment can implement multi-user health management. For example, it uses facial recognition (FaceNet model) or voiceprint recognition (based on the X-Vector model) to distinguish family members, generate family health reports (such as "The average sleep duration of the whole family is 6.5 hours"), and recommend group activities (such as family fitness challenges). To ensure data security, only necessary data is collected (for example, raw video is not stored, only heart rate time series signals are retained).
[0087] The heart rate detection solution in this embodiment is compared with medical-grade equipment (such as an electrocardiograph), and the error is controlled within ±3bpm; the respiratory rate solution is compared with medical-grade equipment, and the error is <±0.5 times / minute in a quiet environment.
[0088] This embodiment achieves a deep integration of contactless health monitoring and personalized health management through technologies such as multimodal sensor fusion, large model-driven intelligent agent decision-making, and natural interaction design. It solves the problems of data isolation, rigid interaction, and single scenario in existing technologies, and has significant technological innovation and practical value.
[0089] Through the above solution, this embodiment achieves deep fusion of multimodal data and improves the comprehensiveness and accuracy of health monitoring. Specifically, the advantages and effects of this embodiment compared with the existing technology are as follows: Multimodal data fusion: By integrating multimodal data from cameras, microphones, infrared sensors, and other devices on the TV, combined with large models for comprehensive analysis, we provide more comprehensive health monitoring (such as heart rate, respiratory rate, posture, and body temperature). Multimodal data is mutually verified, significantly improving monitoring accuracy (for example, the heart rate detection error is controlled within ±3bpm).
[0090] Optimized non-contact monitoring: Multi-region fusion (weighted forehead and cheek signals) and infrared-assisted technology reduce ambient light interference and enhance the reliability of non-contact monitoring. Users can obtain accurate health data without wearing additional equipment, making the experience more convenient. Multimodal data fusion enables more comprehensive health status assessments, providing a reliable basis for personalized health management.
[0091] Large-model-driven intelligent agents: A Transformer-based multimodal fusion model comprehensively analyzes user health data and dynamically generates personalized health recommendations (such as exercise plans and dietary adjustments). Decision strategies are optimized through reinforcement learning (PPO algorithm), and the reward function is designed to improve health indicators and enhance user satisfaction.
[0092] Dynamic Adjustment: The intelligent agent dynamically adjusts monitoring strategies and intervention plans (such as reducing alert frequency or changing recommended content) based on historical user data and real-time feedback. Users receive tailored health management plans, significantly improving compliance (testing shows a 40% increase). The intelligent agent proactively adapts to user needs, providing a more personalized service experience.
[0093] Multimodal interaction support: Supports multiple interaction methods such as voice, gestures, and images, allowing users to interact with the system in natural ways (such as querying health data with voice and switching pages with gestures). The large TV screen provides an immersive interactive experience, suitable for multiple users in a home setting.
[0094] Low-latency interaction: Edge computing layer processing latency is less than 200ms, ensuring real-time interaction. This makes user operation more convenient, especially for elderly users or those unfamiliar with complex operations. In family scenarios, multiple users can share health services through natural interaction, increasing usage frequency and satisfaction.
[0095] Multi-user identification and management: Use facial recognition (FaceNet model) or voiceprint recognition (based on x-vector) to distinguish family members and provide personalized services for each user. Generate family health reports (e.g., average family sleep duration 6.5 hours) and recommend group activities (e.g., family fitness challenges).
[0096] Data aggregation and sharing: Supports comparison and analysis of family members' health data, promoting family health awareness. Family members can share health services, improving the coordination and efficiency of family health management. Group activity recommendations enhance family interaction and health awareness.
[0097] Edge computing and localized processing: Health data is processed and encrypted locally, and only desensitized data is uploaded to the cloud, reducing the risk of privacy leakage.
[0098] Federated learning support: During model training, user data is retained locally and only gradient parameters are uploaded to further ensure data security.
[0099] Data minimization: Only necessary data is collected (e.g., raw video is not stored, only heart rate timing signals are retained), reducing privacy risks. This significantly improves user data security and complies with privacy regulations such as GDPR. This increases user trust in the system and promotes long-term use of health services.
[0100] The technical verification and optimization results of this embodiment are significant: Accuracy verification: Heart rate detection error is controlled within ±3bpm, and respiratory rate error is <±0.5 times / minute, achieving medical-grade accuracy.
[0101] Latency optimization: The edge computing layer processing delay is less than 200ms, meeting real-time interaction requirements.
[0102] User testing: A / B testing showed a 40% increase in user compliance, with satisfaction significantly higher than traditional health apps. The system demonstrated high accuracy, low latency, and high user satisfaction in real-world applications, demonstrating its potential for widespread adoption.
[0103] This embodiment achieves the following technical effects through the above technical solution: This embodiment collects video stream data, audio data, and physiological monitoring data of the target object and pre-processes the collected data to extract features from multimodal health data. Based on the multimodal health big model, the extracted features are fused into multimodal health data. Based on the fused health features, a dynamic intervention strategy can be generated using the big model-driven intelligent agent decision logic. Based on the dynamic intervention strategy, voice reminders and health video playback are performed, and multimodal health monitoring data is dynamically output through a multimodal interactive interface. This embodiment achieves a deep integration of non-contact health monitoring and personalized health management through technologies such as multimodal sensor fusion, big model-driven intelligent agent decision-making, and natural interaction design, thereby improving the autonomous decision-making and personalization capabilities of health monitoring technology.
[0104] Exemplary devices Based on the above embodiments, the present invention further provides a multimodal health monitoring system driven by a large model, comprising: A multimodal data acquisition module is used to collect video stream data, audio data, and physiological monitoring data of the target object, and pre-process the collected data to obtain multimodal health data; A multimodal data fusion module is used to extract features from the multimodal health data and perform multimodal health data fusion processing on the extracted features based on the multimodal health macro model to obtain fused health features; An agent decision module, configured to generate a dynamic intervention strategy based on the fused health features using agent decision logic driven by a large model; The multimodal interaction module is used to provide voice reminders and health video playback based on the dynamic intervention strategy, and dynamically output multimodal health monitoring data through a multimodal interaction interface.
[0105] This embodiment achieves the following technical effects through the above technical solution: This embodiment collects video stream data, audio data, and physiological monitoring data of the target object and pre-processes the collected data to extract features from multimodal health data. Based on the multimodal health big model, the extracted features are fused into multimodal health data. Based on the fused health features, a dynamic intervention strategy can be generated using the big model-driven intelligent agent decision logic. Based on the dynamic intervention strategy, voice reminders and health video playback are performed, and multimodal health monitoring data is dynamically output through a multimodal interactive interface. This embodiment achieves a deep integration of non-contact health monitoring and personalized health management through technologies such as multimodal sensor fusion, big model-driven intelligent agent decision-making, and natural interaction design, thereby improving the autonomous decision-making and personalization capabilities of health monitoring technology.
[0106] Based on the above embodiment, the present invention further provides a terminal, whose principle block diagram can be shown as follows: Figure 2 shown.
[0107] The terminal includes: a processor, memory, interface, display screen and communication module connected via a system bus; wherein the processor of the terminal is used to provide computing and control capabilities; the memory of the terminal includes a computer-readable storage medium and an internal memory; the computer-readable storage medium stores an operating system and a computer program; the internal memory provides an environment for the operation of the operating system and computer program in the computer-readable storage medium; the interface is used to connect to external devices; the display screen is used to display corresponding information; and the communication module is used to communicate with a cloud server or other devices.
[0108] When the computer program is executed by a processor, it is used to implement the operation of a multimodal health monitoring method driven by a large model.
[0109] It will be understood by those skilled in the art that Figure 2 The principle block diagram shown in the figure is only a block diagram of a partial structure related to the solution of the present invention, and does not constitute a limitation on the terminal to which the solution of the present invention is applied. The specific terminal may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0110] In one embodiment, a terminal is provided, which includes: a processor and a memory, wherein the memory stores a multimodal health monitoring program driven by a large model, and when the multimodal health monitoring program driven by a large model is executed by the processor, it is used to implement the operation of the above-mentioned multimodal health monitoring method driven by a large model.
[0111] In one embodiment, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a large model-driven multimodal health monitoring program, and when the large model-driven multimodal health monitoring program is executed by a processor, it is used to implement the operation of the above-mentioned large model-driven multimodal health monitoring method.
[0112] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile storage medium. When executed, the computer program can include the processes in the above-described method embodiments. Any reference to memory, storage, database, or other media used in the various embodiments provided herein may include both non-volatile and volatile memory.
[0113] In summary, the present invention provides a multimodal health monitoring method, system, terminal and storage medium driven by a large model, including: collecting video stream data, audio data and physiological monitoring data of the target object, and preprocessing the collected data to obtain multimodal health data; extracting features from the multimodal health data, and performing multimodal health data fusion processing on the extracted features based on the multimodal health large model; generating dynamic intervention strategies based on the fused health features using the intelligent agent decision logic driven by the large model; performing voice reminders and health video playback based on the dynamic intervention strategy, and dynamically outputting multimodal health monitoring data through a multimodal interactive interface. The present invention realizes the deep integration of non-contact health monitoring and personalized health management through technologies such as multimodal sensor fusion, large model-driven intelligent agent decision-making, and natural interaction design, thereby improving the autonomous decision-making and personalization capabilities of health monitoring technology.
[0114] It should be understood that the application of the present invention is not limited to the above examples. For those skilled in the art, improvements or changes can be made based on the above description. All these improvements and changes should fall within the scope of protection of the claims attached to the present invention.
Claims
1. A multimodal health monitoring method based on large model driving, characterized in that: include: Collecting video stream data, audio data, and physiological monitoring data of the target object, and preprocessing the collected data to obtain multimodal health data; Extracting features from the multimodal health data, and performing multimodal health data fusion processing on the extracted features based on a multimodal health macro model to obtain fused health features; Based on the fused health characteristics, a dynamic intervention strategy is generated using the intelligent agent decision logic driven by the large model; Based on the dynamic intervention strategy, voice reminders and health video playback are performed, and multimodal health monitoring data are dynamically output through a multimodal interactive interface.
2. The multimodal health monitoring method based on large model driving according to claim 1 is characterized in that: The collecting of the video stream data, audio data and physiological monitoring data of the target object includes: Capturing video stream data of the target object through a camera; The microphone array collects ambient sound and uses beamforming technology to focus on the mouth and nose area of the target object to collect audio data of the target object; The physiological monitoring data is obtained by collecting non-contact body temperature through an infrared sensor and synchronizing blood pressure and blood oxygen data through a wearable device.
3. The multimodal health monitoring method based on large model driving according to claim 1 is characterized in that: The preprocessing of the collected data to obtain multimodal health data includes: Performing face detection on the video stream data based on a multi-task convolutional neural network, cropping a region of interest, and obtaining video stream data of the facial region of interest; Based on the posture estimation algorithm, the human body key points in the video stream data are detected in real time to obtain the video stream data of body posture and movement; Based on the independent component analysis method, the audio data is subjected to respiratory sound separation and frame and window processing to obtain preprocessed audio data.
4. The multimodal health monitoring method based on large model driving according to claim 1 is characterized in that: The feature extraction of the multimodal health data includes: Use public health datasets to learn multimodal associations, and combine user personalized data to optimize model parameters to obtain a trained multimodal health model.
5. The multimodal health monitoring method based on large model driving according to claim 3 is characterized in that: The extracting features from the multimodal health data includes: Performing independent component analysis on the RGB channels of the video stream data of the facial region of interest, selecting independent components corresponding to the heartbeat features, removing noise using a Butterworth bandpass filter, and calculating the real-time heart rate based on a dynamic threshold algorithm to obtain the heart rate features; Extracting the fundamental frequency of the respiratory sound signal in the preprocessed audio data, and calculating the respiratory cycle according to the extracted fundamental frequency to obtain a respiratory frequency feature; The posture and action are recognized based on the video stream data of the posture and action to obtain posture and action features.
6. The multimodal health monitoring method based on large model driving according to claim 5 is characterized in that: The multimodal health data fusion processing is performed on the extracted features based on the multimodal health big model to obtain fused health features, including: The heart rate features, the respiratory rate features, and the posture and movement features are input into a multimodal deep learning model, the correlation between different modalities is learned through a cross-attention mechanism, and weighted voting is performed on the multimodal analysis results to obtain the fused health features.
7. The multimodal health monitoring method based on large model driving according to claim 1 is characterized in that: The method of generating a dynamic intervention strategy based on the fused health characteristics and utilizing the agent decision logic driven by the large model includes: Based on the fused health characteristics, the agent decision logic driven by the large model is used to generate health recommendations and intervention measures for exercise plans and dietary adjustments; When the heart rate variability of the target subject is improved, a positive reward signal is given to the intelligent agent to encourage the intelligent agent to continuously explore and generate strategies that can continuously improve the health indicators of the target subject; Acquire feedback information from the target object, integrate usage information of the target object into the intervention measure, and optimize the intervention measure according to the usage information of the target object.
8. A multimodal health monitoring system driven by a large model, characterized in that: include: A multimodal data acquisition module is used to collect video stream data, audio data, and physiological monitoring data of the target object, and pre-process the collected data to obtain multimodal health data; A multimodal data fusion module is used to extract features from the multimodal health data and perform multimodal health data fusion processing on the extracted features based on the multimodal health macro model to obtain fused health features; An agent decision module, configured to generate a dynamic intervention strategy based on the fused health features using agent decision logic driven by a large model; The multimodal interaction module is used to provide voice reminders and health video playback based on the dynamic intervention strategy, and dynamically output multimodal health monitoring data through a multimodal interaction interface.
9. A terminal, characterized in that: include: A processor and a memory, wherein the memory stores a large model-driven multimodal health monitoring program, and when the large model-driven multimodal health monitoring program is executed by the processor, it is used to implement the operation of the large model-driven multimodal health monitoring method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a large-model-driven multimodal health monitoring program, which, when executed by a processor, is used to implement the operation of the large-model-driven multimodal health monitoring method as described in any one of claims 1 to 7.