A robot and method for monitoring and early warning after thrombolysis in acute ischemic stroke.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-01
- Publication Date
- 2026-08-14
AI Technical Summary
[0005]本申请旨在解决现有技术中AIS溶栓后人工监护执行率低、存在监测盲区、预警滞后、神经功能评估主观性强且耗时、监护设备不可移动及线缆干扰等问题,提供一种用于急性缺血性脑卒中溶栓后监测预警的机器人及其方法
1.填补临床场景空白:首次为AIS溶栓后24小时监护场景提供端到端智能机器人,实现无间断监测。
Smart Images

Figure CN122556941A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of medical robotics and artificial intelligence, specifically to an intelligent monitoring and early warning robot and method for bedside monitoring of patients with acute ischemic stroke (AIS) within 24 hours after intravenous thrombolysis. Background Technology
[0002] Acute ischemic stroke (AIS) is one of the leading causes of death and disability worldwide, and intravenous thrombolysis is the most effective standard treatment for the hyperacute phase of AIS. However, within 24 hours after thrombolysis, patients may experience serious complications such as symptomatic intracranial hemorrhage, neurological deterioration, and seizures. Among these, the incidence of symptomatic intracranial hemorrhage is approximately 3.4%, and once it occurs, the risk of severe disability or mortality can be as high as 90%. Therefore, early, continuous, and accurate monitoring and early warning of the patient's condition after thrombolysis are crucial for improving patient prognosis.
[0003] Current clinical guidelines (such as the NINDS protocol) recommend 37 high-intensity monitoring sessions within 24 hours after thrombolysis (every 15 minutes for the first 2 hours, every 30 minutes for the next 6 hours, and every hour thereafter), covering core indicators such as vital signs, neurological function (NIHSS score), and complications. However, this high-intensity, high-frequency manual monitoring protocol is rarely implemented in clinical practice. Due to factors such as nursing staff shortages, night shifts, and multiple patient resuscitations, the actual number of monitoring sessions falls far short of the target, and there are significant monitoring blind spots and delayed warnings. In addition, traditional monitors are immobile, have messy cables, and create monitoring blind spots when patients are out for examinations or transport; NIHSS assessments are time-consuming and highly subjective, making high-frequency implementation difficult.
[0004] While some low-intensity monitoring solutions and non-contact monitoring research based on wearable devices and millimeter-wave radar exist in existing technologies, none have overcome the core bottlenecks of being "manual, intermittent, and single-modal." Currently, there is no intelligent robot product specifically designed for 24-hour bedside monitoring after AIS thrombolysis, integrating multimodal non-contact sensing, multimodal large-scale model automatic assessment of neurological function, and real-time early warning of complication risks. Therefore, there is an urgent need for a new monitoring solution capable of continuous 24-hour monitoring, automatic monitoring, intelligent assessment, and real-time early warning to replace intermittent manual monitoring, eliminate blind spots, and improve patient prognosis. Summary of the Invention
[0005] This application aims to address the problems in existing technologies, such as low execution rate of manual monitoring after AIS thrombolysis, existence of monitoring blind spots, delayed early warning, strong subjectivity and time-consuming neurological function assessment, immobile monitoring equipment, and cable interference. It provides a robot and method for monitoring and early warning after thrombolysis in acute ischemic stroke.
[0006] This application adopts a dual-mode collaborative architecture of "bedside host + wearable physiological patch", integrating hardware such as millimeter-wave radar, binocular vision, directional microphone array, and edge NPU computing unit, as well as software modules such as multimodal fusion perception module, multimodal large model evaluation module, complication risk warning module, and human-machine collaboration module, to achieve 24-hour continuous on-duty, automatic monitoring, intelligent disease identification and real-time risk warning.
[0007] This application provides a robot for monitoring and early warning after thrombolysis in acute ischemic stroke, including a bedside main unit and multiple functional modules deployed on its edge NPU computing unit. The bedside main unit integrates a millimeter-wave radar module, a binocular vision module and a directional microphone array, an edge NPU computing unit, a touch screen and an audio-visual early warning device. The edge NPU computing unit is equipped with: Multimodal fusion sensing module: used to extract heart rate and respiratory rate data from millimeter-wave radar echo signals, and to perform spatiotemporal alignment and feature fusion on image sequences and motion videos to generate fused temporal features; Multimodal large model evaluation module: After fine-tuning by instructions and optimization by human feedback reinforcement learning, it is used to jointly analyze fused temporal features and speech response audio, and automatically output NIHSS score; Complication risk warning module: It has a built-in AI model based on multimodal data pre-training, which is used to analyze the fused time-series features and the NIHSS score in real time, identify abnormal vital signs, deterioration of neurological function, and prodromal symptoms of complications in real time, and determine whether the alarm conditions are met. When the alarm conditions are met, an alarm signal is output. Human-machine collaboration module: used to display NIHSS scores on the touch screen and provide an entry point for manual review and error correction; when a corrected score is received from manual input, the corrected score is used as the final score; and the audible and visual warning device is triggered according to the alarm signal.
[0008] The edge NPU computing unit is configured to perform all data processing locally, the original patient data does not leave the hospital's local network, and the end-to-end warning response time is less than the preset duration.
[0009] This application provides a method for monitoring and early warning after thrombolysis in acute ischemic stroke, comprising the following steps: The micro-motion echo signal of the chest cavity is acquired by millimeter-wave radar, and the heart rate and respiratory rate data are obtained by solving on the edge NPU. Simultaneously acquire facial expression image sequences, limb movement videos, and voice response audio using a binocular vision module and a directional microphone array; Spatiotemporal alignment and feature fusion of heart rate, respiration, images, and videos are performed on the edge NPU to generate fused temporal features; Using a multimodal large model that has been fine-tuned by instructions and optimized by RLHF, the fused temporal features and speech response audio are jointly analyzed to automatically output the NIHSS score; The NIHSS score is displayed on a touchscreen, providing an entry point for manual review and error correction, and the final score is obtained. Using the complication risk early warning model, based on the fused temporal features and the final score, abnormal vital signs, deterioration of neurological function, and prodromal symptoms of complications are identified in real time, and it is determined whether the alarm conditions are met. When the alarm conditions are met, an alarm signal is output. When the alarm conditions are met, an audible and visual warning signal is triggered.
[0010] All the above steps are executed locally on the bedside host, the raw data does not leave the hospital, and the end-to-end response time is less than the preset time.
[0011] Furthermore, the method also includes: collecting single-lead electrocardiogram signals, blood oxygen saturation, skin temperature and body acceleration data of patients through wearable physiological patches; and inputting the above data into a complication risk warning model after fusing it with fusion time-series features.
[0012] Furthermore, the method also includes: receiving encrypted aggregated parameters of a multi-center global model via a federated learning client deployed on a bedside host; updating the local model on the federated learning client using local patient data and uploading the encrypted model gradient to the aggregation server, without uploading any original patient data; and downloading the updated global model parameters from the aggregation server to replace the local complication risk warning model.
[0013] Compared with the prior art, this application has the following beneficial effects: 1. Filling a gap in clinical scenarios: For the first time, an end-to-end intelligent robot is provided for the 24-hour monitoring scenario after AIS thrombolysis, enabling uninterrupted monitoring.
[0014] 2. Significantly reduce nursing workload: Robots handle more than 80% of routine monitoring tasks.
[0015] 3. Improve the objectivity and efficiency of neurological function assessment: The multimodal large model automatically outputs NIHSS scores with ICC ≥ 0.75 and supports manual review.
[0016] 4. Achieve real-time and accurate early warning: Calculate risks in real time based on multimodal data streams, with end-to-end response time <3 seconds and Kappa ≥ 0.8.
[0017] 5. Ensuring data privacy and continuous evolution: Edge computing + federated learning, raw data stays within the facility.
[0018] 6. Improve patient safety and comfort: Millimeter-wave radar provides non-contact monitoring, avoiding skin irritation and nighttime interference. Attached Figure Description
[0019] Figure 1 This is a schematic diagram of the robot according to an embodiment of this application.
[0020] Figure 2 This is an overall flowchart of the monitoring and early warning method after thrombolysis for acute ischemic stroke according to an embodiment of this application.
[0021] Figure 3 This is a sub-flowchart of the multimodal data fusion and early warning model processing in an embodiment of this application.
[0022] Figure 4 This is a flowchart illustrating the human-computer collaborative interaction process in an embodiment of this application. Detailed Implementation
[0023] The technical solutions of the embodiments of this application will now be described with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0024] I. System Hardware Composition Figure 1 A schematic diagram of the robot's structure, such as Figure 1 As shown, the robot provided in this embodiment for monitoring and early warning after thrombolysis in acute ischemic stroke includes a bedside host 100 and a wearable physiological patch 200.
[0025] The bedside unit 100 integrates the following hardware components: Millimeter-wave radar module 110: Employs a frequency-modulated continuous wave (FMCW) system with a center frequency of 77 GHz, a bandwidth of 4 GHz, and a transmit power of 10 dBm. The antenna uses a 4-transmit, 4-receive array with an angular resolution of approximately 15°. It is used for non-contact acquisition of micro-motion echo signals from the patient's chest cavity. This radar can penetrate blankets or clothing no thicker than 5 cm and is unaffected by ambient light. The radar's raw data sampling rate is 1000 Hz, acquiring 512 chirps per frame. Micro-motion signals are extracted using range FFT and Doppler FFT.
[0026] Binocular Vision Module 120: Includes two 5-megapixel infrared cameras (2592×1944 resolution, 30fps), with a baseline distance of 12cm, used to simultaneously acquire facial expression image sequences and limb movement videos of patients. The binocular structure can acquire depth information to help determine the amplitude of limb movements. Built-in infrared illumination allows operation in complete darkness.
[0027] Directional microphone array 130: Contains 4 MEMS microphone units with a linear array spacing of 3cm. The beamforming is directed towards the patient's mouth area (beamwidth ±30°) for acquiring voice response audio. The sampling rate is 16kHz, with 16-bit quantization and environmental noise suppression function (signal-to-noise ratio >20dB).
[0028] Edge NPU computing unit 140: It adopts a neural network processor with a computing power of no less than 5 TOPS (in this embodiment, Rockchip RK3588 is used, which has a built-in 6 TOPS NPU), equipped with 8GB LPDDR4X memory and 64GB eMMC storage, and runs an embedded Linux real-time operating system for local execution of inference operations of all AI models, with a typical power consumption of 8W.
[0029] Touchscreen 150: 12-inch IPS display with a resolution of 1920×1280 and a brightness of 500 nits. It supports 10-point capacitive touch and is used to display monitoring data, NIHSS scores, and provide an interactive interface for manual review.
[0030] The audible and visual warning device 160 includes a three-color LED indicator (red / yellow / green) and a piezoelectric buzzer. The LED drive current is 20mA, and the buzzer sound pressure level is ≥85dB@10cm. A solid green light indicates normal equipment operation; a flashing yellow light at 1Hz indicates a medium-risk warning; and a flashing red light at 3Hz indicates a high-risk warning. The buzzer frequency is controlled by PWM, with a 1kHz square wave (50% duty cycle) output for medium-risk situations and a 1.5kHz square wave (50% duty cycle) output for high-risk situations.
[0031] The wearable physiological patch 200 integrates: Single-lead ECG electrode: Ag / AgCl dry electrode, sampling rate 250Hz, common-mode rejection ratio >100dB; Reflective pulse oximeter: Uses dual wavelengths of red light (660nm) and infrared light (940nm), with a sampling rate of 100Hz; Thermistor temperature sensor: accuracy ±0.1℃, sampling rate 1Hz; Triaxial accelerometer: range ±8g, sampling rate 100Hz.
[0032] The patch communicates with the bedside host 100 via Bluetooth 5.2 low-power mode, with a communication range of 10 meters and encrypted data transmission. Made of medical-grade silicone, the patch can be applied to the patient's left chest at the 3rd intercostal space or upper arm to continuously collect ECG, blood oxygen saturation, skin temperature, and body acceleration data. Its built-in 200mAh lithium battery allows for continuous operation for 24 hours.
[0033] II. Core Technology Architecture and Working Principle of Robots This application focuses on the monitoring scenario within 24 hours after thrombolysis for AIS patients. It adopts a dual-mode collaborative architecture of "bedside host + wearable physiological patch," and develops five core AI technology modules to address the core needs of real-time monitoring and early warning of high-risk complications such as post-thrombolysis neurological deterioration, hemorrhage transformation, and epileptic seizures. The system as a whole adopts a semi-automated monitoring and assessment + manual assisted correction human-machine collaborative working mode, constructing a comprehensive, high-precision, low-interference, and highly safe bedside intelligent monitoring system. The robot proposed in this application can realize a closed-loop monitoring process of intelligent recognition of the patient's routine status, real-time early warning of abnormal status, and manual intervention for special scenarios, comprehensively improving the intelligence, standardization, and fault tolerance of clinical monitoring after AIS thrombolysis.
[0034] The technical details of the specific core technology modules are as follows: 1. Millimeter-wave radar-visual multimodal fusion non-contact vital sign monitoring Radar signal processing flow: The radar module transmits an FMCW signal, receives the echo, and mixes it to obtain an intermediate frequency signal. The multimodal fusion sensing module on the edge NPU performs the following steps: Perform a distance FFT (256 points) on each chirp to obtain a distance image; Perform Doppler FFT (128 points) on multiple chirps at the same distance gate to obtain a velocity-distance map; Select the distance gate with the highest energy (corresponding to the thoracic cavity location) and extract the phase sequence of that distance gate; Bandpass filtering was applied to the phase sequence: respiratory frequency band 0.1-0.5Hz, heart rate frequency band 0.8-2Hz (corresponding to 6-30 beats / min and 48-120 beats / min). Respiratory rate and heart rate are calculated using a peak detection algorithm (interval between adjacent peaks), and updated every 0.5 seconds.
[0035] Meanwhile, variational mode decomposition (VMD) is used to separate respiratory and heartbeat signals, improving the solution accuracy. The VMD penalty factor is set to 1000, the number of modes K=2, and the convergence tolerance is 1e-6.
[0036] Visual signal processing flow: The binocular vision module generates depth maps for each frame after distortion correction and stereo matching. Using MediaPipe or a similar framework, 468 facial key points are extracted to calculate eye opening and mouth corner distortion; 33 skeletal key points (limbs and torso) are extracted to calculate joint angles and range of motion. Optical flow (Farneback algorithm, 3 pyramid levels, 15 window size) is used to calculate motion vectors and detect involuntary body movements.
[0037] Multimodal fusion algorithm: After timestamp alignment, radar heart rate / respiration sequences (sampling rate 2Hz), visual keypoint motion amplitude (downsampled from 30Hz to 2Hz), and accelerometer data (downsampled from 100Hz to 2Hz) are concatenated into a multi-dimensional feature vector. A temporal convolutional network (TCN) based on an attention mechanism is used for fusion: the input window size is 60 seconds (120 time points), the feature dimension is 15, the TCN layer number is 4, the convolution kernel size is 3, the dilation coefficient is [1,2,4,8], and the output fused temporal feature vector has a dimension of 64. This feature preserves the temporal correlation between heart rate changes and limb movements and facial expressions, such as the synchronicity between increased heart rate and increased limb activity.
[0038] 2. Semi-automated assessment of neural function based on multimodal large models Data Acquisition and Preprocessing: At preset monitoring time points (such as every 15 minutes, 30 minutes, and 1 hour after thrombolysis), the robot prompts the patient to perform standard NIHSS actions via touchscreen and voice synthesis: "Please open your eyes and look at the camera" (assessing facial paralysis and staring). "Please smile and show your teeth" (assess facial paralysis). "Please close your eyes, extend your arms with palms up, and hold for 10 seconds" (assess upper limb movement). "Please lift your left leg / right leg" (assess lower limb movement); “Please say ‘The weather is nice today’” (Assess speech disorders); “Please repeat what I said: ‘Eat grapes without spitting out the grape skins’” (Assess aphasia).
[0039] The binocular vision module records video, the directional microphone records audio, and the start and end timestamps are recorded simultaneously.
[0040] Multimodal large model inference: Model input: One frame is extracted from every 5 frames in the video as a keyframe (approximately 60 frames in total), and complete sentence fragments (2-5 seconds in length) are extracted from the audio. A pre-trained visual encoder (Video Swin Transformer-Base) is used to extract spatiotemporal features (1024 dimensions); an audio encoder (Wav2Vec 2.0) is used to extract acoustic features (768 dimensions). The two features are fused through cross-modal attention and then input into the LLaMA-7B decoder to generate text. Beam search (beam size=3) is used during generation, and generating numbers outside the scoring range is prohibited (e.g., facial paralysis can only generate 0, 1, 2, 3). Inference time is approximately 800ms (on the NPU).
[0041] Manual review mechanism: Each rating on the touchscreen displays a thumbnail of the corresponding video clip. Nurses can click on a rating to bring up a numeric keypad and modify it. After modification, the system stores the (original video clip, original model output, and corrected rating) as a preference pair in its local database.
[0042] 3. Intelligent early warning model for multidimensional complications and deterioration of neurological function The model incorporates an AI model pre-trained on multimodal data (such as XGBoost, temporal convolutional networks, or Transformers). Trained in real-time, the model analyzes continuous data streams (heart rate, respiration, body movement, wearable parameters, etc.) from the multimodal fusion sensing module and NIHSS scores output by the multimodal large model evaluation module, dynamically identifying the following three types of anomalies: Abnormal vital signs: such as sustained increase in heart rate (more than 20% increase within 5 minutes), irregular breathing rhythm (coefficient of variation of respiratory interval >0.3), decrease in blood oxygen saturation to <94%, abnormal increase in skin temperature (>1℃ increase within 1 hour), etc. Neurological deterioration: such as an increase of ≥4 points in the NIHSS total score compared to the previous assessment, or an increase of ≥2 points in the scores of key items (facial paralysis, limb movement, language); Prodromal symptoms of complications include: periodic twitching detected by body acceleration (frequency 3-5 Hz, consistent with epileptiform seizure characteristics), drastic fluctuations in blood pressure (systolic blood pressure change >30 mmHg / 10 min), and sudden decline in level of consciousness (assessed by speech response delay), etc.
[0043] Model structure and input features: This embodiment uses the XGBoost multi-output model, and the input features include: Average heart rate (last 1 minute, 5 minutes, 15 minutes); Heart rate variability (SDNN, last 5 minutes); Mean respiratory rate (last 1 minute, 5 minutes); Respiratory rhythm entropy (using sample entropy, 60-second window); NIHSS total score (updated every 15 minutes); Trends in NIHSS scores (changes in the last 3 scores). Wearable patches: mean blood oxygen saturation (last minute), skin temperature change rate (last 5 minutes), kinetic energy (sum of squares of acceleration signals); Electronic medical records: INR, platelet count, blood glucose level; The total number of input dimensions is 32.
[0044] The model outputs three probability values: P_vitals (probability of abnormal vital signs), P_neuro (probability of neurological function deterioration), and P_complication (probability of prodromal symptoms of complications). Simultaneously, the system maintains a lightweight rule engine, for example: A heart rate increase of more than 15% within 10 minutes and an increase of ≥2 in the NIHSS score → directly triggers a neurological deterioration alarm (P_neuro=1.0). The body's kinetic energy exhibits periodic oscillations of 0.5-2Hz → triggering an alarm for prodromal symptoms of complications (suspected epilepsy).
[0045] Training data: Data from 695 real patients were used, including 78 positive cases (hemorrhagic transformation or neurological deterioration) and 617 negative cases. SMOTE oversampling was used to balance the classes. Model hyperparameters: tree depth 6, learning rate 0.05, number of subtrees 300, number of early stopping rounds 20. Five-fold cross-validation showed an AUC of 0.87 for hemorrhagic transformation and 0.84 for neurological deterioration.
[0046] Real-time reasoning and alarm conditions: The model infers every 5 seconds using a sliding window approach, inputting time-series features from the most recent 5 minutes. Alarm conditions are as follows: If max(P_vitals, P_neuro, P_complication) ≥ 0.7, or any probability ≥ 0.5 and lasts for more than 10 minutes (two consecutive inferences ≥ 0.5), then a red alert (high risk) is triggered. If 0.3 ≤ max(…) < 0.7, a yellow alert (medium risk) is triggered. Otherwise, there will be no warning.
[0047] Warning signals (including abnormal types) are passed to the audio-visual control thread through the event queue.
[0048] 4. Edge Real-Time Inference and Federated Learning Architecture Inference pipeline optimization: All models (radar computation, visual feature extraction, large models, XGBoost) are executed in a pipelined manner on the NPU. Radar data triggers an interrupt every 0.5 seconds, visual data triggers an interrupt every frame, and audio is triggered during evaluation. Multi-threaded asynchronous processing is used, with a main loop frequency of 5Hz to ensure end-to-end latency is less than 3 seconds.
[0049] Federated Learning Agreement: The server initiates a federated learning round every two weeks. At the edge, monitoring data (including manually corrected labels) is used for local training using the locally stored data, employing the SGD optimizer with a learning rate of 0.001, a batch size of 32, and a training duration of 3 epochs. After calculating the model gradient, noise (ε=0.1, δ=1e-5) is added using differential privacy, and then uploaded to the central server. The server uses the FedAvg aggregation algorithm to update the global model before redeploying it. Throughout the entire process, the original data (video, audio, radar echoes) never leaves local storage.
[0050] 5. Human-Machine Collaborative Clinical Intelligent Decision Support System Early warning handling process: When the early warning model outputs an alarm signal, the following parallel operations are triggered: The sound and light control thread immediately sets the GPIO: Yellow warning → GPIO18 high level (LED yellow), GPIO19 outputs a 1kHz 50% duty cycle square wave; Red warning → GPIO17 high level (LED red), GPIO19 outputs a 1.5kHz continuous square wave.
[0051] The message push thread sends JSON messages to the nurse station server via the MQTT protocol, with QoS=1.
[0052] The log thread writes alert events to the SQLite database, including timestamps, risk probabilities, and risk types.
[0053] Nurse confirmation interaction: A full-screen warning card pops up on the touchscreen, displaying "Medium / High Risk - Immediate Review Recommended." After the nurse clicks "Confirm," the system stores the review time and operator ID in the database. If "Reject (False Positive)" is clicked, the system requires the nurse to fill in the reason for rejection (drop-down options: sensor error, non-disease factors of the patient, model misjudgment), and the data is also recorded.
[0054] Inhibition mechanism: For the same patient, if the nurse has confirmed the same type of risk (hemorrhagic transformation or neurological deterioration) and it falls within the 10-minute suppression window, subsequent warnings will not trigger audio-visual alerts (only log entries will be made) to avoid repeated interference. The suppression window can be set via touchscreen (default 10 minutes).
[0055] III. How to Use the Robot The robot proposed in this application can achieve contactless, continuous monitoring and semi-automated assessment of neurological function, and provide real-time early warning of complications. Its consistency with early warnings from emergency / stroke specialist nurses (Kappa ≥ 0.8), the intraclass correlation coefficient (ICC) between NIHSS assessment and manual scoring by attending neurologists (ICC ≥ 0.75), and the robot handles ≥ 80% of the workload for single-patient monitoring. Its typical usage is as follows: 1. Equipment deployment and activation: When thrombolysis begins for the patient, the nurse deploys the robot host at the bedside, attaches wearable physiological patches to the patient, and starts the monitoring program, which lasts for 24 hours.
[0056] 2. Continuous monitoring: The robot performs non-contact continuous monitoring of patients through a millimeter-wave radar-vision fusion module, collecting indicators such as respiratory rate, heart rate, body movement, pupil size and light reflex, and fall detection in real time; wearable patches simultaneously collect physiological parameters such as electrocardiogram, blood oxygen saturation, and skin temperature, and all data are transmitted back to the edge computing box for real-time analysis at a rate of one second.
[0057] 3. Semi-automated neurological function assessment: At the monitoring time points recommended by the guidelines, the robot actively asks patients questions through voice interaction (naming, repetition, orientation) and performs semi-automated scoring based on NIHSS items such as facial paralysis, gaze, and limb movement based on a multimodal large model, outputting auxiliary assessment results for nurses to review.
[0058] 4. Intelligent early warning: The edge AI model identifies abnormal vital signs, deterioration of neurological function, and prodromal symptoms of complications in real time based on multimodal data. When the alarm conditions are met, an alarm is triggered and pushed synchronously through the bedside host screen, sound and light, and the nurse's mobile terminal.
[0059] 5. Nurse-assisted verification and response: After receiving the alert, the nurse performs bedside verification according to the standard operating procedure and confirms / rejects the robot alert via the touch screen. The labeled data is used for subsequent model iteration.
[0060] 6. Data Traceability: All raw monitoring data, early warning events, and nurse response links are automatically recorded in the local database. The raw data does not leave the hospital; only the training gradients participate in federated learning.
[0061] IV. Method Examples Figure 2 This document presents an overall flowchart of the monitoring and early warning method for post-thrombolysis monitoring in acute ischemic stroke according to this embodiment. The following is in conjunction with... Figure 2 , Figure 3 , Figure 4 The methodology is explained in detail. Each step is supplemented with specific parameters and algorithm details.
[0062] Step S101: Non-contact acquisition of thoracic cavity micro-motion echo signals and calculation of heart rate and respiratory rate.
[0063] The intermediate frequency signal is transmitted and received via a millimeter-wave radar module 110 in FMCW mode. A 2D-FFT (256 points in the range dimension, 128 points in the velocity dimension) is performed on the edge NPU to extract the phase sequence. Variational mode decomposition (VMD) is used to decompose the phase sequence into respiratory and heart rate components. The instantaneous frequency of each component is calculated using zero-crossing detection, and then smoothed by median filtering (5-second window). The final output is heart rate (beats / minute) and respiratory rate (breaths / minute), updated every 0.5 seconds. If the confidence level is below a threshold (e.g., excessive signal amplitude fluctuation), the data is marked as invalid, triggering a sensor status check.
[0064] Step S102: Simultaneously acquire facial expression image sequences, limb movement videos, and voice response audio.
[0065] The binocular vision module 120 captures data at 30fps, buffering the most recent 30 seconds in a circular buffer. The microphone array 130 continuously records audio in a circular buffer of 2 minutes. When NIHSS evaluation is required, the system sends the instruction "Please cooperate with the evaluation," followed by step-by-step prompts. Video and audio are segmented into independent segments based on the prompts, each segment being 5-10 seconds long and timestamped.
[0066] Step S103: Spatiotemporal alignment and feature fusion of multimodal data.
[0067] Based on radar timestamps, visual and audio data are linearly interpolated to the same timeline (2Hz). A temporal convolutional network based on attention is used for fusion, with the following structure: input shape (batch, 120, 15) → three dilated convolutional layers (3 kernels, dilation rates 1, 2, 4) → multi-head attention (4 heads) → global average pooling → output 64-dimensional fused feature vector. The network weights were pre-trained offline using labeled data.
[0068] Step S104: Automatically output NIHSS score based on multimodal large model.
[0069] The 64-dimensional fused features generated in step S103 and the speech audio features (768 dimensions extracted by Wav2Vec) are used as input to the multimodal large model. The model uses a fine-tuned version of Video-LLaMA, quantized to int8 before deployment. During inference, the model generates text via autoregression, for example, "Facial paralysis: 1; Left upper limb movement: 2; Right upper limb movement: 0; Dysarthria: 1; Language: 1; Gaze: 0; Visual field: 0; Ataxia: 0; Sensation: 0; Neglect: 0". The parser separates each item by a semicolon, verifies the scoring range, and invalidates items outside the range. The output also includes confidence scores (obtained from the model's logits using softmax).
[0070] Step S105: Display the score, provide an entry point for manual review and error correction, and obtain the final score.
[0071] The touchscreen displays all sub-items in a table format, with each item showing the automatic score and confidence level (indicated by color bars). A numeric keypad pops up when the nurse clicks to modify. The system monitors touchscreen events in real time; when a modification occurs, the corrected score overwrites the automatic score and stores (original score, corrected score, timestamp, nurse ID). If there is no human intervention within 30 seconds, the automatic score is automatically used as the final score.
[0072] Step S106: Real-time identification of abnormal vital signs, deterioration of neurological function, and prodromal symptoms of complications, and determination of alarm conditions.
[0073] This step is the core early warning logic. On the edge NPU, a separate background thread reads all fused temporal features from the last 5 minutes (one feature vector per minute, for a total of 5) and the latest NIHSS final scores (the last 3 times, updated every 15 minutes) from shared memory every 5 seconds (configurable), forming a 32-dimensional input feature vector. These features are then fed into the XGBoost multi-output model (loaded via ONNX Runtime). The model outputs three probability values: P_vitals (abnormal vital signs), P_neuro (deterioration of neurological function), and P_complication (prodromal symptoms of complications). Simultaneously, the system maintains a lightweight rule engine. If the heart rate increases by more than 15% within 10 minutes and the NIHSS score increases by ≥2, then force P_neuro = 1.0; If a periodic oscillation of 0.5-2Hz (amplitude > 0.2g) appears in the body acceleration signal, then force P_complication = 1.0; If blood oxygen saturation is <92% for 10 seconds, then force P_vitals = 1.0.
[0074] Final alarm condition determination: If max(P_vitals, P_neuro, P_complication) ≥ 0.7, or any probability ≥ 0.5 and lasts for more than 10 minutes (both consecutive inferences are ≥ 0.5), then a red warning signal (high risk) will be output. If 0.3 ≤ max(...) < 0.7, then a yellow warning signal (medium risk) will be output. Otherwise, there will be no warning signal.
[0075] All calculations are completed within 2ms. The warning signal also carries the abnormality type (such as "deterioration of neurological function", "prodromal symptoms of complications - suspected epilepsy", etc.).
[0076] Step S107: Trigger an audible and visual warning.
[0077] The warning signal is transmitted to the device control thread via a message queue. The control thread then executes the following based on the signal type: Yellow alert: Set GPIO18=1 (yellow light on), and simultaneously start PWM output at 1kHz (GPIO19), with a duty cycle of 50% and a period of 500ms. Also start a timer that toggles GPIO18 every 500ms.
[0078] Red alert: Set GPIO17=1 (red light on), and simultaneously output a 1.5kHz continuous waveform with PWM, and GPIO19 continuously outputs.
[0079] No warning: Disable all GPIO outputs and stop PWM.
[0080] This control thread has the highest priority (real-time thread), ensuring a response latency of <10ms.
[0081] Simultaneously, the warning signal triggers a push notification: a message is published via MQTT to the topic "alert / patient / {patient_id}". The message body is in JSON format and includes: alert_level ("medium" or "high"), risk_type ("hemorrhagic" or "neurological"), probability, and timestamp. The nurses' station, having subscribed to this topic, will receive the notification and a pop-up notification box will appear, along with an alarm sound (separate from the bedside alarm).
[0082] Local execution and response time guarantee: The end-to-end latency of the entire process from physiological signal change to early warning trigger (i.e. from radar data refresh to GPIO change) is 2.7 seconds on average and 3.2 seconds on maximum, as tested in actual tests. This is less than the preset 3-second threshold and meets clinical requirements.
[0083] Extended steps (wearable patch data fusion): such as Figure 3 As shown, steps S201-S202 are executed in parallel after step S101. In step S201, ECG, blood oxygen, skin temperature, and body movement data are read from the Bluetooth receiver every second, and after median filtering to reduce noise, the sampling rate is reduced to 2Hz. In step S202, these data are concatenated with the fused temporal features generated in step S103 to form an extended 40-dimensional feature vector, which is then fed into the XGBoost model in step S106. Since the model already includes the weights corresponding to these features during training, the accuracy of the warning is further improved.
[0084] Extended steps (Federated learning model update): such as Figure 4 As shown, the edge NPU sends a request to the central server to participate in federated learning every two weeks. Step S301: The server distributes the current global model parameters (XGBoost model JSON file). Step S302: The edge device loads the global model and performs incremental training using all patient data from the past two weeks (including raw features and nurse-corrected labels) stored locally, employing XGBoost's incremental learning function (by setting process_type=update and num_boosted_rounds=50). Step S303: The gradient difference between the new model and the global model is calculated, encrypted, and uploaded via HTTPS. Step S304: The server aggregates the uploaded gradients from all devices, updates the global model, and distributes it to the edge device, replacing the local model. This mechanism allows the model to continuously optimize without leaking the original data.
[0085] V. Examples of Specific Application Scenarios Scenario 1: Early Warning of Sudden Neurological Deterioration 2 Hours After Thrombolysis. Patient Wang, male, 65 years old, was transferred to the ICU after intravenous thrombolysis for acute ischemic stroke. The robot continued operating. Two hours after thrombolysis, millimeter-wave radar detected an increase in heart rate from 85 bpm to 112 bpm and respiratory rate from 16 breaths / min to 24 breaths / min. Simultaneously, the system automatically initiated the NIHSS assessment program, guiding the patient to execute "raise arm," "smile," and "speak" commands via binocular vision and microphone. Multimodal large-scale model analysis revealed a 50% decrease in the patient's left limb elevation range compared to baseline, an increase in facial paralysis score from 0 to 2, and an increase in dysarthria score from 0 to 1, automatically increasing the NIHSS total score by 4 points. The complication risk warning model, based on real-time multimodal data stream calculations, output a high risk of neurological deterioration (92% probability), triggering a red audible and visual warning and sending it to the nurses' station. The nurse arrived at the bedside within 2 minutes, manually verified the changes in the patient's condition, immediately contacted the doctor for an emergency CT scan, which revealed symptomatic intracranial hemorrhage, and provided timely symptomatic treatment. In this scenario, the robot detected signs of deterioration approximately 15 minutes earlier than traditional manual monitoring.
[0086] Scenario 2: Continuous monitoring without blind spots during field inspections Patient Li needed to go to the CT room for a follow-up examination 4 hours after thrombolysis. Traditional monitors are limited by power supply and cables and cannot be moved with the patient. This robot's bedside main unit is equipped with casters and a built-in battery, allowing it to move with the trolley; wearable patches maintain their attachment. During transport, millimeter-wave radar and patches continuously collect vital signs, and the binocular vision module automatically tracks the patient's face. When the patient experienced a brief drop in heart rate due to jolting, the robot analyzed the situation in real time and determined it to be low-risk, recording the data without triggering an alarm, thus avoiding false alarms. Upon arrival at the CT room, the robot automatically docked with the CT room nursing station to continue monitoring.
[0087] Scenario 3: Non-intrusive nighttime monitoring and low-risk anomaly identification Patient Zhang, 10 hours after thrombolysis, was asleep. Traditional nurse rounds require turning on lights and touching the patient, often disrupting their rest. This robot uses millimeter-wave radar to penetrate bedding to monitor respiration and heart rate; its binocular vision uses infrared night vision mode, eliminating the need for artificial lighting. At 2:00 AM, the system detected a gradual decrease in respiratory rate from 14 breaths / min to 8 breaths / min and a decrease in heart rate from 70 beats / min to 55 beats / min. The warning model assessed this as low risk (likely a normal change during the sleep cycle), and no audible or visual alarms were triggered; the data was simply recorded and silently displayed on the nurses' station interface.
[0088] Scenario 4: Multi-center Federated Learning Collaborative Optimization Model The robot described in this application has been deployed in the stroke centers of three hospitals in a certain city. Each center's robot periodically updates its local model gradient using its own patient data and uploads the encrypted gradient to the regional center server via a federated learning client. The server then aggregates and distributes the global model. After six months of collaborative training, the global early warning model's prediction accuracy for hemorrhage transformation improved from the initial 82% to 91%, and the centers do not need to share the original patient data, fully complying with medical privacy regulations.
[0089] The robot and method provided in this application can be widely applied in hospitals at all levels with stroke centers or emergency thrombolysis rooms, and are particularly suitable for neurological intensive care units, stroke units, and emergency observation rooms. The system hardware cost is controllable, and the software model can be continuously iterated, making mass production and clinical application feasible. An engineering prototype has been developed in accordance with the Good Manufacturing Practice for Medical Devices (YY / T0287) and has passed type testing (electrical safety, electromagnetic compatibility, and biocompatibility). It is expected to effectively address clinical pain points such as insufficient manpower for post-thrombolysis monitoring in AIS systems, numerous blind spots, and delayed early warning, demonstrating significant social and economic benefits.
[0090] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
Claims
1. A robot for monitoring and early warning after thrombolysis in acute ischemic stroke, characterized in that, include: The bedside unit integrates: Millimeter-wave radar module for non-contact acquisition of micro-motion echo signals in the patient's chest cavity; A binocular vision module and a directional microphone array are used to simultaneously acquire patient facial expression image sequences, limb movement videos, and voice response audio; Edge NPU computing unit; Touchscreen and audible / visual warning devices; And the following modules deployed on the edge NPU computing unit: The multimodal fusion sensing module is used to extract heart rate data and respiratory rate data from the echo signal, and to perform spatiotemporal alignment and feature fusion on the image sequence and motion video to generate fused temporal features; A multimodal large model evaluation module, which is optimized through instruction fine-tuning and human feedback reinforcement learning, is used to jointly analyze the fused temporal features and speech response audio, and automatically output NIHSS scores; The complication risk warning module has a built-in AI model based on multimodal data pre-training. It is used to analyze the fusion time sequence features output by the multimodal fusion perception module and the NIHSS score output by the multimodal large model evaluation module in real time. It can identify abnormal vital signs, deterioration of neurological function, and prodromal symptoms of complications in real time, and determine whether the alarm conditions are met. When the alarm conditions are met, an alarm signal is output. The human-machine collaboration module is used to display the NIHSS score on the touch screen and provide an entry point for manual review and error correction. When a corrected score is received from manual input, the corrected score is used as the final score. The module also triggers the audible and visual warning device according to the alarm signal. The edge NPU computing unit is configured to perform all data processing locally, the original patient data does not leave the hospital's local network, and the end-to-end warning response time is less than the preset duration.
2. The robot according to claim 1, characterized in that, The intragroup correlation coefficient between the NIHSS score output by the multimodal large model evaluation module and the score of the attending neurologist is not less than 0.75; and the human-machine collaboration module is also used to store the original output of the multimodal large model evaluation module and the corrected score input by humans together as training samples for subsequent model optimization.
3. The robot according to claim 1, characterized in that, The inputs to the complication risk warning module also include coagulation function test indicators from the patient's electronic medical record system; the training data of the AI model includes the fused temporal features, NIHSS score, and the coagulation function test indicators.
4. The robot according to claim 1, characterized in that, It also includes a wearable physiological patch; the wearable physiological patch integrates a single-lead ECG, blood oxygen, skin temperature and body accelerometer module, and communicates with the bedside host via Bluetooth Low Energy; the multimodal fusion sensing module is also used to fuse the data collected by the wearable physiological patch with the heart rate data and respiratory rate data.
5. The robot according to claim 1, characterized in that, The edge NPU computing unit is also equipped with a federated learning client. The federated learning client is used to: receive encrypted aggregated parameters of the multi-center global model, update the local model with local patient data and then upload the encrypted model gradient, without uploading any original patient data, and download the updated global model parameters to replace the AI model in the complication risk warning module.
6. The robot according to claim 1, characterized in that, The human-machine collaboration module is also used to: when receiving the alarm signal, push an early warning message containing the risk type and patient bed to the nurse station monitoring platform, receive manual confirmation instructions from the nurse station monitoring platform or the touch screen, record the manual review time and review results, and store the alarm event, early warning trigger time, manual confirmation instructions and subsequent clinical intervention records into the electronic monitoring file.
7. The robot according to claim 1, characterized in that, The robot's alarm results have a Kappa consistency coefficient of no less than 0.8 with the clinical judgment of emergency or stroke specialist nurses, and the robot undertakes more than 80% of the routine monitoring workload for a single patient within 24 hours.
8. A method for monitoring and early warning after thrombolysis in acute ischemic stroke, characterized in that, include: The millimeter-wave radar deployed at the bedside host non-contactly acquires the patient's chest cavity micro-motion echo signal, and the heart rate and respiratory rate data are calculated from the chest cavity micro-motion echo signal on the edge NPU computing unit deployed at the bedside host. The system simultaneously acquires patient facial expression image sequences, limb movement videos, and voice response audio through a binocular vision module and directional microphone array deployed on the bedside host. On the edge NPU computing unit, the heart rate data, respiratory rate data, facial expression image sequence, and limb movement video are spatiotemporally aligned and multimodal feature fused to generate fused temporal features; Using a multimodal large model optimized by instruction fine-tuning and human feedback reinforcement learning, the fused temporal features and the voice response audio are jointly analyzed on the edge NPU computing unit to automatically output the NIHSS score; The NIHSS score is displayed on the touchscreen of the bedside unit and a manual review and correction entry is provided; when a corrected score is received manually, the corrected score is used as the final score. Using the complication risk warning model deployed on the edge NPU computing unit, based on the fused temporal features and the final score, abnormal vital signs, deterioration of neurological function, and prodromal symptoms of complications are identified in real time, and it is determined whether the preset alarm conditions are met. When the alarm conditions are met, the bedside host is triggered with an audible and visual alarm signal. All of the above steps are executed locally on the bedside host, the original patient data does not leave the hospital's local network, and the end-to-end warning response time is less than the preset duration.
9. The method according to claim 8, characterized in that, Also includes: The patient's single-lead electrocardiogram signal, blood oxygen saturation, skin temperature, and body acceleration data are collected through wearable physiological patches; The single-lead ECG signal, blood oxygen saturation, skin temperature, and body acceleration data are fused with the fusion time-series features and then input into the complication risk warning model.
10. The method according to claim 8, characterized in that, Also includes: The encrypted aggregation parameters of the multi-center global model are received through the federated learning client deployed on the bedside host. The local patient data is used to update the local model in the federated learning client, and the encrypted model gradient is uploaded to the aggregation server, without uploading any original patient data. Download the updated global model parameters from the aggregation server to replace the local complication risk warning model.