Intelligent generation system for severe multi-modal data perception with semantic reasoning capability
Patent Information
- Application Number
- CN202610741361.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-27
- Publication Date
- 2026-09-25
AI Technical Summary
[0004]在现有技术中,姿态识别通常依赖RGB图像进行分析,未能充分利用深度信息提升识别精度;且缺乏面部表情与情绪识别
[0019]与现有技术相比,本发明的有益效果是:本发明提出的具备语义推理能力的重症多模态数据感知与智能生成系统,通过多模态数据的高效采集与处理、智能分析及异常行为预警,能够全面监测患者的姿态、情绪及周围环境,不仅解决了现有技术中的同步性差、分析方法局限和异常行为检测不敏感的问题,并提供实时预警,还能够为医护人员提供可视化、直观的数据支持,提升患者护理的质量与安全性。
Smart Images

Figure CN122817682A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image data analysis technology, specifically a critical care multimodal data perception and intelligent generation system with semantic reasoning capabilities. Background Technology
[0002] Image data analysis technology involves extracting useful information from images, covering multiple areas such as image processing, computer vision, image classification, and object detection.
[0003] The ICU (Intensive Care Unit) is a crucial department in a hospital for the centralized monitoring and treatment of critically ill patients. In the ICU environment, patients typically require prolonged bed rest and rely on various monitoring devices to continuously monitor their vital signs so that medical staff can promptly grasp changes in their condition.
[0004] In existing technologies, posture recognition typically relies on RGB image analysis, failing to fully utilize depth information to improve recognition accuracy; and it lacks facial expression and emotion recognition. Meanwhile, existing abnormal behavior detection methods largely depend on physiological signals such as heart rate and blood oxygenation, lacking a multi-dimensional comprehensive analysis mechanism that combines posture changes, behavioral characteristics, and emotional information, resulting in weak real-time recognition capabilities for sudden patient behaviors (such as struggling, strenuous exercise, etc.).
[0005] Furthermore, existing ICU monitoring systems are mostly based on a single data source, such as monitoring methods based on physiological signals or video analysis, lacking the ability to comprehensively understand multimodal data. When inconsistencies or even semantic conflicts occur between different data sources, traditional systems often fail to identify such cross-modal semantic conflicts. Summary of the Invention
[0006] The purpose of this invention is to provide a critical care multimodal data perception and intelligent generation system with semantic reasoning capabilities to solve the problems raised in the prior art.
[0007] To achieve the above objectives, the present invention provides the following technical solution: a critical care multimodal data perception and intelligent generation system with semantic reasoning capabilities, the system comprising a data acquisition module, a preprocessing module, an analysis module, an anomaly detection module, a data storage module, and a visualization management module; The data acquisition module is used to acquire multimodal image data in real time, including RGB images, depth data and thermal imaging images; Among these measures, recording timestamps ensures data synchronization in time sequence; The preprocessing module is used to process the collected multimodal data to ensure data consistency; this includes data synchronization, noise reduction, and standardization. The analysis module is used to analyze the preprocessed data, including posture and body position recognition, facial expression and emotion analysis, and environmental perception analysis. The anomaly detection module is used to detect abnormal behavior based on a preset threshold and issue an early warning in real time; it also automatically notifies nursing staff. The data storage module is used to store the collected multimodal data, analysis results, and intelligent analysis reports.
[0008] The visualization management module is used to visualize the data and provide a user interface.
[0009] Furthermore, specifically: integrating multimodal image acquisition equipment, which includes an RGB camera, a depth camera, and a thermal imager; ensuring that the equipment covers the entire patient bedside area.
[0010] Calibrate the equipment to eliminate data bias and set the acquisition parameters; The acquisition parameters include frame rate, resolution, and sampling frequency; Start the image acquisition device to capture RGB images, depth data, and thermal imaging data in real time; During the data acquisition process, timestamps are recorded to ensure the synchronization of multimodal data.
[0011] The acquired multimodal image data is preprocessed as follows: The preprocessing includes data synchronization, data denoising, and data standardization.
[0012] Specifically, the data denoising includes: RGB images are smoothed and noise is removed by applying Gaussian or bilateral filtering, and color deviations caused by changes in lighting conditions are corrected. Median filtering is used to fill in missing point cloud data and remove outliers from depth data; Thermal imaging data: The temperature distribution map is smoothed to reduce the impact of sensor noise and remove false hotspots that may be caused by reflection or equipment error.
[0013] Specifically, the data standardization refers to: RGB image: The pixel values are normalized and scaled to the range of [0,1]. Depth data: Standardize the range of depth values using z-score to adapt to the input requirements of subsequent algorithms; Thermal imaging data: Temperature values are linearly normalized; Furthermore, specifically: adjust the time axis of different modal data uniformly according to the device timestamp to ensure the timing matching of the data; Based on the equipment installation location and viewing angle, geometric correction is performed using calibration data to map data from different modes to the same coordinate system.
[0014] Further analysis is performed based on the preprocessed data, including posture and body position recognition, facial expression and emotion analysis, and environmental perception analysis, specifically: The posture and body position recognition specifically refers to: Using a standardized RGB image, the user's appearance and facial features are obtained; The user's spatial position and posture changes are obtained through standardized depth data; Pose estimation is performed using a deep learning model, and keypoint locations are extracted. The set of keypoint locations is denoted as {p1, p2, ..., p...}. n}, p i ∈{p1,p2,…,p n}, p i =(x i ,y i ,z i ); where n represents the total number of keypoints, and n is a positive integer; i represents the number of keypoint labels, and i is a positive integer, 1≤i≤n; p i It is an element in the set of key point locations, p i Represents the position coordinates of the i-th key point; (x i ,y i ,z i ) represents the coordinates of the i-th key point in three-dimensional space; x i The x-coordinate of the i-th key point is represented by y. i z represents the ordinate of the i-th key point. i Represents the vertical coordinate of the i-th key point; The deep learning model can be selected from OpenPose, HRNet, etc. The user's body position is determined based on the coordinates of key points in three-dimensional space. The body position includes supine, lateral, and prone positions. The determination of body position is achieved by calculating coordinate differences and angles. Where p1, p2, and p3 are the position coordinates of three key points; θ represents the angle formed by p1, p2, and p3. This method is used to calculate the angles of the head, torso, and limbs in sequence to determine body position; The facial expression and emotion analysis specifically includes: Use standardized RGB images to obtain the user's facial expression information; The three-dimensional structure and feature information of the user's face are obtained by using standardized depth data; Facial feature points are extracted using a facial expression recognition model. These feature points include the geometric shapes of the eyes, eyebrows, and mouth. The set of facial feature points is {f1, f2, ..., f...}. m}, f j ∈{f1,f2,…,f m}, f j =(x j ,y j ); where m represents the total number of facial feature points, m is a positive integer; j represents the number label of facial feature points, j is a positive integer, 1≤j≤m; f j It is an element in the set of facial feature points, f j Represents the position coordinates of the j-th facial feature point; (x j ,y j ) represents the two-dimensional coordinates of the j-th facial feature point in the image; x j The x-coordinate of the j-th facial feature point is represented by y. j This represents the ordinate of the j-th facial feature point; The facial expression recognition model can be FELNet, VGG-Face, etc. A convolutional neural network (CNN) is used to classify facial expressions into emotions, where emotion labels include happiness, anger, sadness, surprise, etc. The output of the emotion classification model is set as follows: Emotion=argmax(softmax(f(features))); Where f(features) represents the features extracted from facial feature points; Emotion represents the emotion classification obtained through a convolutional neural network; Where softmax represents the state transition function; argmax represents the state selection function; Other classification models, such as LSTM or RNN, can also be used to classify facial expressions as emotions. Optionally, the environmental perception analysis specifically includes: Use standardized thermal imaging data to obtain information about the surrounding temperature distribution; The spatial layout of the user's bedside environment is obtained through standardized depth data; In thermal imaging data, pixel values are T(x,y), where T represents the temperature value and (x,y) represents the pixel coordinates. The average temperature and standard deviation of the user's area are calculated. Where N represents the total number of pixel values in the thermal imaging data, and N is a positive integer; This indicates the average temperature of the user's location. ; where σT This represents the standard deviation of temperature in the user's region.
[0015] Furthermore, set thresholds for abnormal behavior, specifically as follows: Install early warning devices; Set the first attitude change threshold θ'; The first posture change threshold represents the angle of the user's posture change, which varies in different situations and should be set according to the user's actual situation. Set the second attitude change threshold t'; The second posture change threshold represents the time a user maintains a posture, and it can be set according to the actual situation. When the angle change between key points exceeds the first posture change threshold θ', the early warning device will automatically issue an alarm to prompt the nursing staff to adjust the posture. When a user maintains a certain posture for more than the second posture change threshold t', the early warning device will automatically issue an alarm to prompt the caregiver to adjust the posture. Set a third temperature threshold range [t1, t2]; The third temperature threshold range is usually taken as [22, 28]; Set a fourth temperature change threshold Δt; The fourth temperature change threshold is usually set at 2-3 degrees Celsius; When the temperature in the user's area is outside the third temperature threshold range [t1, t2], the early warning device will automatically issue an alarm to prompt caregivers to adjust the ambient temperature. When the temperature standard deviation of the user's area exceeds the fourth temperature change threshold Δt, the early warning device will automatically issue an alarm to prompt caregivers to adjust the ambient temperature.
[0016] Furthermore, specifically: First, the emotional data obtained from facial expression analysis, voice analysis, or other physiological signals (such as heart rate, skin conductance response, etc.) needs to be converted into a time series format. The emotional value at each moment (such as happiness, anger, anxiety, etc.) will be assigned a value according to a certain emotional model. The emotion probability distribution is obtained through the emotion model; The emotion model includes the Ekman emotion model and the Plutchik emotion wheel; Specifically, the emotion probability distribution is as follows: Among them, E t It is the probability distribution of emotions at time t, where P represents the probability value of different emotional states; The LSTM model is trained to obtain the user's emotion prediction value by real-time monitoring of the user's facial expression information and the three-dimensional structure and feature information of the user's face. Calculate the amplitude of mood fluctuations: Where t represents time; Time represents the length of the observation time window; This represents the average emotional value. This represents the predicted sentiment value at time t; Volatility represents the fluctuation of sentiment. Calculate the rate of change in mood: ; where ΔE t Indicates the rate of change in mood; This represents the predicted sentiment value at time t-1; Set the threshold for mood changes ΔE; When the rate of change in emotion ΔE t If the emotional change exceeds the threshold ΔE, the warning device will automatically issue an alarm, prompting caregivers to check on the user.
[0017] Furthermore, the real-time monitored multimodal image data is displayed in real time through a visualization screen; the entire process data of posture and body position recognition, facial expression and emotion analysis, and environmental perception analysis is backed up; The warning results will be displayed in real time through a visual screen; Generate intelligent analysis reports, including user emotional fluctuations, posture changes, environmental changes, warning times, and the handling process of relevant personnel.
[0018] Furthermore, the visualization management module includes a disease viewing unit and an intelligent document unit; The condition viewing unit is used to display monitoring results, including multimodal image data recorded according to timestamps; The intelligent document unit is used to display intelligent analysis reports for medical staff to view.
[0019] Compared with existing technologies, the beneficial effects of this invention are as follows: The critical care multimodal data perception and intelligent generation system with semantic reasoning capabilities proposed in this invention can comprehensively monitor the patient's posture, emotions, and surrounding environment through efficient collection and processing of multimodal data, intelligent analysis, and abnormal behavior early warning. It not only solves the problems of poor synchronization, limited analysis methods, and insensitivity to abnormal behavior detection in existing technologies, but also provides real-time early warning and provides medical staff with visualized and intuitive data support, thereby improving the quality and safety of patient care. Attached Figure Description
[0020] Figure 1This is a flowchart illustrating a critical care multimodal data perception and intelligent generation system with semantic reasoning capabilities according to the present invention. Detailed Implementation
[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0022] Example: Figure 1 As shown, the present invention provides a technical solution: a critical care multimodal data perception and intelligent generation system with semantic reasoning capabilities. The system includes a data acquisition module, a preprocessing module, an analysis module, an anomaly detection module, a data storage module, and a visualization management module. The data acquisition module is used to acquire multimodal image data in real time, including RGB images, depth data, and thermal images; Among these measures, recording timestamps ensures data synchronization in time sequence; The preprocessing module is used to process the collected multimodal data to ensure data consistency; this includes data synchronization, noise reduction, and standardization. The analysis module is used to analyze the preprocessed data, including posture and body position recognition, facial expression and emotion analysis, and environmental perception analysis. The anomaly detection module is used to detect abnormal behavior based on a preset threshold and issue an early warning in real time; it also automatically notifies nursing staff. The data storage module is used to store the collected multimodal data, analysis results, and intelligent analysis reports.
[0023] The visualization management module is used to visualize the data and provide a user interface.
[0024] The acquisition of multimodal image data includes: Install a multimodal image acquisition device, which includes an RGB camera, a depth camera, and a thermal imager; ensure that the device covers the entire patient bedside area.
[0025] Calibrate the equipment to eliminate data bias and set the acquisition parameters; The acquisition parameters include frame rate, resolution, and sampling frequency; Start the image acquisition device to capture RGB images, depth data, and thermal imaging data in real time; During the data acquisition process, timestamps are recorded to ensure the synchronization of multimodal data.
[0026] The acquired multimodal image data is preprocessed as follows: The preprocessing includes data synchronization, data denoising, and data standardization.
[0027] Specifically, the data denoising includes: RGB images are smoothed and noise is removed by applying Gaussian or bilateral filtering, and color deviations caused by changes in lighting conditions are corrected. Median filtering is used to fill in missing point cloud data and remove outliers from depth data; Thermal imaging data: The temperature distribution map is smoothed to reduce the impact of sensor noise and remove false hotspots that may be caused by reflection or equipment error.
[0028] Specifically, the data standardization refers to: RGB image: The pixel values are normalized and scaled to the range of [0,1]. Depth data: Standardize the range of depth values using z-score to adapt to the input requirements of subsequent algorithms; Thermal imaging data: Temperature values are linearly normalized; The data synchronization specifically refers to: Based on the device's timestamp, the time axis of different modal data is uniformly adjusted to ensure data timing matching. Based on the equipment installation location and viewing angle, geometric correction is performed using calibration data to map data from different modes to the same coordinate system.
[0029] Analysis is performed based on preprocessed data, including posture and body position recognition, facial expression and emotion analysis, and environmental perception analysis, specifically: The posture and body position recognition specifically refers to: Using a standardized RGB image, the user's appearance and facial features are obtained; The user's spatial position and posture changes are obtained through standardized depth data; Pose estimation is performed using a deep learning model, and keypoint locations are extracted. The set of keypoint locations is denoted as {p1, p2, ..., p...}. n}, p i ∈{p1,p2,…,p n}, p i =(x i ,y i ,z i ); where n represents the total number of keypoints, and n is a positive integer; i represents the number of keypoint labels, and i is a positive integer, 1≤i≤n; p i It is an element in the set of key point locations, p iRepresents the position coordinates of the i-th key point; (x i ,y i ,z i ) represents the coordinates of the i-th key point in three-dimensional space; x i The x-coordinate of the i-th key point is represented by y. i z represents the ordinate of the i-th key point. i Represents the vertical coordinate of the i-th key point; The deep learning model can be selected from OpenPose, HRNet, etc. The user's body position is determined based on the coordinates of key points in three-dimensional space. The body position includes supine, lateral, and prone positions. The determination of body position is achieved by calculating coordinate differences and angles. Where p1, p2, and p3 are the position coordinates of three key points; θ represents the angle formed by p1, p2, and p3. In this embodiment, the pose determination of multimodal image data is implemented to test its accuracy. Specifically: This method is used to calculate the angles of the head, torso, and limbs in sequence to determine body position; Among them, the head: p1=(0,0,1.6); Chest: p2=(0,0,1.4); Waist region: p3=(0,0,1.2); Calculate the direction vector: p2-p1=(0,0,-0.2); p3-p1=(0,0,-0.4); ∥p2-p1∥=0.2; ∥p3-p1∥=0.4; The calculation yields θ = 0°; Since p1>p2>p3, the user is in a supine position. Facial expression and emotion analysis specifically includes: Use standardized RGB images to obtain the user's facial expression information; The three-dimensional structure and feature information of the user's face are obtained by using standardized depth data; Facial feature points are extracted using a facial expression recognition model. These feature points include the geometric shapes of the eyes, eyebrows, and mouth; the set of facial feature points is {f1, f2, ..., f...}. m}, f j ∈{f1,f2,…,f m}, f j =(x j ,y j); where m represents the total number of facial feature points, m is a positive integer; j represents the number label of facial feature points, j is a positive integer, 1≤j≤m; f j It is an element in the set of facial feature points, f j Represents the position coordinates of the j-th facial feature point; (x j ,y j ) represents the two-dimensional coordinates of the j-th facial feature point in the image; x j The x-coordinate of the j-th facial feature point is represented by y. j This represents the ordinate of the j-th facial feature point; Among them, facial expression recognition models can be FELNet, VGG-Face, etc.; A convolutional neural network (CNN) is used to classify facial expressions into emotions, where emotion labels include happiness, anger, sadness, surprise, etc. The output of the emotion classification model is set as follows: Emotion=argmax(softmax(f(features))); Where f(features) represents the features extracted from facial feature points; Emotion represents the emotion classification obtained through a convolutional neural network; Where softmax represents the state transition function; argmax represents the state selection function; Other classification models, such as LSTM or RNN, can also be used to classify facial expressions as emotions. The environmental perception analysis specifically includes: Use standardized thermal imaging data to obtain information about the surrounding temperature distribution; The spatial layout of the user's bedside environment is obtained through standardized depth data; In thermal imaging data, pixel values are T(x,y), where T represents the temperature value and (x,y) represents the pixel coordinates. The average temperature and standard deviation of the user's area are calculated. Where N represents the total number of pixel values in the thermal imaging data, and N is a positive integer; This indicates the average temperature of the user's location. ; where σ T This represents the standard deviation of temperature in the user's region.
[0030] The setting of abnormal behavior thresholds specifically involves: Install early warning devices; Set the first attitude change threshold θ'; The first posture change threshold represents the angle of the user's posture change, which varies in different situations and should be set according to the user's actual situation. Set the second attitude change threshold t'; The second posture change threshold represents the time a user maintains a posture, and it can be set according to the actual situation. When the angle change between key points exceeds the first posture change threshold θ', the early warning device will automatically issue an alarm to prompt the nursing staff to adjust the posture. When a user maintains a certain posture for more than the second posture change threshold t', the early warning device will automatically issue an alarm to prompt the caregiver to adjust the posture. Set a third temperature threshold range [t1, t2]; The third temperature threshold range is usually taken as [22, 28]; Set a fourth temperature change threshold Δt; The fourth temperature change threshold is usually set at 2-3 degrees Celsius; When the temperature in the user's area is outside the third temperature threshold range [t1, t2], the early warning device will automatically issue an alarm to prompt caregivers to adjust the ambient temperature. When the temperature standard deviation of the user's area exceeds the fourth temperature change threshold Δt, the early warning device will automatically issue an alarm to prompt caregivers to adjust the ambient temperature.
[0031] Specifically, the process involves first converting emotional data obtained from facial expression analysis, voice analysis, or other physiological signals (such as heart rate and skin conductance) into a time series format. The emotional value at each moment (such as happiness, anger, anxiety, etc.) will be assigned a value based on a certain emotional model. The emotion probability distribution is obtained through the emotion model; The emotion model includes the Ekman emotion model and the Plutchik emotion wheel; Specifically, the emotion probability distribution is as follows: Among them, E t It is the probability distribution of emotions at time t, where P represents the probability value of different emotional states; The LSTM model is trained to obtain the user's emotion prediction value by real-time monitoring of the user's facial expression information and the three-dimensional structure and feature information of the user's face. Calculate the amplitude of mood fluctuations: Where t represents time; Time represents the length of the observation time window; This represents the average emotional value. This represents the predicted sentiment value at time t; Volatility represents the fluctuation of sentiment. Calculate the rate of change in mood: ; where ΔE t Indicates the rate of change in mood; This represents the predicted sentiment value at time t-1; It should be noted that this emotion detection method is an emotion prediction mechanism based on deep learning. The "numericalization-time-sequencing" of facial expressions, voice, and physiological signals (heart rate, skin conductance, etc.) is a mature technology in the field of affective computing. In this embodiment, the dynamic coordinates of 68 facial key points are extracted by the computer vision tool Dlib, including the degree of mouth corner raising and the intensity of frowning, and mapped to emotions such as "pleasure" and "anger" by the Plutchik model. The 3D facial structure is reconstructed by a depth sensor to supplement the facial muscle movement depth information missing in the planar image and improve the robustness of emotion judgment. The magnitude of emotional fluctuations is essentially the "standard deviation of emotional values within a time window" (the numerator in the formula is the sum of squared deviations, and the denominator is the length of the time window, which is equivalent to the square of the standard deviation), which can quantify emotional stability (e.g., large fluctuations in ICU patients within 1 hour after surgery indicate a risk of pain or agitation). The rate of change of emotion is the "first-order difference of emotion values at adjacent moments", which can capture sudden changes in emotion. The calculation logic of both is simple and does not require a complex model. In this embodiment, it is calculated by Python Pandas scrolling window, which can be embedded in the real-time data stream processing link, and the engineering implementation difficulty is low.
[0032] Set the threshold for mood changes ΔE; When the rate of change in emotion ΔE t If the emotional change exceeds the threshold ΔE, the warning device will automatically issue an alarm, prompting caregivers to check on the user.
[0033] Specifically, the real-time monitored multimodal image data is displayed in real time through a visualization screen; the entire process data of posture and body position recognition, facial expression and emotion analysis, and environmental perception analysis is backed up; The warning results will be displayed in real time through a visual screen; In this implementation, a minimal structure is established for real-time display, including {event type, time window, evidence index set, context (position / emotion / environment / vital signs), associated object (bed / equipment ID)}.
[0034] Generate intelligent analysis reports, including user emotional fluctuations, posture changes, environmental changes, warning times, and the handling process of relevant personnel.
[0035] It also includes: a visual management module comprising a patient condition viewing unit and an intelligent document unit; The disease status viewing unit is used to display monitoring results, including multimodal image data recorded according to timestamps; Intelligent document units include template constraints, evidence citation, and consistency verification; Template constraints must include at least the generation of structured templates for ICU shift handover and nursing records; Evidence citation includes binding an evidence index to each generated conclusion; The evidence index includes a time window, a data source ID, an image segment index, and a record ID; Consistency verification includes checking key values, time points, and event sequences before and after generation; if inconsistencies are found, a rollback is performed. In this implementation, the verification is based on cosine similarity, but Euclidean distance or other methods can also be used; no specific restrictions are imposed. Based on the aforementioned deep learning data, an ICU language model with question-answering capabilities is also provided. This model takes acute and critical care clinical thinking as its core and achieves the transformation from data to value through a three-stage collaboration of understanding (information decoding), reasoning (information matching based on cosine similarity), and generation (results output and display to medical staff). Each stage is deeply adapted to the special characteristics of the ICU (including data fragmentation and time sensitivity). It should be noted that semantic reasoning is a core component of the ICU language model's reasoning (information matching, based on cosine similarity) stage. It receives various types of ICU data processed in the understanding (information decoding) stage and, through reasoning, provides accurate information matching results for the generation (outputting results to medical staff) stage. This is a crucial step in transforming data into value, further converting the complex, multimodal raw information from the ICU into structured clinical semantics, laying the foundation for subsequently presenting relevant information to medical staff. The core objective is to transform the complex, multimodal raw information from the ICU into structured clinical semantics, addressing the issues of information silos and professional barriers. The data processed includes the entire ICU dataset, specifically: Real-time vital signs (time series of heart rate, blood oxygen, blood pressure, etc.), laboratory indicators (complete blood count, blood gas analysis, inflammatory factors, etc.), medication records (dosage / time of antibiotics and vasoactive drugs). Unstructured data: medical records (including information such as patient agitation and decreased SpO2 6 hours after surgery), nursing records (including urine output data), imaging reports, and voice commands (including verbal handover information between medical staff and nurses). Among them, text data is captured in real time from structured texts such as medical records, nursing records, and imaging reports through the electronic medical record system (EMR); voice transcription devices are deployed to collect voice information such as oral handovers and bedside instructions from medical staff and convert them into text streams in a synchronous manner; and the medical order module of the hospital information system (HIS) is connected to collect medication orders (including drug name, dosage, frequency and route of administration), examination orders, and operation orders in real time, while recording the time of order issuance, execution status and executor information. Denoising was performed on unstructured texts such as medical records, and redundant symbols were removed. Medical terminology was split using the medical terminology tool BioNLP and standardized using a medical thesaurus. The processed text data is used to create a database with the patient as the unique identifier. The Hadoop Distributed File System (HDFS) can be used to store the massive amounts of text and medical order data. During interactive visualization with medical staff, information matching based on cosine similarity is used to display relevant medications and treatment plans for the corresponding disease records, providing assistance to medical staff in extremely stressful situations. Data visualization can be achieved through one or more of the following: integrated timeline display, relational graph visualization, and interactive query interface. Integrated Timeline Display: Using time as the horizontal axis, text records, medical order execution nodes, and vital sign curves within a certain time period are displayed synchronously, with outliers highlighted by color. Visualization of the association graph: A force-directed graph is used to display the association network of symptoms-medical orders-drugs. The size of the node represents the similarity weight, and the nodes are connected by thick lines to intuitively present strong association relationships. Interactive query interface: Allows medical staff to filter data by keywords or time range. When a text record is clicked, it automatically highlights several related medical orders and treatment plans with the highest cosine similarity, and displays the similarity value for reference.
[0036] It is important to note that before a database with each patient as a unique identifier is created, medical staff must perform anomaly checks on the medical orders to ensure their accuracy. Cosine similarity is a mathematical metric that measures the directional consistency of two vectors in a vector space. It is widely used in semantic matching, and its application process in this embodiment is as follows: When medical staff focus on a certain piece of information (urine volume data in nursing records) through a visual interface, the model first converts the information into a target vector, and the vector dimension includes key features such as symptoms, signs, and time. By matching the constructed database (including vectors such as medication records, treatment plans, and historical disease courses), the cosine similarity between the target vector and each vector in the database is calculated one by one, and the matched data is pushed out.
[0037] In emergency scenarios, cosine similarity can complete the comparison of all data within milliseconds, quickly locate highly relevant information, help medical staff shorten decision-making time and reduce information retrieval costs under extreme pressure. At the same time, through quantified similarity values, medical staff can intuitively judge the strength of information association and provide a reference for multiple options.
[0038] The intelligent document unit is used to display intelligent analysis reports for medical staff to view.
[0039] Example 1: The multimodal image data consists of an RGB camera (15fps, 1080P resolution), a depth camera (10Hz sampling frequency, depth error ≤2mm), and a thermal imager (5Hz sampling frequency, temperature accuracy ±0.5℃). The time window size is 10 seconds (continuous monitoring window), and the early warning judgment window consists of three consecutive 10-second windows. The thresholds are set as follows: first attitude change threshold θ'=15°, second attitude maintenance threshold t'=30min, temperature threshold range [22℃, 28℃], and temperature change threshold Δt=2℃. The semantic events were extracted by the analysis module as follows: The angle changes of the key points of posture (head p1, chest p2, waist p3) exceeded θ'=15° for three consecutive 10s windows; Thermal imaging data showed that the standard deviation of the patient's trunk temperature was 2.5℃, exceeding Δt=2℃; The fluctuation amplitude of facial expression feature points (corner of the mouth, eyebrows) is Volatility=0.8, the emotion change rate ΔEt=0.6, which exceeds the preset emotion threshold ΔE=0.4; Based on the aforementioned formula, corresponding calculations are performed. When the change in posture angle exceeds the threshold, the temperature fluctuation exceeds the threshold, or the emotional fluctuation exceeds the threshold, an early warning is issued, and medical staff provide care.
[0040] Example 2: Multimodal data includes data from an RGB camera (10fps), a depth camera (8Hz sampling frequency), a vital signs monitor (spO2 blood oxygen, 1Hz sampling frequency), and medical orders (endotracheal intubation orders and sedation drug usage records). The time window size is 30 minutes (shift handover statistics window), and the real-time monitoring window is 5 minutes. The thresholds were set as follows: SpO2 threshold ≤ 92%, facial movement amplitude threshold (related to tube removal: frequency of hand approaching face ≥ 5 times / 5min), and emotion change threshold ΔE = 0.5. The following semantic events were extracted using the analysis module and the ICU language model: SpO2 was below 92% 5 times within 30 minutes, with the lowest value being 88%; the frequency of hand contact with face was 7 times / 5 minutes, exceeding the threshold 5 times / 5 minutes; the rate of mood change ΔEt=0.7, exceeding the threshold 0.5; The system highlights the patient's relevant data on the shift handover interface and simultaneously generates a structured handover document, which is then pushed to the doctor's workstation.
[0041] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims.
Claims
1. A critical care multimodal data perception and intelligent generation system with semantic reasoning capabilities, characterized in that: The system includes a data acquisition module, a preprocessing module, an analysis module, an anomaly detection module, a data storage module, and a visualization management module; The data acquisition module is used to acquire multimodal image data in real time; Among these measures, recording timestamps ensures data synchronization in time sequence; The preprocessing module is used to process the collected multimodal data to ensure data consistency; this includes data synchronization, noise reduction, and standardization. The analysis module is used to analyze the preprocessed data, including posture and body position recognition, facial expression and emotion analysis, and environmental perception analysis. The anomaly detection module is used to detect abnormal behavior based on a preset threshold and issue an early warning in real time; it also automatically notifies nursing staff. The data storage module is used to store the collected multimodal data, analysis results, and intelligent analysis reports; The visualization management module is used to visualize the data and provide a user interface.
2. The critical care multimodal data perception and intelligent generation system with semantic reasoning capability according to claim 1, characterized in that: Specifically, the multimodal image data includes RGB images, depth data, and thermal imaging images; The data acquisition module integrates a multimodal image acquisition device, which includes an RGB camera, a depth camera, and a thermal imager. Equip with calibration equipment and set the acquisition parameters; The acquisition parameters include frame rate, resolution, and sampling frequency; Start the image acquisition device to capture RGB images, depth data, and thermal imaging data in real time; During the data acquisition process, timestamps are recorded to ensure the synchronization of multimodal data.
3. The critical care multimodal data perception and intelligent generation system with semantic reasoning capability according to claim 1, characterized in that: include: Based on the device's timestamp, the time axis of different modal data is uniformly adjusted to ensure data timing matching. Based on the equipment installation location and viewing angle, geometric correction is performed using calibration data to map data from different modes to the same coordinate system.
4. The critical care multimodal data perception and intelligent generation system with semantic reasoning capability according to claim 3, characterized in that: The analysis module is used to analyze the preprocessed data, specifically including posture and body position recognition, facial expression and emotion analysis, and environmental perception analysis. The posture and body position recognition specifically refers to: Using a standardized RGB image, the user's appearance and facial features are obtained; The user's spatial position and posture changes are obtained through standardized depth data; Pose estimation is performed using a deep learning model, and keypoint locations are extracted. The set of keypoint locations is denoted as {p1, p2, ..., p...}. n }, p i ∈{p1,p2,…,p n }, p i =(x i ,y i ,z i ); where n represents the total number of keypoints, and n is a positive integer; i represents the number of keypoint labels, and i is a positive integer, 1≤i≤n; p i It is an element in the set of key point locations, p i Represents the position coordinates of the i-th key point; (x i ,y i ,z i ) represents the coordinates of the i-th key point in three-dimensional space; x i The x-coordinate of the i-th key point is represented by y. i z represents the ordinate of the i-th key point. i Represents the vertical coordinate of the i-th key point; The user's body position is determined based on the coordinates of key points in three-dimensional space. The body position is determined by calculating the coordinate differences and angles.
5. The critical care multimodal data perception and intelligent generation system with semantic reasoning capability according to claim 2, characterized in that: The facial expression and emotion analysis specifically involves: using standardized RGB images to obtain the user's facial expression information; The three-dimensional structure and feature information of the user's face are obtained by using standardized depth data; Facial feature points are extracted using a facial expression recognition model. These feature points include the geometric shapes of the eyes, eyebrows, and mouth. The set of facial feature points is {f1, f2, ..., f...}. m }, f j ∈{f1,f2,…,f m }, f j =(x j ,y j ); where m represents the total number of facial feature points, m is a positive integer; j represents the number label of facial feature points, j is a positive integer, 1≤j≤m; f j It is an element in the set of facial feature points, f j Represents the position coordinates of the j-th facial feature point; (x j ,y j ) represents the two-dimensional coordinates of the j-th facial feature point in the image; x j The x-coordinate of the j-th facial feature point is represented by y. j This represents the ordinate of the j-th facial feature point; We use a convolutional neural network (CNN) to classify facial expressions into emotions, and set the output of the emotion classification model as follows: Emotion=argmax(softmax(f(features))); Where f(features) represents the features extracted from facial feature points; Emotion represents the emotion classification obtained through a convolutional neural network; The environmental perception analysis specifically includes: Use standardized thermal imaging data to obtain information about the surrounding temperature distribution; The spatial layout of the user's bedside environment is obtained through standardized deep data.
6. A critical care multimodal data perception and intelligent generation system with semantic reasoning capability according to any one of claims 4-5, characterized in that: include: Set an abnormal behavior threshold, specifically as follows: Install early warning devices; Set the first attitude change threshold θ'; Set the second attitude change threshold t'; When the angle change between key points exceeds the first posture change threshold θ', the early warning device will automatically issue an alarm to prompt the nursing staff to adjust the posture. When a user maintains a certain posture for more than the second posture change threshold t', the early warning device will automatically issue an alarm to prompt the caregiver to adjust the posture. Set a third temperature threshold range [t1, t2]; Set a fourth temperature change threshold Δt; When the temperature in the user's area is outside the third temperature threshold range [t1, t2], the early warning device will automatically issue an alarm to prompt caregivers to adjust the ambient temperature. When the temperature standard deviation of the user's area exceeds the fourth temperature change threshold Δt, the early warning device will automatically issue an alarm to prompt caregivers to adjust the ambient temperature.
7. The critical care multimodal data perception and intelligent generation system with semantic reasoning capability according to claim 5, characterized in that: Specifically: An emotion probability distribution is obtained through an emotion model; wherein the emotion model includes at least one of an emotion recognition model, a sequence model, and an emotion classification system; the user's emotion prediction value is obtained by training an LSTM model and by real-time monitoring of the user's facial expression information and the three-dimensional structure and feature information of the user's face. Calculate the amplitude of mood fluctuations: Where t represents time; Time represents the length of the observation time window; This represents the average emotional value. This represents the predicted sentiment value at time t; Volatility represents the fluctuation of sentiment. Calculate the rate of change in mood: ; where ΔE t Indicates the rate of change in mood; This represents the predicted sentiment value at time t-1; Set the threshold for mood changes ΔE; When the rate of change in emotion ΔE t If the emotional change exceeds the threshold ΔE, the warning device will automatically issue an alarm, prompting caregivers to check on the user.
8. A critical care multimodal data perception and intelligent generation system with semantic reasoning capability according to claim 7, characterized in that: Specifically: The real-time monitored multimodal image data is displayed in real time through a visualization screen; The entire process data of posture and body position recognition, facial expression and emotion analysis, and environmental perception analysis is backed up; The warning results will be displayed in real time through a visual screen; Generate intelligent analysis reports, including user emotional fluctuations, posture changes, environmental changes, warning times, and the handling process of relevant personnel.
9. A critical care multimodal data perception and intelligent generation system with semantic reasoning capability according to claim 8, characterized in that: Also includes: The visual management module includes a medical condition viewing unit and an intelligent document unit; The condition viewing unit is used to display monitoring results, including multimodal image data recorded according to timestamps; The intelligent document unit is used to display intelligent analysis reports for medical staff to view. The intelligent document unit includes template constraints, evidence citation, and consistency verification. The template constraints include at least the generation of structured templates based on ICU shift handover and nursing records; The evidence citations include an evidence index bound to each generated conclusion; The evidence index includes a time window, a data source ID, an image segment index, and a record ID; The consistency check includes verifying key values, time points, and event sequences before and after generation; if inconsistencies are found, a rollback is initiated.
10. A critical care multimodal data perception and intelligent generation system with semantic reasoning capability according to claim 1, characterized in that: It also includes a three-stage collaborative processing architecture with semantic reasoning capabilities, wherein: The understanding phase is used to decode ICU data collected from the electronic medical record system (EMR), hospital information system (HIS), and speech-to-text devices. It unifies the parsing of real-time vital signs data, nursing records, image reports, and speech-to-text text, and converts them into structured semantic vectors with symptom features, sign features, time features, and event tags. The reasoning phase is used to perform semantic matching on the structured semantic vector. By calculating the cosine similarity in the vector space, the historical medical information, medical records and treatment plans most relevant to the target semantic vector are retrieved from the database constructed with the patient as the unique identifier. The generation phase is used to generate visual information display content for medical staff based on the matching results output by the inference phase.