Intelligent Monitoring and Early Warning Method and System Based on Multidimensional State Awareness
By using multidimensional state perception technology to acquire physiological, behavioral, and contextual data streams, modeling and fusing intermodal interaction relationships, and decoupling latent risk factors, this technology solves the problem of being unable to identify latent compound risks in existing technologies, and enables accurate identification and interpretable early warning of progressive health deterioration processes.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANXI ZHIJIE COMPUTER SOFTWARE ENG CO LTD
- Filing Date
- 2026-02-11
- Publication Date
- 2026-06-02
AI Technical Summary
Existing monitoring technologies cannot effectively identify hidden, progressive complex risks formed by the superposition and synergistic effects of multiple subclinical states. They lack the ability to model the high-order interaction relationships behind different modalities, resulting in blind spots in the perception of progressive health deterioration processes.
By acquiring physiological, behavioral, and contextual data streams, internal modal feature encoding and temporal feature extraction are performed, intermodal interaction relationships are established and fused, latent risk factors are decoupled, and ultimately interpretable early warning signals are generated.
It achieves accurate identification and interpretable early warning of latent risk factors, breaks through the limitations of atomized modeling and linear causal assumptions, and can identify a progressive health deterioration process driven by multiple subclinical states.
Smart Images

Figure CN122135968A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent monitoring, and more specifically, to an intelligent monitoring and early warning method and system based on multi-dimensional state perception. Background Technology
[0002] With the aging population and the increasing demand for chronic disease management, intelligent monitoring and early warning technologies are becoming increasingly important in homes, communities, and medical institutions. Their core objective is to achieve continuous tracking of individual health status and proactive risk identification, thereby shifting from passive response to proactive prevention. However, human health is a complex and comprehensive system; its changes depend not only on a single physiological indicator but also on the combined effects of multiple dimensions, including physiological, behavioral, and environmental contexts. Therefore, constructing an intelligent monitoring and early warning system capable of comprehensively and dynamically perceiving and understanding this multidimensional information is of crucial practical significance and application value for improving the accuracy and timeliness of early warnings and preventing serious health events.
[0003] Currently, existing monitoring technologies have made progress in monitoring specific monomodal data, such as monitoring physiological data like heart rate and blood oxygen through wearable devices, or analyzing behavioral patterns through cameras. However, these technologies generally suffer from the limitation of atomized risk modeling, treating various data streams as independent, decoupled information sources and focusing on identifying explicit, single risk events such as tachycardia or falls. The fundamental flaw of this approach lies in its implicit linear causal assumption, which assumes that the occurrence of a risk is linearly caused by a single indicator exceeding a threshold. This makes existing technologies inadequate when facing implicit, progressive compound risks formed by the superposition and synergistic effects of multiple seemingly normal subclinical states (such as mild dehydration, mood swings, and drug side effects). At its root, the reason existing technologies struggle to effectively identify implicit and compound risks is their lack of ability to model the higher-order interactions behind different modalities. The deterioration of health is often a non-linear process; the risk meaning of a feature (such as a slight increase in heart rate) can fundamentally change depending on the context or when combined with other features (such as gait instability). Current technologies are unable to capture this synergistic effect, resulting in a blind spot in the perception of the gradual health deterioration process.
[0004] Therefore, there is an urgent need for a new method that can overcome the limitations of atomized modeling and linear causal assumptions, and achieve effective decoupling of implicit risk factors and accurate identification and interpretable early warning of complex risks. Summary of the Invention
[0005] To address the aforementioned problems in the existing technology, this application provides an intelligent monitoring and early warning method based on multi-dimensional state perception, which includes: Acquire physiological data streams, behavioral data streams, and contextual data streams; Internal modal feature encoding and temporal feature extraction are performed on physiological data stream, behavioral data stream and contextual data stream to obtain physiological temporal feature vector, behavioral temporal feature vector and contextual temporal feature vector; Modeling and fusing the intermodal interaction relationships of physiological temporal feature vectors, behavioral temporal feature vectors, and contextual temporal feature vectors to obtain the fused contextual vector and intermodal attention weight vector; The implicit risk factors are decoupled from the fused context vector to obtain the implicit risk factor distribution vector; Based on the intermodal attention weight vector, a composite risk synthesis is performed on the distribution vector of latent risk factors to obtain an interpretable early warning signal.
[0006] This application also provides an intelligent monitoring and early warning system based on multi-dimensional state perception, which includes: The multidimensional state data acquisition module is used to acquire physiological data streams, behavioral data streams, and contextual data streams; The multidimensional state data encoding module is used to perform internal modal feature encoding and temporal feature extraction on physiological data streams, behavioral data streams, and contextual data streams to obtain physiological temporal feature vectors, behavioral temporal feature vectors, and contextual temporal feature vectors. The intermodal interaction analysis module is used to model and fuse the intermodal interaction relationships of physiological temporal feature vectors, behavioral temporal feature vectors, and contextual temporal feature vectors to obtain the fused context vector and intermodal attention weight vector. The implicit risk factor decoupling module is used to decouple the implicit risk factors from the fused context vector to obtain the implicit risk factor distribution vector. The early warning signal generation module is used to perform composite risk synthesis on the distribution vector of latent risk factors based on the intermodal attention weight vector to obtain an interpretable early warning signal.
[0007] Compared with existing technologies, this application provides an intelligent monitoring and early warning method and system based on multidimensional state perception. Instead of analyzing isolated single data streams, it treats the physiological, behavioral, and environmental context data of the monitored object as an organic whole for unified input. By deeply modeling the nonlinear interaction relationships and synergistic effects between different modalities, the model can generate a fusion vector that comprehensively reflects an individual's overall health status. This directly overcomes the limitations of existing technologies that treat risk as an atomized risk modeling event. Furthermore, the model actively decouples multiple potential latent risk factors from this fusion vector and quantifies and synthesizes composite risks based on the dynamic contribution of each data modality in the risk formation process. This mechanism abandons the implicit linear causal assumptions of traditional technologies, enabling it to accurately identify and warn of progressive health deterioration processes driven by multiple subclinical states. Ultimately, it outputs interpretable signals containing evidence of risk tracing, thus effectively solving the key technical challenges raised in the background. Attached Figure Description
[0008] The above and other objects, features and advantages of this application will become more apparent from the more detailed description of the embodiments of this application in conjunction with the accompanying drawings.
[0009] Figure 1 This is a flowchart of an intelligent monitoring and early warning method based on multi-dimensional state perception according to an embodiment of this application.
[0010] Figure 2 This is a schematic diagram of data flow in an intelligent monitoring and early warning method based on multi-dimensional state perception according to an embodiment of this application.
[0011] Figure 3 This is a flowchart of step S2 in the intelligent monitoring and early warning method based on multi-dimensional state perception according to an embodiment of this application.
[0012] Figure 4 This is a flowchart of step S5 in the intelligent monitoring and early warning method based on multi-dimensional state perception according to an embodiment of this application.
[0013] Figure 5 This is a block diagram of an intelligent monitoring and early warning system based on multi-dimensional state perception according to an embodiment of this application. Detailed Implementation
[0014] The embodiments of this application will now be described in more detail with reference to the accompanying drawings. It should be understood that the drawings and embodiments of this application are for illustrative purposes only and are not intended to limit the scope of protection of this application.
[0015] To address the shortcomings in the aforementioned technical fields, this application proposes an intelligent monitoring and early warning method based on multi-dimensional state perception. Figure 1This is a flowchart of an intelligent monitoring and early warning method based on multi-dimensional state perception according to an embodiment of this application. Figure 2 This is a schematic diagram of the data flow in an intelligent monitoring and early warning method based on multi-dimensional state perception according to an embodiment of this application. Figure 1 and Figure 2 As shown, the intelligent monitoring and early warning method based on multi-dimensional state perception according to an embodiment of this application includes: Step S1, acquiring physiological data stream, behavioral data stream, and context data stream; Step S2, performing internal modal feature encoding and temporal feature extraction on the physiological data stream, behavioral data stream, and context data stream to obtain physiological temporal feature vector, behavioral temporal feature vector, and context temporal feature vector; Step S3, modeling and fusing the inter-modal interaction relationship of the physiological temporal feature vector, behavioral temporal feature vector, and context temporal feature vector to obtain a fused context vector and an inter-modal attention weight vector; Step S4, decoupling the latent risk factors on the fused context vector to obtain a latent risk factor distribution vector; Step S5, performing composite risk synthesis on the latent risk factor distribution vector based on the inter-modal attention weight vector to obtain an interpretable early warning signal.
[0016] In step S1, physiological data streams, behavioral data streams, and contextual data streams are acquired. It should be understood that the health status of the human body is a complex and dynamic system determined by internal physiology, external behavior, and the environment. The bottleneck of existing monitoring technologies lies in their isolation of these multi-dimensional information, failing to reveal the deep connections between various factors. Health risks, especially complex risks arising from multiple subclinical states, do not evolve from linear changes in a single indicator, but rather are a comprehensive manifestation of the nonlinear coupling of multimodal information. Therefore, to accurately and proactively identify such hidden risks, the primary and most fundamental step is to comprehensively and synchronously capture various types of data reflecting the overall state of an individual. This provides the necessary data input for subsequent intermodal interaction modeling, risk factor decoupling, and the generation of interpretable early warning signals. This is a crucial prerequisite for overcoming the limitations of atomized risk modeling and achieving intelligent early warning from multi-dimensional state perception.
[0017] In one embodiment, step S1 is implemented as follows: To achieve comprehensive perception of the health status of the monitored individual, this embodiment synchronously collects data streams across three dimensions: physiological, behavioral, and contextual. The entire process is conducted using an elderly person living alone with hypertension as the monitored individual. First, to obtain physiological data streams that can finely characterize changes in vital signs, the monitored individual is equipped with a wearable device integrating multifunctional sensors, such as a close-fitting vest or a medical-grade wristband. This device incorporates at least three types of core sensors: a three-lead electrocardiogram (ECG) sensor to capture complete cardiac electrical activity, with a sampling frequency set at 256 Hz; a photoplethysmography (PPG) sensor to monitor indicators such as heart rate and blood oxygen saturation, with a sampling frequency set at 100 Hz; and an electrical conductance per skin (EDA) sensor to reflect autonomic nervous system activity and emotional stress levels, with a sampling frequency set at 4 Hz. The device operates continuously, transmitting the collected high-frequency raw time-series signals, along with timestamps accurate to milliseconds, in real time to the central gateway indoors via Bluetooth Low Energy protocol. The central gateway aligns and packages the received data to form a structured physiological data stream. For example, within the time window of 09:00:00 to 09:00:01 AM, the generated data packet will contain 256 ECG signal sampling points, 100 pulse wave signal sampling points, and 4 skin conductance sampling points. These data constitute the first dimension of input required for subsequent analysis.
[0018] Secondly, to capture behavioral data streams reflecting the monitored individual's daily activity abilities, behavioral patterns, and even potential abnormalities, this embodiment employs a combination of wearable sensing and environmental perception. On one hand, the aforementioned wearable device integrates a six-axis inertial measurement unit (IMU), including a three-axis accelerometer and a three-axis gyroscope, continuously collecting the individual's posture and movement information at a frequency of 50 Hz to identify basic activities such as walking, sitting, and standing, and to monitor gait stability and fall risk. On the other hand, non-invasive millimeter-wave radar sensors are deployed in the main areas of the elderly's activity, such as the living room, bedroom, and bathroom. These sensors can detect the presence, location, posture (such as standing, sitting, and lying down), and respiratory rate of a person in the space without infringing on privacy. Simultaneously, a smart speaker with an integrated microphone array is placed in the center of the living room to collect ambient sounds and snippets of the individual's voice interactions to analyze their social activity and emotional state. This multi-source behavioral data is also transmitted to a central gateway for temporal and spatial fusion. For example, when detecting gait instability data from wearable devices, millimeter-wave radar might simultaneously capture the trajectory of an object wandering near a bathroom, while a microphone might record faint groans. The gateway integrates this information to form a behavioral data stream containing acceleration sequences, spatial coordinate trajectory sequences, and speech fragments, serving as the second dimension of input for analysis.
[0019] Finally, to introduce the background information necessary for in-depth interpretation of physiological and behavioral data—that is, the contextual data stream—this embodiment establishes a data interface and recording mechanism. This data stream is not continuously collected but event-driven. Through an application programming interface (API) connected to the community health center's database, the electronic health record (EHR) of the monitored individual is periodically synchronized, automatically acquiring static information such as their diagnosed medical history (e.g., hypertension, diabetes). Simultaneously, a smart pillbox linked to a central gateway is provided. When the elderly person takes medication, the pillbox automatically records the type, dosage, and time of administration; for example, recording an event at 8:00 AM: "Taking a new type of antihypertensive diuretic, 20 mg." Furthermore, through a simple interactive interface installed on a tablet, family members or caregivers can enter important non-clinical events, such as "3:00 PM, feeling down after talking to children" or "Today's visitor: daughter, stayed for one hour." All these structured event records are accompanied by precise timestamps and are aggregated into a contextual data stream. This data, which includes diagnostic information, medication records, social interactions, and other key events, provides an indispensable reference for accurately assessing the true risk implications of the aforementioned physiological and behavioral data, and constitutes the third dimension of input required for the analysis.
[0020] In step S2, internal modal feature encoding and temporal feature extraction are performed on the physiological data stream, behavioral data stream, and contextual data stream to obtain physiological temporal feature vectors, behavioral temporal feature vectors, and contextual temporal feature vectors. Correspondingly, the original multidimensional data streams obtained in step S1 are diverse in form, uneven in information density, and contain a large amount of noise and redundancy. For example, physiological data is a high-frequency continuous temporal signal, while contextual data is a discrete, sparse record of events. Directly fusing these heterogeneous original data will not only face the problems of the curse of dimensionality and low computational efficiency, but will also interfere with the model's effective capture of core information due to the huge differences between the data. Therefore, before performing cross-modal interaction and fusion, each modality needs to undergo independent and in-depth preprocessing and refinement to transform these original, mixed data streams into a unified, semantically condensed mathematical expression, namely, a temporal feature vector. This process captures the inherent patterns of each modality by encoding its features and extracts its dynamic evolution features using a time series model, thereby generating a fixed-length feature vector that represents the core state of the modality within a specific time window.
[0021] Figure 3 This is a flowchart of step S2 in the intelligent monitoring and early warning method based on multi-dimensional state perception according to an embodiment of this application. Figure 3As shown, in one embodiment, step S2, which involves encoding internal modal features and extracting temporal features from the physiological data stream, behavioral data stream, and contextual data stream to obtain physiological temporal feature vectors, behavioral temporal feature vectors, and contextual temporal feature vectors, includes: step S21, inputting the physiological data stream into a temporal encoder based on a bidirectional gated recurrent unit network to obtain a physiological modality temporal hidden state sequence; step S22, calculating the importance weights of key temporal nodes in the physiological modality temporal hidden state sequence to obtain a physiological modality attention weight vector; and step S23, performing weighted fusion of the physiological modality temporal hidden state sequence based on the physiological modality attention weight vector to obtain the physiological temporal feature vector.
[0022] In the above implementation, step S2 is implemented as follows: First, in step S21: Before input, the original data stream needs to be preprocessed. The processing unit aligns and segments the data with a time window of 60 seconds. For the three types of data with different sampling frequencies, namely ECG, pulse wave, and skin conductance, feature extraction and synchronization are performed in the following way: the 60-second window is divided into 60 time steps of 1 second each. Within each time step (i.e., per second), the heart rate value and the standard deviation of two consecutive heartbeat intervals are calculated for the high-frequency ECG data as a heart rate variability index; the mean blood oxygen saturation is calculated for the pulse wave data; and the mean and standard deviation are calculated for the skin conductance data. Thus, each one-second time step generates a feature vector containing five key indicators. For example, at the time step of 09:00:30 AM, the feature vector may be [heart rate: 85, heart rate variability: 30ms, blood oxygen saturation: 97%, mean skin conductance: 1.5μS, standard deviation of skin conductance: 0.2μS]. After normalization, the 60 seconds of data form a 60-dimensional temporal feature sequence with each element having a dimension of 5, which serves as the input to the temporal encoder. The temporal encoder is a bidirectional gated recurrent unit (Bi-GRU) network. This network consists of a forward-gated recurrent unit layer and a backward-gated recurrent unit layer in parallel, designed to simultaneously capture the historical dependencies and future information of the temporal data. Specifically, the hidden state dimension of the gated recurrent unit is set to 128. When the aforementioned 5-dimensional feature sequence is input, the forward-gated recurrent unit processes sequentially from the first time step (second 0) to the 60th time step (second 59), outputting a 128-dimensional forward hidden state at each time step t. Simultaneously, the backward-gated recurrent unit processes in reverse order from the 60th time step to the first time step, also outputting a 128-dimensional backward hidden state at each time step t. Subsequently, at each time step t, the forward and backward hidden states of that time step are concatenated to form a 256-dimensional composite hidden state. After this process iterates through all 60 time steps, the final output is a sequence containing 60 256-dimensional vectors, which is the physiological modality temporal hidden state sequence. Each vector ht in this sequence contains the physiological characteristics at second t and its contextual information over the entire 60-second window.
[0023] Next, in step S22: not all physiological states contribute equally to risk assessment. For example, a brief arrhythmia is far more indicative than a stable heart rate. The purpose of this step is to quantify the importance of each time point. In one embodiment, step S22, calculating the importance weights of key time nodes in the physiological modality time-series hidden state sequence to obtain the physiological modality attention weight vector, includes: calculating the importance weights of key time nodes in the physiological modality time-series hidden state sequence using the following formula:
[0024]
[0025] in, Let be the physiological modality temporal hidden state at the t-th time step in the physiological modality temporal hidden state sequence. The weight matrix is a learnable matrix. For learnable bias vectors, Let be the intermediate features of the physiological modality time series at time step t. For learnable globally shared context vectors, The length of the physiological modality temporal hidden state sequence. Let be the physiological modality attention weight at time step t in the physiological modality attention weight vector. First, for each hidden state in the sequence... (t ranges from 1 to 60), using the formula Perform the transformation. Here... It is a 256×256 learnable weight matrix. It is a 256-dimensional learnable bias vector. These two parameters are learned through backpropagation during end-to-end training of the entire early warning model using a large-scale labeled monitoring dataset. Their function is to store the original hidden states... Projecting onto a new feature space yields intermediate features. Then, another 256-dimensional globally shared context vector, also learned during training, is introduced. This vector can be understood as a query vector, representing the most important information patterns that the model is interested in. Through calculation... and The dot product is then normalized using the Softmax function, i.e. Obtain the physiological modality attention weights at time step t. This weight It is a scalar value between 0 and 1, with the sum of the weights of all 60 time steps being 1. For example, if an atrial premature beat occurs around the 30th second, its corresponding hidden state... After transformation and dot product operations, a high alignment score is obtained, resulting in a significantly higher attention weight (e.g., 0.15) compared to other stationary time steps, where the weight might only be 0.005. Finally, a physiological modality attention weight vector consisting of 60 attention weights is output. .
[0026] Finally, in step S23: the information within the entire 60-second time window is compressed into a single, fixed-length vector, while highlighting the impact of key time nodes. The fusion process is a weighted summation operation: each hidden state vector in the physiological modality time-series hidden state sequence is weighted... The corresponding physiological modality attention weights Multiply them, and then sum all 60 multiplied vectors together. Specifically, the final physiological time-series feature vector... The calculation method is as follows Where t ranges from 1 to 60, i.e., T is 60. Thus, hidden states with higher attention weights (such as in the previous example) Hidden states with lower weights dominate the summation process, while their influence is weakened accordingly. The final generated physiological time-series feature vector... This is a 256-dimensional vector. This vector not only summarizes the overall physiological state of the monitored subject over the past 60 seconds, but also incorporates the features of the most noteworthy physiological events through an attention mechanism. For example, the numerical distribution of this vector may strongly reflect the characteristics of that particular premature atrial contraction event. Thus, the original, variable-length, multi-channel physiological data stream has been successfully encoded into a condensed, fixed-length physiological temporal feature vector, which can be processed in subsequent steps. Similarly, a similar internal modality encoding and temporal feature extraction process is performed on the behavioral data stream and the contextual data stream, respectively, to obtain the behavioral temporal feature vector and the contextual temporal feature vector, which will not be elaborated further here.
[0027] In step S3, the physiological time-series feature vector, behavioral time-series feature vector, and contextual time-series feature vector are modeled and fused to obtain the fused contextual vector and intermodal attention weight vector. It is understandable that although highly condensed feature vectors have been extracted from each independent data stream through preprocessing, these vectors remain isolated and fail to reflect the inherent coupling and mutual influence between different dimensional states. This is the fundamental problem preventing existing technologies from identifying complex risks. The evolution of health risks, especially latent risks, is often the result of the synergistic effect of multiple factors. The true risk level of an abnormal signal in one modality will dynamically change depending on the state of other modalities. Therefore, after extracting single-modal features, a mechanism is needed to explicitly model the nonlinear interaction between modalities. To this end, this application uses an attention mechanism to dynamically evaluate the relative contributions of the physiological, behavioral, and contextual modalities to risk formation within the current time window, and performs weighted fusion of information based on these contributions to generate a unified representation that comprehensively and accurately reflects an individual's overall risk state. This provides a high-quality input with a fully understood synergistic effect for subsequent risk factor decoupling.
[0028] In one embodiment, step S3, which involves modeling and fusing the intermodal interaction relationships of the physiological time-series feature vector, behavioral time-series feature vector, and contextual time-series feature vector to obtain a fused context vector and an intermodal attention weight vector, includes: step S31, performing nonlinear projection on the physiological time-series feature vector, behavioral time-series feature vector, and contextual time-series feature vector to obtain projected physiological time-series feature vector, projected behavioral time-series feature vector, and projected contextual time-series feature vector; step S32, calculating intermodal importance weights on the projected physiological time-series feature vector, projected behavioral time-series feature vector, and projected contextual time-series feature vector to obtain intermodal attention weights from the first to the third modality, wherein the intermodal attention weights from the first to the third modality constitute the intermodal attention weight vector; and step S33, performing weighted fusion of information on the physiological time-series feature vector, behavioral time-series feature vector, and contextual time-series feature vector based on the intermodal attention weights from the first to the third modality to obtain the fused context vector.
[0029] In the above implementation, step S3 is implemented as follows: Continuing with the example of the elderly person living alone with hypertension, due to the newly taken diuretic taking effect, the elderly person exhibits mild signs of electrolyte imbalance. The meanings implied by the three currently input vectors are specified as follows: a physiological temporal feature vector, whose numerical distribution significantly reflects the characteristics of the atrial premature beat event that occurs at approximately the thirtieth second of the window due to the internal attention mechanism; a behavioral temporal feature vector, whose value summarizes the overall pattern of minor gait instability exhibited by the elderly person after getting up within this minute, such as increased gait frequency variability; and a contextual temporal feature vector, which encodes the key background event that is highly relevant to the current time window, namely 'taking the new antihypertensive diuretic at 8:00 AM'.
[0030] First, in step S31: vectors originating from different modalities and residing in their respective independent feature spaces are mapped to a shared, unified semantic space to facilitate meaningful comparison and correlation assessment. In one embodiment, step S31 involves nonlinearly projecting the physiological temporal feature vector, behavioral temporal feature vector, and contextual temporal feature vector to obtain projected physiological temporal feature vectors, projected behavioral temporal feature vectors, and projected contextual temporal feature vectors. This includes: performing nonlinear projection on the physiological temporal feature vector, behavioral temporal feature vector, and contextual temporal feature vector using the following formula:
[0031] in, These can be physiological time-series feature vectors, behavioral time-series feature vectors, or contextual time-series feature vectors. For learnable projection matrices, is a learnable projection bias vector. for Activation function These are the projected physiological temporal feature vectors, projected behavioral temporal feature vectors, or projected contextual temporal feature vectors. For each modality m (physiological, behavioral, or contextual), there is a dedicated set of learnable parameters, namely the projection matrix. and projection bias vector Specifically, the physiological temporal feature vector is multiplied by a 256×256-dimensional physiological modality projection matrix and then supplemented with a 256-dimensional physiological modality projection bias. The result is then input into the hyperbolic tangent activation function. A nonlinear transformation is performed. All these projection matrices and bias vectors are obtained through iterative learning using gradient descent and backpropagation algorithms during end-to-end optimization of the entire early warning model using a training set containing a large amount of multimodal data and corresponding risk labels. For example, some dimensions in the input physiological time-series feature vector may strongly indicate cardiac arrhythmia. By multiplying it with the physiological modality projection matrix, its expression in the new semantic space is enhanced and aligned with features in other modalities that also indicate risk. After this transformation, the 256-dimensional physiological time-series feature vector is converted into a projected physiological time-series feature vector of the same size (256). Similarly, the behavioral time-series feature vector and the contextual time-series feature vector are also converted into projected behavioral time-series feature vectors and projected contextual time-series feature vectors through their respective projection matrices and bias vectors.
[0032] Next, in step S32: dynamically determine which modality provides the most critical information for assessing potential risks within the specific 60-second window. In one embodiment, step S32 involves calculating intermodal importance weights for the projected physiological temporal feature vector, projected behavioral temporal feature vector, and projected contextual temporal feature vector to obtain the attention weights between the first and third modalities. This includes calculating the intermodal importance weights for the projected physiological temporal feature vector, projected behavioral temporal feature vector, and projected contextual temporal feature vector using the following formula:
[0033] in, This is a context reference vector shared by all learnable modalities. For the number of modes, This is the transpose of the vector. This represents the inter-modal attention weights among the attention weights for the first to third modalities. A key learnable parameter is introduced here: the context reference vector shared by all modalities. .this This is a 256-dimensional vector, also learned during model training. Its role is similar to a query or probe, representing the model's learned, ideal multimodal feature patterns that indicate high risk. For each projected modal feature vector, its relationship with the shared reference vector is calculated. The dot product of the transpose of . The result of this dot product is a scalar that measures the alignment or similarity between the feature representation of the current modality and the high-risk pattern. In the current example, since the model has been trained to learn the high risk of arrhythmia and gait instability in the context of diuretic administration, the high-risk pattern is represented by . The vector will produce a large dot product with the projected physiological temporal feature vector (whose features are highly correlated with arrhythmia) and the projected behavioral temporal feature vector (whose features are highly correlated with gait instability), while the dot product with the projected contextual temporal feature vector (reflecting the medication event itself) may be relatively small. For example, the calculated dot products may be 2.8, 2.6, and 1.2, respectively. Subsequently, these scalar values are passed through an exponential function. The values are converted to positive numbers and then normalized by summing the exponential results of all modalities (i.e., a softmax operation). This yields the attention weights for each modality. These values form a probability distribution that sums to one. For example, through normalization, we might obtain an attention weight of 0.48 for the physiological modality, 0.40 for the behavioral modality, and 0.12 for the contextual modality. The attention weight vector formed by these three weight values between the first and third modalities, i.e., the intermodal attention weight vector, clearly indicates that changes in physiological and behavioral states are the primary focus in the current risk assessment.
[0034] Finally, in step S33: the process follows the weighted summation formula described in the corpus. It is worth noting that the original feature vector output from step S2 is used for weighting. Instead of the projected vector output in step S31 This is done to preserve the most original and richest high-order feature information during the fusion process, while the projected vector... This is used solely for weight calculation to avoid potential information loss during projection. Specifically, the physiological temporal feature vector is multiplied by a weight of 0.48, the behavioral temporal feature vector by a weight of 0.40, and the contextual temporal feature vector by a weight of 0.12. Then, the resulting 256-dimensional vector is summed element-wise. The result of this operation is a completely new, single 256-dimensional vector—the fused context vector. This vector It is a highly condensed and comprehensive description of the monitored subject's status within the time window of 09:00:00 to 09:00:59 AM. Its numerical distribution not only integrates information from heart rhythm, gait, and medication history, but also, through the adjustment of attention weights, ensures that the two most urgent and critical risk signals—arrhythmia indicated by premature atrial contractions and behavioral abnormalities indicated by gait instability—dominate in the final fused representation.
[0035] In step S4, the latent risk factors are decoupled from the fused context vector to obtain the latent risk factor distribution vector. It should be understood that although the fused context vector already contains the synergistic effect of multimodal information and forms a high-level generalization of an individual's overall state, its essence remains a high-dimensional abstract mathematical representation. For recipients of the warning signal, such as medical staff or family members, its inherent clinical meaning is obscure. This vector alone cannot directly determine the nature and root cause of the risk, thus making it difficult to guide specific interventions. This is precisely the limitation of traditional "black box" models. To enable warnings to not only reveal the "what" but also the "why," achieving true interpretability, this holistic risk representation needs to be decomposed into a series of concepts that humans can understand and that have clear clinical implications. Therefore, this application decouples the fused vector and projects it onto a set of predefined, medically significant latent risk dimensions, thereby transforming the machine's abstract perception into a concrete, quantifiable risk factor assessment, laying the foundation for generating interpretable warning signals.
[0036] In one embodiment, step S4, decoupling the latent risk factors from the fused context vector to obtain the latent risk factor distribution vector, includes: step S41, predefining a set of latent risk factors; step S42, inputting the fused context vector into a fully connected neural network layer to obtain a latent risk factor logical vector; and step S43, calculating the independent risk confidence of the latent risk factor logical vector to obtain the latent risk factor distribution vector.
[0037] In the above implementation, step S4 is implemented as follows: First, in step S41: the latent risk factor set is designed based on clinical medical knowledge and common health risks for the target monitored population (such as elderly people with chronic diseases), and consists of a set of terms with clear clinical or physiological significance. Its role is to provide a target semantic space for subsequent risk decoupling, that is, a series of potential problem dimensions that need to be evaluated. In this embodiment, combined with the characteristics of the current scenario, the latent risk factor set is defined as a set containing five core factors, specifically: {cardiac instability, insufficient circulatory volume, neuromuscular control abnormality, emotional load, and potential infection risk}. Among them, cardiac instability corresponds to cardiac problems such as arrhythmia; insufficient circulatory volume corresponds to fluid balance problems such as dehydration and electrolyte imbalance, which is directly related to the use of diuretics; neuromuscular control abnormality corresponds to gait instability, fall risk, etc.; emotional load corresponds to psychological stress; and potential infection risk is a generalized health monitoring dimension. The size of this set, that is, the number of factors, is 5 here, which will determine the output dimension of the neural network layer in the next step.
[0038] Next, in step S42: the fully connected neural network layer acts as a linear transformer, tasked with learning the mapping from the fused, abstract feature space to the predefined, concrete risk factor semantic space. This layer's structure includes a weight matrix W and a bias vector b. Since the input vector C has a dimension of 256, and the target logical vector needs to correspond to the size of the risk factor set (i.e., a dimension of 5), the weight matrix W has a dimension of 5×256, and the bias vector b has a dimension of 5. These two parameters are obtained through end-to-end learning on a dataset containing massive amounts of multimodal data and their corresponding expert-annotated risk factor labels during the entire training phase of the early warning model. The goal of the training process is to adjust W and b so that when the input vector C implies a specific risk pattern, the output logical vector has a higher value in the corresponding risk dimension. In the current example, since vector C represents the combined state of heart problems, diuretic side effects, and gait instability, after multiplying with the learned weight matrix W and adding the bias vector b, the resulting five-dimensional latent risk factor logistic vector is expected to have higher positive values in the three dimensions of cardiac instability, insufficient circulatory volume, and neuromuscular control abnormalities, while having lower or negative values in the other two less relevant dimensions. For example, the output latent risk factor logistic vector might be [2.5, 3.0, 1.8, -0.5, -2.0], and these values are called logistic values, corresponding to the five risk factors mentioned above.
[0039] Finally, in step S43: the Sigmoid activation function is used to process each element of the latent risk factor logistic vector independently. The Sigmoid function maps any real number (logic value) to the interval between zero and one, and its output value can be interpreted as the confidence or probability of the corresponding risk factor. The reason for using the Sigmoid function instead of the Softmax function is that the various latent risk factors are not mutually exclusive, and a monitored subject may have multiple risks at the same time. For example, cardiac instability and neuromuscular control abnormalities can coexist. Softmax normalizes the probabilities of all risks to a sum of one, while Sigmoid scores each risk independently. The specific calculation process is as follows: the Sigmoid function f(x)=1 / (1+exp(-x)) is applied to each element in the latent risk factor logistic vector [2.5,3.0,1.8,-0.5,-2.0]. The calculation results are as follows: the confidence level for cardiac instability is sigmoid(2.5) approximately equal to 0.924; the confidence level for insufficient circulatory volume is sigmoid(3.0) approximately equal to 0.953; the confidence level for neuromuscular control abnormality is sigmoid(1.8) approximately equal to 0.858; the confidence level for emotional load is sigmoid(-0.5) approximately equal to 0.378; and the confidence level for potential infection risk is sigmoid(-2.0) approximately equal to 0.119. The final output is a five-dimensional vector, namely the distribution vector of latent risk factors: [0.924, 0.953, 0.858, 0.378, 0.119]. This vector clearly and quantitatively reveals that, at the current moment, the model is highly confident that the monitored individual is at risk of insufficient circulatory volume (95.3% confidence), unstable cardiac function (92.4% confidence), and abnormal neuromuscular control (85.8% confidence), while the emotional burden and risk of infection are low.
[0040] In step S5, based on the intermodal attention weight vector, the latent risk factor distribution vector is composite risk synthesized to obtain an interpretable early warning signal. That is, although the previous steps have successfully decoupled an abstract multimodal fusion vector into a quantitative distribution vector composed of the confidence levels of specific risk factors, this set of values is still an intermediate product. It reveals what problems might exist, but does not directly provide an overall judgment on the severity of the problems, nor does it explain how this judgment was reached. For monitoring personnel at the terminal, what they need is a clear, direct conclusion containing decision support information, not a set of data requiring secondary interpretation. To transform the complex risk analysis results into an immediately usable early warning instruction with a complete interpretable chain, the decoupled risk factors need to be finally synthesized, rated, and attributed. Therefore, this application integrates the quantified risk distribution vector with the attention weight vector representing the importance of the data source, ultimately generating a structured, interpretable early warning signal containing risk level, core components, and attribution evidence.
[0041] Figure 4 This is a flowchart of step S5 in the intelligent monitoring and early warning method based on multi-dimensional state perception according to an embodiment of this application. Figure 4 As shown, in one embodiment, step S5, based on the inter-modal attention weight vector, performs composite risk synthesis on the latent risk factor distribution vector to obtain an interpretable early warning signal, including: step S51, performing composite risk quantification and early warning level determination on the latent risk factor distribution vector based on the early warning threshold set to obtain an early warning level and a composite risk score; step S52, based on the inter-modal attention weight vector, performing risk attribution evidence extraction and structuring on the early warning level and the latent risk factor distribution vector to obtain a structured evidence package; step S53, performing formatted generation of an interpretable early warning signal on the early warning level, the composite risk score, and the structured evidence package to obtain the interpretable early warning signal.
[0042] In the above implementation, step S5 is implemented as follows: First, in step S51: a single overall risk severity index is extracted from the confidence levels of multiple risk factors, and risk levels are classified accordingly. The first step is risk synthesis. In order to directly reflect the most urgent risk level and avoid diluting the severity of a high-risk item by averaging multiple low-risk items, the maximum value method (Max Operator) is used to calculate the composite risk score. That is, the maximum value in the input latent risk factor distribution vector [0.924,0.953,0.858,0.378,0.119] is used as the composite risk score. Therefore, the composite risk score is max(0.924,0.953,0.858,0.378,0.119), which is 0.953. The second step is warning level determination. This process requires a set of warning thresholds, which is jointly formulated by medical experts and engineering technicians according to the needs of clinical guidelines and actual application scenarios. This set is used to quantify continuous risk scores into discrete levels with clear operational directions. In this embodiment, the warning threshold set is set as {Tl=0.70, Th=0.90}, where Tl is the lower limit of moderate risk and Th is the lower limit of high risk. Based on this warning threshold set, the specific level determination criteria are as follows: if the composite risk score is below 0.70, the warning level is determined to be 0, representing no risk; if the composite risk score is not lower than 0.70 but lower than 0.90, the warning level is determined to be 1, representing moderate risk; if the composite risk score is not lower than 0.90, the warning level is determined to be 2, representing high risk. In this embodiment, the composite risk score of 0.953 calculated in the previous step is compared with this threshold set. Since 0.953 ≥ Th0.90, the warning level is determined to be 2, representing high risk. The third step is process control. Because the determined warning level is 2, which is not equal to 0, the warning process continues. This step ultimately outputs two key pieces of information: the warning level, whose value is an integer 2; and the composite risk score, whose value is a scalar 0.953.
[0043] Next, in step S52: a data package is constructed to explain the reasons for the warning in detail. The first step is the extraction of the main risk components. In order to screen out the core risk factors that constitute this warning, a risk component screening threshold Tf is set. This threshold is defined by domain experts and aims to filter out risk items with low confidence that may be noise. In this embodiment, its value is set to 0.5. Traversing the input latent risk factor distribution vector [0.924, 0.953, 0.858, 0.378, 0.119], all risk factors with a confidence score of not less than 0.5 are screened out. The screening results are: unstable cardiac function (0.924), insufficient circulatory volume (0.953), and neuromuscular control abnormality (0.858). Then, these screened risk factors and their confidence scores are sorted from high to low to form a list, which is stored in the main risk component field of the structured evidence package. The list is [('Insufficient Circulatory Volume', 0.953), ('Unstable Cardiac Function', 0.924), ('Neuromuscular Control Abnormalities', 0.858)]. The second step is the attribution of key data sources. To trace which data modalities contributed most to the judgment of this warning, a modality contribution screening threshold Tm is set. This threshold is used to identify the data sources that played a decisive role, and in this embodiment, its value is set to 0.1. The input inter-modal attention weight vector [0.48, 0.40, 0.12] is traversed, and all data modalities with a weight not less than 0.1 are screened out. The screening results are: physiological (0.48), behavioral (0.40), and contextual (0.12). Similarly, these screened modalities and their weights are sorted from high to low according to their weights to form a list, which is stored in the key data source field of the structured evidence package. This list is [('Physiological', 0.48), ('Behavioral', 0.40), ('Contextual', 0.12)]. At this point, a structured evidence package containing the core content of the early warning and the basis for judgment has been fully constructed.
[0044] Finally, in step S53: all structured information is integrated into a clear, easily human-understandable text or object. The output is a formatted, interpretable warning signal. The processing flow is as follows: The first step is template selection. Based on the input warning level 2, a high-risk warning template corresponding to it is selected from a predefined template library. This template library presets different information frames and tones for different risk levels. The second step is information filling. The data obtained in the previous two steps are filled into the preset placeholders of the template. The warning level 2 is converted into a text description of high risk. The composite risk score 0.953 is formatted into an easily readable percentage form, i.e., 95.3%. Then, the list of main risk components in the structured evidence package is traversed and converted into descriptive text strings, such as the main risk components being: insufficient circulatory capacity (confidence: 95.3%), unstable cardiac function (confidence: 92.4%), and neuromuscular control abnormalities (confidence: 85.8%). Next, the list of key data sources is traversed, and each source is converted into a descriptive text string. For example, this judgment is primarily based on a comprehensive analysis of the following data sources: physiological data (contribution: 48.0%), behavioral data (contribution: 40.0%), and contextual data (contribution: 12.0%). The third step is the construction and output of the signal. All the filled-in content in the template is integrated to form the final interpretable warning signal. This signal can be a JSON object or a structured text. Taking text as an example, the final output signal is as follows: Risk Level: High Risk; Overall Risk Score: 95.3%; Risk Summary: A complex health risk consisting of multiple factors has been detected, with high severity, and immediate attention is recommended; The main risk components identified include insufficient circulating volume (confidence 95.3%), unstable cardiac function (confidence 92.4%), and neuromuscular control abnormalities (confidence 85.8%); This warning judgment is mainly based on a comprehensive analysis of physiological data (contribution 48.0%), behavioral data (contribution 40.0%), and contextual data (contribution 12.0%). This output signal has a clear structure and complete information, clearly indicating the severity level of the risk, listing the specific risk content and its confidence level in detail, and further tracing the key data modalities on which this judgment was based and their contribution, providing comprehensive decision support information for caregivers.
[0045] In summary, the intelligent monitoring and early warning method based on multidimensional state perception, as described in the embodiments of this application, is elucidated. It no longer analyzes a single data stream in isolation, but instead treats the physiological, behavioral, and environmental context data of the monitored object as an organic whole for unified input. By deeply modeling the nonlinear interaction relationships and synergistic effects between different modalities of data, the model can generate a fusion vector that comprehensively reflects an individual's overall health status. This directly overcomes the limitations of existing technologies that treat risk as an atomized risk modeling event. Furthermore, the model actively decouples multiple potential latent risk factors from this fusion vector and quantifies and synthesizes composite risks based on the dynamic contribution of each data modality in the risk formation process. This mechanism abandons the linear causal assumption implicit in traditional technologies, enabling it to accurately identify and warn of progressive health deterioration processes driven by multiple subclinical states, ultimately outputting interpretable signals containing evidence of risk tracing, thereby effectively solving the key technical problems raised in the background art.
[0046] Figure 5 This is a block diagram of an intelligent monitoring and early warning system based on multi-dimensional state perception, according to an embodiment of this application. Figure 5 As shown, the intelligent monitoring and early warning system 100 based on multidimensional state perception according to an embodiment of this application includes: a multidimensional state data acquisition module 110, used to acquire physiological data streams, behavioral data streams, and context data streams; a multidimensional state data encoding module 120, used to perform internal modal feature encoding and temporal feature extraction on the physiological data streams, behavioral data streams, and context data streams to obtain physiological temporal feature vectors, behavioral temporal feature vectors, and context temporal feature vectors; an intermodal interaction analysis module 130, used to model and fuse the intermodal interaction relationships of the physiological temporal feature vectors, behavioral temporal feature vectors, and context temporal feature vectors to obtain fused context vectors and intermodal attention weight vectors; a latent risk factor decoupling module 140, used to decouple latent risk factors on the fused context vectors to obtain latent risk factor distribution vectors; and an early warning signal generation module 150, used to perform composite risk synthesis on the latent risk factor distribution vectors based on the intermodal attention weight vectors to obtain interpretable early warning signals.
[0047] Specifically, in a preferred embodiment, the intelligent monitoring and early warning method based on multi-dimensional state perception of this application includes: Step 1, acquiring physiological data stream, behavioral data stream, and contextual data stream; Step 2, performing internal modal feature encoding and temporal feature extraction on the physiological data stream, behavioral data stream, and contextual data stream to obtain physiological temporal feature vector, behavioral temporal feature vector, and contextual temporal feature vector; Step 3, modeling the intermodal interaction relationship of the physiological temporal feature vector, behavioral temporal feature vector, and contextual temporal feature vector to obtain intermodal attention weight vector; Step 4, performing dynamic risk factor decoupling based on attention gating on the physiological temporal feature vector, behavioral temporal feature vector, contextual temporal feature vector, and intermodal attention weight vector to obtain latent risk factor distribution vector; Step 5, performing composite risk synthesis on the latent risk factor distribution vector based on the intermodal attention weight vector to obtain an interpretable early warning signal. In particular, the implementation process of obtaining data in steps 1, 2, 3, and 5 of this preferred embodiment is the same as the implementation process of the above-described embodiment of the intelligent monitoring and early warning method based on multi-dimensional state perception, and therefore will not be described in detail. Here, the implementation of step 4 will be described in detail.
[0048] It is worth noting that in the basic implementation described above, the strategy of directly mapping by first fusing all modal information into a unified context vector and then inputting it into a fully connected neural network layer for risk decoupling has inherent limitations. This process inevitably discards or dilutes the inter-modal attention weights, which contain rich attribution information and are calculated in the modality fusion step—that is, the dynamic judgment about which data sources are more critical in the current context—making the final decision-making process lack transparency. More importantly, this method forces the model to use a static and shared set of evaluation weights for all risk factors with different properties (e.g., assessing insufficient recurrent capacity versus assessing emotional load), which ignores the differences in the focus of different risk factors on different modal features, thus limiting the model's dynamic adaptability and recognition accuracy. Therefore, to overcome the drawbacks of information bottlenecks and static evaluation, this application proposes a dynamic risk factor gating network based on attention weights. This preferred method aims to construct a dynamic, customized assessment path for each risk factor, and directly integrates intermodal attention weights as the core gating signal into this path, thereby achieving more accurate, dynamic, and interpretable decoupling of latent risks. By learning a dedicated detector for each risk factor and dynamically weighting evidence from different modalities according to the current context, the accuracy of identifying latent and complex risks will be significantly improved.
[0049] Based on this, in one embodiment, step 4 involves performing attention-gated dynamic risk factor decoupling on the physiological time-series feature vector, behavioral time-series feature vector, contextual time-series feature vector, and inter-modal attention weight vector to obtain the latent risk factor distribution vector, including: A predefined set of latent risk factors is first defined to provide a clearly defined and clinically meaningful semantic space for subsequent risk decoupling. This ensures that the risk assessment results output by the model are human-understandable and actionable. This set is determined based on a clinical medical knowledge base and a common health risk profile of the target monitored population. In this embodiment, the set k remains consistent with the previous definition, i.e., k=5, and the set content is: {cardiac instability, insufficient circulatory volume, neuromuscular control abnormalities, emotional burden, potential infection risk}.
[0050] initialization A learnable weight matrix and a set of implicit risk factors corresponding to it. First, a learnable bias vector and a learnable score vector are generated. Second, a dedicated processing module is initialized. This is to build a specialized toolset for the accurate assessment of each risk factor, enabling the model to call different parameters for different risk assessment tasks, achieving a hierarchical information processing approach: first gathering evidence, then making judgments based on that evidence. Specifically, initialization... There are 15 learnable weight matrices, which is 5 × 3 = 15. and the set of implicit risk factors corresponding to There are five learnable bias vectors. and There are 5 learnable rating vectors. . It is a 256×256 dimensional matrix. and All parameters are 256-dimensional vectors, and all of these parameters are learned through backpropagation during the overall training phase of the model.
[0051] Will The temporal feature vectors of each mode and their corresponding modes Multiplying the learnable weight matrices together yields... Modal Each risk-modality-specific feature vector is generated. Next, risk-modality-specific features are extracted to assess different aspects of the same modality required for evaluating different risk factors. For example, assessing circulatory insufficiency requires attention to heart rate and blood pressure fluctuations in physiological data, while assessing cardiac instability focuses more on the specific morphology of the electrocardiogram waveform. This specific transformation effectively solves the problems of information bottlenecks and representation conflicts, extracting the most relevant professional opinions from various general modal features for the assessment of each risk factor. This process follows the formula: Taking the assessment of risk factor k = insufficient circulatory capacity as an example, the 256-dimensional physiological temporal feature vector is multiplied by a dedicated, learnable weight matrix, W{physiological, insufficient circulatory capacity}, to obtain a 256-dimensional vector v{physiological, insufficient circulatory capacity} that specifically reflects the features related to circulatory capacity in the physiological data. Similarly, this process is performed on the behavioral and contextual temporal feature vectors to obtain v{behavioral, insufficient circulatory capacity} and v{contextual, insufficient circulatory capacity}. This process is repeated for all five risk factors, ultimately generating 15 risk-modality-specific feature vectors.
[0052] Based on the intermodal attention weights in the intermodal attention weight vector A learnable bias vector, for Modal Logical values are calculated from risk-modal specific eigenvectors to obtain the set of latent risk factors. The final logical value corresponds to each implicit risk factor. Subsequently, a logical value calculation based on dynamic gating is performed. This not only considers the professional opinions of each modality on specific risks but also dynamically weights the overall credibility of each modality in the current context to make the most reasonable judgment, thereby utilizing the attention weights between modalities. As a gating signal, dynamic and weighted aggregation of professional opinions from various modalities is achieved. If the overall importance of the current physiological modality is high (i.e., the attention weight among first-modes is large), then the importance of the physiological modality in assessing all risk factors will increase accordingly. Furthermore, by introducing a scoring vector... The model can learn how to interpret... The extracted evidence is used to determine the likelihood that these combined characteristics point to risk. This process follows the formula: Continuing with the example of assessing k=circulatory capacity insufficiency, we first calculate the dynamic aggregation vector: 0.48*v{physiological,circulatory capacity insufficiency} + 0.40*v{behavioral,circulatory capacity insufficiency} + 0.12*v{context,circulatory capacity insufficiency}. This weighted sum is a 256-dimensional vector that dynamically integrates opinions from various modalities regarding circular capacity insufficiency, with the physiological and behavioral modalities being dominant due to their higher attention weights. Then, we add the bias vector b{circulatory capacity insufficiency} specific to this risk factor to this aggregation vector, and perform an inner product operation with the specific rating vector μ{circulatory capacity insufficiency} to obtain a final scalar logistic value l{circulatory capacity insufficiency}. This makes the model's judgment on this risk more certain; for example, the calculated logistic value is 3.5. We repeat this process for all five risk factors. Since the model can now better distinguish the relevant characteristics of different risks, the resulting logistic value vector might be [3.1, 3.5, 1.5, -0.8, -2.5].
[0053] Will The final logistic values corresponding to each latent risk factor form a latent risk factor logistic vector, which is then input into the Sigmoid activation function to obtain the latent risk factor distribution vector. Finally, the confidence scores of the latent risk factors are calculated to convert the unbounded logistic values into bounded, standardized confidence scores for easier understanding and subsequent processing. The result is a final, interpretable risk distribution. Each element of the 5-dimensional latent risk factor logistic vector [3.1, 3.5, 1.5, -0.8, -2.5] obtained in the previous step is independently input into the Sigmoid activation function f(x) = 1 / (1 + exp(-x)). After calculation, the final output latent risk factor distribution vector is: [sigmoid(3.1), sigmoid(3.5), sigmoid(1.5), sigmoid(-0.8), sigmoid(-2.5)], approximately equal to [0.957, 0.971, 0.818, 0.310, 0.076]. Compared to the distribution vector [0.924, 0.953, 0.858, 0.378, 0.119] obtained by the basic implementation method, the results obtained by this preferred method show that the model's confidence in the two main risks of cardiac instability and insufficient circulating volume is significantly improved, while the confidence in neuromuscular control abnormalities is reduced, and the confidence in irrelevant risks is further decreased. This indicates that this method, through a dedicated and dynamic assessment path, effectively enhances the identification ability of key risks and suppresses the interference of secondary or irrelevant information, thereby obtaining a more accurate and distinct risk distribution.
[0054] The resulting implicit risk factor distribution vector provides higher-quality input for the subsequent step 5. Due to its higher confidence level for key risks and stronger suppression of irrelevant risks, the signal-to-noise ratio of the risk distribution is significantly improved. This firstly allows the calculated composite risk score to more accurately reflect the peak severity of risks during composite risk quantification and level determination, thus making the judgment of the warning level more decisive and reliable. Secondly, during risk attribution evidence extraction, a more distinct risk distribution allows for more accurate screening of major risk components, effectively eliminating low-confidence interference items, and ultimately resulting in a more focused and pure structured evidence package and interpretable warning signal.
[0055] Here, those skilled in the art will understand that the specific operations of each step in the above-described intelligent monitoring and early warning system based on multi-dimensional state perception have been referenced above. Figures 1 to 4 The method for intelligent monitoring and early warning based on multidimensional state perception has been described in detail, and therefore, its repeated description will be omitted.
Claims
1. An intelligent monitoring and early warning method based on multi-dimensional state perception, characterized in that, include: Acquire physiological data streams, behavioral data streams, and contextual data streams; Internal modal feature encoding and temporal feature extraction are performed on physiological data stream, behavioral data stream and contextual data stream to obtain physiological temporal feature vector, behavioral temporal feature vector and contextual temporal feature vector; Modeling and fusing the intermodal interaction relationships of physiological temporal feature vectors, behavioral temporal feature vectors, and contextual temporal feature vectors to obtain the fused contextual vector and intermodal attention weight vector; The implicit risk factors are decoupled from the fused context vector to obtain the implicit risk factor distribution vector; Based on the intermodal attention weight vector, a composite risk synthesis is performed on the distribution vector of latent risk factors to obtain an interpretable early warning signal.
2. The intelligent monitoring and early warning method based on multi-dimensional state perception according to claim 1, characterized in that, Internal modal feature encoding and temporal feature extraction are performed on physiological data streams, behavioral data streams, and contextual data streams to obtain physiological temporal feature vectors, behavioral temporal feature vectors, and contextual temporal feature vectors, including: The physiological data stream is input into a time encoder based on a bidirectional gated recurrent unit network to obtain a physiological modality time hidden state sequence; The importance weights of key temporal nodes in the physiological modality temporal hidden state sequence are calculated to obtain the physiological modality attention weight vector. Based on the physiological modality attention weight vector, the physiological modality temporal hidden state sequence is weighted and fused to obtain the physiological temporal feature vector.
3. The intelligent monitoring and early warning method based on multi-dimensional state perception according to claim 2, characterized in that, The importance weights of key temporal nodes in the physiological modality temporal hidden state sequence are calculated to obtain the physiological modality attention weight vector. This includes calculating the importance weights of key temporal nodes in the physiological modality temporal hidden state sequence using the following formula: ; ; in, Let be the physiological modality temporal hidden state at the t-th time step in the physiological modality temporal hidden state sequence. The weight matrix is a learnable matrix. For learnable bias vectors, Let be the intermediate features of the physiological modality time series at time step t. For learnable globally shared context vectors, The length of the physiological modality temporal hidden state sequence. Let be the physiological modality attention weight at time step t in the physiological modality attention weight vector.
4. The intelligent monitoring and early warning method based on multi-dimensional state perception according to claim 1, characterized in that, The intermodal interaction relationships of physiological temporal feature vectors, behavioral temporal feature vectors, and contextual temporal feature vectors are modeled and fused to obtain the fused context vector and intermodal attention weight vector, including: Nonlinear projection is performed on the physiological time-series feature vector, behavioral time-series feature vector, and contextual time-series feature vector to obtain the projected physiological time-series feature vector, projected behavioral time-series feature vector, and projected contextual time-series feature vector; Intermodal importance weights are calculated on the projected physiological temporal feature vector, the projected behavioral temporal feature vector, and the projected contextual temporal feature vector to obtain the intermodal attention weights from the first to the third modality. The intermodal attention weights from the first to the third modality constitute the intermodal attention weight vector. Based on the attention weights between the first to third modalities, the physiological temporal feature vector, behavioral temporal feature vector, and contextual temporal feature vector are weighted and fused to obtain the fused contextual vector.
5. The intelligent monitoring and early warning method based on multi-dimensional state perception according to claim 4, characterized in that, Nonlinear projection is performed on the physiological time-series feature vector, behavioral time-series feature vector, and contextual time-series feature vector to obtain the projected physiological time-series feature vector, projected behavioral time-series feature vector, and projected contextual time-series feature vector. This includes performing nonlinear projection on the physiological time-series feature vector, behavioral time-series feature vector, and contextual time-series feature vector using the following formula: ; in, These can be physiological time-series feature vectors, behavioral time-series feature vectors, or contextual time-series feature vectors. For learnable projection matrices, is a learnable projection bias vector. for Activation function These are the projected physiological temporal feature vectors, projected behavioral temporal feature vectors, or projected contextual temporal feature vectors.
6. The intelligent monitoring and early warning method based on multi-dimensional state perception according to claim 4, characterized in that, The intermodal importance weights of the projected physiological temporal feature vector, the projected behavioral temporal feature vector, and the projected contextual temporal feature vector are calculated to obtain the attention weights between the first and third modalities. This includes calculating the intermodal importance weights of the projected physiological temporal feature vector, the projected behavioral temporal feature vector, and the projected contextual temporal feature vector using the following formula: ; in, This is a context reference vector shared by all learnable modalities. For the number of modes, This is the transpose of the vector. This refers to the intermodal attention weights among the attention weights between the first and third modalities.
7. The intelligent monitoring and early warning method based on multi-dimensional state perception according to claim 1, characterized in that, The latent risk factor distribution vector is obtained by decoupling the fused context vector from the latent risk factors, including: A predefined set of implicit risk factors; The fused context vector is input into a fully connected neural network layer to obtain the implicit risk factor logical vector; Independent risk confidence scores are calculated on the logistic vectors of the latent risk factors to obtain the distribution vectors of the latent risk factors.
8. The intelligent monitoring and early warning method based on multi-dimensional state perception according to claim 1, characterized in that, Based on the intermodal attention weight vector, a composite risk synthesis is performed on the distribution vector of latent risk factors to obtain interpretable early warning signals, including: Based on the set of early warning thresholds, the distribution vector of hidden risk factors is used to perform composite risk quantification and early warning level determination to obtain the early warning level and composite risk score. Based on the intermodal attention weight vector, risk attribution evidence is extracted and structured from the distribution vectors of early warning levels and hidden risk factors to obtain a structured evidence package. The interpretable warning signal is obtained by formatting and generating the warning level, composite risk score, and structured evidence package.
9. An intelligent monitoring and early warning system based on multi-dimensional state perception, characterized in that, include: The multidimensional state data acquisition module is used to acquire physiological data streams, behavioral data streams, and contextual data streams; The multidimensional state data encoding module is used to perform internal modal feature encoding and temporal feature extraction on physiological data streams, behavioral data streams, and contextual data streams to obtain physiological temporal feature vectors, behavioral temporal feature vectors, and contextual temporal feature vectors. The intermodal interaction analysis module is used to model and fuse the intermodal interaction relationships of physiological temporal feature vectors, behavioral temporal feature vectors, and contextual temporal feature vectors to obtain the fused context vector and intermodal attention weight vector. The implicit risk factor decoupling module is used to decouple the implicit risk factors from the fused context vector to obtain the implicit risk factor distribution vector. The early warning signal generation module is used to perform composite risk synthesis on the distribution vector of latent risk factors based on the intermodal attention weight vector to obtain an interpretable early warning signal.