Postoperative early complication early warning method based on multi-modal fusion and related device
By constructing a postoperative spatiotemporally coupled attention risk prediction model and combining physiological and clinical data to dynamically assess patient risk, the accuracy and individualization problems of traditional postoperative complication early warning methods are solved, and efficient early postoperative complication early warning is achieved.
Patent Information
- Application Number
- CN202511131006.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-13
- Publication Date
- 2025-11-14
AI Technical Summary
Traditional methods for predicting postoperative complications rely on clinical experience and single-modality data, lacking in-depth integration of multi-source heterogeneous data. This makes it difficult to capture subtle physiological changes and the synergistic effects of multiple factors, resulting in inaccurate predictions and a lack of consideration for individualized risk tolerance and interventions for patients.
By constructing a postoperative spatiotemporal coupled attention risk prediction model, combining physiological temporal features collected by wearable devices and clinical context features from the hospital information system, and employing multi-granular temporal coding, hierarchical semantic context coding, cross-modal attention fusion, and a complication-specific prediction layer, the model dynamically assesses patient risk and provides individualized early warnings.
It significantly improves the accuracy and transparency of early postoperative complication warning, enables dynamic assessment of risk levels and intervention windows, provides reliable intervention recommendations, and enhances the model's interpretability and ability to identify rare complications.
Smart Images

Figure CN120954597A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical auxiliary diagnostic technology, specifically to a method and device for early warning of postoperative complications based on multimodal fusion, and a computing device. Background Technology
[0002] Traditional methods for early warning of postoperative complications mainly rely on the clinical experience of healthcare professionals, regular ward rounds, manual record-keeping, and simple clinical indicator monitoring. Because monitoring methods are often intermittent, they cannot achieve continuous, real-time monitoring of the patient's physiological state, especially during out-of-hospital periods or at night. While some methods utilize wearable devices to collect data, these often represent single-modality physiological data, lacking deep integration with the patient's overall clinical context. Furthermore, by focusing only on threshold abnormalities of a few key indicators, they fail to fully utilize the inherent correlations and temporal evolution characteristics of multi-source heterogeneous data (such as physiological data, clinical records, laboratory tests, and medical imaging), resulting in insufficient early identification of complications and difficulty in capturing subtle physiological changes and risks arising from the synergistic effects of multiple factors.
[0003] While some studies utilize machine learning methods for prediction, most models tend to use single-modal data or simply piece together multimodal data, failing to effectively capture the complex spatiotemporal coupling relationships and potential semantic information between different modalities. Furthermore, they do not adequately consider data gaps, small sample sizes, model interpretability, and uncertainty quantification. Existing algorithms often provide one-way risk warnings, lacking further interpretation of the warning information (such as the source of risk and the intervention window), and fail to fully consider patients' individual risk tolerance, the severity of complications, and their manageability, resulting in inaccurate warnings and potentially leading to over- or under-intervention. Moreover, they offer weak support for healthcare professionals' decision-making, making it difficult to provide standardized and evidence-based intervention recommendations.
[0004] To address the aforementioned issues, this invention proposes a method for early postoperative complication prediction based on multimodal fusion. By constructing a postoperative spatiotemporally coupled attention risk prediction model, it aims to provide early warnings of postoperative complications and improve the accuracy of the warnings. Summary of the Invention
[0005] In view of the above problems, the present invention provides a method and device for early warning of postoperative complications based on multimodal fusion, as well as a computing device.
[0006] According to one aspect of the present invention, a method for early postoperative complication warning based on multimodal fusion is provided, comprising:
[0007] The postoperative physiological time-series characteristics tensor of the patient is collected by wearable devices. The postoperative physiological time-series characteristics tensor includes heart rate, blood oxygen saturation, respiratory rate, body temperature, skin conductance and activity level.
[0008] The patient's clinical context feature tensor is obtained through the hospital information system, electronic medical record system, and medical image archiving and communication system. The clinical context feature tensor includes clinical static data, laboratory test results, medical image reports, and medical records.
[0009] The postoperative physiological temporal feature tensor and the clinical context feature tensor are input into the postoperative spatiotemporal coupled attention risk prediction model to predict the real-time probability of the occurrence of preset complications. The postoperative spatiotemporal coupled attention risk prediction model includes a multi-granularity temporal coding layer, a hierarchical semantic context coding layer, a cross-modal attention fusion layer, and a complication-specific prediction layer. The complication-specific prediction layer includes a multi-task prediction head, a sequence decision layer, and an uncertainty quantification layer.
[0010] Early postoperative complication warnings are provided based on the patient's individualized risk tolerance, the severity of anticipated complications, and their modifiability.
[0011] In an alternative approach, the method further includes:
[0012] The physiological temporal feature tensor and the clinical context feature tensor are mapped to a unified latent space tensor by a variational autoencoder (VAE).
[0013] The missing data of the unified latent space tensor is reconstructed using a cross-modal generator;
[0014] The evidence lower bound loss function of the unified latent space tensor is:
[0015] ;
[0016] in, For physiological data; For clinical data; To unify the latent space tensor; It is the divergence function; Let z be the prior distribution of z; Let z be the posterior probability distribution of the latent variable z when the data are x and y; The conditional probability of generating x and y when z is z.
[0017] In one alternative approach, the complication-specific prediction layer employs a Markov decision process to transfer risk states;
[0018] The risk status includes the current risk level, the remaining intervention window, and the patient's physiological stability indicators.
[0019] The learning formula for the state transition probability of the risk state is:
[0020] ;
[0021] in, To transition to the next risk state after performing action a in the current risk state s. The probability of; To transition to state after performing action a from state s. The number of times; For smoothing coefficients; Let s be the total number of possible states in the risk state space.
[0022] In an alternative approach, during the training of the postoperative spatiotemporally coupled attention risk prediction model, the method further includes:
[0023] A self-supervised task is constructed based on unlabeled postoperative physiological data, which includes temporal data reconstruction and future step size prediction.
[0024] After pre-training, the encoder layer parameters are frozen, and supervised fine-tuning is performed on the labeled dataset.
[0025] Specifically, for small sample complication types, a meta-learning strategy is used to initialize model parameters and achieve rapid convergence for multi-round task adaptation.
[0026] In one alternative approach, the clinical context feature tensor is used to calculate causal effects using the front door adjustment criterion;
[0027] For continuous confounding factors, a propensity score matching method is used to construct a balanced dataset.
[0028] In one alternative approach, the causal contribution formula for the causal effect is:
[0029] ;
[0030] in, Let X be the conditional probability distribution of the target variable Y when the intervention X=x is applied. Let z be the conditional probability that the mediator variable Z takes the value z given that X=x; Let Y be the conditional probability distribution of the target variable Y given that X=x and Z=z; Let X be the intervention function, representing the state of setting variable X to x.
[0031] In one alternative approach, the sequence decision layer includes:
[0032] The risk summary layer displays real-time risk values, complication types, and intervention recommendation levels.
[0033] The evidence traceability layer demonstrates the strength of the association between physiological indicators and clinical factors;
[0034] The intervention protocol layer is used to recommend standardized intervention measures based on the guideline library and localization protocols, and to label the evidence level.
[0035] In one alternative approach, the uncertainty quantification layer randomly activates some or all of the neural network layers in the complication-specific prediction layer during the model inference phase.
[0036] In this process, multiple forward propagations are performed on the same input data to obtain multiple sets of prediction results;
[0037] The mean and variance of the predicted result set are calculated to obtain the risk prediction value and the corresponding uncertainty measure, respectively.
[0038] According to another aspect of the present invention, a postoperative early complication warning device based on multimodal fusion is provided, comprising:
[0039] The physiological time series data acquisition module is used to collect postoperative physiological time series feature tensors of patients through wearable devices. The postoperative physiological time series feature tensors include heart rate, blood oxygen saturation, respiratory rate, body temperature, skin conductance, and activity level.
[0040] The clinical context data acquisition module is used to acquire the patient's clinical context feature tensor through the hospital information system, electronic medical record system, and medical image archiving and communication system. The clinical context feature tensor includes clinical static data, laboratory test results, medical image reports, and medical records.
[0041] A multimodal fusion risk prediction module is used to input the postoperative physiological temporal feature tensor and the clinical context feature tensor into the postoperative spatiotemporal coupled attention risk prediction model to predict the real-time probability of the occurrence of preset complications; wherein, the postoperative spatiotemporal coupled attention risk prediction model includes a multi-granularity temporal coding layer, a hierarchical semantic context coding layer, a cross-modal attention fusion layer and a complication-specific prediction layer, and the complication-specific prediction layer includes a multi-task prediction head, a sequence decision layer and an uncertainty quantification layer;
[0042] The individualized early warning generation module is used to provide early warnings of postoperative complications based on the patient's individualized risk tolerance, the severity of preset complications, and the modifiability of intervention.
[0043] According to another aspect of the present invention, a computing device is provided, comprising: a processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other via the communication bus;
[0044] The memory is used to store at least one executable instruction, which causes the processor to perform the operation corresponding to the above-described method for early postoperative complication warning based on multimodal fusion.
[0045] According to the solution provided by this invention, a wearable device is used to collect postoperative physiological temporal feature tensors of the patient, including heart rate, blood oxygen saturation, respiratory rate, body temperature, skin conductance, and activity level. The patient's clinical context feature tensors are obtained through a hospital information system, electronic medical record system, and medical image archiving and communication system. These clinical context feature tensors include clinical static data, laboratory test results, medical image reports, and medical records. The postoperative physiological temporal feature tensors and the clinical context feature tensors are input into a postoperative spatiotemporally coupled attention risk prediction model to predict the real-time probability of preset complications. The postoperative spatiotemporally coupled attention risk prediction model includes a multi-granularity temporal coding layer, a hierarchical semantic context coding layer, a cross-modal attention fusion layer, and a complication-specific prediction layer. The complication-specific prediction layer includes a multi-task prediction head, a sequence decision layer, and an uncertainty quantification layer. Based on the patient's individualized risk tolerance, the severity of preset complications, and the modifiability of intervention, early postoperative complication warnings are provided. This invention significantly improves the accuracy of early postoperative complication warnings through the postoperative spatiotemporally coupled attention risk prediction model. Specifically, missing data in medical data is addressed by using a variational autoencoder for unified latent space mapping and a cross-modal generator to reconstruct missing data. Risk states are transferred based on Markov decision processes, enabling dynamic assessment of patient risk levels, remaining intervention windows, and physiological stability, thus providing dynamic risk assessment based on the current state. The sequence decision layer, comprising risk summary, evidence tracing, and intervention protocol layers, not only displays risk values but also traces back to key physiological indicators and clinical factors, enhancing the model's transparency and interpretability. A meta-learning method is employed for small sample complication types, allowing the model to quickly adapt to new tasks and improve its ability to identify rare complications. The uncertainty quantification layer calculates the mean and variance of prediction results through multiple forward propagations, providing healthcare professionals with risk prediction values and a measure of their reliability. The causal effect of the clinical context feature tensor is calculated using the front-door adjustment criterion, and a balanced dataset is constructed using propensity score matching, helping to reveal the true causal relationship between complications and clinical factors, making the prediction results more clinically credible.
[0046] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and in order to make the above and other objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description
[0047] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:
[0048] Figure 1 A flowchart illustrating an early postoperative complication warning method based on multimodal fusion according to an embodiment of the present invention is shown.
[0049] Figure 2 A schematic diagram of the framework of a postoperative early complication warning device based on multimodal fusion according to an embodiment of the present invention is shown;
[0050] Figure 3 A schematic diagram of the structure of a computing device according to an embodiment of the present invention is shown. Detailed Implementation
[0051] Exemplary embodiments of the invention will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the invention are shown in the drawings, it should be understood that the invention may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this invention will be thorough and complete, and will fully convey the scope of the invention to those skilled in the art.
[0052] Figure 1 A flowchart illustrating a method for early postoperative complication warning based on multimodal fusion according to an embodiment of the present invention is shown. Specifically, as... Figure 1 As shown, it includes the following steps:
[0053] Step S101: Collect postoperative physiological time-series characteristic tensors of the patient through wearable devices. The postoperative physiological time-series characteristic tensors include heart rate, blood oxygen saturation, respiratory rate, body temperature, skin conductance, and activity level.
[0054] In this embodiment, after postoperative patients are admitted to the ward or recovery room, medical staff assist them in correctly wearing the corresponding wearable devices, ensuring smooth communication between the wearable devices and the data receiving terminals (such as the gateway at the nurse station or edge computing devices). The wearable devices take the form of wristbands, finger clips, or patches, minimizing invasiveness and providing high comfort, facilitating free movement during postoperative recovery and reducing anxiety and discomfort. Simultaneously, the non-invasive nature avoids the risk of infection or skin damage associated with the monitoring itself. Specifically, wrist-worn devices (e.g., medical-grade smartwatches or bracelets) monitor heart rate, partial blood oxygen saturation, activity level, and partial body temperature. Finger clip-on pulse oximeters monitor blood oxygen saturation (SpO2) and pulse rate in real time. Chest patches / sensors monitor respiratory rate, ECG (heart rate), and skin conductance. Smart mattresses / pads monitor respiratory rate, heart rate, and body movement (activity level).
[0055] Step S102: Obtain the patient's clinical context feature tensor through the hospital information system, electronic medical record system, and medical image archiving and communication system. The clinical context feature tensor includes clinical static data, laboratory test results, medical image reports, and medical records.
[0056] In this embodiment, for example, clinical static data is obtained from HIS / EMR: Basic information: Zhang San, male, 75 years old, weight 60kg, height 170cm. Past medical history: hypertension (controlled with medication), type 2 diabetes (insulin therapy), 40-year smoking history (quit 5 years ago). Surgical information: radical gastrectomy for gastric cancer, operation time: 4 hours, intraoperative blood loss: 400ml, anesthesia method: general anesthesia. Medication information: postoperative use of pain pump, antibiotics (cephalosporin), hypoglycemic drugs, and antihypertensive drugs. Laboratory test results are obtained from EMR: Preoperative: hemoglobin 130g / L, C-reactive protein (CRP) 5mg / L, blood glucose 7.2mmol / L, creatinine 80umol / L. Postoperative Day 1: hemoglobin 95g / L (low), CRP 35mg / L (high), blood glucose 10.5mmol / L (high), creatinine 90umol / L. Postoperative Day 2: Blood routine: White blood cell count 15x10^9 / L (elevated), neutrophil percentage 85% (elevated); Coagulation function: PT 15s (prolonged), APTT 38s (prolonged). Medical imaging report obtained from PACS: Preoperative chest X-ray: No obvious exudative shadows seen in both lungs, heart size and shape normal. Postoperative Day 2 chest CT report: Patchy high-density shadows visible in the lower lobe of the right lung, with indistinct borders, bronchial air signs visible in some areas, and a small amount of pleural effusion. Medical records obtained from EMR: Doctor's ward round record (postoperative Day 1): The patient's mental state is good after surgery, with a chief complaint of mild abdominal distension and pain, and a small amount of drainage. Nurse's nursing record (postoperative Day 2): The patient coughed at night, coughing up a small amount of yellow sputum, with a body temperature of 38.2℃, rapid breathing, and a respiratory rate of 25 breaths / min. The incision dressing is dry. Consultation record (postoperative Day 2): Respiratory consultation: Lung infection is suspected, sputum culture is recommended, and antibiotics should be adjusted. The above information was integrated into a clinical context feature tensor. Numerical features included: age (75), weight (60), height (170), operation duration (4), blood loss (400), preoperative hemoglobin (130), postoperative Day 1 hemoglobin (95), CRP (35), blood glucose (10.5), white blood cells (15), PT (15), APTT (38), etc. Categorical features included: gender (male), past medical history (hypertension_Yes, diabetes_Yes, smoking history_Yes), type of surgery (radical gastrectomy), anesthesia method (general anesthesia), etc. Embedding of textual features included: imaging reports (such as "right lower lobe lobe high density shadow, air bronchogram" vector representation obtained after NLP processing), and medical records (such as "cough, yellow sputum, shortness of breath" vector representation obtained after NLP processing).
[0057] Step S103: Input the postoperative physiological temporal feature tensor and the clinical context feature tensor into the postoperative spatiotemporal coupled attention risk prediction model to predict the real-time probability of the occurrence of preset complications; wherein, the postoperative spatiotemporal coupled attention risk prediction model includes a multi-granularity temporal coding layer, a hierarchical semantic context coding layer, a cross-modal attention fusion layer and a complication-specific prediction layer, and the complication-specific prediction layer includes a multi-task prediction head, a sequence decision layer and an uncertainty quantification layer.
[0058] In this embodiment, a multi-granularity temporal coding layer captures the dynamic patterns and trends of physiological data over time, identifying subtle abnormal changes. For example, subtle, persistent changes in heart rate or respiratory rate may indicate early signs of infection or respiratory failure. Spatial correlation (hierarchical semantic context coding layer and cross-modal attention fusion layer) not only understands the complex relationships within a single data modality but also learns the interactions and influences between different modalities through the cross-modal attention fusion layer. The sequence decision layer provides real-time risk summaries, evidence tracing, and intervention recommendations, upgrading from simple prediction to decision support, enhancing the practicality and clinical value of the early warning system. Doctors not only know that there is a risk but also know the source of the risk and how to intervene. The uncertainty quantification layer assesses the confidence level of the model's prediction results; high-confidence predictions enhance doctors' decision-making confidence.
[0059] In an alternative approach, the method further includes:
[0060] The physiological temporal feature tensor and the clinical context feature tensor are mapped to a unified latent space tensor by a variational autoencoder (VAE).
[0061] The missing data of the unified latent space tensor is reconstructed using a cross-modal generator;
[0062] The evidence lower bound loss function of the unified latent space tensor is:
[0063] ;
[0064] in, For physiological data; For clinical data; To unify the latent space tensor; It is the divergence function; Let z be the prior distribution of z; Let z be the posterior probability distribution of the latent variable z when the data are x and y; The conditional probability of generating x and y when z is z.
[0065] In this embodiment, a variational autoencoder maps heterogeneous physiological temporal features and clinical context features to a unified latent space tensor, addressing the inconsistency in data formats, dimensions, and feature types across different modalities. This enables the model to perform feature learning even with incomplete data. (Reconstruction error term) This measures the quality of reconstructing the original data from the latent space, ensuring that the latent representation learned by the model retains important information from the original data. KL divergence term. Measuring the posterior distribution With prior distribution The distance between them encourages the learned latent space distribution to approach a pre-defined simple distribution, thereby ensuring the sampleability of the latent space and helping to generate reasonable missing data.
[0066] In one alternative approach, the complication-specific prediction layer employs a Markov decision process to transfer risk states;
[0067] The risk status includes the current risk level, the remaining intervention window, and the patient's physiological stability indicators.
[0068] The learning formula for the state transition probability of the risk state is:
[0069] ;
[0070] in, To transition to the next risk state after performing action a in the current risk state s. The probability of; To transition to state after performing action a from state s. The number of times; For smoothing coefficients; Let s be the total number of possible states in the risk state space.
[0071] In this embodiment, the remaining intervention window is considered as part of the risk state, enabling the model to perceive and optimize the optimal timing of interventions. This is crucial for early postoperative complications, as many complications can have significantly improved prognoses with timely intervention. Patient physiological stability indicators allow the model to more comprehensively assess the patient's overall health, considering not only the abnormality of a single indicator but also the volatility and trends of physiological parameters and the patient's responsiveness to interventions. The learning formula for state transition probabilities allows the model to learn from historical data the influence of different risk states and intervention actions on risk state transitions. This allows for continuous optimization and adaptation based on actual clinical data, improving its predictive utility in the real world. Compared to traditional black-box models, by analyzing the transition paths between different states and related intervention actions, healthcare professionals can better understand the reasons behind the early warning results and the recommended intervention logic at different risk levels.
[0072] In an alternative approach, during the training of the postoperative spatiotemporally coupled attention risk prediction model, the method further includes:
[0073] A self-supervised task is constructed based on unlabeled postoperative physiological data, which includes time-series data reconstruction and future step size prediction.
[0074] After pre-training, the encoder layer parameters are frozen, and supervised fine-tuning is performed on the labeled dataset.
[0075] Specifically, for small sample complication types, a meta-learning strategy is used to initialize model parameters and achieve rapid convergence for multi-round task adaptation.
[0076] In this embodiment, labeled data for postoperative complications, in particular, is often scarce, while unlabeled physiological time-series data is relatively easy to obtain (e.g., continuous collection by wearable devices). Self-supervised pre-training allows the model to learn general and useful feature representations from a large amount of unlabeled physiological data, thereby alleviating the problem of insufficient labeled data. Two self-supervised tasks—time-series data reconstruction and future step size prediction—force the model to learn the temporal dependence, periodicity, trends, and abnormal patterns of physiological time-series data, enabling the pre-trained encoder to better capture the complex changes in physiological signals over time. Postoperative complications are diverse, but case data for some rare complications are very limited. Traditional supervised learning is prone to overfitting in small sample situations. Meta-learning methods enable the model to learn the ability to quickly adapt to new tasks, and even with only a small amount of labeled data, it can quickly adjust parameters to achieve effective prediction, significantly improving the ability to identify rare complications.
[0077] In one alternative approach, the clinical context feature tensor is used to calculate causal effects using the front door adjustment criterion;
[0078] For continuous confounding factors, a propensity score matching method is used to construct a balanced dataset.
[0079] In this embodiment, the model identifies and quantifies the true causal effect between clinical context features and complications through the front-door adjustment criterion. The model not only learns "what" happened, but also learns more deeply "why" it happened. Furthermore, since clinical data often contains numerous confounding factors (i.e., variables that simultaneously affect exposure and outcome), propensity score matching can effectively balance the distribution of continuous confounding factors in the observational data, making the treatment and control groups comparable in terms of these confounding factors. This reduces the interference of confounding factors on causal inference and improves the accuracy of causal effect estimation.
[0080] In one alternative approach, the causal contribution formula for the causal effect is:
[0081] ;
[0082] in, Let X be the conditional probability distribution of the target variable Y when the intervention X=x is applied. Let z be the conditional probability that the mediator variable Z takes the value z given that X=x; Let Y be the conditional probability distribution of the target variable Y given that X=x and Z=z; Let X be the intervention function, representing the state of setting variable X to x.
[0083] In this embodiment, the front-door adjustment criterion can handle situations with unobservable confounding factors. When traditional back-door adjustment fails due to the inability to fully observe confounding factors, front-door adjustment provides an alternative. By identifying mediating variables, it can fully mediate the causal effect from X to Y and, under certain conditions, bypass the difficulty of directly adjusting for unobservable confounding factors. Front-door adjustment more accurately estimates the true causal effect of a specific intervention (such as medication use or treatment regimen) on postoperative complications (target variable Y), rather than merely a statistical association. This helps avoid spurious correlations caused by confounding factors, leading to more scientific early warning and intervention decisions. Calculating causal contribution not only provides predictive results but also reveals why certain factors lead to complications. By understanding causal relationships, healthcare professionals can better understand the underlying mechanisms, rather than simply making judgments based on correlations. The do(X=x) operator represents an intervention, not just an observation. It can answer the counterfactual question, "What would happen to Y if X were set to a specific value x?" This allows for simulating the impact of different interventions on patient prognosis, thereby selecting the optimal intervention strategy.
[0084] In one alternative approach, the sequence decision layer includes:
[0085] The risk summary layer displays real-time risk values, complication types, and intervention recommendation levels.
[0086] The evidence traceability layer demonstrates the strength of the association between physiological indicators and clinical factors;
[0087] The intervention protocol layer is used to recommend standardized intervention measures based on the guideline library and localization protocols, and to label the evidence level.
[0088] In this embodiment, the evidence traceability layer directly demonstrates the correlation strength between physiological indicators and clinical factors (i.e., the model's input features) and the predicted outcome (risk value). This allows healthcare professionals to understand "why" such a warning is issued, rather than simply accepting the numbers. When healthcare professionals see the underlying evidence supporting the predicted outcome, their trust in the system increases significantly, making them more willing to adopt the system's recommendations.
[0089] In one alternative approach, the uncertainty quantification layer randomly activates some or all of the neural network layers in the complication-specific prediction layer during the model inference phase.
[0090] In this process, multiple forward propagations are performed on the same input data to obtain multiple sets of prediction results;
[0091] The mean and variance of the predicted result set are calculated to obtain the risk prediction value and the corresponding uncertainty measure, respectively.
[0092] In this embodiment, when the predicted risk is high but accompanied by significant uncertainty, doctors may take a more cautious approach, such as conducting further tests for confirmation, rather than immediately performing invasive treatment. Conversely, if the risk value is moderate but the certainty is high, more confident early warning measures can be taken. If the model consistently outputs high uncertainty for certain types of input data (e.g., cases of rare complications, patients with poor data quality), it indicates that the model's learning is insufficient, or that the data itself has inherent ambiguity. Revealing when the model is "uncertain" is itself a valuable insight, suggesting under what circumstances the model might "hesitate" or "be confused," and helps in understanding the boundaries of model behavior. Using Dropout during training and maintaining Dropout activation during inference allows for uncertainty quantification during the inference phase without major modifications to the model structure or training process.
[0093] Step S104: Based on the patient's individualized risk tolerance, the severity of the pre-set complications, and the modifiability of intervention, conduct early postoperative complication warning.
[0094] In this embodiment, for example, if patient A (low-risk tolerance) receives an alert, the model predicts a 40% probability of postoperative hypokalemia and a 5% probability of pulmonary embolism. Traditional alerts might only issue a low-level warning for hypokalemia, but not for pulmonary embolism. For hypokalemia (40% probability, low to moderate severity, high intervention capability), although the probability is not particularly high, considering the patient's low risk tolerance and high intervention capability, early intervention can significantly avoid serious consequences. The alert outputs that patient A has a moderate risk of hypokalemia (predicted probability 40%). Considering their advanced age, multiple illnesses, and poor tolerance, it is recommended to immediately check blood electrolytes and prepare for potassium supplementation. For pulmonary embolism (5% probability, high severity, moderate intervention capability), although the probability is low, the severity of pulmonary embolism is high, and the consequences can be severe if it occurs. The alert outputs that patient A has a low-risk but life-threatening pulmonary embolism possibility (5% predicted probability). It is recommended to strengthen monitoring of lung-related signs, such as dyspnea and chest pain, and consider further imaging examinations if abnormalities occur.
[0095] According to the solution provided by this invention, a wearable device is used to collect postoperative physiological temporal feature tensors of the patient, including heart rate, blood oxygen saturation, respiratory rate, body temperature, skin conductance, and activity level. The patient's clinical context feature tensors are obtained through a hospital information system, electronic medical record system, and medical image archiving and communication system. These clinical context feature tensors include clinical static data, laboratory test results, medical image reports, and medical records. The postoperative physiological temporal feature tensors and the clinical context feature tensors are input into a postoperative spatiotemporally coupled attention risk prediction model to predict the real-time probability of preset complications. The postoperative spatiotemporally coupled attention risk prediction model includes a multi-granularity temporal coding layer, a hierarchical semantic context coding layer, a cross-modal attention fusion layer, and a complication-specific prediction layer. The complication-specific prediction layer includes a multi-task prediction head, a sequence decision layer, and an uncertainty quantification layer. Based on the patient's individualized risk tolerance, the severity of preset complications, and the modifiability of intervention, early postoperative complication warnings are provided. This invention significantly improves the accuracy of early postoperative complication warnings through the postoperative spatiotemporally coupled attention risk prediction model. Specifically, missing data in medical data is addressed by using a variational autoencoder for unified latent space mapping and a cross-modal generator to reconstruct missing data. Risk states are transferred based on Markov decision processes, enabling dynamic assessment of patient risk levels, remaining intervention windows, and physiological stability, thus providing dynamic risk assessment based on the current state. The sequence decision layer, comprising risk summary, evidence tracing, and intervention protocol layers, not only displays risk values but also traces back to key physiological indicators and clinical factors, enhancing the model's transparency and interpretability. A meta-learning method is employed for small sample complication types, allowing the model to quickly adapt to new tasks and improve its ability to identify rare complications. The uncertainty quantification layer calculates the mean and variance of prediction results through multiple forward propagations, providing healthcare professionals with risk prediction values and a measure of their reliability. The causal effect of the clinical context feature tensor is calculated using the front-door adjustment criterion, and a balanced dataset is constructed using propensity score matching, helping to reveal the true causal relationship between complications and clinical factors, making the prediction results more clinically credible.
[0096] Figure 2 A schematic diagram of the framework of a postoperative early complication early warning device based on multimodal fusion according to an embodiment of the present invention is shown. The postoperative early complication early warning device based on multimodal fusion includes:
[0097] The physiological time series data acquisition module 210 is used to acquire the postoperative physiological time series feature tensor of the patient through a wearable device. The postoperative physiological time series feature tensor includes heart rate, blood oxygen saturation, respiratory rate, body temperature, skin conductance and activity level.
[0098] The clinical context data acquisition module 220 is used to acquire the patient's clinical context feature tensor through the hospital information system, electronic medical record system and medical image archiving and communication system. The clinical context feature tensor includes clinical static data, laboratory test results, medical image reports and medical records.
[0099] The multimodal fusion risk prediction module 230 is used to input the postoperative physiological temporal feature tensor and the clinical context feature tensor into the postoperative spatiotemporal coupled attention risk prediction model to predict the real-time probability of the occurrence of preset complications; wherein, the postoperative spatiotemporal coupled attention risk prediction model includes a multi-granularity temporal coding layer, a hierarchical semantic context coding layer, a cross-modal attention fusion layer and a complication-specific prediction layer, and the complication-specific prediction layer includes a multi-task prediction head, a sequence decision layer and an uncertainty quantification layer;
[0100] The individualized early warning generation module 240 is used to provide early warnings of postoperative complications based on the patient's individualized risk tolerance, the severity of preset complications, and the modifiability of intervention.
[0101] Figure 3 The diagram shows a structural schematic of an embodiment of the computing device of the present invention. The specific embodiments of the present invention do not limit the specific implementation of the computing device.
[0102] like Figure 3 As shown, the computing device may include: a processor 302, a communications interface 304, a memory 306, and a communications bus 308.
[0103] The processor 302, communication interface 304, and memory 306 communicate with each other via communication bus 508. Communication interface 304 is used to communicate with other network elements such as clients or other servers. The processor 302 executes program 310, specifically performing the relevant steps in the above-described embodiment of the postoperative early complication warning method based on multimodal fusion.
[0104] Specifically, program 310 may include program code that includes computer operation instructions.
[0105] Processor 302 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention. The computing device includes one or more processors, which may be processors of the same type, such as one or more CPUs; or processors of different types, such as one or more CPUs and one or more ASICs.
[0106] Memory 306 is used to store program 310. Memory 306 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.
[0107] According to the solution provided by this invention, a wearable device is used to collect postoperative physiological temporal feature tensors of the patient, including heart rate, blood oxygen saturation, respiratory rate, body temperature, skin conductance, and activity level. The patient's clinical context feature tensors are obtained through a hospital information system, electronic medical record system, and medical image archiving and communication system. These clinical context feature tensors include clinical static data, laboratory test results, medical image reports, and medical records. The postoperative physiological temporal feature tensors and the clinical context feature tensors are input into a postoperative spatiotemporally coupled attention risk prediction model to predict the real-time probability of preset complications. The postoperative spatiotemporally coupled attention risk prediction model includes a multi-granularity temporal coding layer, a hierarchical semantic context coding layer, a cross-modal attention fusion layer, and a complication-specific prediction layer. The complication-specific prediction layer includes a multi-task prediction head, a sequence decision layer, and an uncertainty quantification layer. Based on the patient's individualized risk tolerance, the severity of preset complications, and the modifiability of intervention, early postoperative complication warnings are provided. This invention significantly improves the accuracy of early postoperative complication warnings through the postoperative spatiotemporally coupled attention risk prediction model. Specifically, missing data in medical data is addressed by using a variational autoencoder for unified latent space mapping and a cross-modal generator to reconstruct missing data. Risk states are transferred based on Markov decision processes, enabling dynamic assessment of patient risk levels, remaining intervention windows, and physiological stability, thus providing dynamic risk assessment based on the current state. The sequence decision layer, comprising risk summary, evidence tracing, and intervention protocol layers, not only displays risk values but also traces back to key physiological indicators and clinical factors, enhancing the model's transparency and interpretability. A meta-learning method is employed for small sample complication types, allowing the model to quickly adapt to new tasks and improve its ability to identify rare complications. The uncertainty quantification layer calculates the mean and variance of prediction results through multiple forward propagations, providing healthcare professionals with risk prediction values and a measure of their reliability. The causal effect of the clinical context feature tensor is calculated using the front-door adjustment criterion, and a balanced dataset is constructed using propensity score matching, helping to reveal the true causal relationship between complications and clinical factors, making the prediction results more clinically credible.
[0108] Those skilled in the art will understand that modules in the device of the embodiments can be adaptively changed and placed in one or more devices different from that embodiment. Modules, units, or components in the embodiments can be combined into a single module, unit, or component, and further, they can be divided into multiple sub-modules, sub-units, or sub-components. Except where at least some of such features and / or processes or units are mutually exclusive, any combination of all features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all processes or units of any method or device so disclosed can be employed. Unless expressly stated otherwise, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) may be replaced by an alternative feature that serves the same, equivalent, or similar purpose. Furthermore, those skilled in the art will understand that although some embodiments herein include certain features included in other embodiments but not others, combinations of features from different embodiments are intended to be within the scope of the invention and form different embodiments. For example, in the following claims, any of the claimed embodiments can be used in any combination. The invention can be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In the unit claims listing several devices, several of these devices may be embodied by the same hardware item. Unless otherwise specified, the steps in the above embodiments should not be construed as limiting the order of execution.
Claims
1. A method for early postoperative complication prediction based on multimodal fusion, characterized in that, include: The postoperative physiological time-series characteristics tensor of the patient is collected by wearable devices. The postoperative physiological time-series characteristics tensor includes heart rate, blood oxygen saturation, respiratory rate, body temperature, skin conductance and activity level. The patient's clinical context feature tensor is obtained through the hospital information system, electronic medical record system, and medical image archiving and communication system. The clinical context feature tensor includes clinical static data, laboratory test results, medical image reports, and medical records. The postoperative physiological temporal feature tensor and the clinical context feature tensor are input into the postoperative spatiotemporal coupled attention risk prediction model to predict the real-time probability of the occurrence of preset complications. The postoperative spatiotemporal coupled attention risk prediction model includes a multi-granularity temporal coding layer, a hierarchical semantic context coding layer, a cross-modal attention fusion layer, and a complication-specific prediction layer. The complication-specific prediction layer includes a multi-task prediction head, a sequence decision layer, and an uncertainty quantification layer. Early postoperative complication warnings are provided based on the patient's individualized risk tolerance, the severity of anticipated complications, and their modifiability.
2. The method for early postoperative complication warning based on multimodal fusion according to claim 1, characterized in that, The method further includes: The physiological temporal feature tensor and the clinical context feature tensor are mapped to a unified latent space tensor by a variational autoencoder (VAE). The missing data of the unified latent space tensor is reconstructed using a cross-modal generator; The evidence lower bound loss function of the unified latent space tensor is: ; in, For physiological data; For clinical data; To unify the latent space tensor; It is the divergence function; Let z be the prior distribution of z; Let z be the posterior probability distribution of the latent variable z when the data are x and y; The conditional probability of generating x and y when z is z.
3. The method for early postoperative complication warning based on multimodal fusion according to claim 1, characterized in that, The complication-specific prediction layer uses a Markov decision process to transfer risk status. The risk status includes the current risk level, the remaining intervention window, and the patient's physiological stability indicators. The learning formula for the state transition probability of the risk state is: ; in, To transition to the next risk state after performing action a in the current risk state s. The probability of; To transition to state after performing action a from state s. The number of times; For smoothing coefficients; Let s be the total number of possible states in the risk state space.
4. The method for early postoperative complication warning based on multimodal fusion according to claim 1, characterized in that, During the training process of the postoperative spatiotemporal coupled attention risk prediction model, the method further includes: A self-supervised task is constructed based on unlabeled postoperative physiological data, which includes time-series data reconstruction and future step size prediction. After pre-training, the encoder layer parameters are frozen, and supervised fine-tuning is performed on the labeled dataset. Specifically, for small sample complication types, a meta-learning strategy is used to initialize model parameters and achieve rapid convergence for multi-round task adaptation.
5. The method for early postoperative complication warning based on multimodal fusion according to claim 1, characterized in that, The clinical context feature tensor is used to calculate causal effects using the front door adjustment criterion. For continuous confounding factors, a propensity score matching method is used to construct a balanced dataset.
6. The method for early postoperative complication warning based on multimodal fusion according to claim 5, characterized in that, The causal contribution formula for the aforementioned causal effect is: ; in, Let X be the conditional probability distribution of the target variable Y when the intervention X=x is applied. Let z be the conditional probability that the mediator variable Z takes the value z given that X=x; Let Y be the conditional probability distribution of the target variable Y given that X=x and Z=z; Let X be the intervention function, representing the state of setting variable X to x.
7. The method for early postoperative complication warning based on multimodal fusion according to claim 1, characterized in that, The sequence decision layer includes: The risk summary layer displays real-time risk values, complication types, and intervention recommendation levels. The evidence traceability layer demonstrates the strength of the association between physiological indicators and clinical factors; The intervention protocol layer is used to recommend standardized intervention measures based on the guideline library and localization protocols, and to label the evidence level.
8. The method for early postoperative complication warning based on multimodal fusion according to claim 1, characterized in that, The uncertainty quantification layer randomly activates some or all of the neural network layers in the complication-specific prediction layer during the model inference stage. In this process, multiple forward propagations are performed on the same input data to obtain multiple sets of prediction results; The mean and variance of the predicted result set are calculated to obtain the risk prediction value and the corresponding uncertainty measure, respectively.
9. A postoperative early complication warning device based on multimodal fusion, characterized in that, include: The physiological time series data acquisition module is used to collect postoperative physiological time series feature tensors of patients through wearable devices. The postoperative physiological time series feature tensors include heart rate, blood oxygen saturation, respiratory rate, body temperature, skin conductance, and activity level. The clinical context data acquisition module is used to acquire the patient's clinical context feature tensor through the hospital information system, electronic medical record system, and medical image archiving and communication system. The clinical context feature tensor includes clinical static data, laboratory test results, medical image reports, and medical records. A multimodal fusion risk prediction module is used to input the postoperative physiological temporal feature tensor and the clinical context feature tensor into the postoperative spatiotemporal coupled attention risk prediction model to predict the real-time probability of the occurrence of preset complications; wherein, the postoperative spatiotemporal coupled attention risk prediction model includes a multi-granularity temporal coding layer, a hierarchical semantic context coding layer, a cross-modal attention fusion layer and a complication-specific prediction layer, and the complication-specific prediction layer includes a multi-task prediction head, a sequence decision layer and an uncertainty quantification layer; The individualized early warning generation module is used to provide early warnings of postoperative complications based on the patient's individualized risk tolerance, the severity of preset complications, and the modifiability of intervention.
10. A computing device, comprising: The processor, memory, communication interface, and communication bus are provided, wherein the processor, memory, and communication interface communicate with each other via the communication bus. The memory is used to store at least one executable instruction, which causes the processor to perform the operation corresponding to the above-described method for early postoperative complication warning based on multimodal fusion.