Multi-mode-based peritoneal compartment syndrome preoperative evaluation method and multi-mode-based peritoneal compartment syndrome preoperative evaluation system

Through the XAI-based multimodal evaluation system, combined with the Transformer self-attention mechanism and multiple physiological data, the preoperative indicators of abdominal compartment syndrome are automatically evaluated, which solves the problem of doctors' multiple comprehensive evaluations and achieves efficient and reliable evaluation results.

CN120636772AInactive Publication Date: 2025-09-12SHANXI CHILDRENS HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510781914.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-09-12
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the existing technology, during the preoperative evaluation of abdominal compartment syndrome, doctors need to conduct multiple comprehensive evaluations in many aspects, resulting in a heavy workload and long periods of mental concentration, and a lack of efficient automated evaluation methods.

Method used

A multimodal evaluation system based on the XAI explainable artificial intelligence model is adopted, which integrates the self-attention mechanism of Transformer, combines multiple physiological data and CT images, and automatically determines whether the surgical indicators are met through deep learning.

Benefits of technology

Automated and interpretable preoperative evaluation of abdominal compartment syndrome has been achieved, which reduces the workload of doctors and improves the reliability and accuracy of the evaluation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120636772A_ABST
    Figure CN120636772A_ABST
Patent Text Reader

Abstract

The invention provides an abdominal compartment syndrome preoperative evaluation method and system based on multiple modes. The method comprises the steps that multiple physiological data of a patient are obtained; inputting the physiological indexes and the physical examination parameters into a pre-established preoperative index judgment system, and judging whether the physiological indexes and the physical examination parameters of the patient accord with operation indexes or not by utilizing the preoperative index judgment system; the preoperative index judgment system is constructed on the basis of an XAI interpretable artificial intelligence model, and is fused with a self-attention mechanism of a transformer. According to the peritoneal compartment syndrome preoperative evaluation method based on the multiple modes, the self-attention mechanism of transformer is ingeniously fused, the model can automatically learn and pay attention to information most critical to diagnosis, and the diagnosis accuracy is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent medical technology, and in particular to a multimodal preoperative assessment method and system for abdominal compartment syndrome. Background Art

[0002] With the advancement of critical care medicine and the increase in critically ill patients, the incidence of intra-abdominal hypertension (IAH) and abdominal compartment syndrome (ACS) is also increasing, and there is an urgent need for more effective and safe treatments.

[0003] When IAH develops into ACS and conservative treatment fails, surgical treatment is recommended. With advances in science and technology, treatments using central venous catheters for peritoneal puncture and drainage to reduce intra-abdominal pressure and thereby treat abdominal compartment syndrome have been developed and implemented.

[0004] The increase in the number of surgeries means a corresponding increase in the workload for doctors reviewing whether patients meet surgical requirements. Preoperative assessments typically involve multiple doctors conducting comprehensive, multifaceted assessments, a significant burden for doctors who face a heavy daily workload and are required to maintain high concentration for extended periods. Summary of the Invention

[0005] In order to solve the above problems, the purpose of the present invention is to provide a multimodal preoperative assessment method and system for abdominal compartment syndrome.

[0006] According to one aspect of the present invention, a multimodal preoperative assessment method for abdominal compartment syndrome is provided, the method comprising: Obtain multiple physiological data of patients; Inputting the plurality of physiological indicators and physical examination parameters into a pre-established preoperative indicator judgment system, and using the preoperative indicator judgment system to judge whether the patient's physiological indicators and physical examination parameters meet the surgical indicators; The preoperative indicator judgment system is built based on the XAI explainable artificial intelligence model and integrates the transformer's self-attention mechanism.

[0007] Optionally, the preoperative indicator judgment system includes: The data input layer is used to receive multiple physiological data of patients as original learning data; A feature extraction layer, used to extract data features from the original learning data; The main model module integrates the Transformer self-attention mechanism to perform deep learning and refinement on the processed data input by the correction module to generate explanatory and predictive data suitable for transmission to subsequent modules; The explanation module is responsible for generating the final explanation data of the model prediction; An interpretation output module is used to output the interpretation data generated by the interpretation module to the user interface for display; The prediction module generates the final prediction data based on the prediction results of the main model; The prediction output module is used to output the prediction data generated by the prediction module to the user interface for display; User interface, which provides an interactive interface for displaying explanation output and prediction output to end users and further optimizing the main model based on the feedback information provided by users.

[0008] Optionally, the feature extraction layer includes a position embedding layer for embedding position information:

[0009] in, and Represents the horizontal and vertical position indexes respectively, and Represent dimensions, represents the dimension of embedding; The new position after embedding is expressed as:

[0010] Among them, z represents the image pixel feature and PE represents the embedded position feature vector.

[0011] Optionally, the correction module integrates features of different modalities through a feature fusion strategy to form a comprehensive feature representation. The process is as follows: Normalize the features of each modality to ensure they are on the same scale; Perform polynomial transformation on the normalized features to capture the nonlinear relationship between features; The polynomial features of different modalities are integrated, and based on the feature integration, the features of different modalities are further weighted to reflect their different contributions to the final task.

[0012] Optionally, the main model calculates the similarity between text features and image features to generate a multimodal attention matrix; The similarity calculation formula is as follows:

[0013] in, Represents text features and image The similarity between the feature maps of the sub-regions, Represents the encoded text features, Represents the first feature vectors.

[0014] Optionally, the prediction module calculates the prediction score using the following formula:

[0015] in, represents the sigmoid activation function, Represents the prototype The weight corresponding to the generated similarity score, Similarity Representing image features and disease prototypes The similarity score between .

[0016] The interpretation module is used to generate interpretation data by combining the two interpretation technologies of Pattern Net and Pattern Attribution.

[0017] Optionally, the physiological data includes: One or more of intra-abdominal pressure, peritoneal perfusion pressure, vital signs, blood routine, blood component data, abdominal CT data, and hemodynamic monitoring data; Vital signs include blood pressure, heart rate, respiration, pulse, etc. What needs to be paid attention to is one or more of the patient's urine volume, jugular vein fullness, and whether the skin of the limbs is cold.

[0018] The present invention also provides a multimodal preoperative evaluation system for abdominal compartment syndrome, comprising a processor and a memory; the memory is used to store program code and transmit the program code to the processor; the processor is used to execute any of the above-mentioned multimodal preoperative evaluations for abdominal compartment syndrome according to the instructions in the program code.

[0019] The multimodal preoperative assessment method and system for abdominal compartment syndrome (ACS) utilizes Explainable AI (XAI) Models to establish a preoperative assessment system. This system not only automatically assesses patient conditions but also, due to its inherent interpretability and traceability, is fully compatible with hospital scenarios. This ensures that assessment results can be traced back to the original data, improving the reliability of the results.

[0020] By inputting the patient's various physiological indicators and CT images and other related information into the preoperative indicator judgment system built based on Explainable AI (XAI) Models and integrating the transformer's self-attention mechanism, a complete preoperative indicator judgment system is formed. The process of determining whether the surgical indicators are met is changed from manual to deep learning.

[0021] Using an XAI interpretable model to determine whether a patient's physical signs are suitable for central venous catheterization for abdominal puncture and drainage surgery reduces physician workload and allows for accurate treatment of abdominal compartment syndrome. This approach can be used for patients with clinical manifestations of abdominal compartment syndrome or intra-abdominal pressure greater than 20 mmHg who have failed or ineffective medical treatment.

[0022] Based on the following detailed description of specific embodiments of the present invention in conjunction with the accompanying drawings, those skilled in the art will become more aware of the above and other objects, advantages and features of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiment below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present invention. The same reference symbols are used throughout the drawings to represent the same components. In the drawings: Figure 1 1 is a flow chart of a multimodal preoperative evaluation method for abdominal compartment syndrome according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the structure of a preoperative indicator judgment system according to an embodiment of the present invention; Figure 3 Schematic diagram of the Transformer attention mechanism framework according to an embodiment of the present invention; Figure 4 Schematic diagram of a multi-head attention model according to an embodiment of the present invention. DETAILED DESCRIPTION

[0024] The embodiments of the present invention are described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to illustrate the present invention and are not intended to limit the present invention.

[0025] The embodiment of the present invention provides a multimodal preoperative evaluation method for abdominal compartment syndrome, such as Figure 1 As shown, the multimodality-based preoperative evaluation method for abdominal compartment syndrome according to an embodiment of the present invention includes the following steps S1 to S2.

[0026] S1, obtain multiple physiological data of the patient; S2: Input the multiple physiological indicators and physical examination parameters into a pre-established preoperative indicator judgment system, and use the pre-operative indicator judgment system to determine whether the patient's physiological indicators and physical examination parameters meet the surgical criteria. The pre-operative indicator judgment system is built based on the XAI interpretable artificial intelligence model and incorporates the transformer's self-attention mechanism.

[0027] The physiological data obtained in the above step S1 can be as follows: Intra-abdominal pressure (IAP): IAP is a key indicator for assessing abdominal compartment syndrome. Normally, IAP is <5 mmHg, and 5-7 mmHg is considered normal in critically ill patients. In obese individuals, IAP can normally rise to 9-14 mmHg.

[0028] Peritoneal perfusion pressure (APP): APP refers to the difference between mean arterial pressure and intra-abdominal pressure. It is a measure of blood perfusion to the abdominal viscera. Normally, APP is >60 mmHg. A decreased APP indicates that visceral blood perfusion may be impaired.

[0029] Vital signs: Focus on observing the patient's vital signs, including blood pressure, heart rate, respiration, pulse, etc. It is necessary to pay attention to the patient's urine volume, jugular vein fullness, whether the skin of the limbs is cold, etc.

[0030] Complete blood test: Mainly used to diagnose whether there is secondary systemic infection.

[0031] Blood and chemical composition: Check whether the patient has abnormal liver and kidney function.

[0032] Others: Patients with acute pancreatitis need to undergo some tests, including lipase, amylase, etc.

[0033] Abdominal CT data: CT findings associated with intra-abdominal hypertension include: round abdomen sign (maximum anteroposterior diameter of the abdominal wall / maximum left-right diameter of the abdominal wall > 0.8), intestinal wall thickening (> 3 mm), elevated diaphragm, inferior vena cava stenosis (< 3 mm), and large amounts of ascites.

[0034] Hemodynamic monitoring data: During ACS, parameters such as central venous pressure, pulmonary capillary wedge pressure, and mean arterial pressure may be unreliable.

[0035] In actual applications, adjustments can be made according to different needs, which is not limited in this embodiment.

[0036] The preoperative indicator judgment system of the embodiment of the present invention may include: The data input layer is used to receive multiple physiological data of patients as original learning data; A feature extraction layer, used to extract data features from the original learning data; The correction module is used to perform feature conversion based on the data features extracted by the feature extraction layer. Data feature conversion refers to mathematical transformation of the extracted features to optimize the feature representation, making them more suitable for subsequent analysis or model training. By combining normalization and polynomial transformation, features from different modalities are integrated to optimize the data representation, making it more suitable for model training.

[0037] The main model module integrates the Transformer self-attention mechanism to perform deep learning and refinement on the processed data input by the correction module, thereby generating explanatory and predictive data suitable for transmission to subsequent modules.

[0038] The explanation module is responsible for generating the final explanation data of the model prediction; An interpretation output module is used to output the interpretation data generated by the interpretation module to the user interface for display; The prediction module generates the final prediction data based on the prediction results of the main model; The prediction output module is used to output the prediction data generated by the prediction module to the user interface for display; User interface, which provides an interactive interface for displaying explanation output and prediction output to end users and further optimizing the main model based on the feedback information provided by users.

[0039] The multimodal intelligent classification method for abdominal compartment syndrome in this embodiment of the present invention innovatively combines the patient's physiological indicators with multimodal information such as CT images, and performs comprehensive evaluation through the XAI model. This breaks through the limitations of traditional surgical indications based solely on physiological indicators or imaging data, and provides a more comprehensive and accurate basis for the diagnosis of abdominal compartment syndrome. The structural diagram of the preoperative indicator judgment system of this embodiment can be found in Figure 2 .

[0040] The multimodal preoperative assessment of abdominal compartment syndrome provided by the embodiments of the present invention is used for the diagnosis of abdominal compartment syndrome. The entire process begins by receiving multimodal patient data, including physiological indicators and CT images. First, a data feature extraction model is used to extract features from this data. Then, a feature fusion strategy designed in the correction module is used to integrate features from different modalities to form a comprehensive feature representation. This feature representation is then input into a deep learning model (i.e., the main model module), which is trained to classify abdominal compartment syndrome and generate predictions. The model's predictions, including classification labels and confidence scores, are sent to the interpretation module, which is responsible for generating explanatory data for the model's decisions, describing why the model made specific classification decisions. The predictions are also displayed to the end user through a user interface, allowing the user to review the predictions and provide feedback. User feedback is used to further optimize the model. Through a feedback loop, the model can self-improve based on user input. The following details the various structures of the preoperative indicator judgment system of the embodiments of the present invention.

[0041] 1. Data Input Layer Collect patient data such as physiological indicators and CT images and input them into the preoperative indicator judgment system.

[0042] 2. Feature Extraction Layer 1. Image processing prediction part First, the feature extraction layer of the deep learning model of the preoperative indicator judgment system is used to extract features from the image data in the collected multimodal data. The image data may include image data such as CT images. In the embodiment of the present invention, the Resnet50 convolutional neural network (CNN) can be used to extract features from images. The function of the text nonlinear projector is to perform a nonlinear transformation on the original features of the text (such as word embedding vectors, etc.). When the original input image enters the Resnet50 encoder, it will go through a series of operations such as convolutional layers, pooling layers, and residual blocks. These operations can automatically learn various features in the image, such as basic information such as edges, textures, and shapes, and then abstract higher-level semantic features layer by layer to ultimately obtain a coded representation of the image.

[0043] Image processing output:

[0044] Among them, z represents the encoded image features, represents a text nonlinear projector, represents the Resnet50 encoder, represents the original input image.

[0045] 2. Position embedding layer, used to embed position information (PE) When processing sequence data (such as text and time series), simply having the features of the elements themselves (such as word embeddings) is not enough, because the order of the elements is crucial for understanding the semantics and structure of the sequence. Position embedding can provide the model with the position information of each element in the sequence, helping the model better capture the order and relative position of the sequence, thereby improving the understanding and processing capabilities of sequence data. In this example, the position embedding formula is as follows:

[0046] Among them, PE(x,y,2i): calculates the value of the 2i-th dimension of the position encoding vector at position (x,y) using the sine function.

[0047] PE(x,y,2i+1): Calculates the value of the 2i+1th dimension of the position encoding vector at position (x,y) using the cosine function.

[0048] PE(x,y,2j+D / 2): Calculates the value of the 2j+2D dimension of the position encoding vector at position (x,y) using the sine function, which is calculated based on the vertical position y.

[0049] PE(x,y,2j+D / 2+1): Calculates the value of the 2j+D / 2+1th dimension of the position encoding vector at position (x,y) using the cosine function, also based on the vertical position y.

[0050] in, and Represents the horizontal and vertical position indexes respectively, and Represent dimensions, Represents the dimension of the embedding.

[0051] In the formula of the position embedding layer, and are all indices used to index a specific dimension in the positional encoding vector. represents the dimension of position embedding, that is, the length of the position encoding vector. When deep learning models process image data, they usually embed position information into image features in the form of vectors. The length of this vector is .

[0052] The new position after embedding can be expressed as:

[0053] Among them, z represents the image pixel feature and PE represents the embedded position feature vector.

[0054] 3. Correction Module By modifying the feature fusion strategy designed by the module, the features of different modalities are integrated to form a comprehensive feature representation. The specific process is as follows: 1. Normalization First, the features of each modality are normalized to ensure that they are on the same scale. The normalization formula is as follows:

[0055] in, represents the original features of the mth mode, represents the normalized features, Represents the maximum value in the original feature of the mth mode, Represents the minimum value in the original feature of the mth mode.

[0056] As mentioned above, physiological data includes multiple types of data, and each data type may be different, such as text data or image data. The embodiment of the present invention normalizes the features of each modality, which means that the feature vectors of different modalities are standardized according to certain rules (such as scaling to a specific range, making them have unit norm, etc.). This makes the features of different modalities comparable and consistent in subsequent fusion, comparison and other operations, avoiding the adverse effects of excessive differences in feature scales between modalities on model performance.

[0057] 2. Polynomial Transformation Next, a polynomial transformation is performed on the normalized features to capture the nonlinear relationship between the features. For the features of the mth mode, the polynomial transformation can be expressed as:

[0058] Where k is the order of the polynomial, represents the polynomial characteristics of the mth mode.

[0059] Polynomial transformation refers to mapping data into a higher-dimensional space through a polynomial function. The purpose of this is to capture nonlinear relationships in the data, thereby improving the model's fitting ability and predictive performance.

[0060] 3. Feature Integration Then, the polynomial features of different modes are integrated. Assuming there are M modes, the integrated features can be expressed as:

[0061] in, Represents the concatenation operation of vectors, is the integrated feature representation.

[0062] Based on feature integration, the features of different modalities can be further weighted to reflect their different contributions to the final task. The weighted features can be expressed as:

[0063] in, It is The weight coefficient of each modal feature.

[0064] 4. Main Model Module The main model module features a deep learning model that has been trained to classify abdominal compartment syndrome and generate predictions. The data is derived from the text and image features extracted by the model's encoder, and the output is a multimodal attention matrix. Specifically, the main model module calculates the similarity between text and image features to generate the multimodal attention matrix. The resulting prediction includes the class label relevant to abdominal compartment syndrome classification (e.g., whether the syndrome is present, what stage of the syndrome is present, and so on). This prediction is based on information such as the similarity between text and image features as reflected in the multimodal attention matrix.

[0065] After completing data feature transformation, including normalization and polynomial transformation feature integration, the next step is to use these integrated features to calculate the similarity between different modalities and generate a multimodal attention matrix. This process is usually performed by the attention mechanism module in the model. Its purpose is to strengthen the model's understanding of the correlation between different modalities, thereby improving the model's overall performance on the task. The formula for calculating similarity is as follows:

[0066] here, Represents text features and image The similarity between the sub-region feature maps, t represents the encoded text features, Represents the first feature vectors.

[0067] This embodiment calculates similarity to capture inter-modal correlations. Data from different modalities may describe the same thing from different perspectives. By calculating similarity, it is possible to identify the interconnected parts of the text and image. In medical image diagnosis scenarios, text descriptions may contain information such as the patient's symptoms and medical history, while images show the patient's internal structure. Calculating the similarity between them can help the model understand the corresponding manifestations of the symptoms mentioned in the text in the image, thereby better integrating multimodal information.

[0068] Calculating similarity can also generate an attention matrix. This similarity calculation is the foundation for generating a multimodal attention matrix. The attention mechanism assigns different weights to different features based on the similarity between different modalities, allowing the model to focus more on task-relevant information. By calculating the similarity between text features and feature maps from different subregions of the image, we can determine the degree of attention the text pays to each part of the image, thereby generating a multimodal attention matrix.

[0069] To get the attention weights, first apply a nonlinear activation function, such as ReLU, to ensure that all weights are non-negative:

[0070] represents the attention weight after processing by the nonlinear activation function, is the threshold of the ReLU function, Represents text features and image The similarity between the feature maps of the sub-regions; Then, by introducing a temperature parameter And applying the softmax function, we get the final attention weights:

[0071] in, Indicates the The attention weight of the feature vector reflects the The importance of each sub-region feature map relative to the text feature.

[0072] Attention weights are mainly used to calculate the weighted combination of different features when fusing multimodal data, so as to better perform subsequent model prediction, classification and other tasks.

[0073] Prototype layer processing: Calculate the Euclidean distance between the feature map latent patch and the disease prototype and convert the distance into a similarity score.

[0074] In the previous calculations, the encoder extracted text and image features, which served as the basis for subsequent calculations. The feature map latent patches in the prototype layer were also further divided or processed based on image features.

[0075] In addition, the previous article calculated the similarity between text features and image feature maps to identify correlations between different modalities and generate a multimodal attention matrix. In prototype layer processing, the Euclidean distance between the latent patch in the feature map and the disease prototype is calculated and converted into a similarity score, which measures the similarity between image features and disease prototypes from another perspective.

[0076] This similarity calculation differs from the previous one in that it focuses more on comparing image features with typical features (prototypes) of the disease, rather than simply comparing text with images. However, both methods aim to enable the model to better understand the information in the data, but with different focuses and angles.

[0077] The prototype layer is responsible for calculating the Euclidean distances between the latent patches in the feature map and the disease prototypes and converting these distances into similarity scores. The main model module can have a prototype network module, and this process is typically performed by the prototype network module in the model. Its function is to sum and average the encoded representations of samples for each category, using this as the prototype representation for that category, thereby enabling classification of new samples. This matrix will be used in the model's subsequent classification tasks.

[0078] The data for this process comes from the image features extracted by the encoder in the model and the encoded representation of text features. The output similarity score will be used in subsequent processing stages, such as classification decisions or feature weighting.

[0079]

[0080] in, represents the similarity score, and are the visual features of image latent patches and disease prototypes, respectively. and are the position embedding features of image latent patches and disease prototypes respectively; and are the hyperparameters of visual feature similarity and position feature similarity respectively.

[0081] 5. Prediction Module 1. The prediction module performs prediction score calculation:

[0082] in, represents the sigmoid activation function, Represents the prototype The weight corresponding to the generated similarity score, Similarity Representing image features and disease prototypes The similarity score between Represents the predicted category label, and x represents the input data. 、 Usually obtained through the training process of the model. It can be automatically adjusted through learning algorithms to reflect the contribution of different prototypes to the final decision. It is usually learned from data and can be a point in the feature space that represents the average or central characteristics of a class of diseases. These prototype vectors can be initialized to random values ​​and then updated by an optimization algorithm during training.

[0083] 6. Explanation Module Determined by the nature of the input data, the weight vector Not always related to the signal direction Align, but try to filter out noise This indicates that in the presence of both signal and interference, the direction of the weight vector in a linear model is primarily determined by the interference. Furthermore, embodiments of the present invention propose two interpretation techniques, Pattern Net and Pattern Attribution, which are theoretically applicable to linear models and can provide improved interpretations for deep networks.

[0084] Combining Pattern Net and Pattern Attribution explains the two techniques, creating a more comprehensive model explanation framework that not only identifies key patterns in the input data, but also quantifies the specific contribution of these patterns to model predictions. Below is an example framework combining these two techniques and its associated formulas.

[0085] First, we use Pattern Net technology to identify and optimize key patterns in the input data. This can be achieved by the following optimization problem:

[0086] in: The model is input Plus Mode The predicted probability after . is the original input data. is the pattern vector to be optimized. is a constraint that ensures that the pattern vector does not become too large. Respectively , Indicates that in the process of finding this pattern vector p, it is necessary to satisfy This constraint, that is, the Euclidean norm of the pattern vector p must be less than or equal to 1, is used to limit the size of the pattern vector p and prevent its value from being too large and causing unreasonable impact on the original input x.

[0087] Next, we use the Pattern Attribution technique to quantify the specific contribution of these patterns to the model prediction. This can be achieved using the following formula:

[0088] in: represents the contribution of input x to the predicted probability of class c.

[0089] The model is input Add mode The predicted probability after .

[0090] is the first input feature elements.

[0091] is the predicted probability relative to the input feature The partial derivative of .

[0092] Finally, combine Pattern Net and Pattern Attribution.

[0093] Combining these two techniques results in a more comprehensive model interpretation framework. First, we identify key patterns through Pattern Net, and then use Pattern Attribution to quantify the contribution of these patterns to model predictions. This process can be expressed as follows:

[0094] in: is input For categories The total predicted probability contribution of . is a weight parameter used to balance the two contributions. It is Prototype The weight corresponding to the generated similarity score. Represents image feature z and disease prototype The similarity score between .

[0095] In this way, we can simultaneously consider key patterns in the input data and the specific contributions of these patterns to model predictions, thereby providing more comprehensive and accurate model interpretations.

[0096] Signal estimation quality metrics :

[0097] This formula is used to measure the signal estimation function quality.

[0098] ρ(S): represents the quality metric of signal estimation. The closer its value is to 1, the better the estimation quality.

[0099] v: represents a vector, usually used to represent a characteristic direction or space of a signal.

[0100] w: represents a weight vector, which is used to perform weighted processing on the input signal to extract specific features of the signal.

[0101] x: represents the input signal vector, which contains the observation data of the signal.

[0102] S(x): represents the signal estimation function, which is the estimation result of the input signal x. .

[0103] : Represents the variance of the vector v^T (xS(x)), which reflects the degree of fluctuation of the estimation error in the direction v.

[0104] : represents the variance of the signal y, which is usually related to the characteristics of the original signal or the target signal and is used to normalize the estimation quality metric.

[0105] Junctional signal estimation :

[0106] in, is achieved by optimizing the above quality metrics To determine the scalar.

[0107] Junctional signal estimation The core function of is to quantify the contribution of each feature in the input data x to the model decision, providing an intuitive numerical basis for model interpretation. In the formula, w is a weight vector that reflects the evaluation of the importance of different features during the model training process; the scalar a is determined by optimizing the quality metric \(\rho(S)\) and is used to adjust the overall weight scale. Through matrix operations, this formula performs a weighted summation of the original input features and the corresponding weights, and the output is The higher the value, the greater the influence of the feature combination in the input x on the model decision, and vice versa.

[0108] Taking the medical imaging diagnosis scenario as an example, the input x contains the patient's CT image features and clinical text information. After model training, the weight vector w is obtained, in which the image features related to the texture and size of the lesion area, as well as the key symptom texts such as "persistent abdominal pain" and "oliguria" have higher weights. After determining the appropriate value of a, the calculated If the value is high, it means that the current input feature combination is highly correlated with the disease diagnosis, and the model tends to make a positive diagnosis; if the value is low, it means that the input feature does not support the diagnosis enough. The contribution of these key patterns can be further quantified, providing a more transparent and explainable basis for medical decision-making.

[0109] Two-component signal estimation :

[0110] : Represents the knotty signal estimation function, which is a linear estimation of the input signal x.

[0111] a: is a scalar coefficient, which is used to optimize the quality metric To determine, it is used to adjust the gain of the estimation function to make the estimation result closer to the real signal.

[0112] : represents the inner product of the weight vector w and the input signal vector x, and extracts the characteristic component of the signal in the w direction.

[0113] In this embodiment, the two-component signal is estimated by calculating In order to more carefully analyze the input signal Specifically, it is based on The value of divides the signal into two cases for processing. When using To estimate the signal, As an adjustment factor, A specific gain adjustment is performed on the part with positive characteristic components in the direction to more accurately reflect the contribution of this part of the signal to the overall estimation; when When using Make an estimate, In this way, the two-component signal estimation can consider the characteristic components of the input signal in different directions respectively, and process them differently according to their positive and negative characteristics, so as to estimate the input signal more comprehensively and accurately. , providing more detailed information for subsequent analysis and decision-making.

[0114] Prediction score :

[0115] in, represents the sigmoid activation function, Represents the prototype The weight corresponding to the generated similarity score.

[0116] Prediction score The current prediction score is displayed in the user interface, which is equivalent to an indicator of confidence.

[0117] Signal estimation quality metrics Used to evaluate the credibility of input multimodal data.

[0118] Junctional signal estimation In

[15] , the weight vector w is designed to extract specific features of the input signal x, such as edges, textures, or other important features of the image. It is used to adjust the extraction focus of the main model in the feedback path.

[0119] The knotty signal estimation provides characteristic support for the prediction score and affects its calculation results; the signal estimation quality metric measures the credibility of the input data by evaluating the knotty signal estimation, and optimizes the parameters in the knotty signal estimation accordingly. At the same time, the credibility of the input data affects the reliability of the prediction score. The three are interrelated and influence each other, and together play an important role in the operation and evaluation of the model.

[0120] As mentioned above, the main module in the preoperative indicator judgment system of this embodiment can integrate the Transformer attention mechanism.

[0121] like Figure 3 As shown in the figure, the Transformer model architecture consists of two main parts: encoder (left side) and decoder (right side). The core components of the model include: Multi-Head Attention, Feed Forward, Residual Connection and Normalization. Among them, the multi-head attention model is actually an upgraded version of the self-attention model. The multi-head self-attention model introduces multiple groups of mappings at the same time, and performs self-attention operations in each group of mappings, and finally combines the feature sequences obtained from each group according to the feature dimension. The multi-head attention mechanism can learn the feature representations of multiple feature subspaces, so that each "head" can pay attention to different value information, such as Figure 4 shown.

[0122] The multimodal preoperative assessment method for abdominal compartment syndrome in an embodiment of the present invention cleverly integrates the transformer's self-attention mechanism, enabling the model to automatically learn and focus on the most critical information for diagnosis, further improving the accuracy of diagnosis. In addition, advanced interpretation technologies such as PatternNet and PatternAttribution are used to provide explainability for the model's decision-making process, which is particularly important in the field of medical diagnosis because it ensures that doctors can understand and trust the model's judgment results, enhancing the feasibility and reliability of the method in practical clinical applications. These innovations together constitute the unique advantages of this method, which enables it to demonstrate great application potential and value in the field of diagnosis of abdominal compartment syndrome, and is expected to lead the diagnostic technology in this field to a new level.

[0123] Based on the same inventive concept, an example of the present invention also provides a multimodal preoperative evaluation system for abdominal compartment syndrome, including a processor and a memory; the memory is used to store program code and transfer the program code to the processor; the processor is used to execute the multimodal preoperative evaluation of abdominal compartment syndrome in the above embodiment according to the instructions in the program code.

[0124] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiment. All technical solutions based on the concept of the present invention are within the scope of protection of the present invention. It should be noted that for those skilled in the art, various improvements and modifications that do not depart from the principles of the present invention should also be considered within the scope of protection of the present invention.

Claims

1. A multimodal preoperative evaluation method for abdominal compartment syndrome, characterized in that: The method comprises: Obtain multiple physiological data of patients; Inputting the plurality of physiological indicators and physical examination parameters into a pre-established preoperative indicator judgment system, and using the preoperative indicator judgment system to judge whether the patient's physiological indicators and physical examination parameters meet the surgical indicators; The preoperative indicator judgment system is built based on the XAI explainable artificial intelligence model and integrates the transformer's self-attention mechanism.

2. The method according to claim 1, characterized in that The preoperative index judgment system includes: The data input layer is used to receive multiple physiological data of patients as original learning data; A feature extraction layer, used to extract data features from the original learning data; A correction module, configured to perform feature conversion based on the data features extracted by the feature extraction layer; The main model module, which integrates the Transformer self-attention mechanism, is used to perform deep learning and refined processing on the data input by the correction module to generate explanatory data and predictive data suitable for transmission to subsequent modules; The explanation module is responsible for generating the final explanation data of the model prediction; An interpretation output module is used to output the interpretation data generated by the interpretation module to the user interface for display; The prediction module generates the final prediction data based on the prediction results of the main model; The prediction output module is used to output the prediction data generated by the prediction module to the user interface for display; User interface, which provides an interactive interface for displaying explanation output and prediction output to end users and further optimizing the main model based on the feedback information provided by users.

3. The method according to claim 2, characterized in that The feature extraction layer includes a position embedding layer for embedding position information: ,in, and Represents the horizontal and vertical position indexes respectively, and Represent dimensions, represents the dimension of embedding; The new position after embedding is expressed as: , where z represents the image pixel feature and PE represents the embedded position feature vector.

4. The method according to claim 2, characterized in that The correction module integrates the features of different modalities through a feature fusion strategy to form a comprehensive feature representation. The process is as follows: Normalize the features of each modality to ensure they are on the same scale; Perform polynomial transformation on the normalized features to capture the nonlinear relationship between features; The polynomial features of different modalities are integrated, and based on the feature integration, the features of different modalities are further weighted to reflect their different contributions to the final task.

5. The method according to claim 2, characterized in that The main model calculates the similarity between text features and image features to generate a multimodal attention matrix; The similarity calculation formula is as follows: ,in, Represents text features and image The similarity between the feature maps of the sub-regions, Represents the encoded text features, Represents the first feature vectors.

6. The method according to claim 2, characterized in that The prediction module calculates the prediction score using the following formula: ,in, represents the sigmoid activation function, Represents the prototype The weight corresponding to the generated similarity score, Similarity Representing image features and disease prototypes The similarity score between .

7. The method according to claim 2, characterized in that The interpretation module is used to generate interpretation data by combining the two interpretation technologies of Pattern Net and Pattern Attribution.

8. The method according to any one of claims 1 to 7, characterized in that The physiological data includes: One or more of intra-abdominal pressure, peritoneal perfusion pressure, vital signs, blood routine, blood component data, abdominal CT data, and hemodynamic monitoring data; Vital signs include blood pressure, heart rate, respiration, pulse, etc. What needs to be paid attention to is one or more of the patient's urine volume, jugular vein fullness, and whether the skin of the limbs is cold.

9. A multimodal preoperative evaluation system for abdominal compartment syndrome, characterized in that: It comprises a processor and a memory; the memory is used to store program code and transmit the program code to the processor; the processor is used to execute the multimodality-based preoperative assessment of abdominal compartment syndrome according to any one of claims 1-8 according to the instructions in the program code.