Prediction method for sepsis in early stage based on multi-dimensional Transform

By using a multi-dimensional Transformer method in early prediction of sepsis, combining horizontal Transformer and vertical Transformer for feature extraction, the shortcomings of the existing methods in variable dimension mining are solved, and the prediction accuracy is significantly improved.

CN120148845APending Publication Date: 2025-06-13CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510206871.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing early prediction methods for sepsis mainly focus on feature extraction and modeling of time series, and failed to fully explore the dependence of clinical data on variable dimensions, resulting in insufficient utilization of feature information and limiting the prediction accuracy.

Method used

The early prediction method of sepsis based on multi-dimensional Transformer is adopted. By obtaining multiple life form feature variable data, laboratory detection variable data and demographic variable data, the shallow feature extraction module is used for preliminary feature extraction, and deformation groups are extracted through multiple cascades, including horizontal Transformer and vertical Transformer feature extraction modules, to realize joint modeling of time dimensions and variable dimensions.

Benefits of technology

Through the deep feature extractor of multi-dimensional Transformer, the interaction relationship between time-change characteristics and variables is comprehensively explored, the potential connections between various data are fully utilized, and the accuracy of early prediction of sepsis is effectively improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148845A_ABST
    Figure CN120148845A_ABST
Patent Text Reader

Abstract

The invention relates to a sepsis early prediction method based on multi-dimensional Transform, comprising: acquiring clinical variable data of a to-be-detected user, the clinical variable data comprising: a plurality of life body feature variable data, a plurality of laboratory detection variable data and a plurality of demographic variable data; performing shallow feature extraction on the clinical variable data through a shallow feature extraction module to obtain shallow features of the clinical variable data; inputting the shallow-layer features of the clinical variable data into a deep-layer feature extractor based on a multi-dimensional Transform for deep-layer feature extraction to obtain deep-layer features of the clinical variable data; and inputting the deep features of the clinical variable data into a result prediction module to obtain an early prediction result of sepsis. According to the method, combined modeling of the time dimension and the variable dimension can be achieved, the time change feature and variable interaction relation is comprehensively mined, potential relations among various data are fully utilized, and the problem that feature information is insufficient in utilization in an existing method is effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of artificial intelligence, and particularly relates to a method for early prediction of sepsis based on multi-dimensional Transformer. Background Art

[0002] Sepsis is a severe clinical syndrome, and early prediction is crucial for improving the prognosis of patients. Currently, early prediction of sepsis mainly relies on machine learning and deep learning technologies.

[0003] In the early prediction of sepsis based on machine learning, the prediction tasks of non-sequential data rely on the computing and analysis capabilities of machine learning models, and early signs of sepsis are identified by extracting static features. Islam et al. conducted a meta-analysis of the applications of various machine learning models such as random forest, support vector machine, and XGBoost in sepsis prediction, confirmed the effectiveness of the models, and emphasized the importance of feature selection and data processing. Hou et al. used the XGBoost model to analyze patient data in the MIMIC-III database and accurately predicted the 30-day mortality rate. Delahanty et al. developed a logistic regression model based on non-sequential features to help emergency department medical staff screen for sepsis risk. Alanazi et al. focused on the early identification of sepsis in ICU patients and proposed a regression synthesis model based on survival analysis. Song et al. developed a support vector machine (SVM) model to improve the accuracy of early identification of high-risk groups by combining static data.

[0004] Early prediction of sepsis based on deep learning has significant advantages in processing and predicting sequential data, and can more effectively capture the dynamic change trends of data. Reyna et al. proposed a deep learning model based on long short-term memory network (LSTM) and convolutional neural network (CNN) for early prediction. Scherpf et al. used recurrent neural network (RNN) to process sequential data in the MIMIC-III database to achieve high-accuracy early detection of sepsis. Lipton et al. systematically studied the application of RNN in clinical time series data and optimized its structure and parameters. Shashikumar et al. proposed the Deep Artificial Intelligence Survival Evaluation (DeepAISE) model, which combines survival analysis and RNN architecture to provide a more interpretable deep learning solution. In recent years, Transformer has made significant progress in sequential data modeling. Wang et al. combined physiological time series data with clinical notes and captured potential associations in the data through the Transformer network to achieve more accurate early prediction of sepsis. Duan et al. introduced a dual-feature fusion mechanism into the deep learning model to improve the generalization ability of the model. Zhao et al. proposed a bidirectional encoding representation model based on BERT (BERT Surv) for prognostic prediction of trauma patients. Lauritsen et al. constructed an LSTM-based model using electronic health record (EHR) event sequences and showed excellent performance in early detection of sepsis. Zhang et al. used an LSTM model combined with an attention mechanism to achieve early prediction of sepsis in the emergency department, improving the accuracy of prediction and the transparency of clinical application.

[0005] However, existing early prediction methods for sepsis mainly focus on feature extraction and modeling of time series, but fail to comprehensively explore the dependency relationships of clinical data in the variable dimension. This limitation may lead to insufficient utilization of feature information, thus restricting the prediction accuracy. Summary of the Invention

[0006] To solve the problems existing in the background technology, one aspect of the present invention provides an early prediction method for sepsis based on multi-dimensional Transformer, including:

[0007] S1: Obtain clinical variable data of the user to be tested, where the clinical variable data includes: multiple vital sign variable data, multiple laboratory test variable data, and multiple demographic variable data;

[0008] S2: Perform shallow feature extraction on the clinical variable data through a shallow feature extraction module to obtain shallow features of the clinical variable data;

[0009] S3: Input the shallow features of the clinical variable data into a deep feature extractor based on a multi-dimensional Transformer for deep feature extraction to obtain the deep features of the clinical variable data. Among them, the deep feature extractor includes: multiple cascaded feature extraction transformation groups; each feature extraction transformation group includes: a horizontal Transformer feature extraction module and a vertical Transformer feature extraction module;

[0010] S4: Input the deep features of the clinical variable data into the result prediction module to obtain the early prediction result of sepsis.

[0011] Another aspect of the present invention provides a sepsis early prediction device based on a multi-dimensional Transformer, including a processor and a memory; the memory is used to store a computer program; the processor is connected to the memory and is used to execute the computer program stored in the memory, so that the sepsis early prediction device based on a multi-dimensional Transformer executes the above-mentioned sepsis early prediction method based on a multi-dimensional Transformer.

[0012] Another aspect of the present invention provides a computer-readable storage medium storing a program, which when executed by a processor, implements the above-mentioned sepsis early prediction method based on a multi-dimensional Transformer.

[0013] The present invention has at least the following beneficial effects

[0014] By obtaining the clinical variable data of the user to be tested, including multiple vital sign variable data, multiple laboratory test variable data, and multiple demographic variable data, the present invention provides a rich data basis for comprehensive analysis. Then, with the help of the shallow feature extraction module for preliminary feature extraction, and then using the deep feature extractor based on a multi-dimensional Transformer, especially the multiple cascaded feature extraction transformation groups, each transformation group includes a horizontal Transformer feature extraction module and a vertical Transformer feature extraction module, which can realize the joint modeling of the time dimension and the variable dimension, comprehensively mine the time-varying features and variable interaction relationships, make full use of the potential connections between various types of data, and effectively solve the problem of insufficient utilization of feature information in existing methods. Brief Description of the Drawings

[0015] Figure 1 It is a schematic flowchart of the method of the present invention;

[0016] Figure 2 It is a schematic diagram of the model framework of the present invention;

[0017] Figure 3Schematic diagram of the model framework of the horizontal Transformer feature extraction module of the present invention;

[0018] Figure 4 Schematic diagram of the model framework of the vertical Transformer feature extraction module of the present invention;

[0019] Figure 5 Comparison chart of the AUC performance curves of the present invention compared with the traditional SEP method on the 2019 PhysioNet Challenge Dataset;

[0020] Figure 6 Comparison chart of the AUC performance curves of the present invention compared with the traditional SEP method on the MIMIC dataset;

[0021] Figure 7 Schematic diagram of the comparison of the AUC performance curve of the present invention with the latest and state-of-the-art SEP method on the 2019 PhysioNet challenge dataset;

[0022] Figure 8 Schematic diagram of the comparison of the AUC performance curve of the present invention with the latest SEP method on the MIMIC dataset. Detailed implementation manners

[0023] The following uses specific specific examples to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention schematically. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0024] Please refer to Figure 1 and Figure 2 The present invention provides a method for early sepsis prediction based on a multi-dimensional Transformer, including:

[0025] S1: Obtain the clinical variable data of the user to be tested, and the clinical variable data includes: a plurality of vital sign variable data, a plurality of laboratory test variable data, and a plurality of demographic variable data;

[0026] In this embodiment, the clinical variable data mainly includes vital signs, laboratory values, and demographic data. For the detailed data representation, please refer to Table 1:

[0027] Table 1 Clinical variable data

[0028]

[0029]

[0030]

[0031] In Table 1, 1-7 represent vital sign variable data, 9-34 represent laboratory test variable data, and 35-39 represent demographic variable data. In this embodiment, some variable data can be selected from the above vital sign variable data, some variable data can be selected from the laboratory test variable data as input data, and some variable data can be selected from the above vital sign variable data as input data. Of course, in addition to the above table, variable data such as plasma heparin-binding protein, procalcitonin, C-reactive protein, arterial blood lactate, SOFA score, APACHE II score, intrathoracic blood volume, pulse pressure variation, cardiac function index, extravascular lung water index, etc. can also be selected as input data.

[0032] S2: Perform shallow feature extraction on the clinical variable data through a shallow feature extraction module to obtain shallow features of the clinical variable data;

[0033] Preferably, the shallow feature extraction module includes a fully connected layer. Performing shallow feature extraction on the clinical variable data through the fully connected layer is expressed as:

[0034] H sha = ReLu(XW 1 +b 1 )

[0035] where X ∈ R T×F represents the input clinical variable data; T represents the time step; F represents the dimension of the clinical variable; W 1 represents the weight matrix of the fully connected layer; b 1 represents the bias matrix of the fully connected layer; ReLU represents the activation function, and H shal ∈ R T×F represents the shallow features of the clinical variable data.

[0036] S3: Input the shallow features of the clinical variable data into a deep feature extractor based on a multi-dimensional Transformer for deep feature extraction to obtain deep features of the clinical variable data; wherein, the deep feature extractor includes: multiple cascaded feature extraction transformation groups; each feature extraction transformation group includes: a horizontal Transformer feature extraction module and a vertical Transformer feature extraction module;

[0037] Preferably, the horizontal Transformer feature extraction module uses a temporal attention mechanism to calculate the similarity between the same clinical variables at different time points, capture the dynamic change patterns of the time series, assign higher weights to important time points, and thus enhance the expressive ability of temporal features; the vertical Transformer feature extraction module analyzes the correlation between different clinical variables at the same time point through a variable attention mechanism, captures the interaction patterns between different clinical variables, and utilizes the potential connections between different clinical variables to enhance the feature expression of multivariate clinical data.

[0038] In this embodiment, the feature extraction process of the horizontal Transformer feature extraction module is as follows:

[0039] S301: Define the input feature of the horizontal Transformer feature extraction module as feature h TA ∈R T×F ;

[0040] S302: Divide feature h TA along the first dimension into T different row vectors

[0041] S303: For each row vector input the row vector into multi-head attention processing to obtain the corresponding attention output feature

[0042] S304: Concatenate the attention output features of all row vectors along the first dimension to obtain feature I TA ∈R T ×F ;

[0043] S305: Perform a residual connection on feature O TA and feature h TA , and after normalization, obtain feature h atten-HT ∈R T×F ;

[0044] S306: Input feature H atten-HT into a feed-forward neural network to further process the hidden state to enhance the feature representation ability and obtain feature FFN(h atten-HT )∈R T×F ;

[0045] S307: Perform a residual connection on feature FFN(h atten-HT ) and feature h atten-HT , and after normalization, obtain the output feature h of the horizontal Transformer feature extraction module HT ∈RT×F 。

[0046] In this embodiment, in the Horizontal Transformer (HT), we propose a Temporal Attention mechanism (TA) to calculate the similarity of the same clinical variables at different times, in order to capture the dependencies and dynamic patterns between time steps in the input data. In laboratory and clinical data, the dynamic changes of time series are usually crucial for diagnosis and prediction. By capturing the relationships of the same clinical variable at different time points, HT can understand the evolution pattern of each clinical indicator over time. Given the input feature h TA ∈R T×F , representing the features of the sample, we input Temporal Attention (TA) to calculate the self-attention between different times. For TA, as Figure 3 shown, first h TA is divided into T non-overlapping row vectors: Then we generate query, key, and value related to as, and where J is the number of multi-head attentions. The multi-head attention mechanism is respectively:

[0047]

[0048] where, and represent the linear projection matrices of query, key, and value in the multi-head attention mechanism respectively. Therefore, the output feature of the j-th head of TA is calculated as follows:

[0049]

[0050] Then we calculate the outputs of all J heads and concatenate the outputs of all J heads to generate Combine all row vectors into a matrix O TA ∈R T×F , which is the output of TA. To stabilize the training results and improve the model performance, we add a residual connection and normalization after TA, and the expression is:

[0051] h atten-HT = LayerNorm(h TA + O TA )

[0052] The hidden state is further processed by a Feed-Forward Network (FFN) to enhance the feature representation ability. As Figure 1 shown, FFN consists of two fully connected layers and a ReLU activation function. Finally, the final output h HT ∈R T×F:

[0053] h HT = LayerNorm(h atten-HT + FFN(h atten-HT ))

[0054] The output of HT not only retains the dimensional structure of the input features, but also enhances the interaction of features in the time dimension through the multi-head self-attention mechanism and the feed-forward network. Therefore, HT can assign weights according to the importance of different time steps and learn the features in the time series more flexibly. The output of HT will be used as the input of the vertical Transformer (VT) to support the next stage of data dimension modeling.

[0055] S4: Input the deep features of the clinical variable data into the result prediction module to obtain the early prediction result of sepsis.

[0056] Preferably, the feature extraction process of the vertical Transformer feature extraction module is as follows:

[0057] S311: Define the input feature of the vertical Transformer feature extraction module as feature h VA ∈R T×F ;

[0058] S312: Divide the feature h VA into F different column vectors along the second dimension

[0059] S313: For each column vector Perform multi-head attention processing on the column vector to obtain the corresponding attention output feature

[0060] S314: Concatenate the attention output features of all column vectors along the second dimension to obtain feature O VA ∈R T ×F ;

[0061] S315: Perform a residual connection on feature O VA and feature h VA , and after normalization, obtain feature h atten-VT ;

[0062] S316: Input feature h atten-V into the feed-forward neural network to further process the hidden state to enhance the feature representation ability and obtain feature FFN(h atten-VT ) ∈R T×F ;

[0063] S317: Residually connect the feature FFN(h atten-VT ) and the feature F atten-VT , and after normalization, obtain the output feature h VT ∈ R T×F .

[0064] In this embodiment, the design goal of the vertical Transformer (VT) is to simulate the dependence and interaction patterns of the input data in the clinical variable dimension. Different from HT which focuses on the relationships between time steps, VT takes the clinical variable dimension as the main modeling object for each feature angle of the input data and explores the correlations between different clinical variables. This design can effectively capture the potential interactions between various clinical data (such as physiological indicators, laboratory test values, etc.) and provide a more comprehensive and detailed variable representation.

[0065] Similar to HT, we use h VA ∈ R T×F to represent the input of variable attention (VA). As Figure 4 shown, h VA is first divided into F non-overlapping column vectors: Then we generate queries, keys, and values related to , denoted as and respectively, where J is the number of heads for the multi-head attention mechanism:

[0066]

[0067] Among them, and represent the linear projection matrices of queries, keys, and values in the multi-head attention mechanism respectively. Therefore, the VA output feature of the j-th head is calculated as follows:

[0068]

[0069] Then we calculate the outputs of all J heads and concatenate the outputs of all J heads to produce a vector Merge all column vectors into a matrix O VA ∈ R T×F , which is the output of VA. Similar to HT, we also apply residual connection and normalization to O VA , and the formula for the final output h VT ∈ R T×F of VT is as follows:

[0070] h atten-VT = LayerNorm(hVA +O VA )

[0071] h VT = LayerNorm(h atten-VT + FFN(h atten-VT ))

[0072] Experimental simulation

[0073] 3. Experimental settings

[0074] 3.1 Dataset and evaluation metrics

[0075] To fairly evaluate the performance of the multi-dimensional Transformer network (MDTN) proposed in the present invention, we conducted experiments using the 2019 PhysioNet Challenge dataset and the MIMIC dataset. First, the sliding window technique was used to process the two datasets, splitting the multivariate time series data into samples of a fixed time length and dividing them into training and test sets at a ratio of 8:2, respectively. During the training process of the model, we ensured that the data preprocessing and splitting methods were consistent with the prior art to ensure the comparability of the results.

[0076] In terms of evaluation metrics, the present invention uses AUC (area under the curve), F1 score, Precision, and Recall as the main evaluation metrics to comprehensively evaluate the prediction performance of the MDTN. Among them, AUC reflects the ability of the model to distinguish positive and negative samples, and the F1 score comprehensively considers the balanced performance of precision and recall. In addition, we also evaluated the number of model parameters (Params) and the number of floating-point operations (FLOPs) to quantify the computational complexity and resource consumption of the model. The selection of the above metrics can comprehensively measure the accuracy, efficiency, and practical application value of the present invention.

[0077] 3.2 Training settings

[0078] The model of the present invention uses the Adam optimizer with an initial learning rate of 1×10 -4 and gradually decays to 1×10 -6 through the cosine annealing algorithm. We set the batch size to 32 and trained a total of 1×10 6 iterations. We set the number of TGs to G = 4, and each TG consists of an HT, a VT, and an ST in sequence. Our MDTN was trained and tested on an NVIDIA RTX3090 GPU and the PyTorch platform. In our MDTN, we use the classical binary cross-entropy loss function:

[0079]

[0080] where y n is the true label of the nth sample (1 represents a positive sample (high risk), 0 represents a negative sample (low risk)), and is the predicted probability of our MDTN for the nth sample.

[0081] 3.3 Experimental Results

[0082] To demonstrate the superiority of the proposed method, we compared the proposed MDTN with representative traditional sequence prediction (SEP) methods LSTM, SVM, RNN, random forest, logistic regression, gradient boosting tree, and Transformer, as well as the latest state-of-the-art SEP methods I-LSTM, GAN+LSTM, decision tree, GBDT, DFSP, and XGBoost. During the comparison with traditional methods, usually only the AUC curve performance is compared. Therefore, Table 2 Figure 4 and Figure 5 show the quantitative comparison of AUC between our MDTN and representative traditional methods. As shown in Table 2 Figure 5 and Figure 6 The proposed MDTN has significantly better AUC values on the two datasets than the representative traditional SEP methods. Especially for the 2019 PhysioNet challenge dataset, the AUC curve area of our MDTN reached 0.989, which is a more significant improvement compared to the second-best method, the Transformer model (0.882). Secondly, as shown in Table 2, our MDTN achieved an AUC value of 0.84 on the MIMIC dataset, showing a significant improvement compared to the second-best method, the gradient boosting tree (0.792).

[0083] Table 2 AUC values of our MDTN and traditional SEP methods on two datasets.

[0084]

[0085] In addition to traditional SEP methods, we also conducted extensive comparative experiments with the latest SEP methods I-LSTM, GAN+LSTM, DecisionTree, GBDT, DFSP, and XGBoost. To ensure the effectiveness of the comparison with the most recent state-of-the-art SEP methods, multiple metrics such as AUC, F1 score, Precision, and Recall are usually used to comprehensively evaluate the performance of all methods. The AUC curve performance diagrams are as shown in Figure 7 , Figure 8 and the comparison results of each metric are shown in Table 3 and Table 4.

[0086] Table 3 Quantitative comparison of our MDTN and the latest state-of-the-art SEP methods on the simulated dataset.

[0087]

[0088] Table 4 Quantitative comparison between our MDTN and the latest state-of-the-art SEP methods on the 2019 PhysioNet Challenge dataset.

[0089]

[0090] From Tables 3 and 4, we can see that for the MIMIC dataset, our MDTN outperforms the current state-of-the-art SEP methods in all metrics (see Table 3). Specifically, the AUC of MDTN reached 0.841, which is 0.009 higher than the second-highest method DFSP (AUC of 0.832), showing better prediction accuracy. At the same time, the F1-score of MDTN is 0.82, which is 0.01 higher than DFSP (F1-score of 0.81), and in terms of recall, MDTN reached 0.83, which is 0.01 higher than DFSP (recall of 0.82), further demonstrating the stability and sensitivity of our MDTN in identifying septic patients. In addition, although the accuracy of MDTN is the same as that of DFSP, both being 0.80, combined with the improvement of other metrics, MDTN shows significant advantages in overall performance, not only improving the detection accuracy but also effectively reducing the risks of false positives and false negatives.

[0091] For the 2019 PhysioNet Challenge dataset, the performance advantage of our MDTN is even more obvious. The F1-score of MDTN reached 0.94, which is 0.04 higher than the second-highest method DFSP (F1-score of 0.90), further confirming its strong performance in key metrics. In terms of recall, MDTN reached 0.95, which is 0.04 higher than DFSP (recall of 0.91), demonstrating its superior ability to identify septic patients. At the same time, the accuracy of MDTN is 0.93, which is 0.02 higher than GAN+LSTM (0.91), further indicating that it can accurately identify patients while maintaining a low false positive rate. Although the decision tree method is slightly higher than our MDTN in the AUC metric, in the other three core metrics, our MDTN far exceeds the decision tree. For example, our MDTN is 0.12, 0.05, and 0.10 higher than the decision tree in terms of F1-score, accuracy, and recall, respectively. Overall, our MDTN has achieved significant improvements in multiple metrics, fully demonstrating its excellent performance as an early predictor of sepsis.

[0092] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0093] In summary, the present invention provides rich data basis for comprehensive analysis by obtaining clinical variable data of a user to be measured, including multiple vital sign variable data, multiple laboratory test variable data, and multiple demographic variable data. Then, through the preliminary feature extraction by the shallow feature extraction module, and then using the deep feature extractor based on multi-dimensional Transformer, especially multiple cascaded feature extraction transformation groups, each of which contains a horizontal Transformer feature extraction module and a vertical Transformer feature extraction module, it can realize the joint modeling in the time dimension and variable dimension, comprehensively mine the time-varying features and variable interaction relationships, make full use of the potential connections between various types of data, and effectively solve the problem of insufficient utilization of feature information in the existing methods.

[0094] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the present technical solution, and they should all be covered by the scope of the claims of the present invention.

Claims

1. An early prediction method for sepsis based on multidimensional Transformer, characterized in that: include: S1: Acquire clinical variable data of the user to be tested, wherein the clinical variable data includes: a plurality of life characteristic variable data, a plurality of laboratory test variable data and a plurality of demographic variable data; S2: shallow feature extraction is performed on the clinical variable data through a shallow feature extraction module to obtain shallow features of the clinical variable data; S3: Inputting shallow features of clinical variable data into a deep feature extractor based on multidimensional Transformer to extract deep features to obtain deep features of clinical variable data; wherein the deep feature extractor comprises: a plurality of cascaded feature extraction deformation groups; each feature extraction deformation group comprises: a horizontal Transformer feature extraction module and a vertical Transformer feature extraction module; S4: The deep features of clinical variable data are input into the outcome prediction module to obtain early prediction results of sepsis.

2. The method for early prediction of sepsis based on multidimensional Transformer according to claim 1, characterized in that: The shallow feature extraction module includes a fully connected layer, and shallow feature extraction of clinical variable data is performed through the fully connected layer as follows: H shall =ReLu(XW1+b1) Where X∈R T×F represents the input clinical variable data; T represents the time step; F represents the dimension of clinical variables; W1 represents the weight matrix of the fully connected layer; b1 represents the bias matrix of the fully connected layer; ReLu represents the activation function, H sha ∈R T×F Represents shallow features of clinical variable data.

3. The method for early prediction of sepsis based on multidimensional Transformer according to claim 1, characterized in that: The horizontal Transformer feature extraction module adopts the temporal attention mechanism to calculate the similarity of the same clinical variable at different time points, captures the dynamic change pattern of the time series, and gives higher weights to important time points, thereby enhancing the expression ability of temporal features; the vertical Transformer feature extraction module analyzes the correlation between different clinical variables at the same time point through the variable attention mechanism, captures the interaction pattern between different clinical variables, and utilizes the potential connection between different clinical variables to enhance the feature expression of multivariate clinical data.

4. The method for early prediction of sepsis based on multidimensional Transformer according to claim 3, characterized in that: The feature extraction process of the horizontal Transformer feature extraction module is as follows: S301: Define the input feature of the horizontal Transformer feature extraction module as feature h TA ∈R T×F ; S302: The feature h TA Split into T different row vectors along the first dimension S303: For each row vector The row vector The input is processed by multi-head attention to obtain the corresponding attention output features S304: Concatenate the attention output features of all row vectors along the first dimension to obtain feature O TA ∈R T×F ; S305: Set feature O TA and feature h TA Perform residual connection and normalize to get feature h atten-HT ∈R T×F ; S306: The feature h atten-HT The input feedforward neural network further processes the hidden state to enhance the feature representation capability and obtain the feature FFN (h atten-HT )∈R T×F ; S307: The feature FFN(h atten-HT ) and feature h atten-HT Perform residual connection and normalization to obtain the output feature h of the horizontal Transformer feature extraction module HT ∈R T×F .

5. The method for early prediction of sepsis based on multidimensional Transformer according to claim 3, characterized in that: The feature extraction process of the vertical Transformer feature extraction module is as follows: S311: Define the input feature of the vertical Transformer feature extraction module as feature h VA ∈R T×F ; S312: The feature h VA Split into F different column vectors along the second dimension S313: For each column vector The column vector Perform multi-head attention processing to obtain the corresponding attention output features S314: Concatenate the attention output features of all column vectors along the second dimension to obtain feature O VA ∈R T×F ; S315: Set feature O VA and feature h VA Perform residual connection and normalize to get feature h atten-VT ; S316: Set feature h atten-VT The input feedforward neural network further processes the hidden state to enhance the feature representation capability and obtain the feature FFN (h atten-VT )∈R T×F ; S317: The feature FFN(h atten-VT ) and feature h atten-VT Perform residual connection and normalization to obtain the output feature h of the vertical Transformer feature extraction module VT ∈R T×F .

6. The method for early prediction of sepsis based on multidimensional Transformer according to claim 1, characterized in that: The result prediction module includes a cascaded fully connected layer and a Softmax classifier.

7. The method for early prediction of sepsis based on multidimensional Transformer according to claim 1, characterized in that: The early prediction results of sepsis include high risk or low risk.

8. An early prediction device for sepsis based on multidimensional Transformer, characterized in that: It comprises a processor and a memory; the memory is used to store a computer program; the processor is connected to the memory and is used to execute the computer program stored in the memory, so that the early prediction device of sepsis based on a multidimensional Transformer executes the early prediction method of sepsis based on a multidimensional Transformer according to any one of claims 1 to 7.

9. A computer-readable storage medium storing a program, characterized in that: When the program is executed by a processor, the early prediction method of sepsis based on a multidimensional Transformer according to any one of claims 1 to 7 is implemented.