Sepsis Risk Prediction Method and System Based on Transfer Learning and Temporal Feature Mining
Through the method based on transfer learning and timing feature mining, the problems of insufficient prediction accuracy and time-consuming calculation in early recognition of sepsis are solved, and high-precision and efficient sepsis risk level prediction is achieved, supporting clinicians' timely treatment decisions.
Patent Information
- Application Number
- CN202510376604.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-03-28
AI Technical Summary
The prior art has problems such as insufficient prediction accuracy and long calculation in the early recognition of sepsis, making it difficult to achieve timely and accurate risk assessment.
Using a method based on transfer learning and timing feature mining, the patient's physiological index data set is obtained for missing value filling, multi-scale timing feature extraction and screening, and the pre-trained timing model is optimized based on the target patient population data to output sepsis risk level prediction results.
It improves the accuracy and calculation efficiency of sepsis risk level prediction, can achieve high-precision and efficient prediction in the clinical environment, and supports timely treatment decisions.
Smart Images

Figure CN119889708B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical data analysis, and specifically relates to a sepsis risk prediction method and system based on transfer learning and time series feature mining. Background Art
[0002] Sepsis is a systemic inflammatory response syndrome caused by the invasion of various pathogenic microorganisms into the blood circulation, with high morbidity and mortality. Early identification and intervention are the keys to improving the prognosis of patients. If not diagnosed and treated in time, sepsis may cause serious complications such as myocarditis, gastrointestinal bleeding, uremia, etc., and even lead to multiple organ failure. The early identification of traditional sepsis mainly relies on the empirical judgment of clinicians, but this method has problems of strong subjectivity and high latency, and it is difficult to achieve timely and accurate prediction.
[0003] To improve the early prediction ability of sepsis, existing technologies have proposed prediction methods based on machine learning. This method collects data on multiple physiological indicators of sepsis patients at different times, constructs a sepsis prediction model; then collects the physiological indicator data of the target patient at the current time, extracts the target feature data, and evaluates the risk level of the target patient's future sepsis based on the prediction model. By extracting the feature data of physiological indicators and using a machine learning model for prediction, this method improves the prediction accuracy to a certain extent. However, existing methods still have significant limitations in practical applications: First, due to the high-dimensionality and time series complexity of physiological indicator data, existing methods fail to fully mine the multi-scale time series features in the data during the feature extraction process, resulting in insufficient prediction accuracy and limiting the prediction performance of the model; Second, the prediction process of existing technologies takes a long time and is difficult to meet the needs of clinical real-time prediction. Summary of the Invention
[0004] To improve the accuracy and calculation efficiency of sepsis risk level prediction, the present invention provides a sepsis risk prediction method and system based on transfer learning and time series feature mining. The specific technical solutions adopted are as follows:
[0005] The technical solution of the first aspect of the present invention provides a sepsis risk prediction method based on transfer learning and time series feature mining, and the method includes:
[0006] Obtain the physiological indicator data set of the patient, and perform filling preprocessing on the missing values based on the time series;
[0007] Use wavelet transform to extract the multi-scale time series features of the physiological indicator data set, and screen the multi-scale time series features based on the feature importance;
[0008] Based on the pre-trained time series model, use transfer learning to combine the data of the target patient group, and optimize the pre-trained time series model according to the screened time series features;
[0009] Based on the optimized pre-trained time series model, output the prediction result of the sepsis risk level of the target patient.
[0010] Furthermore, obtain the physiological index dataset of the patient, and perform preprocessing on filling missing values based on time series, including:
[0011] Detect missing values in the physiological index dataset of the patient, and extract the set of missing time points;
[0012] Classify the missing windows according to the set of missing time points, and fill the missing values in the physiological index dataset based on the category of the missing windows;
[0013] Standardize the filled physiological index data for each patient individual.
[0014] Furthermore, classify the missing windows according to the set of missing time points, and fill the missing values in the physiological index dataset based on the category of the missing windows, including:
[0015] For continuous missing windows, use linear interpolation to fill the missing values in the physiological index dataset, which can be expressed as:
[0016]
[0017] In the formula, represents the filled value of the th item of the th patient at time point ; represents the offset calculated from the start time point of the missing window; represents the number of consecutive missing time points; represents the valid value of the th item of the th patient at time point ; represents the valid value of the th item of the th patient at time point ;
[0018] Furthermore, classify the missing windows according to the set of missing time points, and fill the missing values in the physiological index dataset based on the category of the missing windows, and also include:
[0019] For non - continuous missing windows, use adaptive spline interpolation to fill the missing values in the physiological index dataset, which can be expressed as:
[0020]
[0021] In the formula, Indicates The patient's The indicators at the time point Fill value of Indicates indivual The spline basis function at time point The value of Indicates patient no. The indicator is The spline coefficients on the basis functions; Indicates patient no. Cubic spline interpolation function of the index; express The number of spline basis functions.
[0022] Furthermore, wavelet transform is used to extract multi-scale time series features of the physiological index data set, and the multi-scale time series features are screened based on feature importance, including:
[0023] Perform discrete wavelet transform on the physiological index time series data set according to the preset decomposition level, and output a multi-scale wavelet coefficient matrix;
[0024] Calculate the comprehensive fluctuation intensity of the multi-scale wavelet coefficient matrix;
[0025] The screened multi-scale time series feature set was extracted based on the average value of the comprehensive fluctuation intensity of the patient group.
[0026] Furthermore, the expression for performing discrete wavelet transform on the physiological index time series data set according to the preset decomposition level is:
[0027]
[0028] In the formula, Indicates The patient's The indicator is Layer wavelet decomposition time point The wavelet coefficients of ; Indicates that the wavelet basis function is Layer time point For the original time point The response value, Represents the translation index of the wavelet basis function; Indicates The patient's The indicators are at the original time point The standardized value of Indicates the total number of time points.
[0029] Further, based on the pre-trained time series model, using transfer learning combined with the data of the target patient group, the pre-trained time series model is optimized according to the selected time series features, including:
[0030] Based on maximizing the difference and minimizing the alignment of the feature distributions of the source domain and the target domain;
[0031] Combined with the cross-entropy loss function to adjust the pre-trained time series model and optimize the parameters using the data of the target domain.
[0032] Further, the expression for maximizing the difference and minimizing the alignment of the feature distributions of the source domain and the target domain is:
[0033]
[0034] In the formula, represents the maximum mean difference loss; represents the number of patients in the source domain; represents the number of patients in the target domain; represents the kernel mapping function; represents the th selected time series feature of the th patient in the source domain; represents the th selected time series feature of the th patient in the target domain;
[0035] Further, based on the optimized pre-trained time series model, the prediction result of the sepsis risk level of the target patient is output, including:
[0036] Input the selected time series features of the target patient into the optimized pre-trained time series model, and output the prediction result of the sepsis risk probability of each target patient;
[0037] According to the prediction result of the sepsis risk probability, the target patients are divided into different risk levels.
[0038] The technical solution of the second aspect of the present invention provides a sepsis risk prediction system based on transfer learning and time series feature mining, adopting the sepsis risk prediction method based on transfer learning and time series feature mining described in the technical solution of the first aspect of the present invention. The system includes:
[0039] A data acquisition and preprocessing module, configured to acquire the physiological index data set of the patient and perform preprocessing on the missing values based on the time series;
[0040] A feature extraction module, configured to extract multi-scale time series features of a physiological index data set by using wavelet transform, and screen the multi-scale time series features based on feature importance;
[0041] A transfer learning module, configured to optimize a pre-trained time series model based on the screened time series features by using transfer learning to combine target patient group data based on the pre-trained time series model;
[0042] A sepsis prediction module, configured to output a prediction result of the sepsis risk level of a target patient based on the optimized pre-trained time series model.
[0043] The present invention has the following beneficial effects:
[0044] The sepsis risk prediction method and system based on transfer learning and time series feature mining provided by the present invention extract multi-scale time series features of physiological indexes by using wavelet transform and screen the multi-scale time series features based on feature importance, can capture local and global dynamic change patterns, screen out features highly related to sepsis risk, comprehensively reflect the physiological state of patients while reducing redundant features and reducing the computational complexity of the model; finally, based on the pre-trained time series model, transfer learning is used, combined with target patient group data and the screened time series features, the source domain and target domain feature distributions are aligned by maximizing the difference and minimizing the alignment, and the model is optimized by combining the cross-entropy loss function, so that the model can better adapt to the target patient group, improve the prediction accuracy and generalization ability, and solve the performance bottleneck of the prediction model in the case of limited data. This method realizes high-precision and high-efficiency prediction of sepsis risk levels. Description of the Drawings
[0045] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings required to be used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0046] Figure 1 It is a method flow chart of the sepsis risk prediction method based on transfer learning and time series feature mining provided by an embodiment of the present invention;
[0047] Figure 2 It is a structural schematic diagram of the sepsis risk prediction system based on transfer learning and time series feature mining provided by an embodiment of the present invention. Detailed Embodiments
[0048] To further elaborate on the technical means and effects adopted by the present invention to achieve the intended invention purpose, the following specifically describes, with reference to the accompanying drawings and preferred embodiments, a sepsis risk prediction method and system based on transfer learning and temporal feature mining according to the present invention, including its specific implementation manner, structure, features, and effects. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.
[0049] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs.
[0050] The following specifically describes the specific solution of a sepsis risk prediction method and system based on transfer learning and temporal feature mining provided by the present invention with reference to the accompanying drawings.
[0051] Please refer to Figure 1 , which shows the method flowchart of a sepsis risk prediction method based on transfer learning and temporal feature mining provided by an embodiment of the present invention. The method includes:
[0052] Step S100: Obtain the physiological index dataset of patients and perform preprocessing on filling missing values based on time series; specifically, in this embodiment, the physiological index dataset of all sepsis patients in the past 3 years is obtained from the hospital information system. The physiological index data includes at least: indicators such as body temperature (°C), systolic blood pressure (mmHg), diastolic blood pressure (mmHg), heart rate (beats / minute), respiratory rate (breaths / minute), and blood oxygen saturation (%); the patient population set can be represented as , represents the number of patients; the physiological index time series of each patient can be represented as: , represents the total number of time points; represents the th patient at time - dimensional physiological index, , including body temperature, systolic blood pressure, diastolic blood pressure, heart rate, respiratory rate, and blood oxygen saturation;
[0053] Step S100 specifically includes:
[0054] Step S110: Detect missing values in the physiological index dataset of patients and extract the set of missing time points;
[0055] Step S120: Classify the missing windows according to the set of missing time points, and fill in the missing values in the physiological index data set based on the category of the missing windows; specifically, for continuous missing values, use the sliding window technique. If multiple missing values appear continuously within a certain time range, it is determined as a continuous missing window; for non - continuous missing values, form a set of non - continuous missing time points by recording the time points of the missing values.
[0056] For continuous missing windows, use linear interpolation to fill in the missing values in the physiological index data set. Use the valid data points before and after for linear interpolation, which can be expressed as:
[0057]
[0058] In the formula, represents the filled value of the th item of the th patient's index at time point ; represents the offset calculated from the start time point of the missing window; represents the th item of the th patient's index at time point ; represents the th item of the th patient's index at time point ; represents the number of consecutive missing time points, , that is, at most 5 hours of continuous missing values are filled;
[0059] For the non - continuous missing time points of the patient, in this embodiment, considering the overall trend of the physiological index time series, fit a cubic spline function based on the complete physiological index time series of the patient. The filling expression is:
[0060]
[0061] In the formula, represents the filled value of the th item of the th patient's index at time point ; represents the value of the th spline basis function at time point ; represents the th item of the th patient's index on the th basis function; represents the th item of the The cubic spline interpolation function of the indicators; The number of spline basis functions. For the th patient, the spline coefficients can be solved by minimizing the objective function, and its expression is:
[0062]
[0063] In the formula, represents the set of observed time points of the th patient, that is, the set of all non-missing value time points; represents the smoothing parameter; represents the th patient's th indicator's filled value at the time point ; represents the maximum value of the time range; represents the minimum value of the time range; represents the second derivative of the spline function ;
[0064] Step S130: Standardize the filled physiological indicator data for each patient individually; specifically, in the actual medical scenario, the basic physiological states of different patients are different. To eliminate the influence of this difference on subsequent analysis, in this embodiment, the mean and standard deviation of the physiological indicators are calculated for each patient for the filled physiological indicator data to complete the standardization process, and finally the preprocessed data set is output, which can be expressed as , , ; represents the preprocessed data set, which contains the standardized physiological indicator data of all patients; represents the sub-data set of the th patient, which contains the dimensional standardized physiological indicator data of all time points of this patient; represents the th patient's dimensional standardized physiological indicator vector at the time point ; is the standardized value of the th physiological indicator in this vector.
[0065] In this embodiment, through the system's data acquisition and preprocessing process, the aim is to improve data quality. In the actual medical application scenario, the integrity of the patient's physiological index data is ensured. Through the missing value filling strategy provided in this embodiment, the information loss and deviation caused by data missing are avoided. The normalization process eliminates the interference of differences in the basic physiological states of different patients on the analysis, puts the data of different patients on the same comparable scale, and enhances the consistency and usability of the data. It not only improves the reliability of the data, but also enhances the accuracy and stability of subsequent feature extraction and model training, helps to more accurately predict the sepsis risk, and provides a more reliable decision-making basis for clinicians.
[0066] Step S200: Extract multi-scale time series features of the physiological index data set by using wavelet transform, and screen the multi-scale time series features based on feature importance; specifically, in this embodiment, the physiological change characteristics at different time scales during the development of sepsis are considered; for example, in the early stage of sepsis, some short-term physiological fluctuations may be signals of the onset of the disease; as the disease progresses, medium-term trends and long-term trends will appear. Therefore, this hierarchical decomposition can comprehensively capture the physiological feature changes at different stages.
[0067] Step S200 specifically includes:
[0068] Step S210: Perform discrete wavelet transform on the physiological index time series data set according to the preset decomposition level, and output a multi-scale wavelet coefficient matrix; specifically, the wavelet basis function in this embodiment preferably uses the Daubechies wavelet basis, and the decomposition level preferably uses 3 layers. The purpose is to decompose the physiological index time series data into signals at different time scales through discrete wavelet transform. The 3-layer decomposition level corresponds to the following time scales:
[0069] The first layer is used to capture short-term fluctuations of 0.5 - 1 hour, such as sudden heart rate changes;
[0070] The second layer is used to capture medium-term trends of 1 - 4 hours, such as slow blood decline;
[0071] The third layer is used to capture long-term changes of 4 - 24 hours, such as continuous increase in body temperature;
[0072] The discrete wavelet transform of the physiological index time series data of the patient can be expressed as:
[0073]
[0074] In the formula, represents the th wavelet coefficient of the rd index of the th patient at the time point under the -layer wavelet decomposition, denotes the response value of the wavelet basis function at the time point of the th layer, to the original time point , where represents the translation index of the wavelet basis function; represents the th patient's th indicator's standardized value at the original time point ; represents the total number of time points;
[0075] Finally, a multi-scale wavelet coefficient matrix , can be obtained, where represents the physiological index dimension and 3 represents the decomposition layer number.
[0076] Step S220: Calculate the comprehensive fluctuation intensity of the multi-scale wavelet coefficient matrix; specifically, in this embodiment, the dynamic abnormality degree of the physiological index is quantified by fusing variance and energy, and the variance can be expressed as:
[0077]
[0078]
[0079] In the formula, represents the variance of the th patient's th indicator at the th layer; represents the mean value of the th patient's th indicator at the th layer; The variance reflects the discreteness of the wavelet coefficients of this indicator of this patient at the corresponding layer, and reflects the fluctuation stability of the indicator on the corresponding time scale. For example, when analyzing the heart rate indicator, a larger variance may mean that the heart rate fluctuates violently within this time scale.
[0080] The energy can be expressed as:
[0081]
[0082] In the formula, represents the energy of the th patient's th indicator at the th layer; The energy reflects the signal intensity of this indicator on the corresponding time scale. The higher the energy, the more significant the signal on this time scale.
[0083] The comprehensive fluctuation intensity can be expressed as:
[0084]
[0085] In the formula, Indicates The patient's The indicator is The comprehensive fluctuation intensity under the wavelet decomposition of the layer; the comprehensive fluctuation intensity combines the variance and energy indicators to more comprehensively quantify the dynamic abnormality of physiological indicators. For example, for blood pressure indicators, a high comprehensive fluctuation intensity may indicate that blood pressure has both large fluctuations and strong signal changes within the corresponding time scale, which is more likely to be associated with the risk of sepsis.
[0086] Step S230: extract the screened multi-scale time series feature set according to the average comprehensive fluctuation intensity of the patient group; specifically, this implementation screens the features highly correlated with the risk of sepsis through the average comprehensive fluctuation intensity value of the patient group to avoid manual parameter adjustment; the quantile threshold preferably retains the top 20% of the high comprehensive fluctuation intensity features, adaptively matches the data distribution, and the group average comprehensive fluctuation intensity can be expressed as:
[0087]
[0088] In the formula, Represents the patient population The indicator is The average value of the comprehensive fluctuation intensity under the layer wavelet decomposition, Represents the total number of patients; by calculating the average comprehensive fluctuation intensity of the patient group's item index at the layer, we can understand the fluctuation of the entire patient group on different indicators and time scales. The top 20% of high comprehensive fluctuation intensity features are retained because these features show large fluctuations in the group and are more likely to be highly correlated with sepsis risk. This screening method based on group statistics avoids the subjectivity of manual parameter adjustment and can adaptively match data distribution. The time series features after selection can be expressed as: , To retain the number of features;
[0089] This embodiment uses multi-scale feature extraction and screening, and discrete wavelet transform decomposes the physiological indicator time series data according to different time scales, which can comprehensively capture the changes in physiological characteristics at each stage of sepsis development, from early short-term fluctuations to medium-term and long-term trends, providing rich information for prediction. Then, by calculating the comprehensive fluctuation intensity, integrating variance and energy, the dynamic abnormalities of physiological indicators can be more accurately quantified, which helps to mine potential risk signals. This embodiment is based on feature screening based on patient group statistics, avoiding the subjectivity of manual parameter adjustment, adaptively matching data distribution, screening out features that are highly correlated with sepsis risk, and reducing redundancy. It not only reduces the complexity of subsequent model calculations, but also improves the model prediction accuracy and generalization ability, making the prediction model more stable and accurate in different patient groups and medical environments.
[0090] Step S300: Based on the pre-trained time series model, use transfer learning to combine the data of the target patient group, and optimize the pre-trained time series model according to the selected time series features; specifically, the source domain data in this embodiment is: the multi-scale time series features of historical patients and labels ; the target domain data is the selected time series features of the target patient group ; the pre-trained time series model preferably uses an LSTM model;
[0091] Step S300 specifically includes:
[0092] Step S300 specifically includes:
[0093] Step S310: Maximize the difference and minimize the alignment of the source domain and target domain feature distributions, which can be expressed as:
[0094]
[0095] In the formula, represents the maximum mean difference loss; represents the number of source domain patients, that is, the number of samples of historical patients; represents the number of target domain patients, that is, the number of samples of the target patient group to be predicted currently; represents the kernel mapping function; represents the selected time series feature of the th patient in the source domain; represents the selected time series feature of the th patient in the target domain; represents the reproducing kernel Hilbert space; the kernel mapping function preferably uses a Gaussian kernel function; represents the kernel space norm.
[0096] Step S320: Adjust the pre-trained time series model in combination with the cross-entropy loss function, and optimize the parameters using the target domain data, which can be expressed as:
[0097]
[0098] In the formula, represents the true sepsis risk label of the th patient in the target domain; represents the predicted value of the true sepsis risk probability of the pre-trained model for the th patient in the target domain;
[0099] The optimization objective is:
[0100]
[0101] In the formula, Denote the optimized model parameters; Denote the hyperparameters; during the optimization process, the Adam optimizer is used to update the model parameters.
[0102] In the actual medical scenario, in this embodiment, the feature distributions of the source domain and the target domain are aligned by the maximum mean discrepancy, effectively overcoming the problem of inconsistent data distributions between the source domain and the target domain. This enables the model to fully utilize the knowledge contained in the rich historical patient data in the source domain and successfully transfer it to the target patient population, greatly enhancing the generalization ability of the model among different patient populations. At the same time, the cross-entropy loss function is combined to optimize the model, ensuring the prediction accuracy of the model in the target domain and significantly improving the accuracy of sepsis risk prediction. It provides more timely and accurate support for the early diagnosis and treatment of sepsis, helping clinicians more accurately judge the sepsis risk of patients and formulate more effective treatment plans.
[0103] Step S400: Based on the optimized pre-trained time series model, output the prediction result of the sepsis risk level of the target patient; in actual clinical applications, obtain the multi-scale time series feature data of the target patient population after screening from the hospital's electronic medical record system or data storage platform. These data have been subjected to feature extraction and screening in step S200 and have features highly relevant to sepsis risk; then load the pre-trained time series model optimized in step S300 for prediction; specifically including:
[0104] Step S410: Input the multi-scale time series features of the target patient after screening into the optimized pre-trained time series model, and the model will output the sepsis risk probability of each target patient, which can be expressed as:
[0105]
[0106] In the formula, Denote the optimized time series model; the predicted risk probability values of all target patients can form a vector of length , Denote the sepsis risk probability of the th patient in the target domain;
[0107] Step S420: According to the predicted sepsis risk probability, divide the target patients into different risk levels, which can be expressed as:
[0108]
[0109] Finally, corresponding treatment plans can be formulated according to the risk level. For example, for low-risk patients, according to the established medical process, the unnecessary monitoring frequency can be appropriately reduced, such as adjusting the physiological index monitoring from once per hour to once every 4 hours, while reducing the treatment intensity to avoid over-medical treatment and save medical resources. For example, for medium-risk patients, strengthen the monitoring intensity, increase the monitoring items or frequency, such as increasing the frequency of blood tests, and take preventive measures, such as using antibiotics in advance to prevent the aggravation of infection. For example, for high-risk patients, initiate an emergency treatment plan, organize a multi-disciplinary team for consultation, and formulate a comprehensive and aggressive treatment intervention plan;
[0110] In this embodiment, by inputting the characteristics of the target patient into the pre-trained time series model optimized by transfer learning and target domain data, the high accuracy and reliability of sepsis risk probability prediction are ensured. The model has undergone multi-step optimization in the early stage, fully learning the knowledge of the source domain data and adapting to the characteristics of the target patient group. Further dividing the risk probability into intuitive risk levels greatly facilitates the work of clinicians. Doctors can quickly and clearly understand the sepsis risk degree of each patient, and thus formulate highly targeted treatment plans based on the risk level.
[0111] In summary, the sepsis risk prediction method based on transfer learning and time series feature mining provided by the present invention significantly improves the prediction accuracy and the computational efficiency of the prediction process through multi-stage collaborative optimization. First, based on the filling of missing values in the time series, based on linear interpolation, adaptive spline interpolation, and individual standardization processing, data noise and individual differences are effectively eliminated to ensure a high signal-to-noise ratio of the input features; then, multi-scale time series features are extracted through wavelet transform, and key features are screened by combining the comprehensive fluctuation intensity, accurately quantifying the short-term fluctuations, medium-term trends, and long-term anomalies of physiological indicators, enhancing the sensitivity to early sepsis signals; furthermore, the source domain historical data is used to pre-train the model, the target domain feature distribution is aligned through the maximum mean discrepancy loss, and fine-tuning is combined with the cross-entropy loss to achieve high generalization performance in the scenario of limited data; finally, the optimized lightweight model supports millisecond-level inference, and the risk level (low / medium / high risk) is divided by combining dynamic thresholds, providing interpretable and actionable warning results for clinicians to assist in early intervention. This method combines multi-scale feature fusion, cross-domain knowledge transfer, and an efficient computing architecture, taking into account prediction accuracy, computational efficiency, and clinical interpretability, providing a reliable technical means for the precise prevention and control of sepsis.
[0112] Please refer to Figure 2 , which shows a schematic structural diagram of a sepsis risk prediction system based on transfer learning and time series feature mining provided by an embodiment of the present invention. The system includes:
[0113] A data acquisition and preprocessing module, configured to acquire a physiological index data set of a patient and perform filling preprocessing on missing values based on time series;
[0114] A feature extraction module, configured to extract multi-scale time series features of the physiological index data set by using wavelet transform and screen the multi-scale time series features based on feature importance;
[0115] A transfer learning module, configured to optimize the pre-trained time series model based on the screened time series features by using transfer learning to combine the data of the target patient group based on the pre-trained time series model;
[0116] A sepsis prediction module, configured to output a prediction result of the sepsis risk level of the target patient based on the optimized pre-trained time series model.
[0117] It should be noted that: the above sequence of the embodiments of the present invention is only for description and does not represent the advantages and disadvantages of the embodiments. The processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multi-task processing and parallel processing are also possible or may be advantageous.
[0118] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments.
Claims
1. A sepsis risk prediction method based on transfer learning and temporal feature mining, characterized in that, The method includes: Obtaining a physiological index data set of a patient, and performing filling preprocessing on missing values based on time series; Using wavelet transform to extract multi-scale time series features of the physiological index data set, and screening the multi-scale time series features based on feature importance, including: Performing discrete wavelet transform on the physiological index time series data set according to a preset decomposition level, and outputting a multi-scale wavelet coefficient matrix. The preset decomposition levels include: the first layer for capturing short-term fluctuations of 0.5 - 1 hour, the second layer for capturing medium-term trends of 1 - 4 hours, and the third layer for capturing long-term changes of 4 - 24 hours; Calculating the comprehensive fluctuation intensity of the multi-scale wavelet coefficient matrix, which can be expressed as: In the formula, represents the comprehensive fluctuation intensity of the -th index of the -th patient under the -th layer of wavelet decomposition; represents the variance of the -th index of the -th patient at the -th layer; represents the energy of the -th index of the -th patient at the -th layer; Extracting the screened multi-scale time series feature set according to the average value of the comprehensive fluctuation intensity of the patient group; Based on a pre-trained time series model, using transfer learning to combine with the data of the target patient group, and optimizing the pre-trained time series model according to the screened time series features, including: Based on maximizing the minimum difference to align the feature distributions of the source domain and the target domain, the expression is: In the formula, denotes the maximum mean discrepancy loss; denotes the number of patients in the source domain; denotes the number of patients in the target domain; denotes the kernel mapping function; denotes the screened time series features of the th patient in the source domain; denotes the screened time series features of the th patient in the target domain; denotes the reproducing kernel Hilbert space; denotes the norm of the kernel space; Combining the cross-entropy loss function to adjust the pre-trained time series model, and optimizing the parameters using the target domain data; Based on the optimized pre-trained time series model, outputting the prediction result of the sepsis risk level of the target patient.
2. The sepsis risk prediction method according to claim 1, wherein Obtaining a physiological index data set of a patient, and performing filling preprocessing on missing values based on time series, including: Detecting missing values in the physiological index data set of the patient, and extracting the set of missing time points; Classifying the missing windows according to the set of missing time points, and filling the missing values in the physiological index data set based on the category of the missing windows; Performing standardization processing on the filled physiological index data for each patient individual.
3. The sepsis risk prediction method according to claim 2, wherein Classifying the missing windows according to the set of missing time points, and filling the missing values in the physiological index data set based on the category of the missing windows, including: For continuous missing windows, using linear interpolation to fill the missing values in the physiological index data set, which can be expressed as: In the formula, represents the th index value of the th patient at the time point represents the offset calculated from the start time point of the missing window; represents the number of consecutive missing time points; represents the th index value of the th patient at the time point represents the th index value of the th patient at the time point 4. The sepsis risk prediction method according to claim 3, wherein Classifying the missing windows according to the set of missing time points, and filling the missing values in the physiological index data set based on the category of the missing windows, and also including: For non-continuous missing windows, using adaptive spline interpolation to fill the missing values in the physiological index data set, which can be expressed as: In the formula, represents the th index of the th patient at time point ; represents the value of the th spline basis function at time point ; represents the th spline coefficient of the th patient's th index on the th spline basis function; 5. The sepsis risk prediction method according to claim 1, characterized in that The expression for performing discrete wavelet transform on the physiological index time series data set according to a preset decomposition level is: Wherein, represents the th index of the th patient at the time point under the -layer wavelet decomposition, ; represents the response value of the wavelet basis function at the time point of the th layer to the original time point , represents the translation index of the wavelet basis function; represents the th index of the th patient at the original time point ; represents the total number of time points.
6. The sepsis risk prediction method according to claim 1, wherein Based on the optimized pre-trained time series model, outputting the prediction result of the sepsis risk level of the target patient, including: Inputting the screened time series features of the target patient into the optimized pre-trained time series model, and outputting the prediction result of the sepsis risk probability for each target patient; According to the prediction result of the sepsis risk probability, dividing the target patients into different risk levels.
7. A sepsis risk prediction system based on transfer learning and temporal feature mining, characterized in that Adopting the sepsis risk prediction method based on transfer learning and time series feature mining according to any one of claims 1 to 6. The system includes: A data acquisition and preprocessing module configured to obtain a physiological index data set of a patient, and perform filling preprocessing on missing values based on time series; A feature extraction module configured to use wavelet transform to extract multi-scale time series features of the physiological index data set, and screen the multi-scale time series features based on feature importance; A transfer learning module, configured to optimize a pre-trained time series model based on the screened time series features by using transfer learning to combine the data of the target patient population based on the pre-trained time series model; A sepsis prediction module, configured to output a prediction result of the sepsis risk level of the target patient based on the optimized pre-trained time series model.
Citation Information
Patent Citations
Disease risk prediction method and equipment
CN112669968A
Real-time shock risk early warning and monitoring method and system based on medical internet of things time series data and deep learning algorithm
CN117393153A
Prediction method and device for septicemia
CN118983101A