A preoperative method of predicting probability of postoperative pulmonary complications based on sleep stage data

CN115775630BActive Publication Date: 2026-09-15BEIJING HAISIRUIGE TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202310096690.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-10
Publication Date
2026-09-15
Estimated Expiration
2043-02-10

AI Technical Summary

Technical Problem

另一项前瞻性多中心队列研究专门关注上腹部切口患者的PPCs风险分层,他们的回归模型中确定了五个独立的风险因素,包括麻醉时间、手术类别、呼吸并发症、当前是否吸烟和预测的最大吸氧量,得分低于2.02与PPC的高风险相关[OR(CI)=8.41(3.33–21.26)],但是该模型仍需外部验证

Benefits of technology

[0019] Preferably, the method for forming the prediction function further includes: verifying the prediction model, including: dividing the dataset into 5 equal parts, namely D1, D2, D3, D4 and D5, using D1-D4 as the validation set and D5 as the validation set, calculating the performance parameters of the first fold; repeating this process 4 times to calculate the performance parameters of the second to fifth folds; and calculating the accuracy of the prediction function based on the performance parameters to determine whether the prediction function is suitable.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115775630B_ABST
    Figure CN115775630B_ABST
Patent Text Reader

Abstract

The present disclosure belongs to the field of lung complication prediction, and particularly relates to a preoperative sleep stage data-based postoperative lung complication probability prediction method, comprising: obtaining a preoperative physiological clinical characteristic data set of a predicted person, the preoperative physiological clinical characteristic data set comprising a preoperative physiological characteristic data set and a preoperative clinical characteristic data set; obtaining the preoperative clinical characteristic data set comprises collecting continuous physiological signals of a sleep stage of the predicted person, and extracting preoperative clinical characteristic data of the predicted person based on part of the continuous physiological signals; obtaining a predicted lung postoperative complication probability based on the physiological characteristic data set and the clinical characteristic data with a prediction function; wherein an array formed by input characteristic data of the predicted person x complication probability, θ a feature coefficient vector formed by feature coefficients, x a feature data vector formed by feature data. The probability of lung postoperative complications is conveniently and relatively accurately evaluated before operation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure pertains to the field of pulmonary complication prediction, and more specifically relates to a method for predicting the probability of postoperative pulmonary complications based on preoperative sleep stage data. Background Technology

[0002] Postoperative pulmonary complications (PPCs) are associated with increased postoperative mortality, prolonged hospital stays, and higher healthcare costs, and are a major cause of poor postoperative outcomes in surgical patients. The incidence of PPCs varies significantly among different surgical populations, with cardiac surgeons facing a higher risk. Globally, over 40 million people suffer from mitral or aortic valve disease. Heart valve surgery is one of the riskier cardiac surgeries; the cardiopulmonary bypass during the procedure can easily trigger systemic inflammatory responses and oxidative stress, leading to pulmonary ischemia-reperfusion injury. The main reasons for ICU readmission after heart valve surgery are 40% due to postoperative pulmonary complications and 23% due to postoperative respiratory failure, both common types of PPCs.

[0003] Preoperative assessment of the risk probability of perioperative heart disease (PPCs) is crucial for providing important information on whether and when surgery should be performed. Therefore, effectively predicting the risk of PPCs in the cardiac valve surgery population before surgery, enabling timely warnings and early clinical interventions to reduce adverse PPC outcomes, has become a pressing clinical challenge. Cavayas et al. used diaphragmatic ultrasound to assess preoperative diaphragmatic function in cardiac surgeons, finding that the fraction of maximum diaphragmatic thickness decrease during inspiration (≤38.1%) helped identify the risk of PPCs in this population (OR=4.9; 95% CI, 1.81-13.50; p = 0.002). However, this requires extensive clinical ultrasound experience and expensive ultrasound equipment. In addition, other researchers have developed numerous risk prediction models to identify patients at high risk of PPCs, thereby enabling better perioperative management. The risk of PPCs in Catalan surgical patients was assessed by categorizing them into low, intermediate, and high-risk groups. The regression model included seven independent variables: preoperative peripheral oxygen saturation below 96%, previous month's respiratory infection, age, preoperative anemia (<100 g / dL), surgical site, operative time (>2 hours), and whether it was an emergency surgery. Valentín Mazo et al. validated the external validity of this model in a European population, demonstrating good discriminative ability. Another prospective, multicenter cohort study specifically focused on the risk stratification of PPCs in patients with upper abdominal incisions. Their regression model identified five independent risk factors: anesthesia duration, surgical type, respiratory complications, current smoking status, and predicted maximum oxygen uptake. A score below 2.02 was associated with a high risk of PPCs [OR (CI) = 8.41 (3.33–21.26)], but this model still requires external validation.

[0004] Current risk scoring methods all rely to varying degrees on prior diagnoses (smoking history, COPD history, preoperative sepsis, existing thoracic or upper abdominal surgical incisions, etc.), clinical examination results (pneumonia), and laboratory results (preoperative albumin, hemoglobin, blood urea nitrogen levels, etc.). Furthermore, the discriminatory power of these risk scoring models in the heart valve surgery population remains questionable. In addition, their implementation is not only cumbersome and complex, but also time-consuming and labor-intensive, requiring a high level of expertise from clinicians, which undoubtedly increases the clinical burden under conditions of limited medical resources. Therefore, there is a clinical need for a highly operable, inexpensive, and effective method for predicting PPCs (prolapse, prolapse, and thromboembolism) risk in the preoperative heart valve surgery population. Summary of the Invention

[0005] This disclosure is made based on the aforementioned needs of the prior art. The technical problem to be solved by this disclosure is to provide a method for predicting the probability of postoperative pulmonary complications based on preoperative sleep stage data, so as to conveniently and relatively accurately assess the probability of postoperative pulmonary complications before surgery and provide a reference for the person being predicted.

[0006] To address the aforementioned problems, the technical solutions provided in this disclosure include: This disclosure provides a method for predicting the probability of postoperative pulmonary complications based on preoperative sleep stage data, comprising: acquiring a preoperative physiological and clinical characteristic data set of the subject, wherein the preoperative physiological and clinical characteristic data set includes a preoperative physiological characteristic data set and a preoperative clinical characteristic data set; acquiring the preoperative clinical characteristic data set includes collecting continuous physiological signals during the subject's sleep stage, wherein the continuous physiological signals include: continuous single-lead electrocardiogram signals, continuous chest and abdominal respiratory signals, and continuous sleep state signals; extracting preoperative clinical characteristic data of the subject based on the continuous physiological signals, including a first characteristic data set of heart rate variability calculated based on the NN interval obtained from the continuous single-lead electrocardiogram signals, a second characteristic data set calculated based on the continuous chest and abdominal respiratory signals, and a third characteristic data set calculated based on the continuous sleep state signals; and using a prediction function based on at least one characteristic data from each of the physiological characteristic data set, the first characteristic data set, the second characteristic data set, and the third characteristic data set. The predicted probability of postoperative lung complications was obtained; among them, , , An array formed from the input feature data of the predicted subject x The probability of complications, θ The vector formed by the characteristic coefficients, It is a constant. The feature coefficients corresponding to the nth feature data are: x A vector formed from feature data. , This is the nth feature data.

[0007] This disclosure has revealed the relationship between the aforementioned characteristic data and postoperative complications of lung disease. By obtaining the necessary data and establishing a good predictive model, the probability of postoperative complications can be accurately determined to guide whether surgery should be performed and what adjustments need to be made before surgery can be performed, thus ensuring the success rate of surgery and largely avoiding risks.

[0008] Preferably, the feature data in the first feature data group are: the number of intervals greater than 50ms between two consecutive NN intervals, the average value of the entire NN interval, high-frequency energy, the ratio of low to high frequency, and arrhythmia load.

[0009] Preferably, the characteristic data in the second characteristic data group are: average minute ventilation during sleep, average respiratory rate during sleep, and average inspiratory time during sleep.

[0010] Preferably, the feature data in the third feature data group calculated based on sleep state signals are: the percentage of REM sleep duration in total sleep duration, the percentage of deep sleep duration, and the percentage of effective blood oxygenation duration.

[0011] Preferably, the characteristic data in the preoperative physiological characteristic data group are: surgical method data, age data, and preoperative pulmonary artery diameter data.

[0012] Preferably, the prediction function is determined through the following steps: constructing a first prediction model based on a feature dataset associated with postoperative lung complications, the feature dataset having an array of m feature data items; obtaining the optimal feature coefficients of the first prediction model through iteration; arranging the optimal feature coefficients of the first prediction model in order of size, and deleting the smallest feature coefficient and the feature data item corresponding to the smallest feature coefficient; repeatedly performing the steps of reconstructing, iterating, arranging, and deleting based on the remaining feature data, until the complication prediction value of the Nth prediction model decreases by a first threshold after deleting a certain feature data item; retaining the mN feature data items and the feature coefficients corresponding to the retained feature data as the final feature data and the final feature coefficients, respectively, to obtain the prediction function.

[0013] Preferably, the prediction model includes a logistic regression model, and the likelihood function of the dataset based on the logistic regression model is expressed as: Where m is the number of continuous physiological and clinical parameter arrays in the dataset. For the i-th continuous physiological and clinical parameter array, for The corresponding tags Y The value can be 0 or 1.

[0014] Preferably, the cost function obtained based on the likelihood function is expressed as: Where m is the number of continuous physiological and clinical parameter arrays in the dataset. For the i-th continuous physiological and clinical parameter array, for The corresponding tags Y The value can be 0 or 1.

[0015] Preferably, the feature coefficients are initialized using gradient descent and updated incrementally until the optimal feature coefficients for the feature data in the continuous physiological and clinical parameter array are obtained. , is represented as: ,in, For the first jThe feature coefficients corresponding to each feature data item, where α is the learning rate. Let be the cost function.

[0016] Preferably, the first threshold includes 10% or more.

[0017] Preferably, the feature data extracted from the respiratory signal is obtained by smoothing and filtering the acquired respiratory signal, removing outliers, and then detecting peaks and troughs.

[0018] Preferably, the heart rate variability feature data obtained from the NN interval is obtained by acquiring the electrocardiogram signal and detecting the position of the R wave peak.

[0019] Preferably, the method for forming the prediction function further includes: verifying the prediction model, including: dividing the dataset into 5 equal parts, namely D1, D2, D3, D4 and D5, using D1-D4 as the validation set and D5 as the validation set, calculating the performance parameters of the first fold; repeating this process 4 times to calculate the performance parameters of the second to fifth folds; and calculating the accuracy of the prediction function based on the performance parameters to determine whether the prediction function is suitable.

[0020] Compared with existing technologies, this disclosure determines the probability of postoperative lung complications based on physiological characteristics during sleep through theoretical analysis, empirical summarization, and model calculation. It can obtain highly accurate probabilities of postoperative lung complications simply by testing data during sleep, facilitating accurate assessment for those being predicted. Furthermore, this disclosure constructs a probability prediction function for postoperative lung complications based on four sets of characteristic data. Compared to existing models that rely on theoretical analysis and data verification, this function considers more comprehensive factors and avoids inaccuracies caused by external environmental or psychological factors. Additionally, the selected features have strong correlations and an appropriate amount of data, improving the model's accuracy and computational speed. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the embodiments of this specification. For those skilled in the art, other drawings can be obtained based on these drawings.

[0022] Figure 1 This is a flowchart illustrating the steps of a method for predicting postoperative complications of lung disease based on sleep stage data, as described in this disclosure. Figure 2 This is a flowchart of the method steps for forming a prediction model in an embodiment of this disclosure. Detailed Implementation

[0023] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0024] In the description of the embodiments of this disclosure, it should be noted that, unless otherwise expressly specified and limited, the term "connected" should be interpreted broadly. For example, it can refer to a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium. Those skilled in the art can understand the specific meaning of the above terms in this disclosure according to the specific circumstances.

[0025] Throughout the text, the terms “top,” “bottom,” “above,” “below,” and “on top” refer to the relative positions of components of the device, such as the relative positions of the top and bottom substrates within the device. It is understood that the device is multifunctional and independent of its spatial orientation.

[0026] To facilitate understanding of the embodiments of this application, the following will provide further explanation and description with reference to the accompanying drawings and specific embodiments. These embodiments do not constitute a limitation on the embodiments of this application.

[0027] This embodiment provides a method for predicting postoperative complications of lung disease based on sleep stage data, such as... Figures 1-2 As shown.

[0028] Obtain the preoperative physiological and clinical characteristic data set of the subject to be predicted, which includes the preoperative physiological characteristic data set and the preoperative clinical characteristic data set.

[0029] Preoperative physiological and clinical characteristics of the subjects are considered important physiological indicators reflecting their cardiopulmonary function. Since prostatic pulmonary cancer (PPCs) are closely related to cardiopulmonary function, selecting the preoperative physiological characteristics data set as a factor in assessing the probability of PPC risk can take into account the influence of the subjects' organ parameters on the probability of PPC occurrence. It should be noted that in this specific embodiment, the order of steps in obtaining the preoperative physiological characteristics data set can be flexibly set, as long as the data can be collected when needed, and is not limited to being collected in the first step.

[0030] In this specific embodiment, the following characteristic data may be included: operation, age, New York Heart Association (NYHA) classification, Euro Score, left ventricular diameter, left ventricular enlargement, left atrial diameter, left atrial enlargement, right ventricular diameter, right ventricular enlargement, right atrial diameter, right atrial enlargement, left ventricular end-diastolic volume, left ventricular end-diastolic structural change parameters, left ventricular end-systolic volume, left ventricular end-systolic structural change parameters, left ventricular ejection fraction, left ventricular systolic function change parameters, and pulmonary artery diameter.

[0031] These data are theoretically considered to have a significant impact on the occurrence of PPCs. However, existing studies have not clearly revealed which of these data are strongly correlated with PPCs and which are weakly correlated. On the one hand, theoretically, these data are all closely related to cardiopulmonary function, so the correlation of a specific data point cannot be easily denied. On the other hand, existing methods for judging PPCs are often limited to a certain perspective and cannot comprehensively evaluate the factors that affect PPCs as a whole. The correspondence between these parameters and PPCs is not entirely clear, so it is impossible to accurately assess the magnitude of the influence of these parameters on PPCs.

[0032] In this specific embodiment, preferably, the above data can represent parameters that are strongly correlated with the risk of PPCs in the physiological characteristic data group of the predicted subject. By selecting the above data as representative data of the physiological characteristic data group, the computational load of PPCs prediction can be significantly reduced, the prediction speed can be improved, and the accuracy of the prediction can be guaranteed.

[0033] This is primarily due to the comprehensiveness and rationality of the predictive model considered in this application, as well as the accuracy and scientific rigor in selecting feature data from the feature data set. This will be elaborated upon in the following description of this specific embodiment.

[0034] In addition to the preoperative physiological characteristics of the subjects, the apnea during sleep in the subjects is also considered to be strongly correlated with PPCs.

[0035] Sleep apnea syndrome (OSA) is a clinical syndrome characterized by recurrent apnea or hypoventilation during sleep, primarily caused by obstructive factors, leading to hypoxemia, hypercapnia, and sleep disruption, resulting in a series of pathophysiological changes in the body. OSA patients have a significantly increased risk of postoperative complications during the perioperative period, including respiratory failure, cardiac complications, hypoxemia, prolonged hospital stays, and more frequent ICU transfers. OSA patients have an increased risk of respiratory depression due to their higher sensitivity to sedatives and opioids. Furthermore, abnormalities in respiratory center control after major surgery can reduce the ventilatory response to hypoxia and hypercapnia, further exacerbating respiratory complications in OSA patients. Patients with severe OSA are 2 to 4 times more likely to develop complex arrhythmias than those without OSA. In addition, preoperative hypoxemia in OSA patients reflects cardiac and pulmonary function and is a strong predictor of postoperative pulmonary complications.

[0036] Physiological parameters during sleep are highly correlated with the risk of PPCs. Measuring human physiological parameters during sleep, such as cardiopulmonary function parameters of the person being assessed during sleep, can be of great reference value for judging the probability of PPCs in the person being assessed.

[0037] Therefore, in this specific embodiment, the method for predicting the probability of postoperative pulmonary complications based on sleep stage data includes: collecting continuous physiological signals of the subject's sleep stage, wherein the continuous physiological signals include: continuous single-lead electrocardiogram signals, continuous chest and abdominal respiratory signals, and continuous sleep state signals.

[0038] In this specific embodiment, the single-lead ECG signal can be detected using existing portable ECG detection devices, such as portable ECG detection devices clipped or attached to a predetermined position on the body, to detect the raw single-lead ECG (electrocardiogram) signal. The chest and abdominal respiratory signal can be detected using a strap-type respiratory sensor in the prior art. Sleep state detection can be achieved by using a portable eye-tracking sensor and a portable blood oxygenation sensor in conjunction with a respiratory sensor.

[0039] The aforementioned signal acquisition devices all acquire continuous relevant physiological signals. Preferably, these signals can be continuous signals of the subject throughout the night. Continuous signals can reflect physiological characteristics over a longer period and can reflect changes in these physiological characteristics during the detection period. Therefore, they can be used to better evaluate physiological function indicators related to PPCs.

[0040] However, it should be noted that the above signal detection is only a general detection signal and is not only for the detection of PPC risk. Therefore, in order to more accurately assess the probability of PPC occurrence, it is necessary to process these data and obtain feature data with stronger correlation to PPC based on these continuous signal processing and analysis, namely the preoperative clinical feature data of the predicted person, thereby improving the accuracy and speed of PPC prediction.

[0041] In this specific implementation, based on the summary of theoretical research and experimental data, the continuous physiological signals collected in the previous step are analyzed and processed to obtain the following three sets of preoperative clinical characteristic data of the predicted subjects: the first characteristic data set of heart rate variability calculated based on the NN interval of the continuous single-lead electrocardiogram signal, the second characteristic data set calculated based on the continuous chest and abdominal respiratory signals, and the third characteristic data set calculated based on the continuous sleep state signals.

[0042] The NN interval, as defined in the electrocardiogram, refers to the interval between normal R peaks. The heart rate variability feature dataset obtained from the NN interval may include: the number of intervals greater than 50ms between consecutive NN intervals (nni_50), the number of intervals greater than 20ms between consecutive NN intervals (nni_20), the mean of the entire NN interval (mean_nni), the median of the entire NN interval (median_nni), the standard deviation of the entire NN interval (nni_std), the root mean square of the difference between two adjacent NN intervals (rmssd), ultra-low frequency energy (vlf), low frequency energy (lf), high frequency energy (hf), the ratio of low to high frequency (lf_hf_ratio), the difference between the maximum and minimum NN intervals during sleep (min_max_nni), and the arrhythmia load (ar_burden), calculated as the percentage of intervals with a difference greater than 160ms between consecutive NN intervals out of the total number of NN intervals.

[0043] The aforementioned feature data can be calculated from the raw single-lead ECG (electrocardiogram) signal obtained in the acquisition step and the detection of the R wave peak position. Specifically, the Hamilton method is used to detect the R wave peak position on the raw ECG signal, where the dots represent the detected R wave peak positions. After R wave detection of the ECG signal throughout the night, the NN interval sequence of the entire night is obtained, and the aforementioned feature data is obtained from the RR interval sequence.

[0044] The second feature dataset calculated based on the continuous chest and abdominal respiratory signals includes: mean respiratory rate (br_mean), standard deviation of respiratory rate (br_std), mean inspiratory time (TI_mean), standard deviation of inspiratory time (TI_std), mean percentage of inspiratory time to total respiratory duration (TI_ratio_mean), and mean minute ventilation (min_ven_in_mean). After setting the NN interval signal difference to 2Hz, the power spectral density is calculated using the FFT method, where: high-frequency components (hf) range from 0.15 to 0.4Hz; low-frequency components (lf) range from 0.04 to 0.15Hz; total power (tf) ranges from 0.04 to 0.4Hz; mid-frequency components (hf) range from 0.1 to 0.15Hz; T-low-frequency components (tlf) range from 0.04 to 0.1Hz; and very low-frequency components (vlf) range from 0.0033 to 0.04Hz.

[0045] The aforementioned second feature dataset can be obtained by smoothing and filtering the raw chest and abdominal breathing signals collected throughout the night during sleep, removing outliers, and then detecting peaks and troughs.

[0046] The third feature dataset calculated based on sleep state signals includes: sleep duration (sleep_len), percentage of REM sleep time in total sleep duration (rem_per), percentage of deep sleep duration (deep_per), percentage of light sleep duration (light_per), number of awakenings during sleep (wake_times), percentage of effective blood oxygenation time (used_per), sleep apnea-hypopnea index (ahi), percentage of time with blood oxygenation below 90% during sleep (spo2_90), percentage of time with blood oxygenation below 85% during sleep (spo2_85), and lowest blood oxygenation value during sleep (spo2_min). The calculation method of the third feature data based on the sleep state signal can be implemented using existing methods for evaluating sleep results, or if an evaluation report has already been generated and includes the above data, it can be directly obtained from the existing sleep evaluation report.

[0047] In summary, the physiological characteristic dataset related to sleep contains 47 features, namely: the number of intervals between two consecutive neural interstitial (NN) intervals greater than 50ms (nni_50), the number of intervals between two consecutive NN intervals greater than 20ms (nni_20), the mean of the entire NN interval (mean_nni), the median of the entire NN interval (median_nni), the standard deviation of the entire NN interval (nni_std), the root mean square of the difference between two adjacent NN intervals (rmssd), ultra-low frequency energy (vlf), low frequency energy (lf), high frequency energy (hf), and the ratio of low to high frequency (lf_hf_ratio). The differences between the maximum and minimum NN intervals during sleep (min_max_nni), arrhythmia load (ar_burden) are calculated as the number of NN intervals with a difference greater than 160ms, representing the percentage of all NN intervals, mean respiratory rate (br_mean), standard deviation of respiratory rate (br_std), mean inspiratory time (TI_mean), standard deviation of inspiratory time (TI_std), mean of inspiratory time as a percentage of total respiratory time (TI_ratio_mean), and mean minute ventilation (min_ven_in_mean). After setting the NN interval signal difference to 2Hz, the power spectral density was calculated using the FFT method, where: high-frequency components (hf) ranged from 0.15 to 0.4Hz; low-frequency components (lf) ranged from 0.04 to 0.15Hz; total power (tf) ranged from 0.04 to 0.4Hz; mid-frequency components (hf) ranged from 0.1 to 0.15Hz; T-low frequency components (tlf) ranged from 0.04 to 0.1Hz; very low frequency components (vlf) ranged from 0.0033 to 0.04Hz; and sleep duration (sleep_len) was also considered. The percentage of REM sleep duration in total sleep time (rem_per), percentage of deep sleep duration (deep_per), percentage of light sleep duration (light_per), number of awakenings during sleep (wake_times), percentage of effective blood oxygenation duration (used_per), sleep apnea-hypopnea index (ahi), percentage of time with blood oxygenation below 90% during sleep (spo2_90), percentage of time with blood oxygenation below 85% during sleep (spo2_85), and lowest blood oxygenation value during sleep (spo2_min).

[0048] While theoretically all 47 features mentioned above may be related to PPCs, using so much data as input would impact computation speed. Furthermore, it's crucial to accurately assess which of these 47 data points are strongly correlated with PPCs and which are not strongly correlated with the probability of PPCs occurring. Therefore, in this specific implementation, to filter out the strongly correlated data from the 47 data points and improve the computation speed of the prediction model, the following method is used to form the prediction model.

[0049] A dataset is obtained, comprising an array of continuous physiological and clinical parameters associated with various postoperative complications of lung disease during sleep. The array of continuous physiological and clinical parameters includes acquired physiological feature data and clinical feature data extracted based on single-lead electrocardiogram signals, chest and abdominal respiratory signals, and blood oxygenation signals. The array of physiological and clinical parameters consists of multiple feature data, which are derived from heart rate variability feature data obtained from the NN interval, feature data extracted from respiratory signals, feature data obtained from sleep analysis results, and preoperative physiological feature data.

[0050] Furthermore, the dataset used in this embodiment includes multi-dimensional analysis of sleep patterns, electrocardiogram, respiration, and blood oxygenation throughout the entire sleep period. Features are extracted and filtered to obtain the final dataset. The dataset includes an array of continuous physiological and clinical parameters associated with various postoperative complications of lung surgery during the sleep stage. These physiological and clinical parameter arrays include multiple feature data in four categories: heart rate variability data obtained from the NN interval; feature data extracted from chest and abdominal respiratory signals; feature data extracted from sleep analysis results; and preoperative physiological feature data. Multiple features from these four categories... A prediction model is established based on the dataset, and the final feature data that forms the prediction function and the final feature coefficients corresponding to the final feature data are obtained.

[0051] The prediction model built based on the dataset includes: A first predictive model is constructed based on a feature dataset associated with postoperative complications of lung disease, the feature dataset having an array of M feature data items.

[0052] The outcome based on complications requires the prediction model to output a number between 0 and 1. If the function value is greater than 0.5, it is considered 1; otherwise, it is 0. Simultaneously, the function needs undetermined parameters. Through training with samples, these parameters can be made to accurately predict the data in the training set. Therefore, the prediction model is set as a logistic regression model, and the sigmoid function is used as the prediction function, expressed as: in, An array formed from the input feature data X The probability of complications, Θ The vector formed by the characteristic coefficients, The feature coefficients corresponding to the Nth feature data are: X A vector formed from feature data. , This is the Nth feature data.

[0053] In this specific embodiment, the dataset includes 47 feature data items, i.e., N=47, and the feature coefficients are initially assigned random values.

[0054] The likelihood function of the dataset based on the logistic regression model is expressed as: Where m is the number of continuous physiological and clinical parameter arrays in the dataset. For the i-th continuous physiological and clinical parameter array, for The corresponding tags Y The value can be 0 or 1.

[0055] The cost function obtained based on the likelihood function is expressed as follows: .

[0056] The feature coefficients are initialized using gradient descent and updated incrementally until the optimal feature coefficients for the feature data in the continuous physiological and clinical parameter array are obtained, expressed as: in, For the first j The feature coefficients corresponding to each feature data item, where α is the learning rate. Let be the cost function.

[0057] Taking the above implementation as an example, by using gradient descent, the feature coefficients that minimize the cost function are continuously iterated to approximate the optimal feature coefficients under the first prediction model.

[0058] After obtaining the feature coefficients through the above optimization process, the feature coefficients are arranged in order of magnitude, and the smallest feature coefficient and the feature data corresponding to the smallest feature coefficient are deleted. Specifically, the optimal feature coefficients under the first prediction model are arranged from largest to smallest, and the feature data corresponding to the 47th feature coefficient is removed. That is, the feature data with the smallest weight relative to postoperative pulmonary complications is removed, i.e., the feature data has the least impact on postoperative complications.

[0059] The second prediction model is reconstructed from the dataset after removing one feature data item. Specifically, the above process is repeated on the dataset with one feature data item removed, i.e., the second prediction model is reconstructed based on the dataset with 46 feature data items.

[0060] The second prediction model also uses the sigmoid function. The construction process of the second prediction model is the same as that of the first prediction model. The difference lies in that the second prediction model's dataset did not undergo the feature data deletion process following the first iteration's sorting, resulting in different optimal feature coefficients compared to the first prediction model. Subsequently, the optimal feature coefficients are sorted in order of magnitude, and feature data corresponding to the smallest coefficient are removed from the dataset. This removal process also eliminates factors with the least impact on postoperative pulmonary complications, excluding unnecessary data to simplify the acquisition process, making it more convenient and faster, while also effectively reducing the amount of data.

[0061] The process involves iteratively reconstructing the N'th prediction model based on the remaining feature data and arranging and removing feature data until the complication prediction value of the Qth prediction model decreases by a first threshold after deleting a certain feature data. The first threshold includes a value greater than or equal to 10%.

[0062] The prediction function is obtained by retaining the MQ term feature data and the feature coefficients corresponding to the retained feature data, respectively, as the final feature data and final feature coefficients.

[0063] Compared to models such as neural networks, logistic regression models can provide prediction results more quickly and intuitively while ensuring the accuracy of predicting the probability of complications.

[0064] Based on the above implementation method of this embodiment, the input of the final prediction model is 15 feature data, namely: nni_50, mean_nni, lf_hf_ratio, hf, ar_burden, min_ven_in_mean, br_mean, TI_mean, rem_per, deep_per, used_per, ahi, operation, age, and pulmonary.

[0065] The characteristic data in the heart rate variability characteristic data group obtained from the NN interval are: the number of intervals greater than 50ms between two consecutive NN intervals (nni_50), high-frequency energy (hf), low-to-high frequency ratio (lf_hf_ratio), and arrhythmia load (ar_burden).

[0066] The feature data in the second feature data group calculated based on the continuous chest and abdominal breathing signals are: the average minute ventilation (min_ven_in_mean), the average respiratory rate during sleep (br_mean), and the average respiratory rate during sleep (TI_mean).

[0067] The feature data in the third feature data group calculated based on sleep state signals are: the percentage of REM sleep duration in total sleep duration (rem_per), the percentage of deep sleep duration (deep_per), and the percentage of effective blood oxygenation duration (used_per).

[0068] The characteristic data in the physiological characteristic data group are: operation, age, and pulmonary artery diameter.

[0069] Based on the 15 final feature data items and the corresponding final feature coefficients, a prediction model is obtained. The prediction function of the prediction model is expressed as follows: in, An array formed from the input feature data x The probability of complications, θ The vector formed by the characteristic coefficients, The feature coefficients corresponding to the nth feature data are: x A vector formed from feature data. , This is the nth feature data.

[0070] In the above implementation of this embodiment, the prediction function of the prediction model is expressed as: PPC = 0.0867386 × operation + 0.0024807 × ahi + 0.0145861 × used_per - 0.0569415 × rem_per - 0.000001 × nni_50 - 0.000006 × min_ven_in_mean + 0.0218722 × deep_ per+0.0027144×mean_nni-0.0950354×br_mean-0.0469357×1f_hf_ratio+0.0127289×min_ven _in_cv+0.0181016×TI_mean+0.00003×hf-0.0601069×age+0.0383772×pulmonary+0.00250712.

[0071] After the 15 features were processed by the logistic regression model, the output was a percentage value, which represented the probability of postoperative pulmonary complications based on the Melbourne score.

[0072] To obtain a more reliable and stable model evaluation, a 5-fold cross-validation method is used to generate the training and validation sets. Specifically, this includes: The entire dataset is divided into 5 equal parts: D1, D2, D3, D4, and D5.

[0073] Take D1 to D4 as the training set and D5 as the validation set, and calculate the performance parameters of the first fold.

[0074] Repeat the cycle four times to calculate the model performance evaluation at two to five folds.

[0075] Based on the performance parameters, the accuracy of the prediction function is calculated, and the performance of the prediction model is evaluated.

[0076] Furthermore, the performance parameters obtained by the 5-fold cross-validation method are evaluated by calculating the area under the ROC curve (AUC) to assess the predictive ability of the model. In addition, the accuracy (ACC), f1 score (F1), precision, and recall evaluation metrics are also calculated to evaluate the model performance from multiple aspects.

[0077] Specifically, the following confusion matrix is ​​constructed based on the results of the prediction model output and the true label information: ROC (Relative Correction Curve) is a tool used to measure imbalance in classification. ROC curves and AUC (Area Under Curve) are commonly used to evaluate the performance of a binary classifier. The x-axis represents the False Positive Rate (FPR), and the y-axis represents the True Positive Rate (TPR). AUC is defined as the area under the ROC curve. Because the ROC curve is generally above the y=x line, its value ranges between 0.5 and 1. AUC is used as an evaluation metric because the ROC curve often doesn't clearly indicate which classifier performs better, while a higher AUC value indicates a better classifier. The x and y axes are represented as follows: The accuracy ACC is expressed as: The F1 score is expressed as: The prediction model was evaluated based on the above performance parameters, and the results are as follows: The above verification results demonstrate that the prediction model obtained in this embodiment has good classification performance, and other data illustrate the accuracy and stability of this prediction model. F1 is the harmonic mean of precision and recall, used to measure the overall performance of the classifier. Based on the above data, the prediction model in this embodiment exhibits good overall performance.

[0078] After obtaining the above 15 data points of the predicted patient before surgery, and combining them with the expression for calculating PPC, the probability of postoperative pulmonary complications can be obtained. By using the above method, the probability of postoperative complications can be predicted more accurately, which can guide whether to perform surgery and what adjustments need to be made before surgery, thus ensuring the success rate of surgery and avoiding risks to a large extent.

[0079] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of this application. It should be understood that the above description is only a specific embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A method for predicting the probability of postoperative pulmonary complications based on preoperative sleep stage data, characterized in that, include: Obtain a set of preoperative physiological and clinical characteristic data of the subject to be predicted, wherein the set of preoperative physiological and clinical characteristic data includes a set of preoperative physiological characteristic data and a set of preoperative clinical characteristic data. The characteristic data in the preoperative physiological characteristic data group are: surgical method data, age data, and preoperative pulmonary artery diameter data; The acquisition of preoperative clinical characteristic data includes collecting continuous physiological signals during the sleep stages of the subject. These continuous physiological signals include: continuous single-lead electrocardiogram (ECG) signals, continuous chest and abdominal respiratory signals, and continuous sleep state signals. The single-lead ECG signals are collected using a portable ECG monitoring device, the chest and abdominal respiratory signals are collected using a belt-type respiratory detection sensor, and the sleep state signals are collected using a portable eye movement detection sensor and a portable blood oxygen detection sensor. Based on the continuous physiological signals, the preoperative clinical characteristic data of the predicted subject are extracted, including a first characteristic data group of heart rate variability calculated based on the NN interval obtained from the continuous single-lead electrocardiogram signal, a second characteristic data group calculated based on the continuous chest and abdominal respiratory signals, and a third characteristic data group calculated based on the continuous sleep state signals. The characteristic data in the first characteristic data group are: the number of intervals between two consecutive NN intervals greater than 50ms, the average value of the entire NN interval, high-frequency energy, the ratio of low to high frequency, and arrhythmia load. The characteristic data in the second characteristic data group are: the average minute ventilation during sleep, the average respiratory rate during sleep, and the average inspiratory time during sleep. The characteristic data in the third characteristic data group are: the percentage of REM sleep duration in total sleep duration, the percentage of deep sleep duration, and the percentage of effective blood oxygenation duration. Based on at least one feature data from each of the physiological feature data group, the first feature data group, the second feature data group, and the third feature data group, a prediction function is used. The predicted probability of postoperative lung complications was obtained; among them, , , An array formed from the input feature data of the predicted subject x The probability of complications, θ The vector formed by the characteristic coefficients, It is a constant. The feature coefficients corresponding to the nth feature data are: x A vector formed from feature data. , This is the nth feature data; The prediction function is determined through the following steps: A first predictive model is constructed based on a feature dataset associated with postoperative complications of the lungs, the feature dataset having an array of M feature data items; The optimal feature coefficients of the first prediction model are obtained through iteration; The optimal feature coefficients of the first prediction model are arranged in order of size, and the smallest feature coefficient and the feature data item corresponding to the smallest feature coefficient are deleted. The process of reconstructing, iterating, arranging, and deleting the prediction model based on the remaining feature data is repeated until the complication prediction value of the Q-th prediction model decreases by a first threshold after deleting a certain feature data. The prediction function is obtained by retaining the MQ term feature data and the feature coefficients corresponding to the retained feature data, respectively, as the final feature data and final feature coefficients. The first threshold includes 10% or more.

2. The method for predicting the probability of postoperative pulmonary complications based on preoperative sleep stage data according to claim 1, characterized in that, The prediction model includes a logistic regression model, and the likelihood function of the dataset based on the logistic regression model is expressed as: Where m is the number of continuous physiological and clinical parameter arrays in the dataset. For the i-th continuous physiological and clinical parameter array, for The corresponding tags Y The value can be 0 or 1.

3. The method for predicting the probability of postoperative pulmonary complications based on preoperative sleep stage data according to claim 2, characterized in that, The cost function obtained based on the likelihood function is expressed as follows: Where m is the number of continuous physiological and clinical parameter arrays in the dataset. For the i-th continuous physiological and clinical parameter array, for The corresponding tags Y The value can be 0 or 1.

4. The method for predicting the probability of postoperative pulmonary complications based on preoperative sleep stage data according to claim 3, characterized in that, The feature coefficients are initialized using gradient descent and updated incrementally until the optimal feature coefficients for the feature data in the continuous physiological and clinical parameter array are obtained. , represented as: in, For the first j The feature coefficients corresponding to each feature data item, where α is the learning rate. Let be the cost function.

5. The method for predicting the probability of postoperative pulmonary complications based on preoperative sleep stage data according to claim 1, characterized in that, The feature data extracted from the respiratory signal is obtained by smoothing and filtering the acquired respiratory signal, removing outliers, and then detecting peaks and troughs.

6. The method for predicting the probability of postoperative pulmonary complications based on preoperative sleep stage data according to claim 1, characterized in that, The heart rate variability data obtained from the NN interval are derived from the acquired electrocardiogram signal and the detection of the R wave peak position.

7. The method for predicting the probability of postoperative pulmonary complications based on preoperative sleep stage data according to claim 1, characterized in that, The method for forming the prediction function further includes: validating the prediction model, including: dividing the dataset into 5 equal parts, namely D1, D2, D3, D4 and D5, using D1-D4 as the validation set and D5 as the validation set, calculating the performance parameters of the first fold; repeating this process 4 times to calculate the performance parameters of the second to fifth folds; and calculating the accuracy of the prediction function based on the performance parameters to determine whether the prediction function is suitable.

Citation Information

Patent Citations

  • SBS-based hierarchical feature selection method, system and application

    CN110197706A

  • Disease onset risk prediction device, method, and program

    CN111417337A

  • Logistic regression algorithm-based coal mine power distribution network power failure accident prediction method

    CN111860940A

  • Prediction method and system for senile postoperative systemic inflammatory response syndrome

    CN113707295A