A method for predicting a review time window of coronary CT imaging by fusing multi-source information

By fusing multi-source information and utilizing the LightGBM model and LIME technology, the problem of accurately predicting the time window for coronary CT imaging follow-up examinations was solved, optimizing the allocation of medical resources, reducing unnecessary examinations, and improving the medical efficiency for patients with coronary heart disease.

CN121075710BActive Publication Date: 2026-03-31FIRST AFFILIATED HOSPITAL OF DALIAN MEDICAL UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-05
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing technologies fail to effectively utilize multi-source information from imaging data and electronic medical records when assessing the follow-up time window for coronary CT imaging. This results in an inability to accurately predict the optimal follow-up time, increasing unnecessary examination and medical costs, and potentially causing missed treatment opportunities.

Method used

By collecting multidimensional patient data, including baseline CCTA images, baseline clinical information, and medication use, a multidimensional dataset was constructed. Using the LightGBM model and Stacking method, combined with the locally interpretable technique LIME, the time window for CCTA follow-up examinations was accurately predicted.

Benefits of technology

It enables accurate prediction of CCTA follow-up examination time based on individual characteristics, optimizes the allocation of medical resources, reduces unnecessary examinations, reduces the burden on patients, and provides technical support for precision medicine of coronary heart disease.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121075710B_ABST
    Figure CN121075710B_ABST
Patent Text Reader

Abstract

The application discloses a kind of coronary artery CT imaging review time window prediction methods of fusion multi-source information, comprising: collecting patient multidimensional data, based on multidimensional data confirmation imaging index, based on multidimensional data collation baseline clinical information and collation patient's drug use, secondly induction coronary heart disease patient's CCTA review time, construct multidimensional data set, data set is preprocessed, risk factor is assigned quantization;Through the data constructed after processing LightGBM submodel, the accuracy and generalization ability of final prediction result are improved using Stacking method, the loss function is reduced by introducing model-independent local interpretable technology LIME;Through the above method, the current evaluation method only includes traditional factors such as hypertension and diabetes, ignores the imaging information and rich clinical information in electronic medical record which can be directly quantified and represent individualized plaque characteristics and the interaction between multi-source information, so that it is difficult to accurately predict the best CCTA review time window.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of follow-up prediction technology for patients with coronary heart disease, specifically to a method for predicting the time window of coronary artery CT imaging follow-up examination by integrating multi-source information. Background Technology

[0002] The incidence of coronary artery disease (CAD) continues to rise in my country and globally, and it ranks first in mortality among cardiovascular diseases. Numerous studies have shown that the progression of coronary plaques is a key intermediate step and independent risk factor in the transformation of lesions into cardiac events. Extensive data show that the incidence of cardiac events in patients with plaque progression is as high as 14.3%, significantly higher than the 0.3% in those without progression. This significant difference suggests that plaque progression can serve as a prospective surrogate endpoint for the prevention and treatment of cardiac events, shifting the focus of cardiac event prevention earlier (YUMM, TANGXL, ZHAOX, et al. Plaque progression at coronary artery T angiography links non-alcoholic acid).

[0003] attyliverdiseaseandcardiovascularevents:aprospectivesingle-centerstudy[J].Europeanradiology, 2022,32

[0004] (12):811121.;VANDRIESTFY,BIJNSCM,VANDERGEESTRJ,etal.Utilizing (serial) coronary computed tomography angiography (CCTA) to predict plaque progression and major radverse cardiovascular events (MACE): results, merits and challenges[J]. European radiology,2022,32(5):3408-22.). Therefore, timely detection of plaque progression is helpful for risk stratification in patients with coronary heart disease, and can promote active and effective intervention treatment, which has important clinical value for preventing and delaying the occurrence and development of cardiovascular events. Coronary CT angiography can not only non-invasively and intuitively present information such as the size, location, shape and homogeneity of plaques in the whole heart, but also shows good consistency and reproducibility with invasive intravascular ultrasound (IVUS) in terms of quantifying and monitoring the dynamic evolution of plaques (VANDRIESTFY, BIJNSCM, VANDERGEESTRJ, et al. Utilizing (serial) coronary computed tomography angiography (CCTA) to predict plaque progression and major radiographic cardiac cevents (MACE): results, merits and challenges [J]. European radiology, 2022, 32(5): 3408-22).

[0005] CCTA has been widely used in clinical practice to screen for coronary artery disease and monitor plaque evolution, thus playing an important role in evaluating drug efficacy, optimizing subsequent clinical management, and preventing cardiovascular events (GUH, LUB, GAOY, et al. Prognostic Value of Atherosclerosis Progression for Prediction of Cardiovascular Events in Patients with Nonobstructive Coronary Artery Disease [J]. Academic Radiology, 2021, 28(7): 980-7.; CHENQ, PANT, WANGYN, et al. A Coronary CT Angiography Radiomics Model to Identify Vulnerable Plaque and Predict Cardiovascular Events [J]. Radiology, 2023, 307(2).).

[0006] However, current guidelines do not offer recommendations for the optimal time window for CCTA follow-up examinations for individuals, and there is a lack of effective means in clinical practice to accurately determine the timing of follow-up examinations. Currently, reliance on physicians' subjective experience or limited traditional risk factor assessments often leads to a large proportion of patients receiving excessive or insufficient follow-up examinations. This practice not only increases unnecessary radiation exposure, contrast agent use, and medical costs, but may also miss the optimal time to adjust treatment plans, thereby increasing the risk of cardiovascular events. The reason for this situation is that the factors influencing plaque progression are complex and diverse (WILLIAMS MC. Predictors of Plaque Progression on Coronary Computed Tomography Angiography[J]. Jacc-Cardiovascular Imaging, 2023, 16(4): 505-7; NURMOHAMEDNS, GAILLARDEL, MALKASIANS, et al. Lipoprotein(a) and Long-Term Plaque Progression, Low-Density Plaque, and Pericoronary Inflammation[J]. Jama Cardiology, 2024, 9(9): 826-34.). Current assessment methods only include traditional factors such as hypertension and diabetes, but ignore imaging information and rich clinical information in electronic medical records that can be directly quantified and characterized for individualized plaque features, as well as the interaction between these multi-source information, which makes it very difficult to accurately predict the optimal CCTA follow-up time window. Summary of the Invention

[0007] The purpose of this invention is to address the problem that current assessment methods only incorporate traditional factors such as hypertension and diabetes, while neglecting imaging information that can directly quantify and characterize individualized plaque features, as well as the rich clinical information in electronic medical records, and the interaction between these multi-source information sources. This makes it very difficult to accurately predict the optimal time window for CCTA follow-up examinations.

[0008] To address the aforementioned issues, this invention provides a method for predicting the time window of coronary artery CT imaging follow-up examinations by integrating multi-source information, comprising: Step S1: collecting multi-dimensional patient data; Step S2: determining imaging indicators closely related to coronary plaque progression based on baseline CCTA images; Step S3: organizing baseline clinical information; Step S4: organizing patient medication usage; Step S5: summarizing the CCTA follow-up examination time for patients with coronary heart disease; Step S6: constructing a multi-dimensional dataset based on baseline CCTA images, baseline clinical information, and medication usage; Step S7: dataset preprocessing; Step S8: assigning and quantifying risk factors; Step S9: constructing a LightGBM sub-model; Step S10: improving the accuracy and generalization ability of the final prediction results; Step S11: constructing an interpretable module.

[0009] In the preferred method, the process includes: Step S1: Collecting multidimensional patient data, including: baseline CCTA images, baseline clinical information, medication use, and follow-up CCTA images; Step S2: Based on the baseline CCTA images, determining imaging indicators closely related to coronary plaque progression; Qualitative indicators include: plaque nature, high-risk signs, and degree of coronary artery stenosis; Plaque nature is divided into three types: calcified plaques, non-calcified plaques, and mixed plaques; High-risk signs include: positive remodeling, punctate calcification, low-density plaques, and napkin ring sign; Based on the degree of stenosis, coronary artery stenosis is divided into three levels: stenosis 0-30%, stenosis 30%-50%, and stenosis greater than 50%; Quantitative indicators include: plaque length, minimum luminal area, positive remodeling index, and lesion involvement range; The segment involvement score (SIS score) is used. The SIS score assesses the extent of coronary artery disease. Based on the coronary segmentation method established by the AHA, the coronary arteries are divided into the proximal, mid, and distal segments of the right coronary artery; the proximal, mid, and distal segments of the left ventricular posterior branch, left main coronary artery, and left anterior descending artery; the first and second diagonal branches; the proximal and distal segments of the circumflex artery; the first and second obtuse marginal branches; the posterior branch of the circumflex artery; and the left ventricular posterior branch and posterior descending artery. One point is awarded for each diseased coronary artery segment, with a score range of 0-16. Step S3: Compile baseline clinical information; Qualitative indicators include: gender, lifestyle, smoking status, hypertension, hyperlipidemia, diabetes, cardiovascular or peripheral artery disease, and family history of cardiovascular disease; Lifestyle includes: whether the patient stays up late; Other symptoms include: chest pain, chest tightness, and palpitations; Quantitative indicators include: age, body mass index, creatinine, blood glucose, glycated hemoglobin, total cholesterol, high / low density lipoprotein cholesterol, triglycerides, lipoprotein a, C-reactive protein, interleukin-6, and tumor necrosis factor-α. Step S4: Compile patient medication information; including whether the patient is taking aspirin or not, whether the patient is taking lipid-lowering drugs (statins) or not, whether LDL cholesterol is <1.8 or reduced by half, whether hypertension is controlled, and whether diabetes is controlled; Step S5: Summarize the CCTA follow-up time for coronary artery disease patients, categorized into four types: within 2 years, 2 to 4 years, 4 to 6 years, and more than 6 years; based on the time difference between follow-up CCTA and baseline CCTA, categorized into four types: within 2 years, 2 to 4 years, 4 to 6 years, and 6 years; Step S6: Construct a multi-dimensional dataset based on baseline CCTA images, baseline clinical information, and medication information;

[0010] include: ,in It represents the characteristics of the baseline CCTA image, with a total of 14 dimensions, including 11 dimensions of qualitative indicators and 3 dimensions of quantitative indicators; It represents baseline clinical information, totaling 21 dimensions, including 9 dimensions of qualitative indicators and 12 dimensions of quantitative indicators; The data represents the medication use and control status, comprising 19 dimensions, all of which are qualitative indicators; the date of each patient's next CCTA follow-up examination is used as a label. This represents four different types of review time categories; Step S7: Dataset preprocessing; For qualitative indicators, if missing values ​​exist, directly delete samples containing missing values; For quantitative indicators, if missing values ​​exist, use the mean of the indicator to fill them; Step S8: Assign and quantify risk factors; For qualitative indicators, use one-hot encoding according to the specific type and characteristics. Quantization is performed using both encoding and label encoding. Plaque characteristics are subdivided into three types: calcified plaques, non-calcified plaques, and mixed plaques. One-hot encoding converts these three types into three independent binary features, each corresponding to a plaque type. A feature value of 1 is assigned when a plaque type exists, and 0 otherwise. For binary qualitative indicators (gender: male / female, whether staying up late: yes / no, smoking status: smoking / non-smoking, hypertension: yes / no), label encoding is used for quantization. Label encoding maps each category to a unique integer value of 0 or 1, where 0 represents no / female and 1 represents yes / male. Z-score standardization is used to subtract the mean from each data point and divide by the standard deviation, transforming the data points into a distribution with a mean of 0 and a standard deviation of 1. Step S9: Constructing LightGBM sub-models; training three LightGBM sub-models based on baseline CCTA images, baseline clinical information, and medication usage; LightGBM sub-model based on baseline CCTA images. The formula is: ,in It is the output class probability distribution, that is: , in Let represent the probability that the sample belongs to class k, respectively. These are the parameters of the image feature classification model, namely: ,in This is the learning rate, with a default value of 0.1 and a range of [0.01, 0.1]. This is the depth of the tree. The default value is -1, which means there is no limit to the depth. The value range is set to an integer between [3, 8]. This is the number of leaf nodes, with a default value of 31 and a range of values ​​set to [value range missing]. An integer between; This is the number of trees, with a default value of 100 and a range of integers between [100, 1000]. and This is a regularization parameter, with a default setting of 0 and a value range of [0,10]; it is used to construct a system based on baseline clinical information. Sub-model, formula: , in These are the parameters of the clinical information classification model, including their default values ​​and ranges. Maintain consistency. It is the output class probability distribution, that is: , in Let represent the probability that the sample belongs to category k; construct a system based on drug use. Sub-model, formula: ,in These are the parameters of the drug use classification model, including their default values ​​and ranges. Maintain consistency. It is the output class probability distribution, that is: ,in These represent the probabilities that the sample belongs to class k, respectively; Step S10: Improve the accuracy and generalization ability of the final prediction result; Use the Stacking method to take the class probability distributions output by the first-layer image feature model, clinical information model, and drug usage model as new input features and input them into the second-layer classification model; Concatenate the class probabilities output by the three sub-models into a new input z, as shown in the formula:

[0011]

[0012] Using a new LightGBM model as the second layer, we train it on z to obtain:

[0013]

[0014] in, These are the training parameters for the second layer of the LightGBM model, including their default values ​​and ranges. Maintain consistency. It is the final predicted class probability distribution, i.e., the model. The probability estimate of a given input sample belonging to each category is expressed as:

[0015] ,in Let and represent the probabilities that a sample belongs to class k, respectively. The final class label is selected according to the principle of maximum probability, that is, the class with the highest probability is chosen as the model's prediction result. The formula is:

[0016]

[0017] Step S11: Construct an interpretable module; introduce the model-independent locally interpretable technique LIME, and define the objective function of the interpretable module as:

[0018]

[0019] Further expressed as:

[0020]

[0021] in, This represents a simple model, the linear regression model; It is a collection of simple models, that is, all linear models; Represents the new dataset Compared with the original dataset Distance weights; Representation Model The degree of complexity, choose It is a linear regression model. A function, as a metric, describes how... In the local definition, the simple model Approximating complex models When explaining the complexity of the model Minimize when it is low enough to be understood by humans. The function obtains the optimal solution of the objective function.

[0022] In the preferred method, in step S5, each type is divided into a progressive group and a non-progressive group based on follow-up CCTA; the following three conditions are all judged as progressive: an increase of ≥20% in the degree of coronary artery stenosis, an increase of ≥2 high-risk signs, and an increase of ≥2 SIS scores; otherwise, it is considered non-progressive.

[0023] In the preferred approach, baseline CCTA images include: total cardiovascular wall thickness, degree of stenosis and dilation of the vascular lumen, and various morphological and nature characteristics of the lesions; baseline clinical information includes: patient's age, sex, family history, lifestyle, medical history, symptoms, signs, and laboratory test indicators; medication use includes: antiplatelet drugs, lipid-lowering drugs, antihypertensive drugs, and hypoglycemic drugs currently used by the patient, and the control status of blood glucose, blood lipids, and blood pressure; follow-up CCTA images are used to interpret the evolution of the lesions. The following three conditions are considered as progression of coronary artery disease: an increase in coronary artery stenosis ≥20%, an increase of ≥2 high-risk signs, and an increase of ≥2 SIS scores; otherwise, it is considered as no progression.

[0024] The beneficial effects of this invention are as follows: it can accurately predict the optimal time window for CCTA follow-up examination based on the individual characteristics of patients, thereby optimizing the allocation of medical resources, avoiding unnecessary examinations, reducing the medical burden on patients, providing strong technical support for precision medicine of coronary heart disease, and having broad clinical application prospects. Attached Figure Description

[0025] Figure 1This is a design concept diagram of the present invention;

[0026] Figure 2 This is a schematic diagram of the structure of the present invention. Detailed Implementation

[0027] Example 1:

[0028] A method for predicting the follow-up time window of coronary CT imaging by fusing multi-source information, comprising:

[0029] S1: Collect multi-dimensional patient data, including:

[0030] Baseline CCTA images, baseline clinical information, medication use and dynamic control of risk factors, and follow-up CCTA images;

[0031] Baseline CCTA images can clearly show the anatomical structure of the coronary arteries, including: total cardiovascular wall thickness, degree of stenosis and dilation of the lumen, and various morphological and nature characteristics of lesions;

[0032] Baseline clinical information includes: the patient's age, sex, family history, lifestyle, medical history, symptoms, signs and laboratory test results;

[0033] Medication use and risk factor control include: the patient's current antiplatelet drugs, lipid-lowering drugs, antihypertensive drugs, and hypoglycemic drugs; and the control of blood glucose, blood lipids, and blood pressure. Follow-up CCTA images are used to interpret the evolution of lesions. The following three conditions are considered as progression of coronary artery disease: an increase in coronary artery stenosis ≥20%, an increase of ≥2 high-risk signs, and an increase of ≥2 SIS; otherwise, it is considered as no progression.

[0034] S2: Based on baseline CCTA images, identify imaging indicators closely related to coronary plaque progression, as shown in Table 1;

[0035] Qualitative indicators include: plaque characteristics, high-risk signs, and degree of coronary artery stenosis;

[0036] Plaque properties are classified into three types: calcified plaques, non-calcified plaques, and mixed plaques;

[0037] High-risk signs refer to features on CCTA images that indicate a higher risk of coronary plaque progression, including:

[0038] There are four types of coronary artery stenosis: positive remodeling, punctate calcification, low-density plaque, and napkin sign. Based on the degree of stenosis, coronary artery stenosis is classified into:

[0039] The degree of stenosis is classified into three levels: 0-30%, 30%-50%, and greater than 50%.

[0040] Quantitative indicators include: plaque length, minimum luminal area, positive remodeling index, and extent of lesion involvement;

[0041] Plaque length is recorded in millimeters as the longitudinal length of the plaque. The longer the plaque, the wider the area of ​​coronary artery involvement and the greater the risk of progression.

[0042] The minimum lumen area directly reflects the impact of vascular stenosis on blood flow channels; the smaller the minimum lumen area, the more severe the blood flow obstruction.

[0043] The positive remodeling index assesses the remodeling of the vascular wall by calculating the ratio of the area of ​​the external elastic membrane to the area of ​​the vascular lumen. An elevated positive remodeling index usually indicates abnormal proliferation of the vascular wall, which may be related to the progression of coronary artery plaques.

[0044] The segment involvement score (SIS score) is used to evaluate the extent of coronary artery lesions. Based on the coronary artery segmentation method established by the American Heart Association (AHA), the coronary arteries are divided into the proximal, mid, and distal segments of the right coronary artery; the proximal, mid, and distal segments of the left ventricular posterior branch, left main coronary artery, and left anterior descending artery; the first and second diagonal branches; the proximal and distal segments of the circumflex artery; the first and second obtuse marginal branches; the posterior branch of the circumflex artery; and the left ventricular posterior branch and the left anterior descending artery. One point is awarded for each diseased coronary artery segment, with a score range of 0-16.

[0045] Table 1. Classification of Baseline Image Feature Indicators

[0046]

[0047] S3: Compile baseline clinical information, as shown in Table 2;

[0048] Qualitative indicators include:

[0049] Gender, lifestyle, smoking status, hypertension, hyperlipidemia, diabetes, cardiovascular or peripheral artery disease, family history of cardiovascular disease, other symptoms;

[0050] Multiple studies have confirmed that gender is closely related to the risk of developing coronary heart disease; male patients tend to have an earlier onset of the disease and experience faster disease progression.

[0051] Lifestyle includes:

[0052] Do you stay up late?

[0053] Long-term sleep deprivation may disrupt the biological clock, affect the normal function of the cardiovascular system, and increase the risk of coronary heart disease;

[0054] Smoking has been proven to be an independent risk factor for coronary heart disease. It can cause vasoconstriction, platelet aggregation, and abnormal lipid metabolism, thereby accelerating the development of atherosclerosis.

[0055] The patient's history of chronic diseases such as hypertension, hyperlipidemia, and diabetes was clearly recorded. The presence of cardiovascular or peripheral artery disease suggests that the patient may have systemic atherosclerosis.

[0056] A family history of cardiovascular disease reflects a patient's genetic susceptibility; patients with a family history of cardiovascular disease have a relatively higher risk of developing coronary heart disease.

[0057] Other symptoms include:

[0058] Chest pain, chest tightness, and palpitations may also be related to the activity of coronary heart disease.

[0059] Quantitative indicators include: age, body mass index, creatinine, blood glucose, glycated hemoglobin, total cholesterol, high / low density lipoprotein cholesterol, triglycerides, lipoprotein a, C-reactive protein, interleukin-6 (IL-6), and tumor necrosis factor-α (TNF-α).

[0060] Age is one of the important risk factors for coronary heart disease, and the incidence of coronary heart disease increases with age.

[0061] Body Mass Index (BMI) reflects a patient's weight status, and obesity is an independent risk factor for coronary heart disease.

[0062] A higher BMI may increase the risk of developing coronary heart disease;

[0063] Creatinine levels are an important indicator for assessing kidney function;

[0064] Renal insufficiency may affect drug metabolism and cardiovascular stability;

[0065] Blood glucose and glycated hemoglobin levels reflect a patient's blood glucose control.

[0066] Blood lipid indicators such as total cholesterol, high / low density lipoprotein cholesterol, triglycerides, and lipoprotein a are closely related to the formation of atherosclerosis;

[0067] Abnormal blood lipid levels, especially low-density lipoprotein, are an important risk factor for coronary heart disease;

[0068] The levels of inflammatory factors such as C-reactive protein, IL-6, and TNF-α can reflect the body's inflammatory state. Elevated levels of these inflammatory factors may indicate the activity of coronary heart disease.

[0069] Table 2 Baseline Clinical Information

[0070]

[0071] S4: Compile the patient's medication usage information, as shown in Table 3;

[0072] This includes whether the patient is taking aspirin, whether they are taking lipid-lowering drugs such as statins or others, whether their LDL cholesterol is <1.8 or reduced by half, whether their hypertension is under control, and whether their diabetes is under control.

[0073] Aspirin, as an antiplatelet drug, can effectively prevent platelet aggregation and reduce the risk of thrombosis. It is one of the basic medications for the treatment of coronary heart disease.

[0074] Lipid-lowering therapy is of great significance for controlling blood lipid levels and stabilizing plaques. Statins are currently the most commonly used lipid-lowering drugs, and cholesterol absorption inhibitors and PSK inhibitors are also beginning to be used in clinical practice.

[0075] The dynamic changes in low-density lipoprotein cholesterol (LDL-C) are closely related to the progression of atherosclerosis. Controlling LDL-C to <1.8 mmol / L or reducing it by half is an important goal of lipid management in patients with coronary heart disease.

[0076] Good blood pressure control can reduce the burden on the heart and lower the risk of coronary heart disease;

[0077] Good blood sugar control is also crucial for preventing complications of coronary heart disease in diabetic patients;

[0078] Table 3. Patient Medication Usage Table

[0079]

[0080] S5: Summarize the CCTA follow-up time of coronary artery disease patients, dividing them into four categories: within 2 years, 2 to 4 years, 4 to 6 years, and more than 6 years; based on the time difference between the actual follow-up CCTA and the baseline CCTA, divide them into four categories: within 2 years, 2 to 4 years, 4 to 6 years, and more than 6 years; within each category, divide them into a progression group and a non-progression group based on the follow-up CCTA. The following three conditions are considered as progression: an increase in coronary artery stenosis ≥20%, an increase of ≥2 high-risk signs, and an increase of ≥2 SIS; otherwise, it is considered non-progression; Step S6: Construct a multi-dimensional dataset based on baseline CCTA images, baseline clinical information, and medication use. ;

[0081] include: , Representing multidimensional datasets A multi-dimensional dataset with 54 dimensions. Includes:

[0082] , Indicates the characteristics of the baseline CCTA image. The baseline CCTA image features are represented in 14 dimensions, including 11 qualitative indicators and 3 quantitative indicators.

[0083] , Indicates baseline clinical information, The baseline clinical information consists of 21 dimensions, including 9 dimensions of qualitative indicators and 12 dimensions of quantitative indicators.

[0084] , This indicates the status of medication use and control. The data represents 19 dimensions of drug use and control, all of which are qualitative indicators.

[0085] The date of each patient's next CCTA follow-up examination is used as a label. , The follow-up examination is scheduled for within 2 years. The follow-up examination period is indicated to be 2 to 4 years. The follow-up examination is scheduled for 4 to 6 years. This indicates that the follow-up examination period is more than 6 years.

[0086] Step S7: Dataset preprocessing;

[0087] For qualitative indicators, if there are missing values, samples containing missing values ​​are directly deleted.

[0088] For quantitative indicators, if there are missing values, the mean of the indicator will be used to fill them.

[0089] Step S8: Assign and quantify the risk factors;

[0090] For qualitative indicators, one-hot encoding and label encoding are used for quantification, depending on the specific type and characteristics.

[0091] The plaque property index is subdivided into three types: calcified plaques, non-calcified plaques, and mixed plaques. Through one-hot encoding, the three types of plaque property index are converted into three independent binary features. Each binary feature corresponds to a plaque type. When a certain plaque type exists, the corresponding feature value is 1, otherwise it is 0.

[0092] For the binary qualitative indicators, gender (male, female), whether or not to stay up late (yes, no), smoking status (smoking, non-smoking), and high blood pressure (yes, no), a label coding method is used for quantification. The label coding maps each category to a unique integer value 0 and 1, where 0 represents no, female, and 1 represents yes, male.

[0093] Z-score standardization is used to subtract the mean from each data point and divide by the standard deviation, transforming the data points into a distribution with a mean of 0 and a standard deviation of 1.

[0094] Step S9: Construct the LightGBM sub-model;

[0095] Three LightGBM sub-models were trained based on baseline CCTA images, baseline clinical information, and drug use.

[0096] LightGBM sub-model based on baseline CCTA imagery The formula is:

[0097]

[0098] in, It is the output class probability distribution, that is:

[0099]

[0100] in, Let represent the probability that the sample belongs to class k, respectively. These are the parameters of the image feature classification model, namely:

[0101]

[0102] in, This is the learning rate, with a default value of 0.1 and a range of [0.01, 0.1].

[0103] This is the depth of the tree. The default value is -1, which means there is no limit to the depth. The value range is set to an integer between [3, 8].

[0104] This is the number of leaf nodes, with a default value of 31 and a range of values ​​set to [value range missing]. An integer between; This is the number of trees, with a default value of 100 and a range of integers between [100, 1000]. and This is a regularization parameter, which is set to 0 by default and has a value range of [0, 10].

[0105] Building based on baseline clinical information Sub-model, formula:

[0106]

[0107] in, These are the parameters of the clinical information classification model, including their default values ​​and ranges. Maintain consistency. It is the output class probability distribution, that is:

[0108]

[0109] in, These represent the probabilities that the sample belongs to category k;

[0110] Build a drug use-based Sub-model, formula:

[0111]

[0112] in, These are the parameters of the drug use classification model, including their default values ​​and ranges. Maintain consistency. It is the output class probability distribution, that is:

[0113]

[0114] in, These represent the probabilities that the sample belongs to category k;

[0115] Step S10: Improve the accuracy and generalization ability of the final prediction results;

[0116] The Stacking method is used to input the class probability distributions output by the first-layer image feature model, clinical information model, and drug usage model as new input features into the second-layer classification model; the class probabilities output by the three sub-models are concatenated into a new input z, as shown in the formula:

[0117]

[0118] Using a new LightGBM model as the second layer, we train it on z to obtain:

[0119]

[0120] in, These are the training parameters for the second layer of the LightGBM model, including their default values ​​and ranges. Maintain consistency. It is the final predicted class probability distribution, i.e., the model. The probability estimate of a given input sample belonging to each category is expressed as:

[0121]

[0122] in, These represent the probabilities that the sample belongs to class k;

[0123] The final category label is selected based on the principle of maximizing probability, that is, the category with the highest probability is chosen as the model's prediction result. The formula is:

[0124]

[0125] Step S11: Construct an interpretable module;

[0126] Introducing the model-independent locally interpretable technique LIME, the objective function of the interpretable module is defined as:

[0127]

[0128] Further expressed as:

[0129]

[0130] in, This represents a simple model, the linear regression model;

[0131] It is a collection of simple models, that is, all linear models;

[0132] Represents the new dataset Compared with the original dataset Distance weights; Representation Model The degree of complexity, choose It is a linear regression model. A function, as a metric, describes how... In the local definition, the simple model Approximating complex models When explaining the complexity of the model Minimize when it is low enough to be understood by humans. The function obtains the optimal solution of the objective function;

[0133] This optimal solution is known as a locally interpretable model, which clearly demonstrates which features significantly influence the prediction results and how these features affect them. This local interpretability not only enhances the model's transparency but also provides valuable reference information for clinicians, enabling them to better understand the model's decision-making rationale and thus apply the model's predictions with greater confidence in clinical practice.

[0134] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the present invention. Various changes and modifications can be made to the present invention without departing from its spirit and scope. All such changes and modifications fall within the scope of the present invention as claimed, which is defined by the appended claims and their equivalents.

Claims

1. A method for predicting a review time window for coronary CT angiography using fusion of multi-source information, characterized in that, Comprise: Step S1: Collecting patient multi-dimensional data; Step S2: According to baseline CCTA image, determine the imaging index closely related to coronary plaque progression; Step S3: Organize baseline clinical information; Step S4: Organize the drug use of patients; Step S5: Summarize the CCTA review time of coronary heart disease patients; Divided into: 2 years, 2 to 4 years, 4 to 6 years, more than 6 years four types; According to the time difference between follow-up CCTA and baseline CCTA, divided into 2 years, 2 to 4 years, 4 to 6 years, 6 years four types; Step S6: According to baseline CCTA image, baseline clinical information, drug use, construct multi-dimensional data set; Step S7: Data set preprocessing; Step S8: Quantitative assignment of risk factors; Step S9: Constructing LightGBM sub-models; training three LightGBM sub-models based on baseline CCTA images, baseline clinical information and drug use; LightGBM sub-model based on baseline CCTA images The formula is: wherein, is the output class probability distribution, i.e.: wherein, respectively represent the probability that the sample belongs to class k, is the parameter of the image feature classification model, i.e.: wherein, is a learning rate, with a default value of 0.1 and a range set between [0.01, 0.1]; is the depth of the tree, the default value is -1, representing no limit to the depth, and the value range is set to an integer between [3, 8]; is the number of leaf nodes, the default value is 31, and the value range is set to an integer between ; is the number of trees, the default value is 100, and the value range is set to an integer between [100, 1000]; and are regularization parameters, the default settings are both 0, and the value range is [0, 10]; Constructing sub-models based on baseline clinical information sub-models, formula: wherein, is a parameter of the clinical information classification model, the default value, value range of the parameter are consistent with , is the output category probability distribution, that is: wherein, respectively represent the probability that the sample belongs to class k. Constructing sub-models based on drug usage The formula is: wherein, are parameters of the drug use case classification model, the default values, value ranges of the parameters are consistent with , is the output category probability distribution, that is: wherein, respectively represent the probability that the sample belongs to class k. Step S10: improve the accuracy and generalization ability of the final prediction result; concatenate the class probabilities output by the three sub-models into a new input , train a new LightGBM model as the model of the second layer, and obtain the final prediction class probability distribution . According to the maximum probability principle, select the final class label, that is, select the class with the highest probability as the prediction result of the model; S11: construct an interpretable module.

2. The method of claim 1, wherein the method further comprises: Comprise: Step S1: Collecting patient multi-dimensional data, including: baseline CCTA image, baseline clinical information, drug use, follow-up CCTA image; Step S2: According to baseline CCTA image, determine the imaging index closely related to coronary plaque progression; Qualitative indicators include: plaque properties, high-risk signs, and coronary artery stenosis degree; Plaque properties are divided into: calcified plaque, non-calcified plaque and mixed plaque three types; High-risk signs include: positive remodeling, punctate calcification, low-density plaque and napkin ring sign four; According to the different degree of stenosis, the coronary artery stenosis is divided into: stenosis degree 0-30%, stenosis degree 30%-50% and stenosis degree greater than 50% three levels; Quantitative indicators include: plaque length, minimum lumen area, positive remodeling index, lesion involvement range; Use lesion segment score, SIS score, SIS score evaluates the involvement of coronary artery lesions, based on the coronary segment division method formulated by AHA, the coronary artery is divided into right coronary artery proximal segment, middle segment, distal segment, left posterior branch, left main stem, anterior descending branch proximal segment, middle segment, distal segment, first and second diagonal branch, circumflex artery proximal segment, distal segment, first and second obtuse margin branch, circumflex artery left posterior branch and posterior descending branch, 1 point for 1 lesion coronary artery segment, the score range is 0-16 points; Step S3: Organize baseline clinical information; Qualitative indicators include: gender, lifestyle, whether smoking, hypertension, hyperlipidemia, diabetes, cardiovascular or peripheral arterial disease, family history of cardiovascular disease; Lifestyle includes: whether to stay up late; Other symptoms include: chest pain, chest tightness, palpitation; Quantitative indicators include: age, body mass index, creatinine, blood glucose, glycosylated hemoglobin, total cholesterol, high / low density lipoprotein cholesterol, triglyceride, lipoprotein a, C-reactive protein, interleukin-6, tumor necrosis factor-α; Step S4: Organize the drug use of patients; Including patients taking aspirin and not taking aspirin, taking lipid-lowering agents (statins) and not taking lipid-lowering agents (statins), whether low-density lipoprotein is <1.8 or reduced by half, whether hypertension is controlled, whether diabetes is controlled; Step S6: According to baseline CCTA image, baseline clinical information, drug use, construct multi-dimensional data set; Comprises: wherein denotes baseline CCTA image features, 14 dimensions in total, including 11 qualitative indicators and 3 quantitative indicators; Baseline clinical information, 21 dimensions, including 9 qualitative indicators and 12 quantitative indicators; representing drug use and control situation, 19 dimensions, all are qualitative indicators; Next CCTA review time for each patient as label , indicating four different types of review time categories; Step S7: Data set pre-processing; For qualitative indicators, there are missing values, directly delete samples containing missing values; For quantitative indicators, there are missing values, fill in the mean of the indicator; Step S8: Assign values to risk factors; For qualitative indicators, according to the specific type and characteristics, use One-Hot Encoding and Label Encoding to quantize the data; The plaque property index is divided into calcified plaque, non-calcified plaque and mixed plaque, and through One-Hot Encoding, the three types of plaque property index are converted into three independent binary features, each binary feature corresponds to a plaque type, and the feature value is 1 when a plaque type exists, otherwise 0; For binary qualitative indicators, gender: male, female, whether to stay up late: yes, no, smoking status: smoking, non-smoking, hypertension: yes, no, use label encoding for quantization, label encoding maps each category to a unique integer value 0 and 1, 0 represents no, female, and 1 represents yes, male; Use Z-score standardization to subtract the mean and divide by the standard deviation of each data point, and convert the data points to a distribution with a mean of 0 and a standard deviation of 1; Step S10: Improve the accuracy and generalization ability of the final prediction result; Use the Stacking method to input the class probability distribution output by the first layer of image feature model, clinical information model and drug use model into the second layer of classification model; The class probability output by the three sub-models is spliced into a new input z, and the formula is: Use a new LightGBM model as the second layer model to train z, and get: wherein, are the training parameters of the second layer LightGBM model, the default values, value ranges of the parameters are consistent with , is the final predicted class probability distribution, i.e., the model For the probability estimate of the given input sample belonging to each class, it is expressed as: wherein, respectively represent the probability that the sample belongs to class k; According to the maximum probability principle, select the final class label, that is, select the class with the highest probability as the prediction result of the model, and the formula is: Step S11: Build an interpretable module; Introduce a model-independent local interpretable technique LIME, and the objective function of the interpretable module is defined as: Further represented as: in, This represents a simple model, the linear regression model; It is a collection of simple models, that is, all linear models; Represents the new dataset Compared with the original dataset Distance weights; Representation Model The degree of complexity, choose It is a linear regression model. A function, as a metric, describes how... In the local definition, the simple model Approximating complex models When explaining the complexity of the model Minimize when it is low enough to be understood by humans. The function obtains the optimal solution of the objective function.

3. The method of claim 2, wherein the method further comprises: In step S5, each type is divided into progression group and non-progression group according to follow-up CCTA; The following three conditions are judged as progression: coronary artery stenosis increases by ≥20%, high-risk signs increase by ≥2, and SIS score increases by ≥2; Otherwise, it is not progression.

4. The method of claim 2, wherein the method further comprises: Baseline CCTA images include: whole heart blood vessel wall thickness, blood vessel lumen stenosis and dilation degree, various morphological and property conditions of lesions; Baseline clinical information includes: patient's age, gender, family history, lifestyle, medical history, symptoms, signs and laboratory examination indicators; Drug use includes: antiplatelet drugs, lipid-lowering drugs, antihypertensive drugs, hypoglycemic drugs being used by patients, blood glucose, blood lipid and blood pressure control; Follow-up CCTA images are used to judge the evolution of lesions, and the following three conditions are judged as coronary artery lesion progression: coronary artery stenosis increases by ≥20%, high-risk signs increase by ≥2, and SIS score increases by ≥2; Otherwise, it is not progression.