Coronary artery disease early prediction model construction method, system, device and medium
By constructing an early prediction method for coronary artery disease based on lipidomics analysis and multivariate logistic regression model, the problem of insufficient identification ability of traditional biomarkers is solved, and early accurate screening and risk quantification are achieved, providing a basis for personalized diagnosis and treatment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- RENJI HOSPITAL AFFILIATED TO SHANGHAI JIAO TONG UNIV SCHOOL OF MEDICINE
- Filing Date
- 2026-04-29
- Publication Date
- 2026-05-29
AI Technical Summary
Existing technologies are insufficient for early and accurate screening of coronary artery disease. Traditional biomarkers and blood lipid indicators are not capable of identifying the disease in its early stages and cannot meet the needs of early clinical diagnosis.
By collecting clinical covariates and fasting plasma samples from the coronary artery group and the healthy control group, a systematic lipidomics analysis was conducted. Combining the pathophysiological mechanism of coronary artery disease (CAD) and progressive lipid feature screening, a multivariate logistic regression model was constructed. The model was evaluated using the training set and validation set, and finally a risk estimation model was formed.
It enables early and efficient screening and precise risk stratification of coronary artery disease, providing an objective basis for personalized clinical diagnosis and treatment, and quantifying the risk of disease with only a single blood test.
Smart Images

Figure CN122117431A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, system, device, and medium for constructing an early prediction model for coronary artery disease. Background Technology
[0002] Coronary artery disease (CAD) is a prevalent cardiovascular disease with high mortality and disability rates worldwide. Its pathogenesis is closely related to abnormal lipid metabolism, and early screening of high-risk individuals is of great significance in reducing the risk of cardiovascular events.
[0003] Currently, clinical diagnosis of coronary artery disease (CAD) mainly relies on imaging examinations (such as coronary CT angiography and coronary angiography) and serum biomarkers. However, existing biomarkers (such as troponin and myoglobin) usually only rise significantly after myocardial injury occurs, limiting their ability to identify the early stages of the disease. Traditional lipid indicators (such as total cholesterol, low-density lipoprotein cholesterol, and triglycerides), while associated with cardiovascular disease risk, lack specificity and predictive ability, making accurate early screening difficult and failing to meet the needs of early clinical diagnosis. Therefore, there is an urgent need for an effective method for constructing an early prediction model for coronary artery disease to address these issues. Summary of the Invention
[0004] In view of the above problems, the present invention is proposed to provide a method, system, device and medium for constructing an early prediction model for coronary artery disease that overcomes or at least partially solves the above problems.
[0005] To achieve the above and other related objectives, the present invention provides a method for constructing an early prediction model for coronary artery disease, the method comprising: Clinical covariates and fasting plasma samples were collected from subjects in the coronary artery group and the healthy control group. Systematic lipidomics analysis was performed based on a pre-set quality control system to obtain lipid characteristic data for each subject. The clinical covariates included gender, age, and BMI. The lipid feature data is preprocessed to obtain the corresponding standardized lipid variables, and combined with the clinical covariates of the corresponding subjects to jointly construct the model input variables; The model input variables of each subject were divided into training set and validation set according to a set ratio. The set of input lipid features for the training set was determined by combining the pathophysiological mechanism of CAD and the progressive lipid feature screening strategy. A multivariate logistic regression model was constructed by fitting the set of lipid features and clinical covariates in the training set. The model performance was independently evaluated using a validation set. After successful validation, the model parameters were solidified in the form of a model parameter table, and finally a risk estimation model for calculating the CAD risk probability of the test sample was constructed.
[0006] Optionally, the systematic lipidomics analysis based on a pre-established quality control system to obtain lipid characteristic data for each subject includes: A unique sample code was established for fasting plasma samples from subjects in the coronary artery group and the healthy control group, and a label was used to identify the testing status of each plasma sample; wherein, the testing status includes pending testing, under testing, completed testing, and retained sample; After inserting QC samples into each testing batch, fasting plasma samples are selected for lipidomics testing and data collection based on their unique sample codes and testing status. Data correction and quality control screening are completed based on the QC samples to obtain lipid characteristic data for each subject.
[0007] Optionally, the preprocessing of the lipid feature data to obtain the corresponding standardized lipid variables includes: Perform a logarithmic transformation on each lipid feature data to obtain the logarithmic lipid value; Based on the logarithmic lipid values of all fasting plasma samples from the coronary artery group and the healthy control group, the mean and standard deviation of each lipid feature were calculated, and the corresponding standardized lipid variables were obtained by Z-score standardization.
[0008] Optionally, the step of combining the pathophysiological mechanism of CAD and the progressive lipid feature screening strategy to determine the set of input lipid features for the training set includes: For the standardized lipid variables in the training set, combined with the core pathophysiological mechanisms of CAD cholesterol accumulation, foam cell formation and plaque instability, the t-test was used to analyze the differences between the standardized lipid variables of the coronary artery group and the healthy control group, and lipid metabolites that showed significant differences between the two groups and were related to the pathological process of CAD were screened to construct a candidate lipid set. For the candidate lipid set, LASSO regression was used for sparsification and feature dimensionality reduction, and redundant variables were eliminated by combining CAD pathophysiological mechanism constraints, retaining key lipid features with clear pathological significance, thus forming a key lipid feature set. PLS-DA analysis was used to verify the discriminative power and visualize the set of key lipid features, calculate the variable importance projection value, and help evaluate the classification stability and pathological relevance of the features. Finally, the set of lipid features to be used for model construction was determined.
[0009] Optionally, the step of constructing a multivariate logistic regression model using the input lipid feature set and clinical covariates from the training set, and independently evaluating the model performance using a validation set, includes: Using the set of lipid features and clinical covariates in the training set as independent variables and the corresponding CAD disease labels of the subjects as dependent variables, a multivariate logistic regression model was constructed; wherein, the CAD disease status labels include CAD disease status and healthy control status. After the multivariate logistic regression model has converged, the model intercept and the regression coefficients corresponding to each input variable are obtained. The lipid feature set of the validation set and clinical covariates are input into the converged model to obtain the CAD risk prediction results of the validation set. Based on the CAD disease label and CAD risk prediction results of each subject in the validation set, pre-set evaluation indicators were calculated to complete the independent evaluation of model performance; wherein, the pre-set evaluation indicators include accuracy, sensitivity, recall, F1 score and AUC value.
[0010] Optionally, after the verification is passed, the model parameters corresponding to the model are solidified in the form of a model parameter table, and a risk estimation model for calculating the CAD risk probability of the sample to be tested is finally constructed, including: When the evaluation indicators meet the preset threshold, the model is deemed to have passed the validation. The intercept and regression coefficients of each input variable of the validated model are extracted, organized into a structured model parameter table, and then solidified. Based on the solidified model parameter table, a risk estimation model is constructed to calculate the CAD risk probability of the sample to be tested.
[0011] Optionally, after the step of finally constructing a risk estimation model for calculating the CAD risk probability of the sample to be tested, the method further includes: Clinical covariates and fasting plasma samples were collected from the subjects to be tested. Lipomics analysis was performed on the fasting plasma samples to obtain the corresponding input lipid data. Based on the clinical covariates and the lipid data of the model, the risk estimation model and the model parameter table are used to make predictions and estimates, and the CAD risk probability of the subject to be tested is calculated. By using preset probability threshold classification rules, the CAD risk probability is mapped to the corresponding risk level to assist in clinical diagnosis and follow-up management.
[0012] Secondly, the present invention also provides a system for constructing an early prediction model for coronary artery disease, the system comprising: The collection module is used to collect clinical covariates and fasting plasma samples from subjects in the coronary artery group and the healthy control group, and to perform systematic lipidomics analysis based on a pre-set quality control system to obtain lipid characteristic data for each subject; wherein, the clinical covariates include gender, age and BMI; The preprocessing module is used to preprocess the lipid feature data to obtain the corresponding standardized lipid variables, and combine them with the clinical covariates of the corresponding subjects to jointly construct the model input variables; The determination module is used to divide the model input variables of each subject into training set and validation set according to a set ratio, and to determine the set of input lipid features for the training set by combining the pathophysiological mechanism of CAD and the progressive lipid feature screening strategy. The training module is used to fit and construct a multivariate logistic regression model using the set of lipid features and clinical covariates in the training set. The model performance is independently evaluated using a validation set. After successful validation, the model parameters are fixed in the form of a model parameter table, and finally a risk estimation model for calculating the CAD risk probability of the test sample is constructed.
[0013] Thirdly, the present invention provides an electronic device comprising: a memory and a processor; the memory for storing a computer program; and the processor for executing the computer program stored in the memory to enable the electronic device to perform the method for constructing an early prediction model for coronary artery disease as described above.
[0014] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by an electronic device, implements the method for constructing an early prediction model for coronary artery disease as described above.
[0015] Fifthly, the present invention provides a computer program product including computer program code, which, when run on a computer, causes the computer to implement the method described above.
[0016] The above-described one or more technical solutions provided by this invention can have the following advantages or at least achieve the following technical effects: This invention simultaneously collects clinical covariates and fasting plasma lipidomics data from subjects, relying on a quality control system to ensure the accuracy and reliability of the test data. Lipid characteristics are preprocessed and standardized, and unified model input variables are constructed in conjunction with clinical indicators. Training and validation sets are divided, and core lipid characteristics are screened based on CAD pathological mechanisms, redundant variables are eliminated, and the correlation between characteristics and coronary artery lesions is strengthened. A multivariate logistic regression model is constructed using the training set, and evaluated using the validation set to ensure the model's discriminative efficacy and generalization ability. Key model parameters are solidified and a parameter table is generated, establishing a standardized CAD risk estimation model. In clinical application, only a single blood test combined with basic demographic information is needed to quickly quantify the early risk of disease, achieving efficient screening and accurate risk stratification, providing objective evidence for early warning of coronary artery disease and personalized clinical diagnosis and treatment. Attached Figure Description
[0017] Figure 1 The diagram shows a flowchart of a method for constructing an early prediction model for coronary artery disease in one embodiment of the present invention.
[0018] Figure 2 The image shown is a schematic diagram of coronary angiography of a healthy control group and a coronary artery group in one embodiment of the present invention;
[0019] Figure 3 This is a schematic diagram illustrating the predictive ability of different variables in the training set for coronary artery disease risk in one embodiment of the present invention.
[0020] Figure 4 This diagram illustrates a comparison of the relative concentrations of three lipid metabolites in a healthy control group and a CAD group in one embodiment of the present invention.
[0021] Figure 5 The diagram shown is a schematic of a coronary artery prediction nomogram model constructed based on a training set in one embodiment of the present invention.
[0022] Figure 6 The diagram shows a calibration curve for the predictive capability of the training set model in one embodiment of the present invention.
[0023] Figure 7 The diagram shows the ROC curves of five machine learning models on the validation set in one embodiment of the present invention.
[0024] Figure 8 The diagram shows the performance of five machine learning models on six key metrics in one embodiment of the present invention.
[0025] Figure 9 This is a schematic diagram of the functional modules of a coronary artery disease early prediction model construction system according to an embodiment of the present invention;
[0026] Figure 10The diagram shown is a schematic representation of an electronic device according to an embodiment of the present invention. Detailed Implementation
[0027] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, unless otherwise specified, the following embodiments and features described therein can be combined with each other.
[0028] It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Therefore, the drawings only show the components related to the present invention and are not drawn according to the actual number, shape and size of the components in the actual implementation. In the actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.
[0029] In the following description, numerous details are explored to provide a more thorough explanation of embodiments of the invention. However, it will be apparent to those skilled in the art that embodiments of the invention may be practiced without these specific details. In other embodiments, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring embodiments of the invention.
[0030] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate for the embodiments of this disclosure described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion.
[0031] Unless otherwise stated, the term "multiple" means two or more.
[0032] In this embodiment of the disclosure, the character " / " indicates that the objects before and after it are in an "or" relationship. For example, A / B means: A or B.
[0033] The term "and / or" describes an association between objects, indicating that three relationships can exist. For example, A and / or B means: A or B, or A and B.
[0034] The technical solutions of the present invention will now be described in detail with reference to the accompanying drawings.
[0035] Please see Figure 1 An embodiment of the present invention provides a method for constructing an early prediction model for coronary artery disease, the method comprising the following steps S10-S40: Step S10: Collect clinical covariates and fasting plasma samples from subjects in the coronary artery group and the healthy control group, and perform systematic lipidomics analysis based on a pre-set quality control system to obtain lipid characteristic data for each subject; wherein, the clinical covariates include gender, age and BMI (Body Mass Index).
[0036] The coronary artery group (CAD group) was used to characterize the set of subjects clinically diagnosed with coronary artery disease. All included subjects met the pre-set diagnostic criteria for coronary artery disease and served as the target observation cohort for this study.
[0037] The healthy control group was a set of subjects who did not have coronary artery disease or other serious cardiovascular and cerebrovascular diseases. After baseline feature matching, they were balanced and comparable with the coronary artery group and served as the control cohort for this study.
[0038] Please see Figure 2 , Figure 2 The left side (A) shows a normal coronary artery with a smooth vessel outline, natural lumen transition, no stenosis or plaque, smooth blood flow, and sufficient blood supply to the myocardium. Figure 2 The right side (B) shows a serious lesion in the coronary artery. The white arrow points to a localized severe stenosis caused by the rupture of atherosclerotic plaques with thrombus formation, which obstructs blood flow and is the responsible lesion for angina or myocardial infarction.
[0039] The quality control system includes inserting quality control samples (QC samples), retention time correction, peak area normalization, outlier removal, and batch effect correction to ensure the reliability and reproducibility of lipidomics analysis data.
[0040] Clinical covariates are indicators used to comprehensively reflect the baseline characteristics of subjects, including at least gender, age, and BMI. They are used to balance the baseline of the cohort, eliminate the influence of confounding factors, and provide a standardized variable basis for subsequent lipid characterization analysis and model construction.
[0041] Lipid characterization data refers to the qualitative and quantitative detection results of various lipid metabolites in plasma, covering information on the types, relative contents or concentrations of various lipids, and is used to reflect the lipid metabolism profile characteristics in the subject's body.
[0042] In practice, clinical information of each subject in the coronary artery disease group and the healthy control group can be collected and organized, including age, gender, BMI, history of hypertension, history of diabetes, and family history of coronary heart disease, to form clinical covariates for each subject. These clinical covariates include at least gender, age, and BMI. Based on the aforementioned clinical covariates, baseline matching is performed between the two groups of subjects to ensure that the coronary artery disease group and the healthy control group are comparable and to eliminate the influence of confounding factors. Fasting plasma samples are collected from each group of subjects simultaneously. Subsequently, high-throughput lipidomics detection technology can be used to qualitatively and quantitatively detect various lipid metabolites in each fasting plasma sample, including but not limited to phospholipids, triglycerides, cholesterol esters, and free fatty acids, to comprehensively obtain lipid characteristic data for each subject, providing reliable data support for subsequent lipid characteristic screening and CAD risk estimation model construction.
[0043] Step S20: The lipid feature data is preprocessed to obtain the corresponding standardized lipid variables, and combined with the clinical covariates of the corresponding subjects to jointly construct the model input variables.
[0044] Standardized lipid variables refer to regularized lipid indicators obtained after preprocessing the original lipid characteristic data. This can eliminate dimensional differences, unify data distribution, and is suitable for modeling.
[0045] The model input variables are a comprehensive feature set composed of standardized lipid variables and clinical covariates of the subjects, providing multidimensional feature support for coronary artery disease risk analysis and model training.
[0046] In practice, lipid characteristic data from the two groups of subjects can be preprocessed by normalization, dimensionless removal, and noise reduction to obtain standardized lipid variables. Subsequently, the standardized lipid variables are fused with corresponding clinical covariates to construct model input variables. This approach can eliminate differences in data dimensions and detection interference, unify data distribution specifications, and integrate multi-dimensional features, effectively improving the accuracy and stability of subsequent CAD disease risk model training and analysis.
[0047] Step S30: Divide the model input variables of each subject into training set and validation set according to a set ratio, and determine the set of input lipid features for training set by combining the pathophysiological mechanism of CAD and the progressive lipid feature screening strategy.
[0048] Among them, the pathophysiological mechanism of CAD refers to the pathogenesis and metabolic mechanism of coronary artery disease, which is used to guide lipid feature screening from a medical perspective and ensure that the features are clinically relevant.
[0049] The progressive lipid feature screening strategy refers to screening lipid indicators in a hierarchical manner, eliminating invalid and redundant features, simplifying dimensions, and preventing model overfitting.
[0050] The set of lipid features for modeling refers to the core lipid feature combination that is finally selected and used for model training and risk assessment after combining pathological mechanisms and step-by-step screening.
[0051] In practical implementation, the model input variables of the two groups of subjects can be divided into training and validation sets according to a preset ratio (e.g., training set: validation set = 7:3). The training set is used for model learning and feature fitting, while the validation set is used for subsequent model performance evaluation. Data splitting distinguishes between modeling and evaluation data, suppresses overfitting, and improves the model's generalization ability. Subsequently, for the model input variables of the two groups of subjects in the training set, a preliminary screening can be performed based on the pathophysiological mechanism of coronary artery disease. Based on the mechanism of coronary artery disease lesions and lipid metabolism disorders, lipid indicators with clinical and pathological relevance are retained. Simultaneously, a progressive lipid feature screening strategy is adopted to eliminate redundant and irrelevant interference features step by step, and simplify the feature dimensions. Through bidirectional screening, the set of lipid features suitable for modeling is determined, providing accurate, simplified, and pathologically relevant core feature support for subsequent model training, parameter optimization, and disease risk prediction.
[0052] Step S40: A multivariate logistic regression model is constructed by fitting the set of lipid features and clinical covariates in the training set. The model performance is independently evaluated using a validation set. After successful validation, the model parameters are solidified in the form of a model parameter table, and finally a risk estimation model for calculating the CAD risk probability of the sample to be tested is constructed.
[0053] The model parameter table is a standardized table formed by unifying and solidifying the core operational parameters of the multivariate logistic regression model. It mainly includes the lipid characteristics of each input model, clinical covariates and their corresponding regression coefficients, as well as key operational parameters such as the model intercept term and risk judgment threshold. It is used to unify and reproduce the calculation logic of CAD risk probability, ensuring that the CAD risk calculation process is reproducible and the results are traceable, and providing a clear basis for the risk assessment of subsequent samples to be tested.
[0054] Table 1 shows the model parameters for multivariate regression analysis.
[0055]
[0056] Table 1 shows the model parameters for the multivariate regression analysis, which aims to assess the probability of various lipid metabolites and clinical indicators on the risk of coronary artery disease (CAD). As can be seen from Table 1, cholesterol esters CE (20:4) has the strongest predictive ability (OR=35.687, P<0.001), indicating that its elevated level is highly associated with the risk of CAD. Other significant risk factors such as CE (20:3) (OR=3.571, P=0.023) and TAG (52:5)_FA20:3 (OR=3.725, P=0.018) are also significant risk factors, and their elevated levels significantly increase the risk of CAD.
[0057] In the model parameter table, the estimated values represent logistic regression coefficients; positive values indicate increased risk, and negative values indicate decreased risk. The p-value is used to determine the reliability of the results; typically, p < 0.05 indicates that the variable's impact is statistically significant. The odds ratio (OR) measures the magnitude of the risk impact; OR > 1 indicates a risk factor (the larger the value, the higher the risk), and OR < 1 indicates a protective factor (the smaller the value, the stronger the protective effect). The 95% confidence intervals, corresponding to the 2.50%-97.50% quantiles in the table, are used to verify the reliability of the odds ratio (OR) estimate (an interval not containing 1 is considered significant).
[0058] The risk estimation model is a logistic regression prediction model based on lipid characteristics and clinical indicators. It relies on fixed parameter calculations to accurately output the probability of CAD disease risk for the tested sample; it can be expressed as: (1) In the formula, P(CAD) is the CAD disease risk probability (range 0~1); R is the weighted score of comprehensive risk factors; exp(-R) is the natural exponential function, used to map the linear score (i.e., the risk score, R) to the 0~1 interval.
[0059] The risk score (R) can be represented as follows: (2) In the formula, This is the intercept term (the baseline offset of the model, a fixed coefficient). The regression coefficients for sex, age, and BMI are fixed coefficients, representing the weights of the corresponding factors on the risk of CAD; sex, age, and BMI are the measured values of the subjects' sex, age, and BMI, respectively (clinical covariates).
[0060] S is the simplest lipid feature set; The regression coefficient for the Lth lipid is a fixed coefficient representing the weight of the lipid's impact on CAD risk. The regression coefficient is the standardized quantitative value (dimensionless lipidomics detection result) of the Lth lipid. The weighted sum of all lipid features included in the model (liposome risk contribution).
[0061] In practical implementation, a multivariate logistic regression model can be constructed by using the training set samples as a basis and combining the determined set of lipid features for model input with clinical covariates. Then, an independent validation set is divided for external validation, and evaluation indicators such as accuracy, AUC value, F1 score, precision, specificity, and sensitivity are used to comprehensively evaluate the model's generalization performance and predictive stability from dimensions such as discriminative power and good fit. Subsequently, after the model validation effect meets the standards, the core parameters of the model, such as regression coefficients, intercept, and threshold, are extracted and fixed in the form of a standardized model parameter table. Based on the fixed model parameters, an algorithm logic that can be automatically calculated is built, and finally, a risk estimation model that can quantitatively calculate the probability of CAD disease risk in the test sample is constructed.
[0062] In this embodiment, clinical covariates and fasting plasma lipidomics data of subjects were collected simultaneously, and the accuracy and reliability of the test data were ensured by relying on a quality control system. Lipid characteristics were preprocessed and standardized, and unified model input variables were constructed in conjunction with clinical indicators. Training and validation sets were divided, and core lipid characteristics were screened based on the pathological mechanisms of coronary artery disease (CAD), redundant variables were eliminated, and the correlation between characteristics and coronary lesions was strengthened. A multivariate logistic regression model was constructed using the training set, and evaluated using the validation set to ensure the model's discriminative efficacy and generalization ability. Key model parameters were solidified and a parameter table was formed, establishing a standardized CAD risk estimation model. In clinical applications, only a single blood test combined with basic demographic information is needed to quickly quantify the early risk of disease, achieving efficient screening and accurate risk stratification, providing objective evidence for early warning of coronary artery disease and individualized clinical diagnosis and treatment.
[0063] Based on the foregoing embodiments, a second embodiment of the method for constructing an early prediction model for coronary artery disease according to the present invention is proposed. In this embodiment, step S10 may include the following sub-steps S101~S102: Sub-step S101: Establish unique sample codes for fasting plasma samples from subjects in the coronary artery group and the healthy control group, and use labels to identify the detection status of each plasma sample; wherein, the detection status includes pending detection, under detection, completed detection, and retained sample.
[0064] Among them, the exclusive sample code is a unique and non-repeating identification code assigned to each fasting plasma sample. It is used to distinguish different subjects and different groups of samples (such as healthy or coronary artery samples), realize the traceability of the entire sample process, avoid sample confusion, false detection or false detection, and ensure that the test data corresponds one-to-one with the subject.
[0065] The testing status refers to the stage of a fasting plasma sample in the testing process, which includes at least the stages of pending testing, in progress testing, completed testing, and retained sample.
[0066] In practice, each group of subjects can be assigned a unique sample code and a testing status marker for their fasting plasma samples. The testing status includes pending testing, testing in progress, testing completed, and sample retention, so as to achieve full-process traceability and status differentiation of fasting plasma samples.
[0067] In sub-step S102, after inserting QC samples into each testing batch, fasting plasma samples are selected based on their unique sample codes and testing status for lipidomics detection and data collection. Data correction and quality control screening are completed based on the QC samples to obtain lipid characteristic data for each subject.
[0068] Among them, QC (Quality Control) samples are mixed quality control samples with uniform properties, interspersed in each test batch; they are used to monitor the stability of the instrument and the test batch, correct test and batch deviations, complete the quality control screening and deviation correction of omics data, and ensure the accuracy and reliability of lipidomics test results.
[0069] In practice, QC (Quality Control) samples are uniformly added to each batch of samples for lipidomics testing to monitor the stability of the testing system and reduce batch errors. Then, based on the unique sample codes and testing status of each fasting plasma sample set in the early stage, fasting plasma samples that meet the testing conditions are screened out. High-throughput lipidomics testing is carried out on the qualified samples after screening, and raw test data is collected simultaneously. Subsequently, the raw data is subjected to deviation correction, noise reduction and quality control screening using the QC samples in the batch to remove abnormal data and unqualified test results, and complete the data standardization process to finally obtain accurate and reliable lipid characteristic data for each subject.
[0070] In this embodiment, by assigning unique codes to plasma samples and labeling them with multiple testing statuses, accurate sample differentiation, end-to-end traceability, and standardized management are achieved, preventing sample confusion and misdetection. Each testing batch is mixed with QC samples, and the samples to be tested are screened by combining unique sample codes and testing statuses. In conjunction with QC samples, data correction and quality control screening are completed, effectively reducing batch deviation, eliminating abnormal data, and ensuring the accuracy and reliability of lipid characteristic data. This provides a compliant and high-quality data foundation for subsequent lipid analysis and the construction of coronary disease risk models.
[0071] Based on the foregoing embodiments, a third embodiment of the method for constructing an early prediction model for coronary artery disease according to the present invention is proposed. In this embodiment, step S20 may include the following sub-steps S201~S202: Sub-step S201: Perform a logarithmic transformation on each lipid feature data to obtain the logarithmic lipid value.
[0072] Logarithmic lipid values refer to the derived index obtained by logarithmically transforming the original lipid feature detection data; they are used to improve the skewed distribution of data, weaken the influence of extreme values, and adapt to the requirements of subsequent statistical analysis and model building.
[0073] In practice, logarithmic transformation can be performed on each lipid characteristic data of the two groups of subjects to obtain the logarithmic lipid values of each subject; the distribution pattern of lipid indicators can be optimized to make it closer to a normal distribution, reduce the influence of heteroscedasticity, and effectively improve the adaptability and computational stability of subsequent data standardization, feature screening and model fitting.
[0074] In sub-step S202, based on the logarithmic lipid values of all fasting plasma samples from the coronary artery group and the healthy control group, the mean and standard deviation of each lipid feature are calculated, and Z-score standardization is used to obtain the corresponding standardized lipid variables.
[0075] In practical implementation, the logarithmic lipid values of fasting plasma samples from all subjects in the coronary artery group and the healthy control group can be used as the basis for calculation. For a single lipid feature, the mean and standard deviation within the group are calculated separately. Based on the statistical parameters of this group, the logarithmic lipid values are uniformly standardized using the Z-score standardization algorithm. This method uses the overall data of the same group as a reference benchmark to eliminate the differences in dimensions and numerical magnitudes between different lipid indicators, unify the data fluctuation range and distribution scale, and avoid analytical biases caused by uneven range of indicator values. At the same time, it builds on the data optimization effect of the previous logarithmic transformation, further standardizes the data format, weakens individual differences and residual detection errors, and finally generates standardized lipid variables with uniform format and regular dimensions. This provides standardized and horizontally comparable feature data for subsequent feature fusion, variable screening, and risk estimation model construction.
[0076] In this embodiment, the distribution of lipid data is improved by logarithmic transformation to suppress extreme values and heteroscedasticity. Then, Z-score standardization is implemented by combining the group mean and standard deviation to eliminate differences in the dimensions and magnitudes of the indicators. Through two levels of data optimization, the data format is standardized, experimental errors are reduced, and standardized lipid variables are obtained, providing stable, comparable, and high-quality input data for subsequent multi-feature fusion and coronary artery disease risk model construction.
[0077] Based on the foregoing embodiments, a fourth embodiment of the method for constructing an early prediction model for coronary artery disease according to the present invention is proposed. In this embodiment, step S30 may include the following sub-steps S301 to S303: In sub-step S301, based on the standardized lipid variables in the training set, and combined with the core pathophysiological mechanisms of CAD cholesterol accumulation, foam cell formation, and plaque instability, a t-test is used to analyze the differences between the standardized lipid variables of the coronary artery group and the healthy control group, and lipid metabolites that are significantly different between the two groups and are related to the pathological process of CAD are screened out to construct a candidate lipid set.
[0078] Cholesterol accumulation is a pathological state in which lipid metabolism in the coronary endothelium and vessel wall is disordered during CAD pathology, resulting in the abnormal deposition and continuous accumulation of cholesterol and other lipids.
[0079] Foam cell formation is a pathological process in which monocytes invade the vascular endothelium, differentiate into macrophages, engulf large amounts of oxidized lipids and cholesterol esters, causing excessive accumulation of intracellular lipids, and then transforming into foam-like cells.
[0080] Plaque instability refers to the thinning of the fibrous cap and enlargement of the lipid core in coronary atherosclerotic plaques, accompanied by pathological changes such as inflammation and matrix degradation, which makes them prone to damage and rupture, leading to adverse cardiovascular events such as thrombosis and myocardial ischemia.
[0081] The candidate lipid set refers to the collection of lipid metabolites that are significantly different and highly correlated with coronary artery lesions, which are initially retained after screening by intergroup difference test and CAD pathophysiological mechanism; it provides a preliminary pool of candidate features for further feature screening and model construction.
[0082] In practical implementation, for the standardized lipid variables in the training set, using the CAD group and the healthy control group as the grouping criteria, and combining the core pathophysiological mechanisms such as CAD cholesterol accumulation, foam cell formation, and plaque instability, the independent samples t-test (Student's test) can be used. t -test, Student's t-test) to conduct intergroup difference analysis; set a uniform statistical test level, compare the expression levels of lipid indicators in the two groups one by one, and screen lipid metabolites with significant differences between groups; at the same time, combine with the pathogenesis of coronary artery disease to conduct correlation analysis, eliminate indicators that only have statistical differences but are not related to the pathological process of CAD, and finally integrate the screened effective lipid indicators to construct a candidate lipid set.
[0083] In sub-step S302, for the candidate lipid set, LASSO regression is used for sparsification and feature dimensionality reduction, and redundant variables are eliminated by combining CAD pathophysiological mechanism constraints, retaining key lipid features with clear pathological significance, thus forming a key lipid feature set.
[0084] Among them, the key lipid feature set refers to the core lipid features that are highly correlated and of high value that are retained after LASSO (Least Absolute Shrinkage and Selection Operator) regression dimensionality reduction and redundancy removal, combined with CAD pathological screening, based on the candidate lipid set, and provide core input variables for the CAD risk estimation model.
[0085] In the specific implementation, a LASSO regression model is constructed using all lipid indicators in the candidate lipid set as independent variables and CAD disease status (CAD group or healthy control group) as the outcome variable. Then, the model coefficients are compressed by introducing a penalty coefficient, and variable sparsity processing is carried out to automatically remove redundant lipid variables with strong collinearity and low contribution, thus completing the dimensionality reduction of high-dimensional features. Subsequently, based on the algorithm screening, the pathophysiological mechanisms of CAD such as lipid metabolism disorders, foam cell formation, and plaque lesions are combined for constraint screening to remove redundant variables without pathological support and irrelevant to the pathogenesis, retaining only lipid features with clear pathological mechanisms and high disease correlation, and integrating them to form a key lipid feature set.
[0086] In sub-step S303, PLS-DA analysis is used to verify the discriminative power and visualize the set of key lipid features, calculate the variable importance projection value, assist in evaluating the classification stability and pathological relevance of the features, and finally determine the set of lipid features to be used for model construction.
[0087] Among them, the variable importance projection (VIP) value is an evaluation index for multivariate statistical models such as PLS-DA, used to quantify the contribution of each characteristic variable to intergroup differences and model fit; the higher the VIP value, the stronger the discriminative power and biological relevance of the variable, and it is often used to screen key differential lipid biomarkers.
[0088] In the specific implementation, the key lipid feature set obtained from the above screening is used as the analysis variable, and the CAD group and the healthy control group are used as the grouping variables. A discriminant model can be constructed using PLS-DA (Partial Least Squares-Discriminant Analysis). The sample distribution is visualized through the score map output by the model, and the separation degree and discrimination effect of the two groups of samples are evaluated intuitively. Based on the constructed PLS-DA model, the VIP value of each lipid indicator is calculated. Then, the contribution of each lipid indicator to the group classification is quantified by the VIP value. The higher the VIP value, the stronger the effect of the feature on the distinction between groups. Based on the VIP value and combined with pathological correlation, lipid features with excellent classification efficiency are screened, and finally the lipid feature set for the model is determined.
[0089] In this embodiment, the differences in lipid indicators between the CAD group and the healthy control group were analyzed using a t-test. A candidate lipid set was obtained by combining the core pathological mechanisms of CAD, completing the initial feature screening and narrowing the screening scope. Then, LASSO regression was used to reduce the dimensionality of candidate features, eliminate collinearity and data redundancy, and pathological constraints were used to form a key lipid feature set, simplifying feature dimensions and improving lesion targeting. Finally, PLS-DA was used to verify the feature discrimination efficiency, and the contribution of features and pathological relevance were evaluated by combining VIP values. Inefficient indicators were eliminated, and the lipid feature set for the model was determined. Multi-dimensional, layer-by-layer quality control ensured excellent feature discrimination ability and sufficient pathological basis, providing high-quality core features for model construction.
[0090] Based on the foregoing embodiments, a fifth embodiment of the method for constructing an early prediction model for coronary artery disease of the present invention is proposed. In this embodiment, step S40 may include the following sub-steps S401 to S404: Sub-step S401: Using the set of lipid features and clinical covariates in the training set as independent variables and the corresponding CAD disease label of the subjects as dependent variables, a multivariate logistic regression model is fitted and constructed; wherein, the CAD disease label includes CAD disease status or healthy control status.
[0091] Among them, the CAD disease label is a sample grouping attribute used for binary classification in the model, including the coronary artery disease status (i.e., CAD disease status) and the healthy control status, which are used as the basis for binary classification in the multivariate logistic regression model.
[0092] Please see Figure 3 This forest plot primarily demonstrates the predictive power of demographic characteristics (such as age, sex, and BMI) and lipid metabolites in the training set for coronary artery disease risk. The results show that specific lipid metabolites (especially CE and TAG types) exhibit significant value in predicting coronary artery disease risk. Among them, CE(20:4) showed the strongest risk association, while TAG(52:5)_FA20:3 (OR=3.89) and CE(20:3) (OR=3.57) were robust and significant risk factors.
[0093] Please see Figure 4 The box plot compared the relative concentrations of three lipid metabolites in the healthy control group and the CAD group. The results showed that the relative concentrations of CE(20:3), CE(20:4), and TAG(52:5)_FA20:3 in the CAD group were significantly higher than those in the healthy control group (P<0.0001), indicating that these metabolites are closely related to the disease state and are potential auxiliary diagnostic biomarkers.
[0094] Please see Figure 5This nomogram model provides an intuitive quantitative tool for assessing the risk of coronary artery disease by integrating clinical covariates (age, sex, BMI) and multiple lipid metabolites. It clearly demonstrates how the specific values of each predictor variable are converted into corresponding scores, which are then summed to obtain a total score, which in turn corresponds to a specific odds ratio, thereby assessing an individual's relative risk of disease.
[0095] In practical implementation, the training set samples can be used as the modeling object. The lipid feature set obtained after screening and the clinical covariates are used as independent variables. The CAD disease label (CAD disease status / healthy control status) corresponding to the sample is defined as a binary dependent variable, in which the CAD group is assigned a value of 1 and the healthy control group is assigned a value of 0. Then, the maximum likelihood estimation method is used to fit the independent variables and the dependent variables to construct a multivariate logistic regression model. Subsequently, the influence weights of lipid features and clinical indicators on CAD disease status are learned through the model to complete the initial construction of the risk estimation model.
[0096] Sub-step S402: After the multivariate logistic regression model has converged, obtain the model intercept and the regression coefficients corresponding to each input variable.
[0097] Please see Figure 6 The figure shows the calibration curve for the model's predictive ability on the training set. The model's calibration degree is evaluated by comparing the predicted probability with the actual observed probability. The results show that the bias-corrected prediction curve highly coincides with the ideal reference line, proving that the model fits well on the training set and that the predicted risk has excellent consistency with the actual risk.
[0098] In practice, the maximum likelihood estimation method can be used to iteratively train the constructed multivariate logistic regression model until the model loss function converges and the training process terminates. Then, by solving the model training, the intercept term and the regression coefficients corresponding to each input lipid feature and clinical covariate can be obtained, thus completing the solution and determination of the core operational parameters of the model.
[0099] Sub-step S403 involves inputting the set of lipid features from the validation set and clinical covariates into the converged model to obtain the CAD risk prediction results for the validation set.
[0100] Among them, the CAD risk prediction results of the validation set refer to the CAD disease risk probability and binary classification prediction results of each sample after inputting the lipid characteristics and clinical covariates of the validation set samples into the convergent model; these are used to calculate evaluation indicators and test the discriminative ability and generalization performance of the model.
[0101] In practice, the lipid features and clinical covariates corresponding to the validation set samples can be uniformly input into the multivariate logistic regression model after training and convergence. The model's fixed intercept and regression coefficients of each variable are used for calculation to obtain the CAD risk probability of each validation set sample, thereby outputting the overall CAD risk prediction result of the validation set.
[0102] Sub-step S404: Based on the CAD disease label and CAD risk prediction results of each subject in the validation set, calculate the preset evaluation indicators to complete the independent evaluation of model performance; wherein, the preset evaluation indicators include accuracy, sensitivity, recall, F1 score and AUC value.
[0103] Among them, the preset evaluation index refers to the statistical index used to pre-set and quantify the predictive performance of the model, including accuracy, sensitivity, recall, F1 score and AUC value, which can comprehensively measure the discriminative power and generalization ability of the model.
[0104] In practice, a confusion matrix can be constructed and statistical calculations can be performed based on the CAD disease labels of each subject in the validation set and the CAD risk prediction results output by the model. Pre-set evaluation indicators such as the model's accuracy, sensitivity, specificity, F1 score, and AUC value can then be obtained. Subsequently, the model's prediction accuracy, discriminant efficacy, and generalization ability can be comprehensively evaluated based on the above indicators to complete the independent external validation of the model's performance.
[0105] Please see Figure 7 The figure shows the receiver operating characteristic (ROC) curves of five machine learning models (logistic regression, random forest, support vector machine, XGBoost, and super learner) on the validation set. The AUC values of each model are all greater than 0.8, with the support vector machine (SVM) model performing the best (AUC=0.854), indicating that these models have good risk prediction discrimination.
[0106] Furthermore, to further verify the model performance, various algorithms, including Random Forest, Support Vector Machine, XGBoost, and Super Learner, were used for testing. The results show that the AUC values of each model on the validation set are all greater than 0.80, demonstrating excellent predictive performance. This strongly proves that the risk estimation model possesses superior stability, generalization ability, and reliability.
[0107] Please see Figure 8The radar charts were used to further compare the overall performance of five machine learning models (logistic regression, random forest, support vector machine, XGBoost, and super learner) on six key metrics (AUC, F1 score, precision, specificity, sensitivity, and accuracy). The results show that all models perform well and are relatively balanced across all metrics, with support vector machine (SVM) and super learner exhibiting slightly larger polygon areas, indicating slightly better overall performance.
[0108] Furthermore, in one embodiment, step S40 may further include the following sub-steps S405-S406: Sub-step S405: When the evaluation index meets the preset threshold, the model is deemed to have passed the validation. The intercept and regression coefficients of each input variable of the validated model are extracted, organized into a structured model parameter table, and then solidified. Sub-step S406: Based on the solidified model parameter table, construct a risk estimation model for calculating the CAD risk probability of the sample to be tested.
[0109] In practical implementation, when all evaluation indicators reach preset thresholds, the model is deemed validated. Then, the intercept, lipid characteristics of each input model, and regression coefficients corresponding to clinical covariates of the optimal model are extracted and compiled into a standardized structured table, completing the archiving and solidification of the model's core parameters. Subsequently, based on the solidified model parameter table, the fixed model intercept and regression coefficients of each input variable are retrieved. Combining the rules of multivariate logistic regression, the lipid characteristics and clinical covariate data of the test sample are substituted, and iterative calculations using a unified algorithm formula are performed to calculate the CAD disease risk probability of the test sample in real time, thereby constructing a standardized risk estimation model that can be directly applied. Thus, by selecting qualified models based on evaluation indicators, extracting and solidifying the model intercept and regression coefficients to form a standardized parameter table, and constructing a risk estimation model based on the solidified parameters, a unified calculation rule is established to stably calculate the CAD disease probability of the test sample, ensuring consistent and reliable evaluation results and improving clinical applicability and scalability.
[0110] In this embodiment, a multivariate logistic regression model is constructed by jointly modeling the lipid features of the training set with clinical covariates and combining them with binary disease labels. After iterative training, the model intercept and regression coefficients are stably output. Independent prediction is performed using the validation set, and the model performance is comprehensively verified by combining multiple evaluation indicators. This avoids the defects of single indicator evaluation, effectively ensures the model's fitting stability and generalization ability, and improves the accuracy and objectivity of CAD disease risk prediction, laying a solid technical foundation for model solidification and clinical risk screening.
[0111] Based on the foregoing embodiments, a sixth embodiment of the method for constructing an early prediction model for coronary artery disease according to the present invention is proposed. In this embodiment, after step S40, the method may further include the following steps A10 to A30: Step A10: Collect clinical covariates and fasting plasma samples from the subjects to be tested, and perform lipidomics detection on the fasting plasma samples to obtain the corresponding input lipid data.
[0112] Among them, the input lipid data are quantitative lipid detection data obtained from fasting plasma samples of the subjects to be tested, which correspond one-to-one with the input lipid features of the training set. The indicators of each dimension are kept consistent and are used to input the risk estimation model to complete the calculation.
[0113] In practice, clinical covariates such as age, gender, and BMI of the subjects to be tested can be collected, and fasting plasma samples of the subjects to be tested can be collected simultaneously and properly preserved and preprocessed. Then, lipidomics detection technology is used to quantitatively analyze the fasting plasma samples, detect the relevant concentrations of target lipid biomarkers, obtain target input lipid data that meet the model input requirements, and complete the standardization and organization of detection data.
[0114] Step A20: Based on the clinical covariates and the lipid data used in the model, the risk estimation model and the model parameter table are used to make predictions and estimates, and the CAD risk probability of the subject to be tested is calculated.
[0115] In practice, clinical covariates and lipid data of the test subjects can be extracted to ensure compliance with model input specifications. Then, the fixed model parameter table (including model intercept and regression coefficients of each variable) and risk estimation model are called, and the two types of data are simultaneously input into the model. Subsequently, according to the multivariate logistic regression operation rules, the model intercept and regression coefficients of each variable in the model parameter table are combined to calculate and finally output the CAD risk probability of the test subjects, thus completing the estimation.
[0116] Step A30: Using a preset risk level threshold classification rule, the risk probability is mapped to the corresponding risk level to assist in clinical diagnosis and follow-up management.
[0117] Among them, the probability threshold division rule is a risk cutoff value pre-set by combining clinical diagnostic criteria and validation set data, used to determine the CAD risk probability output by the model; based on this cutoff value, low, medium and high probability intervals are divided, thereby realizing the stratified determination of the disease risk of the subject to be tested.
[0118] In practice, the CAD risk probability output by the model can be compared with the threshold boundary value according to the pre-set probability threshold division rules. Then, based on the interval to which the CAD risk probability value belongs, it can be mapped to low, medium and high risk levels to achieve standardized risk stratification. Subsequently, by quantifying the stratification results, an intuitive and quantifiable reference can be provided for clinical disease diagnosis, stratified intervention and individualized follow-up management.
[0119] In this embodiment, by integrating multidimensional clinical covariates and lipidomics detection data of the subjects to be tested, and relying on fixed model parameters and risk estimation models, standardized quantitative calculations are performed to accurately output the probability of disease risk; combined with preset thresholds, risk stratification and classification are achieved, making the assessment results quantitative, intuitive and easy to interpret clinically.
[0120] Based on the same inventive concept, the seventh embodiment of the present invention also provides a coronary artery disease early prediction model construction system corresponding to the coronary artery disease early prediction model construction method of the foregoing embodiments. Since the principle of the system in the seventh embodiment of the present invention for solving the problem is similar to the coronary artery disease early prediction model construction method of the foregoing embodiments of the present invention, the implementation of the system can refer to the implementation of the method, and repeated details will not be elaborated further. Please refer to... Figure 9 The present invention provides a system for constructing an early prediction model for coronary artery disease, the system comprising: The collection module 10 is used to collect clinical covariates and fasting plasma samples from subjects in the coronary artery group and the healthy control group, and to perform systematic lipidomics analysis based on a pre-set quality control system to obtain lipid characteristic data for each subject; wherein, the clinical covariates include gender, age and BMI; Extraction module 20 is used to preprocess the lipid feature data to obtain the corresponding standardized lipid variables, and combine them with the clinical covariates of the corresponding subjects to jointly construct the model input variables; Training module 30 is used to divide the model input variables of each subject into training set and validation set according to a set ratio, and to determine the set of input lipid features for training set by combining the pathophysiological mechanism of CAD and the progressive lipid feature screening strategy. Simplified module 40 is used to fit and construct a multivariate logistic regression model using the set of input lipid features and clinical covariates in the training set, and to independently evaluate the model performance using a validation set. After successful validation, the model parameters corresponding to the model are solidified in the form of a model parameter table, and finally a risk estimation model for calculating the CAD risk probability of the sample to be tested is constructed.
[0121] In addition, the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described method for constructing an early prediction model for coronary artery disease.
[0122] Figure 10 This is a schematic block diagram of the electronic device provided in an embodiment of this application. Figure 10As shown, the electronic device includes at least one processor 401, a memory 402, at least one network interface 403, and a user interface 405. The various components in the electronic device are coupled together via a bus system 404. It is understood that the bus system 404 is used to implement communication between these components. In addition to a data bus, the bus system 404 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in… Figure 10 The general will label all buses as bus systems.
[0123] The user interface 405 may include a monitor, keyboard, mouse, trackball, clicker, button, touchpad, or touch screen.
[0124] It is understood that memory 402 can be volatile memory or non-volatile memory, or both. Non-volatile memory can be read-only memory (ROM) or programmable read-only memory (PROM), used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM) and synchronous static random access memory (SSRAM). The memories described in the embodiments of this invention are intended to include, but are not limited to, these and any other suitable categories of memory.
[0125] In this embodiment of the invention, the memory 402 is used to store various types of data to support the operation of the electronic device 400. Examples of this data include: any executable program for operation on the electronic device 400, such as the operating system 4021 and application programs 4022; the operating system 4021 contains various system programs, such as the framework layer, core library layer, driver layer, etc., for implementing various basic services and handling hardware-based tasks. The application program 4022 may contain various applications, such as media players, browsers, etc., for implementing various application services. The method for constructing an early prediction model for coronary artery disease provided in this embodiment of the invention can be included in the application program 4022.
[0126] The methods disclosed in the above embodiments of the present invention can be applied to or implemented by processor 401. Processor 401 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in processor 401 or by instructions in the form of software. The processor 401 may be a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Processor 401 can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. General-purpose processor 401 may be a microprocessor or any conventional processor, etc. The steps of the method for constructing an early prediction model for coronary artery disease provided in the embodiments of the present invention can be directly manifested as execution by a hardware decoding processor, or execution by a combination of hardware and software modules in the decoding processor. The software modules may be located in a storage medium, which is located in a memory. The processor reads the information in the memory and combines it with its hardware to complete the steps of the aforementioned method.
[0127] In an exemplary embodiment, the electronic device 400 may be used by one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), or complex programmable logic devices (CPLDs) to perform the aforementioned method.
[0128] In summary, this invention simultaneously collects clinical covariates and fasting plasma lipidomics data from subjects, ensuring the accuracy and reliability of the test data through a quality control system. Lipid characteristics are preprocessed and standardized, and unified model input variables are constructed in conjunction with clinical indicators. Training and validation sets are divided, and core lipid characteristics are screened based on CAD pathological mechanisms, eliminating redundant variables and strengthening the correlation between characteristics and coronary artery lesions. A multivariate logistic regression model is constructed using the training set, and evaluated using the validation set to ensure the model's discriminative efficacy and generalization ability. Key model parameters are solidified and a parameter table is generated, establishing a standardized CAD risk estimation model. In clinical applications, only a single blood test combined with basic demographic information is needed to quickly quantify the early risk of disease, achieving efficient screening and accurate risk stratification, providing objective evidence for early warning of coronary artery disease and personalized clinical diagnosis and treatment.
[0129] The above embodiments are merely illustrative of the principles and effects of the present invention and are not intended to limit the invention. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in the present invention should still be covered by the claims of the present invention.
Claims
1. A method for constructing an early prediction model for coronary artery disease, characterized in that, The method includes: Clinical covariates and fasting plasma samples were collected from subjects in the coronary artery group and the healthy control group. Systematic lipidomics analysis was performed based on a pre-set quality control system to obtain lipid characteristic data for each subject. The clinical covariates included gender, age, and BMI. The lipid feature data is preprocessed to obtain the corresponding standardized lipid variables, and combined with the clinical covariates of the corresponding subjects to jointly construct the model input variables; The model input variables of each subject were divided into training set and validation set according to a set ratio. The set of input lipid features for the training set was determined by combining the pathophysiological mechanism of CAD and the progressive lipid feature screening strategy. A multivariate logistic regression model was constructed by fitting the set of lipid features and clinical covariates in the training set. The model performance was independently evaluated using a validation set. After successful validation, the model parameters were solidified in the form of a model parameter table, and finally a risk estimation model for calculating the CAD risk probability of the test sample was constructed.
2. The method according to claim 1, characterized in that, The systematic lipidomics analysis based on the pre-established quality control system obtains lipid characteristic data for each subject, including: A unique sample code was established for fasting plasma samples from subjects in the coronary artery group and the healthy control group, and a label was used to identify the testing status of each plasma sample; wherein, the testing status includes pending testing, under testing, completed testing, and retained sample; After inserting QC samples into each testing batch, fasting plasma samples are selected for lipidomics testing and data collection based on their unique sample codes and testing status. Data correction and quality control screening are completed based on the QC samples to obtain lipid characteristic data for each subject.
3. The method according to claim 1, characterized in that, The preprocessing of the lipid feature data to obtain the corresponding standardized lipid variables includes: Perform a logarithmic transformation on each lipid feature data to obtain the logarithmic lipid value; Based on the logarithmic lipid values of all fasting plasma samples from the coronary artery group and the healthy control group, the mean and standard deviation of each lipid feature were calculated, and the corresponding standardized lipid variables were obtained by Z-score standardization.
4. The method according to claim 1, characterized in that, The method, which combines the pathophysiological mechanisms of CAD with a progressive lipid feature screening strategy, determines the set of lipid features for the training set, including: For the standardized lipid variables in the training set, combined with the core pathophysiological mechanisms of CAD cholesterol accumulation, foam cell formation and plaque instability, the t-test was used to analyze the differences between the standardized lipid variables of the coronary artery group and the healthy control group, and lipid metabolites that showed significant differences between the two groups and were related to the pathological process of CAD were screened to construct a candidate lipid set. For the candidate lipid set, LASSO regression was used for sparsification and feature dimensionality reduction, and redundant variables were eliminated by combining CAD pathophysiological mechanism constraints, retaining key lipid features with clear pathological significance, thus forming a key lipid feature set. PLS-DA analysis was used to verify the discriminative power and visualize the set of key lipid features, calculate the variable importance projection value, and help evaluate the classification stability and pathological relevance of the features. Finally, the set of lipid features to be used for model construction was determined.
5. The method according to claim 1, characterized in that, The multivariate logistic regression model is constructed by fitting the input lipid feature set and clinical covariates from the training set. The model performance is independently evaluated using a validation set, including: Using the set of lipid features and clinical covariates from the training set as independent variables and the corresponding CAD disease labels of the subjects as dependent variables, a multivariate logistic regression model was constructed; wherein, the CAD disease labels include CAD disease status and healthy control status. After the multivariate logistic regression model has converged, the model intercept and the regression coefficients corresponding to each input variable are obtained. The lipid feature set of the validation set and clinical covariates are input into the converged model to obtain the CAD risk prediction results of the validation set. Based on the CAD disease label and CAD risk prediction results of each subject in the validation set, pre-set evaluation indicators were calculated to complete the independent evaluation of model performance; wherein, the pre-set evaluation indicators include accuracy, sensitivity, recall, F1 score and AUC value.
6. The method according to claim 1, characterized in that, After the verification is passed, the model parameters corresponding to the model are solidified in the form of a model parameter table, and a risk estimation model for calculating the CAD risk probability of the sample to be tested is finally constructed, including: When the evaluation indicators meet the preset threshold, the model is deemed to have passed the validation. The intercept and regression coefficients of each input variable of the validated model are extracted, organized into a structured model parameter table, and then solidified. Based on the solidified model parameter table, a risk estimation model is constructed to calculate the CAD risk probability of the sample to be tested.
7. The method according to claim 1, characterized in that, After the step of finally constructing a risk estimation model for calculating the CAD risk probability of the sample to be tested, the method further includes: Clinical covariates and fasting plasma samples were collected from the subjects to be tested. Lipomics analysis was performed on the fasting plasma samples to obtain the corresponding input lipid data. Based on the clinical covariates and the lipid data of the model, the risk estimation model and the model parameter table are used to make predictions and estimates, and the CAD risk probability of the subject to be tested is calculated. By using preset probability threshold classification rules, the CAD risk probability is mapped to the corresponding risk level to assist in clinical diagnosis and follow-up management.
8. A system for constructing an early prediction model for coronary artery disease, characterized in that, The system includes: The collection module is used to collect clinical covariates and fasting plasma samples from subjects in the coronary artery group and the healthy control group, and to perform systematic lipidomics analysis based on a pre-set quality control system to obtain lipid characteristic data for each subject; wherein, the clinical covariates include gender, age and BMI; The preprocessing module is used to preprocess the lipid feature data to obtain the corresponding standardized lipid variables, and combine them with the clinical covariates of the corresponding subjects to jointly construct the model input variables; The determination module is used to divide the model input variables of each subject into training set and validation set according to a set ratio, and to determine the set of input lipid features for the training set by combining the pathophysiological mechanism of CAD and the progressive lipid feature screening strategy. The training module is used to fit and construct a multivariate logistic regression model using the set of lipid features and clinical covariates in the training set. The model performance is independently evaluated using a validation set. After successful validation, the model parameters are fixed in the form of a model parameter table, and finally a risk estimation model for calculating the CAD risk probability of the test sample is constructed.
9. An electronic device, characterized in that, The electronic device includes a memory and a processor, wherein the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory to cause the processor to perform the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program that, when executed by a processor, implements the method of any one of claims 1 to 7.