A model for identifying subtypes of patients with high residual cholesterol and its construction method
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-12
- Publication Date
- 2026-08-14
AI Technical Summary
[0005]有鉴于此,本申请的目的是提供一种鉴别高残余胆固醇患者亚型的模型及其构建方法,用于解决现有残余胆固醇风险评估同质化、分层精度不足、筛查工具落地性差的问题
[0016]综上所述,本申请提供一种鉴别高残余胆固醇患者亚型的模型构建方法,该方法以残余胆固醇高于0.518mmol/L的人群作为目标人群;对目标人群的多维临床生化表型数据依次进行标准化处理和降维分析,并依托无监督聚类算法对高残余胆固醇人群进行亚型划分;根据划分亚型采用岭回归算法基于标准化的血脂指标构建多分类预测模型,最后通过十折交叉验证优化多分类预测模型的参数,得到鉴别高残余胆固醇患者亚型的模型。本申请通过引入临床常规血脂指标构建高残余胆固醇亚型鉴别模型,整体检测成本低廉,无需新增特殊检测项目,可适配临床常规筛查与大规模人群普查场景;此外,本申请还通过对多维临床生化表型数据进行标准化数据预处理、无监督聚类分型及岭回归建模,精准区分出代谢特征迥异的肝型、肠型高残余胆固醇亚型,有效提升了残余胆固醇人群风险评估的精细化程度,同时大幅优化了模型的稳定性、泛化能力,规避了传统建模易过拟合的缺陷。
Smart Images

Figure CN122388822B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of medical diagnostic technology, and in particular to a model for identifying subtypes of patients with high residual cholesterol and a method for constructing the model. Background Technology
[0002] Lipid metabolism abnormalities are widely recognized as a significant risk factor for the development and progression of cardiovascular disease. For a long time, clinical and epidemiological studies have primarily focused on the role of low-density lipoprotein cholesterol (LDL-C) in atherosclerotic cardiovascular disease, with LDL-C reduction serving as a core strategy for primary and secondary prevention. However, a growing body of research has found that even in populations with effectively controlled LDL-C levels, the risk of cardiovascular events remains significant; this phenomenon is termed "residual cardiovascular risk." Residual cholesterol (RC) is an important lipid metabolism indicator that has received widespread attention in recent years. Clinically, it is measured independently of the four main lipid markers and is typically calculated from routine lipid indicators, reflecting the cholesterol content of triglyceride-rich lipoproteins and their remnants. Numerous studies have shown that residual cholesterol is closely related to the progression of atherosclerosis, inflammatory responses, and the occurrence of adverse cardiovascular events, making it a crucial residual risk factor for cardiovascular disease. Compared to LDL-C, residual cholesterol is prevalent in the general population, and its elevation often occurs early in the disease process. Therefore, early identification and intervention of residual cholesterol are of significant public health importance for primary prevention of cardiovascular disease, especially for early risk screening in grassroots populations.
[0003] Currently, in clinical practice and epidemiological studies, the assessment of residual cholesterol mainly employs a simple calculation method based on routine blood lipid test results. This involves subtracting low-density lipoprotein cholesterol (LDL-C) and high-density lipoprotein cholesterol (HDL-C) from total cholesterol (TC) to obtain the residual cholesterol level. In risk assessment, individuals are typically categorized into "normal residual cholesterol" or "elevated residual cholesterol" based on whether their residual cholesterol exceeds a certain fixed threshold, and their cardiovascular disease risk is assessed accordingly. However, the existing technical approach has significant shortcomings: First, current methods generally treat elevated residual cholesterol as a homogeneous metabolic abnormality, assuming that all individuals with elevated residual cholesterol share similar pathological mechanisms and cardiovascular risks, while ignoring the objective fact that lipid metabolism itself is influenced by multiple physiological and pathological factors. Numerous studies have shown that blood lipid levels are regulated by multiple factors, including dietary intake, intestinal lipid absorption, hepatic lipid synthesis and clearance capacity, and insulin resistance status. The sources and mechanisms of elevated residual cholesterol may differ significantly among individuals. Secondly, while existing research suggests a link between residual cholesterol and cardiovascular disease risk, systematic studies and definitive conclusions are lacking regarding whether elevated residual cholesterol from different sources or with different metabolic characteristics corresponds to different cardiovascular risk levels. In particular, whether elevated residual cholesterol in natural population cohorts can be further subdivided into subtypes with different cardiovascular risk characteristics remains unclear. Thirdly, there is currently a lack of a technical means to perform early, simple, and objective typing of individuals with elevated residual cholesterol without increasing testing costs. Some existing studies have attempted to finely stratify lipid abnormalities using complex lipoprotein typing, metabolomics, or molecular biology detection methods; however, these methods often rely on expensive equipment and have complex testing procedures, making them difficult to promote and apply in primary healthcare institutions or large-scale population screening.
[0004] In summary, existing technologies for residual cholesterol risk assessment have the following main gaps: First, current methods generally treat individuals with high residual cholesterol as a homogeneous group, lacking effective means to identify and classify metabolic heterogeneity within the population; second, the lipid metabolism characteristics of different residual cholesterol subtypes and their differences in association with long-term cardiovascular disease risk have not been clearly defined; third, existing assessment tools mostly rely on complex detection indicators, lacking simple classification tools based on routine blood lipid indicators that are suitable for primary healthcare institutions and large-scale early screening of the population. Based on these gaps and deficiencies in existing technologies, this application proposes a high residual cholesterol subtype identification model and its construction method. Summary of the Invention
[0005] In view of this, the purpose of this application is to provide a model for identifying subtypes of patients with high residual cholesterol and a method for constructing the model, in order to solve the problems of homogeneity, insufficient stratification accuracy, and poor implementation of screening tools in existing residual cholesterol risk assessments.
[0006] To achieve the above-mentioned technical objectives, this application provides a method for constructing a model to identify subtypes of patients with high residual cholesterol, comprising the following steps: Step S1: The target population is individuals with residual cholesterol levels higher than 0.518 mmol / L. Step S2: The multidimensional clinical and biochemical phenotypic data of the target population are standardized and dimensionality reduced sequentially. An unsupervised clustering algorithm is used to classify the target population into subtypes, resulting in the liver type high residual cholesterol subtype and the intestinal type high residual cholesterol subtype. Step S3: Using the liver-type and intestinal-type high residual cholesterol subtypes as classification labels, a multi-class prediction model is constructed based on standardized blood lipid indicators using the ridge regression algorithm. The parameters of the multi-class prediction model are optimized through 10-fold cross-validation to obtain a model for identifying the subtypes of patients with high residual cholesterol.
[0007] Furthermore, blood lipid indicators include total cholesterol, triglycerides, high-density lipoprotein cholesterol, and low-density lipoprotein cholesterol.
[0008] Furthermore, multidimensional clinical biochemical phenotypic data include liver function indicators, kidney function indicators, lipid profile parameters, glucose metabolism indicators, and bilirubin metabolism indicators.
[0009] Furthermore, the standardization process includes at least one of Z-score standardization, min-max standardization, and decimal scaling standardization.
[0010] Furthermore, the dimensionality reduction analysis includes at least one of the UMAP algorithm, PCA principal component analysis, and t-SNE algorithm; the unsupervised clustering algorithm includes at least one of the K-means algorithm, hierarchical clustering, DBSCAN density clustering, and Gaussian mixture model (GMM) clustering.
[0011] Furthermore, the dimensionality reduction analysis was performed using the UMAP algorithm, with parameters set as follows: neighborhood parameter of 15–50 and minimum distance parameter of 0.1–0.5. This preserved the local similarity structure of the samples and the overall data distribution characteristics, enabling automatic spatial clustering and separation of populations with different metabolic phenotypes.
[0012] Furthermore, the criteria for determining the liver-type high residual cholesterol subtype are: at least three of the four indicators—alanine aminotransferase, aspartate aminotransferase, gamma-glutamyl transferase, and the ratio of alanine aminotransferase to aspartate aminotransferase—must be higher than the average level of the general population or the indicator level of the intestinal-type high residual cholesterol subtype. The criteria for identifying the intestinal type of high residual cholesterol subtype are: elevated levels of total bilirubin and indirect bilirubin; and normal liver function indicators.
[0013] Furthermore, the liver-type high residual cholesterol subtype corresponds to a high-risk population for liver disease, while the intestinal-type high residual cholesterol subtype corresponds to a high-risk population for cardiovascular disease and intestinal metabolic disease.
[0014] Furthermore, the model for identifying subtypes of patients with high residual cholesterol is as follows: η k =0.35-0.38×TC+0.54×TG-0.14×LDL-0.81×HDL; Predictive probability value of liver-type high residual cholesterol subtype = exp(η) k ) / [exp(η k )+exp(-η k )]; Predictive probability of intestinal type high residual cholesterol subtype = 1 - Predictive probability of liver type high residual cholesterol subtype; The subtype with the highest predicted probability value was used as the result for determining the high residual cholesterol subtype; among them, TC, TG, LDL, and HDL are all standardized blood lipid indicators; TC is the standardized total cholesterol indicator, TG is the standardized triglyceride indicator, LDL is the standardized low-density lipoprotein cholesterol indicator, and HDL is the standardized high-density lipoprotein cholesterol indicator.
[0015] This application provides a model for identifying subtypes of patients with high residual cholesterol by substituting standardized blood lipid indicators into the following formula: η k =0.35-0.38×TC+0.54×TG-0.14×LDL-0.81×HDL; Predictive probability value of liver-type high residual cholesterol subtype = exp(η) k ) / [exp(η k )+exp(-η k )]; Predictive probability of intestinal type high residual cholesterol subtype = 1 - Predictive probability of liver type high residual cholesterol subtype; The subtype with the highest predicted probability value was used as the result for determining the high residual cholesterol subtype; among them, TC, TG, LDL, and HDL are all standardized blood lipid indicators; TC is the standardized total cholesterol indicator, TG is the standardized triglyceride indicator, LDL is the standardized low-density lipoprotein cholesterol indicator, and HDL is the standardized high-density lipoprotein cholesterol indicator.
[0016] In summary, this application provides a model construction method for identifying subtypes of patients with high residual cholesterol. This method targets individuals with residual cholesterol levels higher than 0.518 mmol / L. The multidimensional clinical and biochemical phenotypic data of the target population are sequentially standardized and subjected to dimensionality reduction analysis. An unsupervised clustering algorithm is then used to classify the high residual cholesterol population into subtypes. Based on the subtype classification, a ridge regression algorithm is employed to construct a multi-class prediction model using standardized blood lipid indicators. Finally, the parameters of the multi-class prediction model are optimized through 10-fold cross-validation to obtain a model for identifying subtypes of patients with high residual cholesterol. This application constructs a subtype identification model for high residual cholesterol by introducing routine clinical lipid indicators. The overall detection cost is low, and no new special tests are required. It can be adapted to routine clinical screening and large-scale population surveys. In addition, this application also accurately distinguishes liver-type and intestinal-type high residual cholesterol subtypes with different metabolic characteristics by standardizing data preprocessing, unsupervised clustering and ridge regression modeling of multidimensional clinical biochemical phenotypic data. This effectively improves the precision of risk assessment for residual cholesterol populations, while significantly optimizing the stability and generalization ability of the model and avoiding the overfitting defects of traditional modeling.
[0017] Compared with existing technologies, this invention significantly improves the precision and clinical applicability of residual cholesterol-related cardiovascular risk assessment without increasing the complexity or cost of detection. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a UMAP dimensionality reduction clustering scatter plot provided in an embodiment of this application; Figure 2 Multidimensional clinical biochemical phenotypic radar chart provided for embodiments of this application; Figure 3 The subject operating characteristic curve provided in the embodiments of this application; Figure 4 Forest plot showing the association between liver / intestinal high residual cholesterol subtypes and the risk of liver, intestinal, and cardiovascular diseases, provided in the embodiments of this application. Detailed Implementation
[0020] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. Based on the embodiments in this application specification, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection claimed in this application.
[0021] In the description of this application, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," indicating orientation or positional relationships, are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation on this application. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0022] Unless otherwise expressly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to fixed connections, detachable connections, or integral connections; they can refer to mechanical connections or electrical connections; they can refer to direct connections or indirect connections through an intermediate medium; and they can refer to the internal connection between two components. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances.
[0023] The raw materials used in this invention are not particularly restricted in their source; they can be purchased on the market or prepared using conventional methods known to those skilled in the art.
[0024] This application provides a method for constructing a model to identify subtypes of patients with high residual cholesterol, comprising the following steps: Step S1 targets individuals with residual cholesterol levels above 0.518 mmol / L. It should be noted that the value of 0.518 mmol / L comes from the article "Low Remnant Cholesterol and In-Hospital Bleeding Risk After Ischemic Stroke or Transient Ischemic Attack," which states that "in the Chinese population, residual cholesterol levels above 0.518 mmol / L can lead to serious cardiovascular disease." Step S2: The multidimensional clinical and biochemical phenotypic data of the target population are standardized and dimensionality reduced sequentially. An unsupervised clustering algorithm is used to classify the target population into subtypes, resulting in the liver type high residual cholesterol subtype and the intestinal type high residual cholesterol subtype. Step S3: Using the liver-type and intestinal-type high residual cholesterol subtypes as classification labels, a multi-class prediction model is constructed based on standardized blood lipid indicators using the ridge regression algorithm. The parameters of the multi-class prediction model are optimized through 10-fold cross-validation to obtain a model for identifying the subtypes of patients with high residual cholesterol.
[0025] In some embodiments, blood lipid indicators include total cholesterol, triglycerides, high-density lipoprotein cholesterol, and low-density lipoprotein cholesterol.
[0026] It should be noted that this application can construct a model for identifying subtypes of patients with high residual cholesterol using only four conventional blood lipid indicators, without the need for additional biochemical tests, genetic tests, or imaging data. It can achieve further subtyping of people with elevated residual cholesterol without increasing testing costs, and has outstanding practicality and promotional value.
[0027] In some embodiments, multidimensional clinical biochemical phenotypic data include liver function indicators, kidney function indicators, lipid profile parameters, glucose metabolism indicators, and bilirubin metabolism indicators.
[0028] In some embodiments, the standardization process includes at least one of Z-score standardization, min-max standardization, and fractional scaling standardization; the dimensionality reduction analysis includes at least one of UMAP algorithm, PCA principal component analysis, and t-SNE algorithm; and the unsupervised clustering algorithm includes at least one of K-means algorithm, hierarchical clustering, DBSCAN density clustering, and Gaussian mixture model (GMM) clustering.
[0029] In some preferred embodiments, the dimensionality reduction analysis is performed using the UMAP algorithm, with the parameters set as follows: neighborhood parameter of 15 to 50, and minimum distance parameter of 0.1 to 0.5.
[0030] It should be noted that the UMAP algorithm can effectively preserve the local correlation features and overall distribution structure of high-dimensional data, solving the problems of information loss and poor clustering effect in traditional dimensionality reduction analysis. Furthermore, the parameters provided in this embodiment are optimal, accurately mapping the original high-dimensional clinical feature data to a two-dimensional visual feature space. The dimensionality-reduced data allows individuals with similar metabolic characteristics and biochemical phenotypes to automatically cluster in the two-dimensional space, while individuals with different metabolic patterns are spatially separated. This provides a clear and stable geometric structure basis for subsequent unsupervised clustering and typing, maximizing the preservation of potential subtype differences in individuals with high residual cholesterol.
[0031] In some embodiments, the criteria for determining the hepatic type of high residual cholesterol subtype are: at least three of the four indicators—alanine aminotransferase, aspartate aminotransferase, gamma-glutamyl transferase, and the ratio of alanine aminotransferase to aspartate aminotransferase—are higher than the average level of the general population or the indicator levels of the intestinal type of high residual cholesterol subtype; the criteria for determining the intestinal type of high residual cholesterol subtype are: elevated levels of total bilirubin and indirect bilirubin; and normal liver function indicators.
[0032] In some embodiments, the hepatic type of high residual cholesterol subtype corresponds to a high-risk population for liver disease, while the intestinal type of high residual cholesterol subtype corresponds to a high-risk population for cardiovascular disease and intestinal metabolic disease.
[0033] It should be noted that existing technologies generally assume that individuals with elevated residual cholesterol have homogeneous metabolic characteristics and cardiovascular disease risk, and therefore often conduct risk assessments based solely on a single residual cholesterol threshold. However, it has been found that lipid metabolism involves multiple steps, including intestinal lipid absorption, hepatic lipid synthesis and clearance, and lipoprotein transport, resulting in significant heterogeneity in the pathological sources and metabolic mechanisms of elevated residual cholesterol among different individuals. This invention subdivides individuals with high residual cholesterol into subtypes to accurately identify subgroups with differentiated risk characteristics, thereby achieving refined risk assessment. Among them, the hepatic subtype has a higher risk of developing liver disease in the future, while the intestinal subtype has a higher risk of developing cardiovascular disease and intestinal metabolic diseases in the future. Further research has verified that the proportion of abnormal triglycerides, low-density lipoprotein cholesterol, and high-density lipoprotein cholesterol in the intestinal subtype with high residual cholesterol is lower than that in the general hyperlipidemia population, and their hyperlipidemia detection rate is not significantly different from that of the healthy population, but this population has a significantly increased risk of developing cardiovascular disease in the future. The results indicate that traditional risk assessment methods relying on a single blood lipid index significantly underestimate the cardiovascular disease risk in this population. This model can serve as an early screening tool based on routine blood lipid testing, is suitable for widespread adoption at the grassroots level, and provides reliable support for precise prevention and tiered management of cardiovascular diseases.
[0034] In some embodiments, the model for identifying patient subtypes with high residual cholesterol is: η k =0.35-0.38×TC+0.54×TG-0.14×LDL-0.81×HDL; Predictive probability value of liver-type high residual cholesterol subtype = exp(η) k ) / [exp(η k )+exp(-η k )]; Predictive probability of intestinal type high residual cholesterol subtype = 1 - Predictive probability of liver type high residual cholesterol subtype; The subtype with the highest predicted probability value was used as the final result for determining the high residual cholesterol subtype. Among them, TC, TG, LDL, and HDL are all standardized blood lipid indicators. TC is the standardized total cholesterol indicator, TG is the standardized triglyceride indicator, LDL is the standardized low-density lipoprotein cholesterol indicator, and HDL is the standardized high-density lipoprotein cholesterol indicator.
[0035] This application provides a model for identifying subtypes of patients with high residual cholesterol by substituting standardized blood lipid indicators into the following formula: η k =0.35-0.38×TC+0.54×TG-0.14×LDL-0.81×HDL; Predictive probability value of liver-type high residual cholesterol subtype = exp(η) k ) / [exp(η k )+exp(-η k )]; Predictive probability of intestinal type high residual cholesterol subtype = 1 - Predictive probability of liver type high residual cholesterol subtype; The subtype with the highest predicted probability value was used as the final result for determining the high residual cholesterol subtype. Among them, TC, TG, LDL, and HDL are all standardized blood lipid indicators. TC is the standardized total cholesterol indicator, TG is the standardized triglyceride indicator, LDL is the standardized low-density lipoprotein cholesterol indicator, and HDL is the standardized high-density lipoprotein cholesterol indicator.
[0036] It should be noted that the model calculation process of this invention is simple and easy to implement, and it is suitable for primary medical institutions and large-scale population screening. While achieving accurate identification of subtypes of residual cholesterol elevation, it provides reliable technical support for early risk stratification and precise prevention of cardiovascular diseases.
[0037] The applicant further provides the following specific embodiments to describe the present invention. It should be noted that these embodiments are merely descriptive and do not limit the present invention in any way.
[0038] Example 1: This invention provides a method for constructing a model to identify subtypes of patients with high residual cholesterol, comprising the following steps: Step S1, Source of research subjects and data collection: This embodiment is based on a natural population cohort and a clinical oral glucose tolerance test (OGTT) validation cohort in some parts of China to carry out model construction and validation. All research protocols comply with medical ethics, and all research subjects in the clinical OGTT validation cohort signed informed consent forms. Among them, the research subjects provided by the natural population cohort in some parts of China are from general community natural populations in some parts of China, including but not limited to Guangdong Province, Guangxi Zhuang Autonomous Region, Fujian Province, and Hainan Province. A total of 102,932 research subjects were recruited through a nationwide multi-center collaborative recruitment method, which served as the core cohort for model training and validation. The population of the clinical OGTT validation cohort came from the Chashan Town Community Health Service Center in Dongguan City, Guangdong Province, and a total of 640 research subjects were recruited for model validation. All research subjects underwent standardized data collection, and the collected information included the following three categories: (1) general demographic information, including age, gender, etc.; (2) clinical information, including past medical history, cardiovascular disease outcomes during follow-up, etc.; (3) laboratory test information, including clinical biochemical characteristic data including four blood lipid indicators. The four lipid indicators included were: 1) Total cholesterol (TC); 2) Triglycerides (TG); 3) High-density lipoprotein cholesterol (HDL-C); and 4) Low-density lipoprotein cholesterol (LDL-C). All laboratory tests were performed using a standardized quality control process to ensure the accuracy and comparability of the data.
[0039] Step S2, calculation of residual cholesterol and identification of individuals with high residual cholesterol: Based on the blood lipid data of four indicators collected in step S1 from natural population cohorts in some parts of China, the residual cholesterol (RC) level of all study subjects was calculated using the formula: RC = TC - (LDL - C) - (HDL - C). 0.518 mmol / L was used as the optimal threshold for determining high residual cholesterol; that is, an individual was included in the high residual cholesterol group when their RC level was higher than this threshold. It should be noted that the value of 0.518 mmol / L comes from the article "Low Remnant Cholesterol and In-Hospital Bleeding Risk After Ischemic Stroke or Transient Ischemic Attack," which states that "in the Chinese population, residual cholesterol levels above 0.518 mmol / L can lead to serious cardiovascular disease."
[0040] Step S3, Dimensionality reduction mapping of multidimensional clinical features based on the UMAP algorithm: Clinical biochemical characteristic data, processed using the Z-score normalization method, including liver function enzyme indicators, renal function metabolic indicators, whole blood lipid profile parameters, and glucose and bilirubin metabolism-related indicators, were used as model input variables. The Unified Manifold Approximation and Projection Algorithm (UMAP) was introduced to perform dimensionality reduction mapping of the high-dimensional feature space. By setting the neighborhood parameter to 15–50 and the minimum distance parameter to 0.1–0.5, the original high-dimensional data was mapped to a two-dimensional feature space, thus preserving the local similarity structure between individuals while also considering the overall data distribution characteristics. The two-dimensional clinical biochemical characteristic data representation obtained after UMAP dimensionality reduction allows individuals with similar metabolic characteristics and biochemical patterns to naturally cluster in space, providing a clear geometric structural basis for subsequent subtyping.
[0041] Step S4, Precise subtype classification of high residual cholesterol populations based on unsupervised clustering: Based on the two-dimensional clinical biochemical feature data obtained in step S3, an unsupervised clustering algorithm was used to automatically classify the high residual cholesterol target population. In this embodiment, the K-means clustering algorithm was selected for subtyping analysis. Simultaneously, to determine the optimal number of clusters, this embodiment used either the silhouette coefficient method or the elbow method to jointly evaluate the cluster compactness, separation, and model stability under different numbers of clusters (k=2, 3, 4, 5). The two-dimensional clinical biochemical feature data were finally clustered into two independent clusters with clear boundaries and no overlap, such as... Figure 1 As shown, significant spatial shifts were observed between the two population groups along the first and second principal axes of the UMAP, confirming an essential biological difference in their overall clinical metabolic phenotypes. Figure 2 It can be seen that individuals within the same cluster show a high degree of consistency in multiple clinical biochemical indicators, while individuals in different clusters exhibit statistically significant differences in liver function enzyme indicators and bilirubin metabolism-related indicators, thus forming an interpretable and definitive biological classification basis. Based on the differences between the two groups, the target population for high residual cholesterol analysis is precisely divided into hepatic high residual cholesterol subtype and intestinal high residual cholesterol subtype. The specific criteria and phenotypic characteristics are as follows: (1) Hepatic type high residual cholesterol subtype: This subtype is characterized by a significant increase in liver damage and metabolic stress-related indicators. Specifically, at least three of the four indicators—alanine aminotransferase, aspartate aminotransferase, gamma-glutamyl transferase, and the alanine aminotransferase-to-aspartate aminotransferase ratio—are higher than the average level of the general population or the level of the intestinal type high residual cholesterol subtype. The formation mechanism of this subtype is usually closely related to the disorder of liver lipid synthesis and secretion, suggesting that the increase in residual cholesterol may mainly be due to the accumulation of lipid residues caused by abnormal synthesis or secretion of very low density lipoprotein in the liver. From the perspective of pathogenesis, the core cause of the increase in residual cholesterol in this subtype is the disorder of liver lipid synthesis and secretion, mainly caused by excessive synthesis and abnormal secretion of very low density lipoprotein in the liver, which leads to the accumulation of a large number of lipid residues and ultimately results in an abnormal increase in residual cholesterol. The core pathological target is abnormal liver metabolic function.
[0042] (2) Intestinal type high residual cholesterol subtype: The core characteristic of this subtype is a specific elevation of bilirubin metabolism indicators without obvious liver damage phenotype. Specifically, it is characterized by elevated bilirubin metabolism-related indicators, while liver damage enzyme indicators are within the normal range or significantly lower than those of the hepatic type. Specifically, the total bilirubin and indirect bilirubin levels in this subtype are significantly higher than those in the hepatic type high residual cholesterol subtype, while no elevation is observed in the four indicators of alanine aminotransferase, aspartate aminotransferase, gamma-glutamyl transferase, and the ratio of alanine aminotransferase to aspartate aminotransferase. This phenotype suggests that the elevated residual cholesterol is more likely related to abnormal enterohepatic circulation and intestinal lipid absorption processes, such as excessive absorption or delayed clearance of intestinal chylomicron debris, rather than simply caused by hepatocyte damage or decreased liver metabolic function.
[0043] Step S5, Construction of a multi-class prediction model based on ridge regression analysis: While the UMAP method can effectively reveal the subtyping structure of the target population with high residual cholesterol, its computational process is complex and relies on multidimensional nonlinear dimensionality reduction, making it difficult to promote and apply in primary healthcare institutions or actual clinical settings. To improve the operability and practicality of the model, this invention further employs ridge regression analysis to construct a simplified multi-class prediction model. During model construction, TC, TG, HDL-C, and LDL-C in the target population with high residual cholesterol from natural population cohorts in some regions of China are first processed using Z-score standardization to eliminate the influence of dimensional differences on the model. Using the high residual cholesterol subtype results obtained by the UMAP method as reference labels, a multi-class prediction model is constructed using ridge regression analysis. Then, the regularization parameters in the multi-class prediction model are optimized using ten-fold cross-validation to achieve a balance between model fit and generalization ability, ultimately yielding the following ridge regression model (a model for identifying subtypes of patients with high residual cholesterol): η k =0.35-0.38×TC+0.54×TG-0.14×LDL-0.81×HDL; Predictive probability value of liver-type high residual cholesterol subtype = exp(η) k ) / [exp(η k )+exp(-η k )]; Predictive probability of intestinal type high residual cholesterol subtype = 1 - Predictive probability of liver type high residual cholesterol subtype; Where, η k The values represent the linear predictive latent variables for blood lipids; TC, TG, HDL, and LDL are the standardized blood lipid indicators; TC is the standardized total cholesterol indicator, TG is the standardized triglyceride indicator, LDL is the standardized low-density lipoprotein cholesterol indicator, and HDL is the standardized high-density lipoprotein cholesterol indicator.
[0044] When using this model, standardized blood lipid data from different cohorts are input into the trained ridge regression model. Predicted probabilities for the liver-type and intestinal-type high residual cholesterol subtypes are calculated, with probabilities ranging from 0 to 1. For the same individual, the subtype with the highest predicted probability is selected as the final determination of the high residual cholesterol subtype.
[0045] Verification Example 1: To verify the stability and generalization ability of the high residual cholesterol genotyping model proposed in this invention based on natural population cohorts from various regions of China under different population backgrounds, this invention further selected an independent OGTT validation cohort with clear clinical phenotypic information for external validation analysis, while using natural population cohorts from various regions of China as validation controls. The OGTT validation cohort was not involved in the model construction stage, effectively avoiding the risk of overfitting, thus objectively evaluating the application value of the method in real-world clinical scenarios. Furthermore, OGTT has a clear metabolic state stratification and standardized detection procedure, which can reflect the true characteristics of individuals with metabolic abnormalities in clinical populations, thus providing an important basis for evaluating the stability of the model in clinical applications.
[0046] Validation method: The blood lipid indicators of the subjects in the natural population cohort and the clinical OGTT validation cohort in some regions of China were standardized. The standardized blood lipid indicators were substituted into the model for identifying the subtype of patients with high residual cholesterol, and the probability of individual subtype classification was calculated and the subtype was identified. Receiver operating characteristic (ROC) curves were plotted based on the discrimination results, and the model's discrimination ability was evaluated through curve analysis.
[0047] Validation results: The ridge regression model and the UMAP classification results show high consistency. For example... Figure 3 As shown, in natural population cohorts in some parts of China, the area under the receiver operating characteristic (AUC) of the model's discriminative ability was approximately 81.6% (95% CI: 79.3%–83.7%); in the OGTT validation cohort, the AUC further improved to approximately 89.8% (95% CI: 81.6%–97.3%). These results demonstrate that the ridge regression model provided by this invention exhibits good discriminative ability and stability in different population contexts.
[0048] Furthermore, in natural population cohorts and clinical OGTT validation cohorts in selected regions of China, the risk differences in different diseases among individuals with hepatic hyperresidual cholesterol and those with intestinal hyperresidual cholesterol were assessed. The analysis results are as follows: Figure 4As shown, in two independent population cohorts, individuals identified as having an elevated residual cholesterol (RCC) subtype (hereinafter referred to as RCC individuals) consistently exhibited an increased risk of cardiovascular and gut-related diseases. In a general population cohort in parts of China, RCC individuals showed a slightly elevated risk of cardiovascular disease, with an odds ratio (OR) of approximately 1.08–1.10; in the clinical OGTT validation cohort, this risk was further increased, with an OR of approximately 1.6–2.0. These results suggest that the RCC subtype exhibits a stable trend of increased cardiovascular risk across different populations. Furthermore, analysis of gut-related diseases revealed a positive correlation between the RCC subtype and the risk of gut diseases. In a general population cohort in parts of China, the hazard ratio (HCR) for gut-related diseases in RCC individuals was approximately 1.04–1.06; in the clinical OGTT validation cohort, the HCR was approximately 1.7–2.2, suggesting that this subtype has a more significant susceptibility to gut diseases in individuals with clinical metabolic abnormalities. Individuals with the hepatic type of high residual cholesterol (referred to as hepatic individuals) primarily experience liver-related disease risks. In general population cohorts in some parts of China, hepatic individuals have a slightly increased risk of developing liver-related diseases (OR approximately 1.03); however, in clinical OGTT-validated cohorts, this risk is significantly increased (OR approximately 1.3–1.8), indicating a more significant association between the hepatic type of high residual cholesterol and liver damage in the context of metabolic abnormalities. In contrast, the association between the hepatic type of high residual cholesterol and cardiovascular and intestinal-related diseases is relatively low.
[0049] Therefore, the above analysis results based on natural populations and clinical OGTT populations in parts of China demonstrate that the model construction method for identifying subtypes of patients with high residual cholesterol proposed in this invention, and the resulting ridge regression model, can stably distinguish subtypes with differentiated disease risk profiles. Specifically, the intestinal type of high residual cholesterol is mainly associated with cardiovascular and intestinal-related disease risks, while the hepatic type is closely associated with liver-related diseases. This classification method can achieve precise stratification of multi-system disease risks, possessing good clinical predictive value and clear biological rationale.
[0050] In summary, the model and its construction method for identifying subtypes of patients with high residual cholesterol provided in this application have the following advantages: (1) To achieve objective typing of individuals with high residual cholesterol in the natural population and reveal the subtype-differentiated disease risk: This invention can perform unbiased objective typing of individuals with elevated residual cholesterol in the natural population, accurately identify two subtypes of high residual cholesterol with significant differences in lipid metabolism characteristics; by clarifying the risk association differences between the two subtypes and cardiovascular diseases and other multi-system diseases, it can achieve accurate stratified identification of high-risk groups and provide a scientific basis for early disease intervention. At the same time, this model is constructed based on only four conventional blood lipid indicators: total cholesterol, triglycerides, high-density lipoprotein cholesterol, and low-density lipoprotein cholesterol. No additional biochemical, genetic or imaging tests are required, and the testing cost is not increased. It is suitable for primary medical institutions and large-scale population epidemiological screening scenarios. The overall calculation logic of the model is simple, highly reproducible, and easy to promote and implement. It can quickly complete the identification of residual cholesterol subtypes and provide key technical support for early prevention and individualized stratified management of cardiovascular diseases.
[0051] (2) Ridge regression modeling is adopted to improve model stability and cross-population generalization ability: This invention introduces the ridge regression algorithm in the model construction stage. By using regularization constraints, the interference of multicollinearity among lipid indicators on the modeling results is reduced, which effectively improves the stability and generalization ability of the model in different population cohorts. This design is specifically adapted to the actual application scenario where clinical lipid indicators are highly correlated, ensuring that the model can be stably applied across cohorts. This is the core technical guarantee for improving the clinical applicability of the model.
[0052] (3) Objective subtyping based on standardized blood lipid indicators and probability maximization rules: This invention establishes a standardized and reproducible subtype discrimination mechanism: First, the original blood lipid indicators are standardized and then input into the trained model to calculate the predicted probability of an individual belonging to different subtypes. The subtype corresponding to the maximum probability is used as the final judgment result. This method eliminates the need for manual subjective threshold setting and effectively ensures the objectivity, consistency and repeatability of subtype classification.
[0053] (4) Model adaptability to primary healthcare and large-scale public health screening scenarios: The model for identifying subtypes of patients with high residual cholesterol constructed in this invention has the characteristics of few input variables, simple calculation and low dependence on software and hardware. It can be widely used in primary healthcare institutions, physical examination centers and large-scale public health screening, and has strong clinical application value and promotion potential.
[0054] The above are merely preferred embodiments of this application and are not intended to limit the present invention. Although this application has been described in detail with reference to examples, those skilled in the art can still modify the technical solutions described in the foregoing examples or make equivalent substitutions for some of the technical features. However, any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A method for constructing a model to identify subtypes of patients with high residual cholesterol, characterized in that, Includes the following steps: Step S1: The target population is individuals with residual cholesterol levels higher than 0.518 mmol / L. Step S2: The multidimensional clinical biochemical phenotypic data of the target population are standardized and dimensionality reduced sequentially. An unsupervised clustering algorithm is used to classify the target population into subtypes, resulting in the liver type high residual cholesterol subtype and the intestinal type high residual cholesterol subtype. Step S3: Using the liver-type high residual cholesterol subtype and the intestinal-type high residual cholesterol subtype as classification labels, a multi-class prediction model is constructed based on standardized blood lipid indicators using the ridge regression algorithm. The parameters of the multi-class prediction model are optimized through ten-fold cross-validation to obtain a model for identifying the subtype of patients with high residual cholesterol. The multidimensional clinical biochemical phenotypic data include liver function indicators, kidney function indicators, lipid profile parameters, glucose metabolism indicators, and bilirubin metabolism indicators. The criteria for determining the hepatic type of high residual cholesterol subtype are: the ratio of alanine aminotransferase to aspartate aminotransferase, and at least three of the four indicators (amylase, aspartate aminotransferase, and gamma-glutamyl transferase) must be higher than the average level of the general population or the indicator level of the intestinal type of high residual cholesterol subtype. The criteria for identifying the intestinal type of high residual cholesterol subtype are: elevated levels of total bilirubin and indirect bilirubin, and no abnormalities in liver function indicators.
2. The method for constructing a model to identify subtypes of patients with high residual cholesterol according to claim 1, characterized in that, The blood lipid indicators include total cholesterol, triglycerides, high-density lipoprotein cholesterol, and low-density lipoprotein cholesterol.
3. The method for constructing a model to identify subtypes of patients with high residual cholesterol according to claim 1, characterized in that, The standardization process includes at least one of Z-score standardization, min-max standardization, and decimal scaling standardization.
4. The method for constructing a model to identify subtypes of patients with high residual cholesterol according to claim 1, characterized in that, The dimensionality reduction analysis includes at least one of the UMAP algorithm, PCA principal component analysis, and t-SNE algorithm; the unsupervised clustering algorithm includes at least one of the K-means algorithm, hierarchical clustering, DBSCAN density clustering, and Gaussian mixture model (GMM) clustering.
5. The method for constructing a model to identify patient subtypes with high residual cholesterol according to claim 4, characterized in that, The dimensionality reduction analysis was performed using the UMAP algorithm, with the following parameters: neighborhood parameter of 15–50 and minimum distance parameter of 0.1–0.
5.
6. The method for constructing a model to identify patient subtypes with high residual cholesterol according to claim 1, characterized in that, The liver-type high residual cholesterol subtype corresponds to high-risk groups for liver disease, while the intestinal-type high residual cholesterol subtype corresponds to high-risk groups for cardiovascular disease and intestinal metabolic disease.
7. The method for constructing a model to identify subtypes of patients with high residual cholesterol according to claim 1, characterized in that, The model for identifying subtypes of patients with high residual cholesterol is as follows: or k =0.35-0.38×TC+0.54×TG-0.14×LDL-0.81×HDL; Predictive probability value of liver-type high residual cholesterol subtype = exp(η) k ) / [exp(η k )+exp(-η k )]; Predictive probability of intestinal type high residual cholesterol subtype = 1 - Predictive probability of liver type high residual cholesterol subtype; The subtype with the highest predicted probability value was used as the result for determining the high residual cholesterol subtype; among them, TC, TG, LDL, and HDL are all standardized blood lipid indicators; TC is the standardized total cholesterol indicator, TG is the standardized triglyceride indicator, LDL is the standardized low-density lipoprotein cholesterol indicator, and HDL is the standardized high-density lipoprotein cholesterol indicator.
Citation Information
Patent Citations
Method for constructing coronary artery disease risk discrimination model based on conventional physical examination indexes
CN115132363A
Data-driven model for predicting progression risk of future diabetes related diseases in early stage of diabetes and construction method thereof
CN121393898A