Diabetes vascular calcification early warning method based on molecular marker detection
By detecting the expression levels of specific genes in diabetic patients and combining them with machine learning models, the accuracy problem of vascular calcification detection in existing technologies has been solved, enabling early warning and personalized management, and reducing the risk of cardiovascular events.
Patent Information
- Application Number
- CN202511055681.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-11-04
AI Technical Summary
Current technologies cannot efficiently and accurately detect vascular calcification in diabetic patients, leading to frequent missed or misdiagnosed cases, which affects the early warning and intervention of cardiovascular events.
By acquiring blood data, blood biochemical index data, and medical data from diabetic patients, plasma was separated and ribonucleic acid was extracted. The gene expression levels of bone sialic acid protein gene, sclerosingin gene, fibronectin gene, type I collagen alpha chain gene, and integrin-binding sialic acid protein gene were detected. Fusion analysis was performed using random forest model and multilayer perceptron model to predict the risk of vascular calcification.
It significantly improves the diagnostic accuracy of diabetic vascular calcification, enables early detection of disease risks, reduces the probability of serious cardiovascular complications, and provides personalized treatment recommendations.
Smart Images

Figure CN120895244A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of early diagnosis of diabetic vascular complications, and relates to a diabetic vascular calcification early warning method based on molecular marker detection. BACKGROUND
[0002] Diabetes mellitus (DM) is a group of metabolic diseases characterized by high blood sugar. High blood sugar is caused by a deficiency in insulin secretion or impaired biological action, or both. The incidence of DM in today's society is increasing year by year, and the phenomenon of vascular calcification (VC) is widespread among DM patients. The incidence of vascular calcification in diabetic patients is much higher than that in non-diabetic patients, and is a key cause of cardiovascular events in diabetic patients. According to statistics, the incidence of vascular calcification in diabetic patients is 2-3 times higher than that in non-diabetic patients, and early diagnosis and timely intervention can reduce the risk of cardiovascular events by 30%-50%.
[0003] In related technologies, detection is performed based on imaging detection methods such as X-ray and computed tomography, but this method is subject to a resolution threshold and is difficult to identify early microcalcification foci in the blood vessel wall. Or a detection method based on a single biomarker, but the above method lacks sufficient diagnostic specificity and sensitivity, and is prone to misdiagnosis or missed diagnosis in actual application.
[0004] Therefore, how to efficiently and accurately detect whether the blood vessels of a diabetic patient are calcified has become a problem to be solved. SUMMARY
[0005] Therefore, the embodiments of the present application provide a diabetic vascular calcification early warning method based on molecular marker detection, which at least solves the problem that related technologies cannot efficiently and accurately detect whether the blood vessels of a diabetic patient are calcified.
[0006] According to a first aspect of the embodiments of the present application, a diabetic vascular calcification early warning method based on molecular marker detection is provided, comprising: obtaining blood data, blood biochemical index data and diagnosis and treatment data of a diabetic patient; and separating the blood data to obtain separated plasma; placing the plasma in a sterile cryogenic tube and storing it in a refrigerator at a preset temperature; and extracting ribonucleic acid from the cryogenic tube in the refrigerator; when the concentration and purity of the ribonucleic acid meet the preset conditions, obtaining gene expression level data of target genes based on the ribonucleic acid; the target genes are bone sialoprotein gene, sclerostin gene, fibronectin gene, collagen type I alpha chain gene and integrin-binding sialoprotein gene, respectively. The gene expression level data, blood biochemical index data and diagnosis and treatment data are fused to obtain fusion data, and the fusion data are input into a random forest model and a multilayer perception model respectively to obtain a first prediction result and a second prediction result; The first prediction result and the second prediction result are weighted and summed to obtain a risk value of calcification of blood vessels of the diabetic patient, and corresponding early warning is performed through the risk value.
[0007] According to a second aspect of the embodiment of the present application, an electronic device is provided, comprising a processor, a memory, a communication interface and a communication bus, the processor, the memory and the communication interface complete communication with each other through the communication bus; the memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform the operation corresponding to the method according to the first aspect.
[0008] According to a third aspect of the embodiment of the present application, a computer storage medium is provided, and the computer storage medium stores a computer program, and the program is executed by a processor to realize the method according to the first aspect.
[0009] According to the scheme provided by the embodiment of the present application, blood data, blood biochemical index data and diagnosis and treatment data of a diabetic patient are acquired; the blood data is separated to obtain separated plasma; the plasma is placed in a sterile cryogenic tube and stored in a refrigerator at a preset temperature; ribonucleic acid is extracted from the cryogenic tube in the refrigerator; when the concentration and purity of the ribonucleic acid meet preset conditions, gene expression level data of target genes are obtained based on the ribonucleic acid; the target genes are bone sialoprotein gene, sclerostin gene, fibronectin gene, collagen type I alpha chain gene and integrin-binding sialoprotein gene; the gene expression level data, blood biochemical index data and diagnosis and treatment data are fused to obtain fused data, and the preprocessed fused data is input into a random forest model and a multilayer perception model to obtain first prediction results and second prediction results; the first prediction results and the second prediction results are weighted and summed to obtain a risk value of vascular calcification of the diabetic patient, and corresponding early warning is performed through the risk value. In this process, by detecting the expression levels of the five key genes, i.e., bone sialoprotein gene, sclerostin gene, fibronectin gene, collagen type I alpha chain gene and integrin-binding sialoprotein gene, the molecular signal of disease occurrence can be captured in the early stage of vascular calcification of a diabetic patient, the risk can be discovered in advance, and the risk of serious diabetic cardiovascular complications can be significantly reduced. The risk of vascular calcification of a diabetic patient is evaluated by integrating gene expression data, blood biochemical index data and diagnosis and treatment data, compared with a traditional single detection method, the diagnostic accuracy is greatly improved, and the uncertainty of clinical decision-making is reduced. The random forest model and the multilayer perception model are combined to predict the final risk value, and the prediction accuracy is improved. BRIEF DESCRIPTION OF DRAWINGS
[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor. Figure 1 A flowchart of a diabetic vascular calcification early warning method based on molecular marker detection provided by the embodiment of the present application; Figure 2 A structural schematic diagram of an electronic device provided by the embodiment of the present application. DETAILED DESCRIPTION
[0011] In order to make the objects, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. The following embodiments are used to describe the present application, but not to limit the scope of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work belong to the scope of protection of the present application.
[0012] In the following description, "some embodiments" are related to a subset of all possible embodiments, but it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.
[0013] It should be noted that the terms "first", "second", "third" involved in the embodiments of the present application are only to distinguish similar objects, and do not represent the specific order of the objects. It can be understood that "first", "second", "third" can be interchanged with specific order or sequence as allowed, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.
[0014] Those skilled in the art can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as generally understood by those skilled in the art to which the embodiments of the present application belong. It should also be understood that terms such as those defined in general dictionaries should be understood as having meanings consistent with those in the prior art, and should not be interpreted in an idealized or overly formal sense unless specifically defined as such.
[0015] Figure 1 A flowchart of a diabetes vascular calcification early warning method based on molecular marker detection provided by the embodiments of the present application, the diabetes vascular calcification early warning method based on molecular marker detection provided by the embodiments of the present application can be executed by an electronic device, which can be a computer, a server, etc.
[0016] As shown in Figure 1 The diabetes vascular calcification early warning method based on molecular marker detection comprises the following steps. S101, obtaining blood data, blood biochemical index data and diagnosis and treatment data of a diabetes patient; and separating the blood data to obtain separated plasma.
[0017] In the embodiment of the present application, 5-10 mL of blood is taken from a diabetic patient by venous blood sampling, and blood data is obtained in a blood collection tube containing ethylenediamine tetraacetic acid anticoagulant. The blood plasma is separated by centrifugation at 3000 rpm for 10 minutes within 30 minutes after collection. At the same time, blood biochemical index data is collected, including but not limited to calcium, phosphorus, alkaline phosphatase, blood glucose, blood lipids, etc. The patient's diagnosis and treatment data are recorded, including clinical symptoms, diabetes duration, treatment history, and other information.
[0018] S102, the blood plasma is placed in a sterile cryogenic tube and stored in a refrigerator at a preset temperature; and ribonucleic acid is extracted from the cryogenic tube in the refrigerator.
[0019] In the embodiment of the present application, the blood plasma is transferred to a sterile cryogenic tube and sealed, and then quickly placed in a refrigerator at a preset temperature (such as -20℃ or -80℃) for long-term storage to prevent nucleic acid degradation. When needed, the cryogenic tube is taken out from the low-temperature environment, thawed, centrifuged to remove impurities, and then a special nucleic acid extraction kit or automatic equipment is used to separate and purify ribonucleic acid from the blood plasma.
[0020] S103, when the concentration and purity of the ribonucleic acid meet the preset conditions, the gene expression level data of five target genes are obtained based on the ribonucleic acid; the five target genes are bone sialoprotein gene, sclerostin gene, fibronectin gene, collagen type I alpha chain gene and integrin-binding sialoprotein gene.
[0021] In the embodiment of the present application, the extracted ribonucleic acid is detected for concentration and purity by a spectrophotometer. When the concentration and purity meet the preset conditions, it is proved that the extracted ribonucleic acid is qualified, and the gene expression level data corresponding to the bone sialoprotein gene, sclerostin gene, fibronectin gene, collagen type I alpha chain gene and integrin-binding sialoprotein gene are obtained in the qualified ribonucleic acid.
[0022] Specifically, the qualified ribonucleic acid is used for reverse transcription, and is converted into complementary deoxyribonucleic acid by using a reverse transcription kit. The complementary deoxyribonucleic acid is used as a template, and a specific primer sequence is designed according to the sequence of each gene by using a professional primer design software. In the quantitative reverse transcription polymerase chain reaction, the reaction system comprises the complementary deoxyribonucleic acid template, the upstream and downstream primers, the SYBRGreen polymerase chain reaction mixture and the nuclease-free water. The reaction condition is set as follows: 95 degrees Celsius pre-denaturation for 3 to 5 minutes, then 40 cycles of reaction, each cycle comprising 95 degrees Celsius denaturation for 10 to 15 seconds, 60 degrees Celsius annealing for 20 to 30 seconds and 72 degrees Celsius elongation for 20 to 30 seconds. The glyceraldehyde-3-phosphate dehydrogenase is used as an internal reference gene to correct the loading amount of the complementary deoxyribonucleic acid in each sample. The expression amount of the target gene relative to the internal reference gene is further calculated by comparing the cycle threshold value, so as to obtain the gene expression level data of the five genes.
[0023] In the embodiment of the present application, the gene expression level data, the blood biochemical index data and the diagnosis and treatment data can be fused by splicing to obtain fusion data, and the preprocessed fusion data is input into the trained random forest model and the trained multilayer perception model to obtain the first prediction result and the second prediction result.
[0024] In the embodiment of the present application, the gene expression level data, the blood biochemical index data and the diagnosis and treatment data can be fused by splicing to obtain fusion data, and the preprocessed fusion data is input into the trained random forest model and the trained multilayer perception model to obtain the first prediction result and the second prediction result.
[0025] Wherein, before the fusion data is input into the random forest model and the multi-layer perception model, the fusion data needs to be preprocessed, specifically, for gene expression data, if there are a small amount of missing values (missing proportion < 10%), a multiple imputation method is used. For missing values of continuous variables (such as blood glucose, blood lipids, etc.), the mean imputation method is used, that is, the average value of the variable in all samples is calculated for imputation; for categorical variables (such as complication type), the mode is used for imputation. For abnormal value detection and processing, the Z-score method is used to detect abnormal values, and the local weighted regression scatter smoothing method is used for correction, and the predicted value is used to replace the abnormal value. For related biochemical indicators, abnormal values are identified by drawing box plots, and for abnormal values, the 95th percentile (P95) or the 5th percentile (P5) of the same group of data is used for replacement. Then, the blood biochemical indicator data and other continuous variables are subjected to Z-score standardization processing to eliminate the dimensional difference and facilitate subsequent analysis. Since the original data has the problems of high dimension and much noise, principal component analysis (PCA) algorithm is used for dimension reduction processing, and finally the preprocessed fusion data is obtained. PCA converts the original data into a set of uncorrelated principal components through linear transformation, and these principal components can retain most of the information of the original data while removing redundancy and noise. In the dimension reduction process, the number of principal components to be retained is determined according to the cumulative variance contribution rate, usually making the cumulative variance contribution rate reach 85%-95%, so as to reduce the data dimension while retaining the data characteristics to the greatest extent.
[0026] S105, the first prediction result and the second prediction result are weighted and summed to obtain a risk value of the blood vessels of the diabetic patient from calcification, and a corresponding warning is given through the risk value.
[0027] In the embodiments of the application, the random forest model and the multi-layer perception model set corresponding weights according to the test index values in the test stage, and the first prediction result and the second prediction result are weighted and summed to obtain a target prediction result, that is, a risk value of the blood vessels of the diabetic patient from calcification. And through the risk value, the final warning is given to the diabetic patient.
[0028] Wherein, the warning includes a risk level and corresponding treatment measures for the diabetic patient.
[0029] It can be understood that in the embodiments of the present application, the blood data, blood biochemical index data and diagnosis and treatment data of the diabetic patient are obtained; the blood data is separated to obtain separated plasma; the plasma is placed in a sterile cryopreservation tube and stored in a refrigerator at a preset temperature; the ribonucleic acid is extracted from the cryopreservation tube in the refrigerator; when the concentration and purity of the ribonucleic acid meet the preset conditions, the gene expression level data of the target gene is obtained based on the ribonucleic acid; the target gene is bone sialoprotein gene, sclerostin gene, fibronectin gene, collagen type I alpha chain gene and integrin-binding sialoprotein gene; the gene expression level data, blood biochemical index data and diagnosis and treatment data are fused to obtain fusion data, and the fusion data are input into a random forest model and a multilayer perception model to obtain a first prediction result and a second prediction result; the first prediction result and the second prediction result are weighted and summed to obtain a risk value of vascular calcification of the diabetic patient, and corresponding warning is performed through the risk value. In this process, by detecting the gene expression level data of the five key genes of bone sialoprotein gene, sclerostin gene, fibronectin gene, collagen type I alpha chain gene and integrin-binding sialoprotein gene, the molecular signal of disease occurrence can be captured in the early stage of diabetic vascular calcification, the risk can be found in advance, and the risk of serious diabetic cardiovascular complications can be significantly reduced. The gene expression data, blood biochemical index data and diagnosis and treatment data are integrated, and scientific data preprocessing methods and prediction models are used for deep analysis to evaluate the risk of diabetic vascular calcification of the patient. Compared with the traditional single detection method, the diagnostic accuracy is greatly improved, and the uncertainty of clinical decision-making is reduced. The random forest model and the multilayer perception model are combined to predict the final risk value, which improves the accuracy of prediction.
[0030] In some embodiments of the present application, the corresponding warning through the risk value in S105 can be implemented through S1051 to S1052, which is illustrated as follows.
[0031] S1051, compare the risk value with a preset risk threshold, obtain a risk level according to the comparison result, and perform different degrees of warning according to the risk level.
[0032] S1052, when the risk level represents high risk, obtain medical image data of the diabetic patient, and develop a treatment plan through the medical image data.
[0033] In some embodiments of the present application, the risk levels include low risk, medium risk and high risk, and risk thresholds such as 0.3 and 0.7 are set, and the risk value is compared with the preset risk threshold to obtain the risk level of the diabetic patient's blood vessel calcification. Each risk level sets a corresponding warning. When the risk level shows that the diabetic patient's blood vessel calcification is high risk, the medical image data of the diabetic patient is obtained, and a personalized comprehensive treatment plan is formulated according to the examination results. When it is low risk, regular monitoring and health management is still needed.
[0034] Exemplarily, as shown in Table 1 below, the warnings in Table 1 include risk levels and main intervention measures, as follows: Among them, the prediction results of the risk value and the warning are output in the form of a visual report, the report content is rich and intuitive. It contains the basic information of the diabetic patient (name, age, gender, diabetes duration, etc.), the molecular marker expression atlas (the relative expression amounts of the five genes of bone sialoprotein gene, sclerostin gene, fibronectin gene, collagen type I alpha chain gene and integrin-binding sialoprotein gene are displayed in the form of column chart or line chart, and compared with the normal reference value), the specific values and change trends of glycosylated hemoglobin, blood lipid indicators and risk level identification. At the same time, according to the risk level, detailed clinical suggestions are given, for example, for low-risk patients, it is recommended to review the relevant indicators once every 3-6 months; for medium-risk patients, it is recommended to adjust the dose of hypoglycemic and lipid-lowering drugs, and to detect blood glucose and blood lipid every month; for high-risk patients, it is recommended to immediately perform further imaging examination, and to formulate a personalized comprehensive treatment plan according to the examination results. The visual report can be directly pushed to the clinician through the electronic medical record system, which facilitates the doctor to quickly obtain patient information and make decisions, realizes early warning and personalized management of diabetic vascular calcification.
[0035] In some embodiments of the present application, S104 is further implemented before S201-S205, which is described by the following steps.
[0036] S201, obtaining blood data samples, blood biochemical index data samples and diagnosis and treatment data samples of the diabetic patient, and obtaining gene expression level data samples based on the blood data samples.
[0037] S202, fusing the gene expression level data samples, the blood biochemical index data samples and the diagnosis and treatment data samples to obtain fused data samples; the fused data samples include training samples.
[0038] In some embodiments of the present application, S201 and S202 are similar to the processes of S101 and S104, which are not described in detail here. After S202, the fused data samples are obtained, which include not only the training samples but also the test samples. The fused data samples can be randomly divided into training samples and test samples according to a ratio of 7:3 or 8:2.
[0039] S203, according to the self-sampling method, a sample subset is extracted from the training samples with replacement, and a preset number of decision trees in the random forest model is trained according to the sample subset, to obtain a risk value corresponding to each decision tree.
[0040] In some embodiments of the present application, when constructing the random forest model, first, the self-sampling method is used to extract multiple sample subsets from the original training samples with replacement, and each subset is used to train a decision tree. In this way, each tree is trained based on a slightly different training sample distribution, thereby enhancing the diversity of the model. For each decision tree, after training on its corresponding sample subset, its performance in the training process can be evaluated, and a corresponding risk value can be calculated.
[0041] Among them, the decision tree in the random forest model sets the maximum depth (such as 5-8 layers) and the minimum number of samples (such as 5-10) of the node. The output of the random forest model finally adopts the voting method, and the classification result of the majority decision tree is taken as the output.
[0042] S204, the risk value is added to the training samples to obtain new training samples, and the random forest model is trained using the new training samples until the trained random forest model is obtained.
[0043] S205, the training sample is used to train the trained multi-layer perceptron model until the trained multi-layer perceptron model is obtained.
[0044] In some embodiments of the present application, after the risk value is obtained, the risk value can be added to the original training samples as a new feature to obtain new training samples, and the random forest model is trained using the new training samples, which is iteratively optimized until the trained random forest model is obtained. While training the random forest model, the multi-layer perceptron model can be trained using the training samples until the trained multi-layer perceptron model is obtained. The training samples can be new training samples or original training samples.
[0045] The network architecture design of the multilayer perceptron model is as follows: the number of neurons in the input layer is determined based on the number of new features after merging. Hidden layers adopt a hierarchical stacked structure, initially with 3-4 hidden layers. The number of neurons in each hidden layer can be adjusted appropriately from 64-128, such as setting it to 128-256, to adapt to the complexity of the new feature set. The ReLU function is used as the activation function for the hidden layers, and a dropout layer technique is introduced between layers with a probability set to 0.3-0.5 to enhance the model's generalization ability. The activation function of the output layer is selected according to the prediction target. For binary classification tasks, the Sigmoid function is used to output the calcification probability (0-1), and for regression tasks, a linear activation function is used to output the risk value. Model training optimization: the Adam optimization algorithm is used to update network parameters, and the batch size (e.g., 32, 64) and the number of training epochs (e.g., 100-300 epochs) are set appropriately. During training, 10-20% of the data from the training samples is used as validation samples, and the loss function value on the validation samples is monitored in real time (cross-entropy loss is used for binary classification, and mean squared error loss is used for regression tasks). When the validation sample loss stops decreasing or begins to increase after 10-20 consecutive rounds, an early stopping strategy is triggered, halting training. Model evaluation and iteration: The multilayer perceptron model is evaluated using test samples, and its performance is quantified using metrics such as accuracy, recall, F1 score (F1 metric), AUC (for classification tasks) or mean squared error (MSE) and mean absolute error (MAE) (for regression tasks). If the results are unsatisfactory, the network architecture, activation function, algorithm parameters are optimized, or data preprocessing is performed again to iteratively optimize the model.
[0046] To adapt to the dynamic changes in the condition of diabetic patients, the constructed random forest model and multilayer perceptron model have dynamic update capabilities. When a certain amount of new sample data is accumulated, the new sample data is merged with historical data, and a complete data preprocessing, feature selection, random forest model and multilayer perceptron model training, and model fusion process are carried out again. By continuously incorporating the latest clinical information and constantly optimizing the model parameters and structure, the combined model can adapt to the changes in the condition of diabetic patients in real time, achieving dynamic and accurate prediction of vascular calcification risk.
[0047] In some embodiments of the present invention, S205 is further followed by S301 to S302, which will be described by the following steps.
[0048] S201. Test the random forest model and the multilayer perceptron model using test samples to obtain the corresponding first index value and second index value.
[0049] S202. Based on the first index value and the second index value, the random forest model and the multilayer perceptron model are weighted to obtain the first weight and the second weight.
[0050] In some embodiments of the present invention, the fused data sample also includes a test sample. The test sample is used to test the random forest model and the multilayer perceptron model respectively to obtain the corresponding first index value and second index value. The index value can be accuracy, error rate, and area under the curve (AUC), etc. The first index value and the multilayer perceptron model are weighted according to the first index value and the second index value to obtain the first weight and the second weight.
[0051] For example, if the AUC value of the random forest model on the test set is 0.8 and the AUC value of the multilayer perceptron model is 0.85, then the weight of the prediction result of the random forest model can be set to 0.4 and the weight of the prediction result of the MLP model can be set to 0.6.
[0052] In some embodiments of the present invention, the weighted summation of the first prediction result and the second prediction result in S105 to obtain the risk value of vascular calcification in the diabetic patient can be achieved through S105A to S105B, as described in the following steps.
[0053] S105A. Multiply the first weight and the first prediction result to obtain the first result, and multiply the second weight and the second prediction result to obtain the second result.
[0054] S105B: Sum the first result and the second result to obtain the risk value.
[0055] In some embodiments of the present invention, the first weight and the first prediction result are multiplied, the second weight and the second prediction result are multiplied, and the first result and the second result are summed to obtain the risk value.
[0056] Example 1: Patient Case: Mr. Zhang, male, 62 years old I. Sample Collection and Data Acquisition Mr. Zhang, a 62-year-old male, had been diagnosed with type 2 diabetes for 8 years, meeting the study requirements (disease duration > 5 years). He routinely controlled his blood sugar with oral hypoglycemic medication. He participated in the early warning detection of vascular calcification in diabetes as a study subject. Blood was drawn venously in the morning on an empty stomach, with 8 mL of blood collected into an EDTA anticoagulant tube. The blood sample was processed within 25 minutes after collection, awaiting subsequent testing.
[0057] In the gene expression detection stage, total ribonucleic acid (RNA) was extracted from frozen plasma samples using the triazole method. Strict procedures were followed to ensure the quality of the extracted RNA. The extracted RNA was then reverse transcribed into complementary deoxyribonucleic acid (DNA), and real-time quantitative polymerase chain reaction (qPCR) was used to detect the expression of five genes closely related to vascular calcification in diabetes: Spp1 (bone sialic acid protein gene), Sost (sclerothelin gene), Fn1 (fibronectin gene), Ibsp (type I collagen alpha chain gene), and Col1a1 (integrin-binding sialic acid protein gene). Glyceraldehyde-3-phosphate dehydrogenase was used as an internal reference gene. The expression levels of each gene relative to the internal reference gene were calculated. To ensure the accuracy and reliability of the detection results, multiple technical replicates were set for each sample, and the coefficient of variation (CV) of the technical replicates was <5%. The final expression levels of each gene relative to the internal reference gene were: Spp1=9.8, Sost=8.5, Fn1=10.2, Ibsp=9.0, and Col1a1=11.5.
[0058] Meanwhile, blood biochemical indicators and medical data were collected from the patient: the blood calcium level was 2.45 mmol / L, slightly higher than the average blood calcium level of the study group (2.35±0.2 mmol / L); the blood phosphorus level was 1.15 mmol / L, within the normal fluctuation range of blood phosphorus levels in the study group (1.2±0.3 mmol / L); and the glycated hemoglobin (HbA1c) level was 8.2%, higher than the average level of the study group (7.8±1.5%), indicating that the patient's blood glucose control was poor recently.
[0059] II. Data Preprocessing The expression levels of 5 genes relative to the internal reference gene in this patient were fused with 12 clinical indicators to construct fused data. Outlier handling was performed, and then Z-score standardization was applied to all continuous variables to eliminate dimensional differences. Next, PCA algorithm was used for dimensionality reduction, ultimately retaining 6 principal components with a cumulative variance of 92%. This approach reduced data dimensionality while preserving data features to the greatest extent possible, resulting in preprocessed fused data.
[0060] III. Risk Assessment 1. Evaluation of the Random Forest Model The preprocessed fusion data of this patient were incorporated into a pre-trained random forest model for analysis. The model output the probability of vascular calcification in this patient, predicting that the probability of vascular calcification in this patient was 68%.
[0061] 2. Evaluation of Multilayer Perceptron Model The preprocessed fused data was used as input to a multilayer perceptron model. After further analysis, the multilayer perceptron model predicted a 75% probability of vascular calcification in this patient.
[0062] 3. Model Fusion and Final Risk Assessment A weighted fusion strategy was adopted. Based on the performance of the Random Forest model and the Multilayer Perceptron model on the test set, weights were assigned to the two models: the Random Forest model had a weight of 0.3, and the MLP model had a weight of 0.7. The predicted probabilities of the two models were then summed according to their weights to obtain the final risk value: 0.68×0.3+0.75×0.7=0.729.
[0063] Based on the established risk grading criteria (low risk <0.3, medium risk 0.3≤risk value <0.7, high risk ≥0.7), the patient's final risk value of 0.729 was determined to be high risk, indicating a high risk of vascular calcification.
[0064] IV. Risk Warning and Intervention Recommendations Based on the model evaluation results, the hospital generated a detailed visual risk assessment report for Zhang and pushed it to the attending physician through the electronic medical record system. The report clearly displayed the patient's basic information, expression profiles of five molecular markers (intuitively presenting the comparison between the expression levels of each gene and normal reference values), data on various indicators (such as specific values and trends of blood glucose, blood lipids, blood calcium, and blood phosphorus), and high-risk level indicators.
[0065] Meanwhile, based on the high-risk level, targeted clinical recommendations were developed for the patient: immediate further imaging examinations, such as coronary CT angiography and lower extremity arterial ultrasound, were arranged to clarify the specific degree and location of vascular calcification; the patient was advised to adjust their current medication regimen, strengthen the control of blood sugar, blood pressure, and blood lipids, and use combination medications if necessary; a personalized diet and exercise intervention plan was developed for the patient, including strict control of sugar, fat, and salt intake and increased dietary fiber intake; close monitoring of changes in the patient's condition was maintained, with weekly outpatient follow-up visits arranged to adjust the treatment plan in a timely manner, achieving early intervention and management of diabetic vascular calcification.
[0066] Reference Figure 2 The diagram shows a structural schematic of an electronic device according to an embodiment of the present invention. The specific embodiments of the present invention do not limit the specific implementation of the electronic device.
[0067] like Figure 2 As shown, the electronic device may include: a processor 502, a communications interface 504, a memory 506, and a communications bus 508.
[0068] in: The processor 502, communication interface 504, and memory 506 communicate with each other via communication bus 508.
[0069] Communication interface 504 is used to communicate with other electronic devices or servers.
[0070] The processor 502 is used to execute program 510, specifically the relevant steps in the above method embodiments.
[0071] Specifically, program 510 may include program code that includes computer operation instructions.
[0072] Processor 502 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention. The smart device may include one or more processors of the same type, such as one or more CPUs; or it may include processors of different types, such as one or more CPUs and one or more ASICs.
[0073] Memory 506 is used to store program 510. Memory 506 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.
[0074] Specifically, program 510 can be used to cause processor 502 to perform the operations corresponding to the methods described in the above method embodiments.
[0075] The specific implementation of each step in program 510 can be found in the corresponding descriptions of the steps and units in the above method embodiments, and will not be repeated here. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the devices and modules described above can be referred to the corresponding process descriptions in the foregoing method embodiments, and will not be repeated here.
[0076] It should be noted that, depending on the implementation needs, the various components / steps described in the embodiments of the present invention can be broken down into more components / steps, or two or more components / steps or parts of the operation of components / steps can be combined into new components / steps to achieve the purpose of the embodiments of the present invention.
[0077] The methods described above according to embodiments of the present invention can be implemented in hardware, firmware, or as software or computer code that can be stored in a recording medium (such as a CD-ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or as computer code originally stored on a remote recording medium or a non-transitory machine-readable medium and subsequently stored on a local recording medium, downloaded via a network. Thus, the methods described herein can be processed by software stored on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an ASIC or FPGA). It is understood that the computer, processor, microprocessor controller, or programmable hardware includes storage components (e.g., RAM, ROM, flash memory, etc.) capable of storing or receiving software or computer code, which, when accessed and executed by the computer, processor, or hardware, implements the methods described herein. Furthermore, when a general-purpose computer accesses code used to implement the methods shown herein, the execution of the code transforms the general-purpose computer into a dedicated computer for executing the methods shown herein.
[0078] Those skilled in the art will recognize that the units and method steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of the embodiments of the present invention.
[0079] The above embodiments are only used to illustrate the embodiments of the present invention, and are not intended to limit the embodiments of the present invention. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the embodiments of the present invention. Therefore, all equivalent technical solutions also fall within the scope of the embodiments of the present invention, and the patent protection scope of the embodiments of the present invention should be defined by the claims.
Claims
1. A method for early warning of vascular calcification in diabetes based on molecular marker detection, characterized in that, include: Obtain blood data, blood biochemical index data, and medical data from diabetic patients; and separate the blood data to obtain separated plasma. The plasma was placed in a sterile cryovial and then stored in a refrigerator at a preset temperature. Ribonucleic acid was extracted from the cryovials in the refrigerator; When the concentration and purity of the ribonucleic acid meet the preset conditions, gene expression level data of the target genes are obtained based on the ribonucleic acid; the target genes are bone sialic acid protein gene, sclerosingin gene, fibronectin gene, type I collagen alpha chain gene and integrin-binding sialic acid protein gene; The gene expression level data, blood biochemical index data, and diagnosis and treatment data are fused to obtain fused data. The preprocessed fused data is then input into a random forest model and a multilayer perceptron model to obtain a first prediction result and a second prediction result. The first prediction result and the second prediction result are weighted and summed to obtain the risk value of vascular calcification in the diabetic patient, and the risk value is used to issue a corresponding warning.
2. The method according to claim 1, characterized in that, The provision of corresponding early warnings based on the risk value includes: The risk value is compared with a preset risk threshold, the risk level is obtained based on the comparison result, and different levels of warnings are issued based on the risk level. When the risk level is characterized as high risk, medical imaging data of the diabetic patient is acquired, and a treatment plan is formulated based on the medical imaging data.
3. The method according to claim 1, characterized in that, Before inputting the preprocessed fused data into the random forest model and the multilayer perceptron model respectively to obtain the first prediction result and the second prediction result, the method further includes: Blood data samples, blood biochemical index data samples, and diagnosis and treatment data samples of diabetic patients were obtained, and gene expression level data samples were obtained based on the blood data samples. The gene expression level data sample, the blood biochemical index data sample, and the diagnostic and treatment data sample are fused to obtain a fused data sample; the fused data sample includes training samples. The bootstrap sampling method is used to extract a subset of samples with replacement from the training samples, and the random forest model is trained with a preset number of decision trees based on the subset of samples to obtain the risk value corresponding to each decision tree. The risk value is added to the training sample to obtain a new training sample, and the random forest model is trained using the new training sample until the trained random forest model is obtained. The training samples are used to train the multilayer perceptron model to be trained until the trained multilayer perceptron model is obtained.
4. The method according to claim 3, characterized in that, The fused data sample also includes test samples; After training the multilayer perceptron model to be trained using the training samples until the trained multilayer perceptron model is obtained, the method further includes: The random forest model and the multilayer perceptron model are tested using the test samples to obtain the corresponding first index value and second index value. The first weight and the second weight are obtained by assigning weights to the random forest model and the multilayer perceptron model based on the first index value and the second index value.
5. The method according to claim 4, characterized in that, The step of weighted summing of the first prediction result and the second prediction result to obtain the risk value of vascular calcification in the diabetic patient includes: Multiply the first weight and the first prediction result to obtain the first result, and multiply the second weight and the second prediction result to obtain the second result; The risk value is obtained by summing the first result and the second result.