Construction method and application of retinal artery occlusion disease risk assessment model

By constructing a retinal artery occlusion risk prediction model and utilizing specific protein combinations, the invasiveness and low sensitivity issues of existing diagnostic methods were resolved, achieving high-accuracy early diagnosis and molecular typing of retinal artery occlusion.

CN120636788APending Publication Date: 2025-09-12RENMIN HOSPITAL OF WUHAN UNIVERSITY (HUBEI GENERAL HOSPITAL)
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510642631.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing diagnostic methods for retinal artery occlusion rely on invasive examinations, imaging features have a weak correlation with disease staging, insufficient sensitivity for early diagnosis, lack of specific biomarkers, and inability to achieve molecular typing.

Method used

A risk prediction model for retinal artery occlusion was constructed. Logistic regression analysis of the combination of integrin subunit αM, sex hormone-binding globulin, S100 calcium-binding protein A7, immunoglobulin heavy chain δ, and synaptobrevin 1 was used to construct a risk assessment model for retinal artery occlusion. The prediction accuracy reached 88%, and the AUC was as high as 0.944.

Benefits of technology

It provides a highly accurate risk assessment for retinal artery occlusion, improves the sensitivity and specificity of early diagnosis, realizes molecular typing, and enhances the clinical application value of diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120636788A_ABST
    Figure CN120636788A_ABST
Patent Text Reader

Abstract

The invention relates to a construction method and application of a retinal artery occlusion disease risk assessment model. Specifically, according to the retinal artery occlusion disease risk assessment model, a retinal artery occlusion disease risk assessment result of a subject is obtained based on the RAO protein marker level of a biological sample of the subject; wherein the RAO protein marker is a combination of integrin subunit alpha M, sex hormone binding globulin, S100 calcium binding protein A7, immune globulin heavy chain delta and synaptic-like protein 1. The prediction accuracy of the retinal artery occlusion disease risk prediction model can reach 88%, the area (AUC) under the working characteristic curve of a subject reaches 0.944, and extremely high distinguishing efficiency and clinical application value are shown.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of biomedical detection technology, and in particular to a method for constructing a retinal artery occlusion risk assessment model and its application. Background Art

[0002] Retinal artery occlusion (RAO) is an acute, blinding eye disease. Current diagnosis mainly relies on fundus fluorescein angiography and optical coherence tomography (OCT), which have the following limitations: invasive examinations lead to poor patient compliance; imaging features are weakly correlated with disease staging, and the sensitivity of early diagnosis is insufficient (approximately 65%); there is a lack of specific biomarkers, making molecular typing impossible. Summary of the Invention

[0003] The present invention discovered 66 differentially expressed proteins in blood samples from patients with retinal artery occlusion and healthy controls. Annotation enrichment analysis revealed 23 key proteins among these differentially expressed proteins. The present invention constructed a retinal artery occlusion diagnostic model by randomly combining these 23 key proteins for logistic regression analysis. The optimal retinal artery occlusion risk prediction model was identified, achieving a prediction accuracy of 88% and an AUC of 0.944, demonstrating excellent discriminatory power. Based on this retinal artery occlusion risk prediction model, the present invention provides a method for constructing a retinal artery occlusion risk assessment model and its application.

[0004] In a first aspect, embodiments of the present invention provide the use of a RAO protein marker or a reagent for detecting the level of a RAO protein marker in the preparation of any of the following products:

[0005] (1) Products used for risk assessment of retinal artery occlusion;

[0006] (2) Products for the prognostic risk assessment of retinal artery occlusion;

[0007] The RAO protein marker is a combination of integrin subunit αM, sex hormone-binding globulin, S100 calcium-binding protein A7, immunoglobulin heavy chain δ, and synaptophysin 1.

[0008] In combination with the first aspect, in one embodiment, the product is a system, a kit, an instrument or a chip.

[0009] In combination with the first aspect, in one embodiment, the product includes a calculation module, which is used to calculate the probability of retinal artery occlusion based on the RAO protein marker level and its logistic regression formula, and output the probability of retinal artery occlusion.

[0010] The Logistic regression formula is In this formula, p represents the probability of retinal artery occlusion, ITGAM represents the level of integrin subunit αM, SHBG represents the level of hormone-binding globulin, S100A7 represents the level of S100 calcium-binding protein A7, P0DOX3 represents the level of immunoglobulin heavy chain δ, SYPL1 represents the level of synaptophysin 1, β0 is the intercept term, β1 is the coefficient of ITGAM, β2 is the coefficient of SHBG, β3 is the coefficient of S100A7, β4 is the coefficient of PODOX3, and β5 is the coefficient of SYPL1.

[0011] Preferably, the intercept term β0 = -25.77, the coefficient β1 of ITGAM = 0.47, the coefficient β2 of SHBG = -0.44, the coefficient β3 of S100A7 = 0.69, the coefficient β4 of PODOX3 = 0.56, the coefficient β5 of SYPL1 = 0.25, and the logistic regression formula is

[0012]

[0013] In combination with the first aspect, in one embodiment, the RAO protein marker is from a blood sample; preferably, the blood sample is plasma or serum.

[0014] In a second aspect, an embodiment of the present invention provides a retinal artery occlusion risk assessment system, comprising:

[0015] an acquisition module for obtaining the level of RAO protein markers in a biological sample of a subject; the RAO protein markers are a combination of integrin subunit αM, sex hormone-binding globulin, S100 calcium-binding protein A7, immunoglobulin heavy chain δ, and synaptophysin 1;

[0016] The calculation module is used to calculate the probability of retinal artery occlusion based on the RAO protein marker level and its logistic regression formula, and output the probability of disease:

[0017] The acquisition module is connected to the calculation module wirelessly and / or wiredly.

[0018] In conjunction with the second aspect, in one embodiment, the Logistic regression formula is In this formula, p represents the probability of retinal artery occlusion, ITGAM represents the level of integrin subunit αM, SHBG represents the level of hormone-binding globulin, S100A7 represents the level of S100 calcium-binding protein A7, P0DOX3 represents the level of immunoglobulin heavy chain δ, SYPL1 represents the level of synaptophysin 1, β0 is the intercept term, β1 is the coefficient of ITGAM, β2 is the coefficient of SHBG, β3 is the coefficient of S100A7, β4 is the coefficient of PODOX3, and β5 is the coefficient of SYPL1.

[0019] In a third aspect, an embodiment of the present invention provides a retinal artery occlusion risk assessment instrument including the above-mentioned retinal artery occlusion risk assessment system.

[0020] In a fourth aspect, an embodiment of the present invention provides a computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, it is configured to: obtain a risk assessment result of retinal artery occlusion in a subject based on the level of RAO protein markers in a biological sample of the subject;

[0021] The RAO protein marker is a combination of integrin subunit αM, sex hormone-binding globulin, S100 calcium-binding protein A7, immunoglobulin heavy chain δ, and synaptophysin 1.

[0022] In a fifth aspect, an embodiment of the present invention provides a computer-readable storage medium storing computer program instructions, which, when executed, implement: obtaining a risk assessment result of retinal artery occlusion in a subject based on the level of RAO protein markers in a biological sample of the subject;

[0023] The RAO protein marker is a combination of integrin subunit αM, sex hormone-binding globulin, S100 calcium-binding protein A7, immunoglobulin heavy chain δ, and synaptophysin 1.

[0024] In a sixth aspect, an embodiment of the present invention provides a method for constructing a retinal artery occlusion risk assessment model, comprising:

[0025] A data-independent acquisition strategy was used to obtain plasma proteome profiles of the diseased and control groups;

[0026] The differential proteins between the diseased group and the control group were screened based on the plasma proteome profile;

[0027] The total samples of the diseased group and the control group were randomly divided into a training set and a test set in a ratio of 7:3. A Lasso regression model was used in the training set to perform 5-fold cross-validation to screen several important proteins from the differentially expressed proteins.

[0028] The important proteins screened out were subjected to logistic regression analysis by random combination, and multiple logistic regression formulas were constructed. The discrimination index AUC and calibration index Brier score calculated by 5-fold cross-validation were used to screen out the optimal logistic regression formula as the retinal artery occlusion risk assessment model.

[0029] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include:

[0030] The retinal artery occlusion risk prediction model provided by the present invention has a prediction accuracy of up to 88% and an AUC value of up to 0.944, demonstrating extremely strong discriminatory efficacy and clinical application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 : Workflow for diagnostic model construction.

[0032] Figure 2: Analysis of differences in protein expression levels between the disease group and the control group; (A) Quantitative protein quantity distribution; (B) Quantitative peptide quantity distribution; (C) Volcano plot of differentially expressed proteins; each point represents a protein, where red represents upregulated proteins and blue represents downregulated proteins; the critical values ​​in the figure are |log2FC| ≥ 1.0 (vertical dashed line) and P < 0.05 (horizontal dashed line); (D) Sample PCA diagram of differential protein characterization, where each point represents a sample; (E) Quantitative heat map of differential proteins between samples; rows represent different differential proteins, columns represent different samples, and colors represent the expression levels of differential proteins.

[0033] Figure 3: Selected results of enrichment analysis of DEPs (hypergeometric test, P < 0.05). (A) GO-based DEP enrichment analysis (two-tailed hypergeometric test; p < 0.05). GO terms were ranked by p-value, and the top 5 entries in each category were selected for presentation. (B) KEGG-based DEP enrichment analysis (two-tailed hypergeometric test; p < 0.05). KEGG terms were ranked by p-value, and the top 15 entries were selected for presentation. (C) DO-based DEP enrichment analysis (two-tailed hypergeometric test; p < 0.05). DO terms were ranked by p-value, and the top 15 entries were selected for presentation.

[0034] Figure 4: PPI network analysis results for differentially expressed proteins. (A) Protein interaction network diagram of DEPs. Circles represent proteins, the thickness of the lines connecting the circles represents the confidence level of the protein-protein interaction evidence, and the color of the circles represents the up-regulated (red) / down-regulated (blue) trend when comparing the two groups. (B) Core proteins derived from the PPI network diagram using the DEGREE algorithm. (C) Core proteins derived from the PPI network diagram using the DMNC algorithm; (D) Core proteins derived from the PPI network diagram using the MCC algorithm. (E) Core proteins derived from the PPI network diagram using the MNC algorithm; (F) Venn diagram of core proteins derived from the four algorithms.

[0035] Figure 5: Model construction process and model parameter results. (A) Graph showing mean squared error (MSE) versus Log(λ) during Lasso regression protein screening. The upper horizontal axis represents the number of predicted proteins screened, with the number of proteins with the minimum MSE representing the number of proteins selected for subsequent analysis by Lasso regression. (B) Graph showing the regression coefficient versus Log(λ) during Lasso regression protein screening. (C) Graph showing the optimal logistic regression model parameters. "Estimate" represents the regression coefficient, "Std.Error" represents the standard error of the regression coefficient, "z value" represents the test statistic for the Z-test of the regression coefficient, and "Pr(>|z|)" represents the P-value for the Z-test of the regression coefficient.

[0036] Figure 6: Diagnostic model evaluation results. (A) Receiver-operating characteristic (ROC) curves for the training and test sets. The horizontal axis represents 1-specificity, and the vertical axis represents sensitivity. (B) Confusion matrix of prediction results in the test set. The horizontal axis represents the actual class, and the vertical axis represents the predicted class. (C) Calibration curve of the diagnostic model in the test set. The dashed line represents the ideal curve, the solid red line represents the actual calibration curve, and the solid blue line represents the calibration curve after 1000 bootstrap resamplings. (D) Decision curve of the diagnostic model in the test set. The horizontal axis represents the threshold probability, and the vertical axis represents the net benefit. The black line represents the net benefit of 0 when no one is treated, the gray line represents the inverse curve with a negative slope when everyone is treated, and the red line represents the net benefit of the model. (E) Nomogram display of the logistic regression model. The individual scores "Points" corresponding to different values ​​of the model predictor variables are summed to form the total score "Total Points", which is then mapped to "Risk", the probability of the outcome event occurring. DETAILED DESCRIPTION

[0037] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0038] Example

[0039] 1. Research Sample

[0040] Participants included 50 patients with retinal artery occlusion and 48 healthy subjects. A total of 147 plasma samples were obtained, including 50 arterial plasma samples from patients with retinal artery occlusion and 49 venous plasma samples. The plasma samples of 99 patients with retinal artery occlusion were classified as the diseased group, and the plasma samples of 48 healthy subjects were classified as the control group.

[0041] 2. Differential Protein Analysis

[0042] The plasma proteome profiles of all plasma samples (n=147) were obtained using a data-independent acquisition (DIA) strategy. A total of 1,820 proteins and 14,126 peptides were quantified in all plasma samples. Figure 2A After removing proteins with deletion ratios exceeding 70% in both groups, a total of 973 proteins were included in the subsequent analysis. Compared with the control group, 66 proteins were significantly differentially expressed in the diseased group (|log2FC| ≥ 1.0, p < 0.05; FC, fold change), of which 51 proteins were upregulated and 15 proteins were downregulated ( Figure 2C ).

[0043] The 66 differentially expressed proteins (DEPs) were plotted using principal component analysis (PCA). Different points represent different samples. The results showed that the RAOG1 group (representing proteins overlapping between arterial and venous blood of RAO patients) and the HC-VG0 group (representing proteins of the control group) could be clearly distinguished by DEPs. Figure 2D ). Using the quantitative information of differentially expressed proteins to draw a heat map, it was found that the expression levels of DEPs were different between the two groups ( Figure 2E ), rows represent different differentially expressed proteins, columns represent different samples, and colors represent the levels of differentially expressed proteins.

[0044] 3. Annotation Enrichment Analysis

[0045] To investigate the biological functions of these DEPs, functional annotation and pathway enrichment analysis were performed for the 66 identified DEPs using the Gene Ontology (GO), Kyoto Encyclopedia of Genes and Genomes (KEGG), and Disease Ontology (DO) databases. A P value < 0.05 using the hypergeometric test was considered statistically significant. The differentially expressed proteins were primarily enriched in cellular components including "vesicles," "organelles," and "extracellular organelles." Their molecular functions included "RNA binding," "cytoskeletal protein binding," and "structural components of the cytoskeleton." Their biological processes included "cell differentiation," "cellular component organization," and "regulation of organelle organization." KEGG enrichment analysis revealed that the differentially expressed proteins were primarily enriched in pathways including "actin cytoskeleton regulation," "pathogenic Escherichia coli infection," and "cGMP-PKG signaling pathway." Disease Ontology (DO) enrichment analysis revealed that the differentially expressed proteins were primarily enriched in diseases including "brain disorders," "hypertension," and "epilepsy."

[0046] In order to explore the interactions between DEPs, protein interaction information was obtained from the STRING database, and the protein-protein interaction (PPI) network diagram was drawn using CytoScape software ( Figure 4A ), four algorithms (DEGREE, DMNC, MCC and MNC) in the software plug-in CytoHubba were used to calculate the top 10 core proteins in the PPI network ( Figure 4B -E) and draw a Venn diagram ( Figure 4F ), the core proteins that appear simultaneously in three or more algorithms are called hub proteins, and the hub proteins are represented by gene names.

[0047] 4. Disease risk prediction

[0048] All plasma samples were randomly divided into training and test sets at a ratio of 7:3. In the training set, 23 important proteins were screened out from 66 DEPs using Lasso regression model with 5-fold cross validation, including integrin subunit αM (ITGAM), sex hormone binding globulin (SHBG), S100 calcium binding protein A7 (S100A7), immunoglobulin heavy chain δ (P0DOX3), synaptobrevin-like protein 1 (SYPL1), inhibin βE (INHBE), sialyltransferase 1 (SIAT1), connective tissue growth factor (CCN2), trypsin 3 (TRY3), aspartic acid protein (ASPN), periostin (POSTN ), stromal protein-like 3 (STML3), phospholipase A1A (PLA1A), proteolipid protein 2 (PLP2), nucleoside diphosphate kinase A (NDKA), alpha-1 acid glycoprotein 2 (A1AG2), spectrin beta 2 (SPTB2), laminin G3 binding protein (LG3BP), immunoglobulin heavy chain variable region 6-1 (IGHV6-1), polymeric immunoglobulin receptor (PIGR), hemoglobin delta chain (HBD), fibrinogen gamma chain (FGG), insulin-like growth factor binding protein 2 (IBP2), respectively ( Figure 5A -B). These 23 important proteins were randomly combined for logistic regression analysis to construct a diagnostic model, and the optimal diagnostic model was screened by the discrimination index AUC and calibration index Brier score calculated by 5-fold cross validation ( Figure 5C ), the Logistic regression formula is

[0049] 5. Model Evaluation

[0050] The receiver operating characteristic (ROC) curve, Hosmer-Lemeshow goodness-of-fit test, calibration curve, and decision curve were used to evaluate the discrimination, calibration, and clinical utility of the model in the training and test sets. The area under the ROC curve (AUC) was used to evaluate the discrimination of the model. The AUC for the training set was 0.944 (0.891-1), and the AUC for the test set was 0.94 (0.868-1) ( Figure 6A ). It is generally believed that 0.90≤AUC, the model has excellent discrimination ability; 0.75≤AUC<0.90, the model has good discrimination ability; 0.60≤AUC<0.75, the model has certain discrimination ability, but it is not recommended; AUC<0.60, the model has poor discrimination ability. The confusion matrix represents the classification of the model and is used to evaluate the model prediction accuracy. The model accuracy in the training set is 88%, and the model accuracy in the test set is 80% ( Figure 6B The Hosmer-Lemeshow goodness-of-fit test P values ​​for the training set and the test set were 0.8222785 and 0.783488 respectively (if P>0.05, the model fit is good). In addition, if the test set ( Figure 6C ) is very close to the standard 45° diagonal line, indicating that the predicted probability is very consistent with the actual frequency, proving that the model calibration is high. The DCA curve is used to evaluate the clinical application value of the model. The horizontal axis of the curve represents the threshold probability, and the vertical axis represents the net rate of return. The black line refers to the net rate of return of 0 when no one intervenes, the gray line refers to the inverse curve with a negative slope of the net rate of return when all people receive intervention, and the red line refers to the net rate of return of the model. The height of the DCA curve represents the size of the clinical utility. If the red line is higher than the gray line, it means that using this model to determine whether to intervene in a population is more clinically useful than intervening in the entire population. Figure 6D The DCA curve of the training set is shown. The Nomograph of the predicted probability of the outcome event is drawn based on the results of multivariate logistic regression ( Figure 6E ). On this scoring table, each variable has a corresponding score. The scores of each variable are added together to get a total score, and then the total score is used to determine the probability of the outcome event.

[0051] The reagents and consumables, sample processing methods, mass spectrometry detection methods, database search methods, data preprocessing methods, and statistical analysis methods used in the following examples to obtain plasma proteome profiles are as follows:

[0052] (1) Reagents and consumables

[0053] Table 1

[0054] Chinese name factory Item No. Sodium deoxycholate Sigma-Aldrich D6750-25G Trifluoroacetic acid Sigma-Aldrich 299537-500G Tris(2-carboxyethyl)phosphine hydrochloride Sigma-Aldrich C4706-10G Chloroacetamide Sigma-Aldrich C0267-100G Trypsin SignalChem T575-31N-10 Tris(hydroxymethyl)aminomethane Sigma-Aldrich T1503-1KG Bradford protein quantitative detection kit Bioengineering C503031-1000 Water(LC-MS) ThermoFisherScientific W6-4 Acetonitrile (LC / MS) ThermoFisherScientific A955-4 Formic acid (LC-MS) Fluka 60-006-17

[0055] (2) Sample processing

[0056] Samples were enriched using superparamagnetic iron oxide nanoparticles. A 20-μl sample was diluted with loading buffer (10 mM Tris-Cl, 1 mM EDTA, 150 mM KCl, 0.05% CHAPS), mixed evenly with 1 mg of nanoparticle suspension, and incubated at 37°C for 1 hour. The beads were washed twice with loading buffer and once with loading buffer without CHAPS (10 mM Tris-Cl, 1 mM EDTA, 150 mM KCl). The beads were adsorbed on a magnetic rack and the supernatant was discarded to obtain the protein-enriched nanoparticles. Lysis buffer (1% SDC / 100 mM Tris-HCl, pH 8.5 / 10 mM TCEP / 40 mM CAA) was added to the sample and incubated at 60°C for 30 minutes for reductive alkylation. Add an equal volume of ddH2O to dilute the SDC concentration to below 0.5%. Add 1 μg of trypsin and incubate at 37°C with shaking overnight for enzymatic digestion. The next day, add TFA to terminate the digestion. Desalt the supernatant using an SDB-RPS desalting column, vacuum dry, and freeze at -20°C until ready for use.

[0057] (3) Mass spectrometry

[0058] The samples were detected by mass spectrometry using an UltiMate 3000RSLCnano nanoliter liquid phase (Thermo) coupled to a timsTOFPro mass spectrometer (Bruker). The peptide samples were injected via an autosampler and bound to a C18 trapping column (75 μm*2 cm, 3 μm particle size, pore size, Thermo), and then enter the analytical column (75μm*15cm, 1.7μm particlesize, pore size, IonOpticks). Analytical gradients were established using mobile phase A (0.1% formic acid) and mobile phase B (0.1% formic acid in ACN). Mass spectrometry data were acquired in diaPASEF mode. The capillary voltage was set to 1500 V. The scan range for MS1 and MS2 spectra was set to 100-1700 m / z. The ion mobility range was set to 0.6-1.6 Vs / cm 2 The accumulation time and ramp time were set to 100ms. The diaPASEF acquisition window was set using timsControl software according to the distribution law of mass-to-charge ratio-ion mobility. The collision energy was set from 1 / K0 = 1.6 Vs / cm 2 The 59eV is linearly reduced to 1 / K0 = 0.6Vs / cm 2 20eV.

[0059] (4) Search the database

[0060] DIA raw data were analyzed using DIA-NN (1.9.2) software in library-free mode. The search database used was the HUMAN protein sequence database (20250113) downloaded from Uniprot. The search parameters were primarily default, with key parameters described as follows: the Precursor ion generation option was enabled for theoretical spectral library prediction; Trypsin / P was used, with a maximum of one missed cleavage site allowed; Carbamidomethyl (C) modification was set as fixed; Oxidaton (M) and Acetylation (protein N-terminal) were set as variable; MS1 and MS2 mass tolerances were set to 15 ppm; MBR was enabled; Heuristic protein inference was enabled; and the FDR was set to 1%. Protein quantification information was normalized using the MaxLFQ algorithm.

[0061] (5) Data preprocessing

[0062] Contaminating proteins were removed from the quantitative protein matrix and logarithmically transformed. Missing values ​​were then imputed using values ​​representing a normal distribution near the detection limit of the mass spectrometer. To this end, the present invention determined the mean and standard deviation of the actual intensity distribution and used it to create a new distribution with a downward shift of 1.8 standard deviations and a width of 0.25 standard deviations. These values ​​were used to estimate missing values ​​in the protein matrix, excluding proteins with a missing ratio exceeding 50% in both groups. The remaining proteins were retained for subsequent analysis.

[0063] (6) Statistical analysis

[0064] Clinical variables were described as follows: for continuous variables that conformed to a normal distribution, mean ± standard deviation was used; for continuous variables that did not conform to a normal distribution, quartiles were used; and for categorical variables, percentages were used. Quantitative protein data were tested for intergroup differences. A t-test was used if both groups conformed to a normal distribution; otherwise, a Wilcoxon rank-sum test was used. After screening for DEPs using fold change (FC) and difference tests (two-sided P < 0.05), Lasso regression and cross-validation were used to identify the final biomarkers from the DEPs and to develop a predictive model. 70% of the total sample data served as the training set, and 30% of the data served as the testing set. Multiple diagnostic models were constructed using multivariate logistic regression using randomly selected protein combinations from the protein biomarkers within the training set. The optimal diagnostic model was selected using the area under the receiver operating characteristic curve (AUC) and Brier score calculated using cross-validation. The Brier score is a statistical measure of predictive accuracy. It reflects the degree of model calibration by calculating the mean squared error between the predicted probability and the actual probability. Finally, the model's discrimination, calibration, and clinical utility were evaluated using receiver operating characteristic (ROC) curves, Hosmer-Lemeshow goodness-of-fit tests, calibration curves, and decision curves in the training and test sets, respectively. All statistical analyses were performed using R (version 4.4.3) and Python (version 3.10.12).

[0065] (7) Bioinformatics analysis process

[0066] The present invention performs functional annotation and enrichment analysis on the identified differentially expressed proteins to understand the functional properties and molecular mechanisms associated with the differentially expressed proteins. A hypergeometric test P < 0.05 indicates that the differentially expressed proteins are significantly enriched in the pathway.

[0067] Functional annotation databases include:

[0068] Gene Ontology (GO, http: / / www.geneontology.org / ) is a database that categorizes gene functions into three main categories: biological process (BP), cellular component (CC), and molecular function (MF). Cellular component describes subcellular structure, location, and macromolecular complexes; molecular function describes the functions of genes and gene products; and biological process describes the biological processes in which gene-encoded products participate.

[0069] The Kyoto Encyclopedia of Genes and Genomes (KEGG, https: / / www.kegg.jp / kegg / ) is a database that systematically describes biological systems at the molecular level. KEGG pathways primarily include metabolism, genetic information processing, environmental information processing, cellular processes, human diseases, and drug development.

[0070] The Disease Ontology (DO, https: / / disease-ontology.org / ) knowledge base annotates genes involved in human diseases. Protein interaction information from the STRING database was used to annotate differentially expressed proteins and explore interaction relationships among these differentially expressed proteins. The STRING database scores protein-protein relationships based on experimental data, text mining data, gene adjacency data, gene co-expression data, and other protein-related databases from multiple public databases. Protein-protein interaction (PPI) networks of the differentially expressed proteins were mapped using CytoScape software. Core proteins in the PPI network were identified using the DEGREE, DMNC, MCC, and MNC algorithms in the "CytoHubba" plugin. Proteins that appear in three or more of these algorithms are considered hub proteins.

[0071] The above are only preferred embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention description and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.

Claims

1. Use of RAO protein markers or reagents for detecting RAO protein marker levels in the preparation of any of the following products: (1) Products used for risk assessment of retinal artery occlusion; (2) Products for the prognostic risk assessment of retinal artery occlusion; The RAO protein marker is a combination of integrin subunit αM, sex hormone-binding globulin, S100 calcium-binding protein A7, immunoglobulin heavy chain δ, and synaptophysin 1.

2. The use according to claim 1, characterized in that: The product is a system, a kit, an instrument or a chip.

3. The use according to claim 1, characterized in that: The product includes a calculation module, which is used to calculate the probability of retinal artery occlusion based on the RAO protein marker level and its logistic regression formula, and output the probability of retinal artery occlusion.

4. The use according to claim 3, characterized in that: The Logistic regression formula is In this formula, p represents the probability of retinal artery occlusion, ITGAM represents the level of integrin subunit αM, SHBG represents the level of hormone-binding globulin, S100A7 represents the level of S100 calcium-binding protein A7, P0DOX3 represents the level of immunoglobulin heavy chain δ, SYPL1 represents the level of synaptophysin 1, β0 is the intercept term, β1 is the coefficient of ITGAM, β2 is the coefficient of SHBG, β3 is the coefficient of S100A7, β4 is the coefficient of PODOX3, and β5 is the coefficient of SYPL1.

5. A retinal artery occlusion risk assessment system, characterized in that: include: an acquisition module for obtaining the level of RAO protein markers in a biological sample of a subject; the RAO protein markers are a combination of integrin subunit αM, sex hormone-binding globulin, S100 calcium-binding protein A7, immunoglobulin heavy chain δ, and synaptophysin 1; A calculation module is used to calculate the probability of retinal artery occlusion based on the RAO protein marker level and its logistic regression formula, and output the probability of retinal artery occlusion; The acquisition module is connected to the calculation module wirelessly and / or wiredly.

6. The retinal artery occlusion risk assessment system according to claim 5, characterized in that: The Logistic regression formula is In this formula, p represents the probability of retinal artery occlusion, ITGAM represents the level of integrin subunit αM, SHBG represents the level of hormone-binding globulin, S100A7 represents the level of S100 calcium-binding protein A7, P0DOX3 represents the level of immunoglobulin heavy chain δ, SYPL1 represents the level of synaptophysin 1, β0 is the intercept term, β1 is the coefficient of ITGAM, β2 is the coefficient of SHBG, β3 is the coefficient of S100A7, β4 is the coefficient of PODOX3, and β5 is the coefficient of SYPL1.

7. A retinal artery occlusion risk assessment instrument, characterized by: The method comprises the retinal artery occlusion risk assessment system according to claim 5 or 6.

8. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the following steps are achieved: obtaining a risk assessment result of retinal artery occlusion of the subject based on the level of RAO protein marker in the biological sample of the subject; The RAO protein marker is a combination of integrin subunit αM, sex hormone-binding globulin, S100 calcium-binding protein A7, immunoglobulin heavy chain δ, and synaptophysin 1.

9. A computer-readable storage medium storing computer program instructions, wherein when executed, the computer program instructions are configured to: obtain a risk assessment result of retinal artery occlusion in a subject based on the level of RAO protein markers in a biological sample of the subject; in, The RAO protein marker is a combination of integrin subunit αM, sex hormone-binding globulin, S100 calcium-binding protein A7, immunoglobulin heavy chain δ, and synaptophysin 1.

10. A method for constructing a risk assessment model for retinal artery occlusion, characterized in that: include: A data-independent acquisition strategy was used to obtain plasma proteome profiles of the diseased and control groups; The differential proteins between the diseased group and the control group were screened based on the plasma proteome profile; The total samples of the diseased group and the control group were randomly divided into a training set and a test set in a ratio of 7:

3. A Lasso regression model was used in the training set to perform 5-fold cross-validation to screen several important proteins from the differentially expressed proteins. The important proteins screened out were subjected to logistic regression analysis by random combination, and multiple logistic regression formulas were constructed. The discrimination index AUC and calibration index Brier score calculated by 5-fold cross-validation were used to screen out the optimal logistic regression formula as the retinal artery occlusion risk assessment model.

Citation Information

Cited By

  • Abdominal aortic aneurysm serum protein fingerprint detection and analysis method and system based on LASSO regression algorithm

    CN121142060A