A gene marker of gfpt2, lum, tnxb, thbs4 alone or in combination and application thereof in cardiovascular disease risk warning and prognosis evaluation
By detecting the expression levels of gene markers GFPT2, LUM, TNXB, and THBS4, and combining them with a risk assessment model, early warning and precise intervention for the risk of heart failure in cardiovascular diseases were achieved. This solved the problems of delayed warning and etiological specificity in existing technologies, and enabled highly accurate prediction and treatment guidance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HENAN CANCER HOSPITAL
- Filing Date
- 2026-03-06
- Publication Date
- 2026-05-29
AI Technical Summary
Existing technologies are insufficient to identify individuals at high risk of heart failure before irreversible fibrotic remodeling of the myocardium occurs, and lack the ability to provide universal and mechanistic early warning of myocardial injury across different etiologies.
By using genetic markers such as GFPT2, LUM, TNXB, and THBS4, either individually or in combination, and by detecting the expression levels or protein concentrations of these genes in biological samples, combined with a risk assessment model, early warning and prognostic assessment of the risk of heart failure in cardiovascular diseases can be achieved.
It provides a universal early warning across different cardiovascular diseases, can identify high-risk patients before obvious abnormalities appear in imaging or function, has extremely high accuracy in predicting the risk of heart failure transformation, and provides biomarker basis for the use of antifibrotic drugs.
Smart Images

Figure CN122104891A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of biomedical technology and relates to a gene biomarker of GFPT2, LUM, TNXB, and THBS4, alone or in combination, and its application in cardiovascular disease risk warning and prognosis assessment. Specifically, it relates to a gene biomarker composition, detection kit, and system for early warning, diagnosis, and prognosis assessment of heart failure risk caused by cardiovascular diseases (including but not limited to hypertensive heart disease, coronary heart disease, cardiomyopathy, myocarditis, valvular heart disease, etc.). Background Technology
[0002] Currently, the management of cardiovascular diseases and the early warning of heart failure mainly rely on clinical symptoms, imaging examinations (such as echocardiography to measure left ventricular ejection fraction LVEF) and blood biomarkers (such as BNP and NT-proBNP). These methods are mostly used for the diagnosis and assessment of the middle and late stages of the disease.
[0003] The existing technology has the following drawbacks: Delayed early warning: Current technologies are insufficient to effectively identify individuals at high risk of heart failure before irreversible fibrotic remodeling of the myocardium occurs, and indicators such as LVEF have insufficient sensitivity in the early stages.
[0004] Poor universality of etiology: Many existing biomarkers (such as BNP) mainly reflect cardiac pressure load, and lack universal and sensitive early warning capabilities for myocardial damage driven by different etiologies (such as hypertension, ischemia, and metabolic abnormalities) but ultimately converge into a common fibrotic pathway.
[0005] Lack of mechanistic correlation: Existing biomarkers are not strongly associated with the core common pathways driving the progression of heart failure—myocardial fibrosis and metabolic reprogramming—and cannot provide guidance for early intervention targeting the pathological mechanisms.
[0006] A search revealed Chinese patent literature disclosing reagents, kits, and methods for detecting serum LUM expression in liver fibrosis [Application No.: CN202011563336.0; Publication No.: CN112763716A]. This invention provides reagents, kits, and methods for detecting serum LUM expression in liver fibrosis, including LUM capture antibodies and LUM detection antibodies for labeling. This invention has significant translational value for early warning and diagnosis of severe liver disease. Furthermore, an enzyme-linked immunosorbent assay (ELISA) diagnostic reagent based on the double-antibody sandwich principle has been developed, enabling rapid and accurate quantitative detection of LUM for rapid auxiliary diagnosis of the degree of fibrosis in various liver diseases.
[0007] A search revealed the application of Chinese patent literature, such as the detection of biomarkers for myocardial ischemia injury, in the preparation of tools for diagnosing and / or treating myocardial ischemia injury [Application No.: CN202110346604.1; Publication No.: CN115141881A]. Verification showed that the expression levels of IFIT3, IFIT2, XAF1, DDX60, IFI44L, UBA7, CTSK, LUM, NT5E, ASPN, and BCL2L1 were significantly correlated with myocardial ischemia injury, and these genes can serve as biomarkers. Original samples and animal models verified that abnormal expression levels of IFIT2, IFIT3, IFI44L, BCL2L1, and CTSK are closely related to myocardial ischemia. Therefore, these biomarkers can be used to prepare adjunctive diagnostic and therapeutic agents for myocardial ischemia injury, screen drugs for treating myocardial ischemia injury, and develop potential therapeutic targets and drugs for myocardial ischemia, possessing significant clinical application value.
[0008] A search revealed that Chinese patent literature discloses biomarkers for cardiovascular diseases [Application No.: CN200880126603.9; Publication No.: CN101946009A]. This invention relates to biomarkers for diagnosing and predicting cardiovascular diseases in patients, methods for diagnosing and predicting cardiovascular diseases in subjects, kits for performing such methods, and microarrays and diagnostic reagents useful for such methods. Specifically, the diagnosed cardiovascular disease involves ischemic heart disease. The invention further relates to methods for treating subjects with cardiovascular diseases (increased risk of developing such diseases) and pharmaceutical compositions suitable for such treatment methods.
[0009] Analysis revealed that the above patent documents only disclosed the application of THBS4 as a genetic marker in cardiovascular diseases, but did not provide detailed experimental evidence or propose a basis for a universal molecular marker composition that can be used to diagnose and assess different cardiovascular diseases.
[0010] A search revealed an article published on July 28, 2021, in the journal *Science, Technology and Engineering*, titled "Bioinformatics Analysis of Phylogenetic and Functional GFPT2 Gene," authored by Wei Si'ang, Ding Zhiwen, and Feng Yan.
[0011] According to the search results, the journal "Research Progress of Platelet-Reactive Protein in Cardiovascular Diseases" published on August 26, 2025 in the Journal of Cardiovascular and Pulmonary Diseases, published on CNKI, was found. The authors are Zhang Jiangtao, Xu Fei, Zhuang Zhong, and Chai Shoudong.
[0012] Analysis revealed that the above journal articles only disclosed the application of GFPT2 and THBS4 as gene markers in cardiovascular diseases, but did not provide detailed experimental evidence or propose a universal molecular marker combination that can be used to diagnose and assess different cardiovascular diseases.
[0013] The present invention aims to overcome the deficiencies of the prior art and provide a universal molecular marker composition that is closely related to the common core pathological mechanism that ultimately leads to heart failure (i.e., myocardial fibrosis). Summary of the Invention
[0014] The purpose of this invention is to address the aforementioned problems in existing technologies by proposing a gene biomarker of GFPT2, LUM, TNXB, and THBS4, either alone or in combination, and its application in cardiovascular disease risk warning and prognostic assessment. The technical problem this invention aims to solve is: how to achieve accurate early warning and prognostic assessment of heart failure risk in the early stages of irreversible myocardial fibrosis, across different initial causes of cardiovascular diseases, providing a time window and targeted basis for early intervention.
[0015] The objective of this invention can be achieved through the following technical solutions: A gene marker for GFPT2, LUM, TNXB, or THBS4, alone or in combination, characterized in that it includes any one or more combinations.
[0016] The application of a genetic marker, GFPT2, LUM, TNXB, or THBS4, alone or in combination, in the preparation of diagnostic products for risk warning and prognostic assessment of the risk of heart failure caused by hypertensive heart disease, coronary heart disease, cardiomyopathy, myocarditis, or valvular heart disease.
[0017] A gene biomarker composition for diagnosing the risk of heart failure in patients with cardiovascular disease, characterized in that the composition comprises a reagent for detecting any one or more genes of GFPT2, LUM, TNXB and THBS4 in a biological sample, wherein the reagent is used to detect the expression level or protein level of the gene.
[0018] A diagnostic kit for early warning and prognostic assessment of heart failure risk in patients with cardiovascular disease, characterized in that it comprises: Primer pairs and / or probes for the specific detection of GFPT2 mRNA, LUM mRNA, TNXB mRNA and THBS4 mRNA; Alternatively, antibodies for the specific detection of GFPT2, LUM, TNXB, and THBS4 proteins; And necessary auxiliary reagents, including buffers, enzymes, and substrates.
[0019] A cardiovascular disease heart failure risk warning and prognostic assessment system, characterized in that it includes: The detection kit is used to measure the expression level or protein concentration of four genes, GFPT2, LUM, TNXB, and THBS4, in biological samples. The data processing unit has a pre-stored risk assessment model, which is configured to receive measurement data of four gene markers and calculate a comprehensive risk score based on their expression levels (GFPT2 upregulated, LUM, TNXB, and THBS4 downregulated). The output unit is used to output the risk score and the corresponding heart failure risk level conclusion (such as low risk, medium risk, high risk).
[0020] How the system works: Sample collection: Collecting blood, plasma, or serum samples from individuals who have or are suspected of having cardiovascular diseases such as hypertension, coronary heart disease, or cardiomyopathy.
[0021] Biomarker detection: Using reagents in the detection kit, the expression levels or concentrations of four gene biomarkers in the sample are quantitatively detected.
[0022] Data analysis and risk assessment: The detection data is input into the risk assessment model. The model integrates the values of the four gene markers based on a pre-trained algorithm (built on multivariate regression or machine learning model) to calculate the probability of the individual progressing to heart failure or the severity level of current myocardial fibrosis within a specific time in the future.
[0023] Output results: The system outputs risk assessment conclusions to assist clinicians in early and proactive intervention and management of high-risk patients.
[0024] The four genes GFPT2, LUM, TNXB, and THBS4 form a functionally synergistic network in the common pathway of myocardial fibrosis, independent of specific initial causes. GFPT2, as an upstream driver gene, acts as an amplifier for pro-fibrotic signal activation when upregulated. LUM, TNXB, and THBS4, as key molecules for extracellular matrix (ECM) homeostasis, are directly reflected in myocardial structural damage and fibrotic scar formation when downregulated. Changes in the expression of these four genes together constitute a "universal molecular tag" reflecting the degree of myocardial fibrosis activity.
[0025] Compared with existing technologies, the gene markers GFPT2, LUM, TNXB, and THBS4, alone or in combination, and their application in cardiovascular disease risk warning and prognostic assessment have the following advantages: This invention identifies four genes—GFPT2, LUM, TNXB, and THBS4—alone or in combination, as universal biomarkers for predicting the common terminal pathway (myocardial fibrosis) of heart failure across different etiologies and with specificity. This invention establishes a mathematical model for predicting heart failure risk based on the expression levels of these four genes, applicable to a broad cardiovascular population.
[0026] Highly applicable: For the first time, it provides a universal early warning biomarker for heart failure risk that can be widely used in a variety of cardiovascular diseases (such as hypertensive heart disease, ischemic heart disease, etc.), overcoming the limitations of etiological specificity.
[0027] Early warning window moved forward: By capturing the core molecular events of the common pathway of myocardial fibrosis, patients at high risk of heart failure can be identified before obvious abnormalities appear on imaging or in function, achieving true early warning.
[0028] High accuracy: The multi-marker joint detection model showed extremely high accuracy in predicting the risk of heart failure conversion (AUC>0.95).
[0029] Guiding precise intervention: This biomarker combination directly targets the fibrosis mechanism, and its evaluation results can provide key biomarker evidence for the use of antifibrotic drugs (such as inhibitors targeting the GFPT2 pathway) or the adjustment of existing treatment regimens. Attached Figure Description
[0030] Figure 1 This is a schematic diagram illustrating how TGF-β significantly induces lactate secretion and fibrosis in MCFs in Embodiment 1 of the present invention; wherein, Figure 1 A showed that TGF-β significantly induced an increase in lactate secretion from MCFs. Figure 1 B showed that TGF-β significantly induced an accelerated fibrosis process.
[0031] Figure 2 This is a schematic diagram illustrating how Gfpt2 dysfunction alleviates lactic acid secretion and fibrosis in MCFs in Embodiment 1 of the present invention; wherein, Figure 2 A showed that knocking down Gfpt2 effectively reduced lactate secretion in MCFs. Figure 2 B shows that knocking down Gfpt2 effectively slows down the fibrosis process.
[0032] Figure 3 This is a schematic diagram illustrating how Lum, Tnxb, and Thbs4, individually and in combination, alleviate lactic acid secretion and fibrosis in MCFs in Embodiment 1 of the present invention; wherein, Figure 3 A showed that overexpression of Lum, Tnxb, and Thbs4 effectively reduced lactate secretion in MCFs. Figure 3 B showed that overexpression of Lum, Tnxb, and Thbs4 effectively slowed down the fibrosis process. Figure 3 C showed that combined overexpression of Lum, Tnxb, and Thbs4 effectively reduced lactate secretion in MCFs. Figure 3 D showed that co-expression of Lum, Tnxb, and Thbs4 effectively slowed down the fibrosis process.
[0033] Figure 4This is a schematic diagram illustrating how the loss of Gfpt2 function combined with the functions of Lum, Tnxb, and Thbs4 in Embodiment 1 of the present invention alleviates lactic acid secretion and fibrosis in MCFs; wherein, Figure 4 A study showed that Gfpt2 knockdown combined with overexpression of any one of Lum, Tnxb, or Thbs4 effectively reduced lactate secretion in MCFs. Figure 4 B showed that Gfpt2 knockdown combined with overexpression of any one of Lum, Tnxb, or Thbs4 effectively slowed down the fibrosis process. Figure 4 C showed that Gfpt2 knockdown combined with overexpression of any two of Lum, Tnxb, and Thbs4 effectively reduced lactate secretion from MCFs. Figure 4 D showed that Gfpt2 knockdown combined with overexpression of any two of Lum, Tnxb, and Thbs4 effectively slowed down the fibrosis process.
[0034] Figure 5 This is a schematic diagram illustrating how the loss of Gfpt2 function combined with the functions of Lum, Tnxb, and Thbs4 in Embodiment 2 of the present invention alleviated cardiac injury and fibrosis in mice; wherein, Figure 5 A study showed that Gfpt2 knockdown combined with overexpression of any one of Lum, Tnxb, or Thbs4 effectively reduced cardiac damage in mice. Figure 5 B showed that Gfpt2 knockdown combined with overexpression of any one of Lum, Tnxb, or Thbs4 significantly increased the ejection fraction (EF) and fractional shortening (FS) of the mouse heart, bringing them within the normal range, and the mouse's cardiac contractile function was good.
[0035] Figure 6 This is a schematic diagram illustrating how, in Embodiment 2 of the present invention, loss of Gfpt2 function combined with the functions of Lum, Tnxb, and Thbs4 slowed down abnormal serum lactate secretion and cardiac fibrosis in mice; wherein, Figure 6 A study showed that Gfpt2 knockdown combined with overexpression of any one of Lum, Tnxb, or Thbs4 effectively reduced lactate secretion in mice. Figure 6 B showed that Gfpt2 knockdown combined with overexpression of any one of Lum, Tnxb, or Thbs4 effectively slowed the progression of cardiac fibrosis in mice.
[0036] Figure 7 This is a nomograph for heart failure prognosis based on the expression levels of four genes (GFPT2, LUM, TNXB, THBS4) in Example 3 of the present invention.
[0037] Figure 8 This is the time-dependent ROC curve of the four-gene prognostic model in the training queue in Embodiment 3 of the present invention. Detailed Implementation
[0038] The following are specific embodiments of the present invention, which are described in conjunction with the accompanying drawings. However, the present invention is not limited to these embodiments.
[0039] To better illustrate the implementation process of this invention, the specific research process of this invention is disclosed as follows: 1. Data Acquisition The GEO database, short for GENE EXPRESSION OMNIBUS, is a gene expression database created and maintained by the National Center for Biotechnology Information (NCBI) in the United States. Download the single-cell data file for GSE145154 from the NCBI GEO public database, which includes 4 control samples and 23 disease-related heart tissue samples. Download the Series Matrix File data file for GSE57338 from the NCBI GEO public database, with annotation file GPL11532, which includes expression profile data from 313 samples, including 137 in the non-heart failure group and 177 in the heart failure group. Download the methylation data for GSE197670 from the NCBI GEO public database, with annotation file GPL13534, which includes 19 samples, including 3 control samples and 16 in the disease group.
[0040] 2. Single-cell data quality control First, we read the expression profiles using the Seurat package, filtering cells based on the total UMI (Unique Mitochondrial Indices) of each cell, the number of genes expressed, and the mitochondrial expression percentage. The mitochondrial gene expression percentage refers to the percentage of total mitochondrial gene expression relative to the total expression of all genes. Cells with high mitochondrial gene expression percentages have low RNA expression levels, indicating they are entering a cell death process. We performed quality control using the median absolute deviation (MAD). Generally, if a variable deviates from the median by more than 3 MADs, it is considered an outlier and needs to be removed. Then, we used DoubletFinder (V2.0.4) to filter the two-cell samples for each cell, thus completing the cell quality control.
[0041] 3. Dimensionality reduction, clustering, and annotation of single-cell data We employed the global normalization method LogNormalize, adjusting the total expression level of each cell to 10,000 by multiplying by a coefficient s0, and then performing logarithmic normalization. CellCycleScoring was used to calculate cell cycle scores. FindVariableFeatures was used to identify hypervariable genes. The ScaleData function was used to remove gene expression fluctuations caused by differences in mitochondrial gene expression ratios, ribosomal gene expression ratios, and cell cycle stages. RunPCA was used for linear dimensionality reduction of the expression matrix, and principal components were selected for subsequent analysis. Harmony was used to remove batch effects, and the RunUMAP function was used for UMAP nonlinear dimensionality reduction to display the global structural relationships between cells in a two-dimensional format. Cell types and corresponding marker genes in the corresponding tissues were identified primarily by querying the CellMarker and PangaoDB databases and literature, supplemented by automated annotation using SingleR software.
[0042] 4. Differential Expression Analysis The Limma package is an R software package used for differential expression analysis of expression profiles to identify genes with significant differential expression between groups. The Limma package was used to analyze the differences in molecular mechanisms between disease data, identifying differentially expressed genes between control and disease samples. The differential gene screening criteria were P-value < 0.05 and |logFC| > 0.585. Preprocessing was performed using the Champ package: probe filtering and standardization (including BMIQ correction) were performed on methylation data; differential methylation sites were screened using differential methylation analysis methods, with an adjusted P-value < 0.05 and an absolute value of methylation difference > 0.1 as the significance threshold, distinguishing between high and low methylation sites in the disease group. The methylation patterns of significant DMPs in the two groups were visualized using heatmaps. Finally, probes corresponding to significant DMPs were annotated to genes to obtain differentially methylated genes, providing a foundation for subsequent research.
[0043] 5. Gene set differential analysis (GSVA) To assess the potential changes in biological function among different samples, this study employed Gene Set Variation Analysis (GSVA) to perform gene set enrichment analysis on transcriptome data. GSVA is a non-parametric, unsupervised analysis method that translates changes in gene expression levels into changes at the pathway level. We downloaded the required gene sets from the Molecular Signatures Database (MSigDB) and used the GSVA algorithm to calculate the comprehensive score for each sample on each gene set, thereby comparing the differences in biological function among different samples.
[0044] 6. Analysis of immune cell infiltration To infer the relative proportions of immune cells in the samples, this study used the CIBERSORT algorithm to perform immune infiltration analysis on the transcriptome expression matrix. CIBERSORT, based on the principle of support vector regression (SVR), uses a feature matrix containing 547 immune marker genes to deconvolve mixed expression signals, thereby distinguishing 22 human immune cell subtypes, including T cells, B cells, plasma cells, and various myeloid cell subsets.
[0045] 7. Transcriptional Regulation Analysis of Key Genes The prediction of transcription factor regulatory relationships of key genes was performed using the RcisTarget R package. This method first calculates the area under the curve (AUC) of each motif based on the gene set's recovery curve, and then obtains a normalized enrichment score (NES) using the AUC distribution of all motifs in the database. Subsequently, annotation is performed by combining motif similarity and gene sequence information to identify the set of transcription factors significantly associated with key genes.
[0046] 8. Nomogram Model Construction To assess the clinical application potential of key genes in disease risk prediction, this study constructed a nomogram model based on multivariate regression analysis. Each variable was scored according to its regression coefficient, and the scores of all variables were summed to obtain a total score, from which the corresponding predicted probabilities were calculated. This model can intuitively present the contribution of different variables to the risk of disease occurrence.
[0047] 9. Correlation Analysis To investigate the expression relationship between key genes and genes related to lactation and fibrosis in heart failure, this study calculated Pearson correlation coefficients in transcriptomic data, setting the screening criteria as |R|>0.5 and P<0.05. Lactic acidification and fibrosis genes significantly associated with key genes were identified.
[0048] Example 1: like Figures 1-4 As shown, this embodiment verifies the effect of gene knockdown / overexpression at the cellular level, specifically involving single genes and any combination of two or more genes. The experimental analysis charts and conclusions for each experimental group and control group are presented below. The detailed steps of the overall experimental process at the cellular level are as follows.
[0049] I. Extraction and Culture of Primary Mouse Cardiac Fibroblasts (MCFs) 1. Tissue Acquisition and Preprocessing Sample collection: One-day-old suckling mice were euthanized by cervical dislocation and their entire bodies were disinfected with 75% alcohol for 5 minutes. The hearts were quickly removed in a laminar flow hood and placed in pre-cooled sterile PBS.
[0050] Cleaning and trimming: Rinse the heart repeatedly with PBS until the blood is washed away. Remove non-target tissues such as the aorta and atria, retaining the ventricles and trimming them thoroughly (each piece should be approximately 1 mm³ in size).
[0051] 2. Cell isolation Transfer tissue fragments to centrifuge tubes and add a mixed digestion solution (0.1% trypsin + 0.1% type II collagenase, containing 30 mmol / L taurine, 10 mmol / L HEPES, and 0.5 μg / ml DNas). Digest in a 37°C water bath with shaking for 12 minutes. After standing for a short time, collect the supernatant containing cells and immediately add serum-containing culture medium to stop digestion. Add digestion solution back to the remaining pellet and repeat the digestion step 5 times until the tissue fragments are completely dissolved. Combine all digestion solutions and filter through a 200-mesh cell strainer to remove large impurities. Centrifuge at 300g for 10 minutes, discard the supernatant, resuspend the cell pellet in complete culture medium, and perform cell counting.
[0052] 3. Purification After obtaining cells via enzymatic digestion and seeding, culture for 2 hours. At this point, most fibroblasts have adhered to the culture vessel, while other cell types (such as cardiomyocytes) remain mostly in suspension. The supernatant is aspirated to remove unadhered cells, and fresh complete culture medium is added for continued culturing.
[0053] 4. Cultivation and Identification Culture conditions: Use DMEM high-glucose medium containing 10% fetal bovine serum. Culture environment: 37℃, 5% CO2. Change the medium every 2 days.
[0054] Passaging: Passage the cells when they reach 90% confluence. Discard the old culture medium, wash with PBS, and then digest with 0.25% trypsin (containing EDTA). When the intercellular spaces become larger and rounder under a microscope, add serum-containing culture medium to stop the digestion. After centrifugation and resuspending, culture in separate flasks.
[0055] II. Induction of MCF fibrosis Second-generation fibroblasts with good growth and approximately 70% confluence were selected. The culture medium for the induction group was replaced with fresh complete medium containing TGF-β1 (10 ng / mL). The control group received physiological saline. The expression levels of α-SMA, type I collagen (COL1A2), and fibrosis-related proteins (FN) were measured after 48 hours.
[0056] III. Determination of Lactic Acid Content The cell culture supernatant and intracellular lactate content were detected using a colorimetric assay kit. For the culture supernatant, it was collected and centrifuged at 3000g for 10 minutes to remove the precipitate, following the manufacturer's instructions. For the cells, the culture supernatant was discarded, and the cells were gently washed twice with pre-chilled PBS. 1 mL of lactate extraction buffer was added per 5 million cells, and the cells were sonicated on ice (300W, 3 seconds, 10-second interval, repeated 10 times). The lysis buffer was then centrifuged at 12000g for 10 minutes at 4°C, and the supernatant was collected as the intracellular lactate sample. The absorption peak at 530 nm was detected using a microplate reader.
[0057] IV. MCFs Transfection Use second-generation, well-grown MCFs. One day before transfection, plate the cells in 6-well plates using antibiotic-free complete medium to achieve 50% cell confluence.
[0058] Purchase the mouse siRNA sequence of Gfpt2 (HY-RS19199) to knock down the gene expression, and prepare and use it according to the instructions.
[0059] Using the pcDNA3.1 overexpression vector, recombinant plasmids for overexpression of mouse Lum, Tnxb, and Thbs4 were synthesized, and the plasmids were transfected into MCFs using Lipofectamine 3000 according to the instructions.
[0060] V. Protein Extraction and Expression Detection Discard the culture medium and gently wash the cells twice with pre-chilled PBS. Add pre-chilled RIPA lysis buffer at a ratio of 100 μL per 100 mm culture dish, ensuring even coverage of the cells. Incubate on ice for 30 minutes, shaking the culture dish occasionally. Quickly scrape the cells off with a cell scraper and transfer the lysate to a pre-chilled centrifuge tube. Centrifuge the lysis buffer at 12000g for 15 minutes at 4°C. Carefully aspirate the supernatant (i.e., the total protein solution) and transfer it to a new pre-chilled centrifuge tube. Determine the protein concentration using the BCA method, then mix with 5× loading buffer in the correct proportion and boil for 10 minutes to denature the protein before use.
[0061] Using a 10% SDS-PAGE gel concentration, the protein expression levels of α-SMA, type I collagen, and fibronectin (FN) were detected through three steps: electrophoresis, membrane transfer, and exposure.
[0062] Example 2: like Figures 5-6 As shown, this embodiment verifies the effect of protein expression in an organism (live mouse). The number of experimental samples and the observation period meet the requirements for experimental proof. Specifically, it involves a single protein and two or more arbitrary combinations. The experimental analysis charts and conclusions of each experimental group and control group are described below. The detailed steps of the biological experiment are as follows.
[0063] 1. Grouping of experimental animals Male C57BL / 6J mice aged 6-8 weeks were used, with 10 mice in each group.
[0064] TGF-β1 control group: Received AAV9 empty vector + TGF-β1 slow-release pump.
[0065] Virus treatment + TGF-β1 group: Received mixed virus (AAV9-Gfpt2-shRNA+AAV9-Lum+AAV9-Tnxb+AAV9-Thbs4) + TGF-β1 slow release pump.
[0066] 2. Preparation and injection of AAV virus (1) Virus construction: Gfpt2 knockout: Construct an AAV9 vector carrying a heart-specific promoter (cTnT) and an effective shRNA targeting Gfpt2.
[0067] Overexpression of Lum, Tnxb, and Thbs4: AAV9 vectors carrying the cTnT promoter and corresponding mouse cDNA were constructed, respectively.
[0068] (2) Virus injection: Route: Tail vein injection. AAV9 can be effectively targeted to the heart after intravenous injection.
[0069] Dosage: Total viral dose of 1×10 12 vg / unit. The ratio of knockdown to overexpression of the virus was 1:1.
[0070] Timing: Viral injection was performed 3 weeks before induced fibrosis (implantation of TGF-β1 sustained-release pump).
[0071] 3. Osmotic micropump implantation surgery Mice were anesthetized with isoflurane inhalation, their backs were shaved and disinfected, a small incision was made, and a pump (0.5 μL / h) was aseptically implanted subcutaneously, followed by suturing. The pump was filled with recombinant mouse TGF-β1 dissolved in PBS at a dose of 1.0 μg / kg / day for 14 days. The control group's pump was filled with an equal volume of PBS.
[0072] 4. Sampling at the endpoint Mouse weight was recorded, and mice were deeply anesthetized for echocardiography to assess cardiac function (ejection fraction EF%, fractional shortening FS%). Blood was collected via the abdominal aorta, and the heart was perfused with pre-cooled PBS. The heart was quickly removed, placed on ice, weighed, and the heart weight / body weight ratio (HW / BW) was calculated. The apical portion of the heart was immediately flash-frozen in liquid nitrogen and then stored at -80°C for protein and RNA extraction. Transverse sections of the mid-ventricle were fixed in 4% paraformaldehyde for 24 hours, embedded in paraffin, and used for histological analysis.
[0073] 5. Histological assessment of cardiac fibrosis Masson trichrome staining provides a visual representation of collagen fiber deposits in the myocardial interstitium and around blood vessels.
[0074] 6. Detection of expression levels of fibrosis-related proteins Western blot was used to detect the expression levels of α-SMA, COL1A2, and FN fibrosis-related proteins.
[0075] Example 3: like Figures 7-8 As shown, this embodiment provides a cardiovascular disease heart failure risk warning and prognostic assessment system, specifically involving a heart failure prognostic assessment model based on a specific combination of gene markers (GFPT2, LUM, TNXB, THBS4) and its construction method, as well as the application of this model in assessing patient risk stratification, predicting disease progression and guiding individualized treatment. It is of great significance for achieving early warning, risk stratification and targeted intervention of heart failure.
[0076] This invention first identified four genes—GFPT2, LUM, TNXB, and THBS4—that play key roles in the development and progression of heart failure through multi-omics data integration and analysis. Specific findings include: 1. Integrated findings: In single-cell transcriptome analysis, tissue-level transcriptome differential analysis, and whole-genome DNA methylation analysis, GFPT2 showed a consistent activated state of "upregulated expression and increased methylation" in the disease group; while LUM, TNXB, and THBS4 showed a consistent inhibited state of "downregulated expression and decreased methylation".
[0077] 2. Functional correlation: The expression levels of the above-mentioned genes are significantly correlated with the degree of myocardial fibrosis, lactate metabolism, and immune microenvironment characteristics. At the single-cell level, these genes are enriched in fibroblasts, and their expression changes are closely related to cellular lactation and fibrosis scores.
[0078] 3. Experimental verification: Using mouse primary cardiac fibroblasts (MCFs) and mouse in vivo models, it was confirmed that intervention with four genes (knockdown of the pro-fibrotic gene GFPT2, and overexpression of the protective genes LUM, TNXB, and THBS4) can significantly improve TGF-β1-induced abnormal lactate secretion and fibrosis process, and improve cardiac function.
[0079] Based on the above findings, this invention constructs an assessment model for evaluating an individual's risk or prognosis of heart failure. The core of this model is to calculate a comprehensive risk score (RS) based on the expression levels of the four key genes mentioned above.
[0080] I. Construction and Validation of Prognostic Assessment Model (1) Training cohort and data source: Transcriptome expression profile data and corresponding clinical information of 313 samples (137 non-heart failure group and 176 heart failure group) were obtained from the NCBI GEO public database (accession number: GSE57338) as the training cohort.
[0081] (2) Key gene expression level acquisition: Extract the expression values of the four genes GFPT2, LUM, TNXB and THBS4 for each sample from the above expression profile data (e.g., standardized FPKM value or chip signal intensity value).
[0082] (3) Construction of a multivariate Cox proportional hazards regression model: Using the patient's important clinical endpoints (such as all-cause mortality, readmission due to heart failure, etc.) as the dependent variable and the expression levels of the above four genes as the independent variables, a multivariate Cox regression model was constructed. Through regression analysis, the independent contribution weight of each gene expression level to the prognosis (i.e., regression coefficient, β) was determined.
[0083] (4) Establishment of the risk scoring formula: Based on the results of the Cox regression model, the individualized risk score (RS) calculation formula is established as follows: RS = (β_GFPT2 × Exp_GFPT2) + (β_LUM × Exp_LUM) + (β_TNXB × Exp_TNXB) + (β_THBS4 × Exp_THBS4); Where β represents the coefficient of each gene in the multivariate Cox regression, and Exp represents the expression level of that gene in the sample. A higher RS value predicts a poorer clinical prognosis.
[0084] (5) Risk Stratification and Nomogram Visualization: Patients were divided into "high-risk" and "low-risk" groups based on the median RS values of the training cohort or the optimal cut-off value determined by X-tile software. Using R software packages such as "rms," the results of the Cox model were plotted as a nomogram, enabling risk visualization and individualized probability prediction. In the initial construction of this invention, the four-gene model showed extremely high discriminative power in the training cohort, with an area under the ROC curve (AUC) reaching 0.961.
[0085] (6) Model performance verification: Statistical validation: Kaplan-Meier survival analysis compared the survival curves of the high-risk and low-risk groups. The Log-rank test showed a significant difference in prognosis between the two groups (P<0.001). Time-dependent ROC curve analysis showed that the model maintained high predictive accuracy at different time points.
[0086] Biological justification: Gene set enrichment analysis (GSVA) of the high-risk group (i.e., high GFPT2 expression, low LUM / TNXB / THBS4 expression pattern) showed significant enrichment in pathways closely related to inflammation and fibrosis, such as TNFA signaling, IL-6-JAK-STAT3 signaling, and epithelial-mesenchymal transition (EMT). Immune infiltration analysis also showed that this group had specific immune microenvironment characteristics (such as increased infiltration of M2 macrophages and neutrophils), which is highly consistent with the pathogenic mechanism on which the model is based.
[0087] Reverse validation of experimental function: The unique advantage of this invention lies in the fact that the gene markers upon which the model relies have been rigorously validated through functional experiments. As shown in "Experimental Methods and Results," in cell and animal models, reversing the high-risk expression pattern (GFPT2↑, LUM / TNXB / THBS4↓) to the target intervention pattern (GFPT2↓, LUM / TNXB / THBS4↑) effectively inhibits lactate production, reduces fibrosis, and improves cardiac function. This demonstrates, from a causal perspective, that the molecular characteristics upon which the model is based directly drive adverse prognoses, thereby greatly enhancing the model's biological credibility and clinical guidance value.
[0088] II. Application Methods of Prognostic Assessment Models The prognostic assessment model of this invention can be applied to the assessment of clinical or research samples through the following steps: (1) Sample acquisition and processing: Obtain biological samples such as myocardial tissue biopsy samples, peripheral blood mononuclear cells or circulating extracellular vesicles rich in cardiac origin from the individual to be evaluated.
[0089] (2) Biomarker detection: The expression levels of GFPT2, LUM, TNXB and THBS4 in the sample were measured using quantitative polymerase chain reaction (qPCR), high-throughput sequencing (RNA-seq) or protein detection technology (such as immunohistochemistry, ELISA, provided that the protein level and mRNA level expression trends are consistent and verified).
[0090] (3) Risk score calculation: Substitute the measured gene expression levels into the risk score (RS) formula established in this invention to calculate the individual's comprehensive risk score.
[0091] (4) Risk stratification and prognosis: The calculated RS value is compared with the pre-set cutoff value to classify the individual as "high risk of heart failure" or "low risk". At the same time, the expression levels of each gene can be input into Nomograph to directly read the predicted probability of the occurrence of a specific clinical endpoint event.
[0092] (5) Treatment guidance: This model can be used not only for prognostic assessment but also to provide guidance for targeted therapy. For example, patients assessed as "high-risk" can be recommended as candidates for key intervention populations such as those targeting GFPT2 inhibitors or LUM / TNXB / THBS4 agonists / complementary therapies, thereby achieving precision treatment based on molecular stratification.
[0093] The risk assessment model constructed in this invention has the following advantages: High precision and strong mechanism: The model is based on multi-omics integrated analysis from single cells to tissues, from epigenetics to transcriptional expression. The biomarker screening process is rigorous and the biological basis is solid.
[0094] Functional validation support: The core gene markers of the model have been validated by gain-of-function and loss-of-function experiments at the cellular level and in mouse animal models, clarifying their causal role in fibrosis and lactic acid metabolism, thus upgrading the model from "correlation prediction" to "mechanistic prediction".
[0095] Dual-function potential: This model possesses both "prognostic assessment" and "therapeutic targeting" functions. It can not only identify high-risk patients, but its assessment results (i.e., the expression patterns of the four genes) directly point to actionable combination therapy strategies (inhibiting GFPT2 and enhancing LUM / TNXB / THBS4), thus achieving a closed loop between diagnosis and treatment.
[0096] Flexible clinical applications: The model can be adapted to different technology platforms (qPCR, sequencing, etc.) for detection, and can achieve personalized and visualized risk assessment through nomograph, making it easy to promote and use in clinical settings.
[0097] This invention constructs a disease prediction model based on the expression levels of four key genes, and visualizes the regression analysis results using a nomogram. The model shows that each gene contributes to the scoring system to varying degrees, indicating that these four genes have potential independent predictive value in disease risk assessment. The cumulative score can intuitively predict an individual's disease probability, providing a reference for clinical stratification management and personalized intervention. Further ROC curve analysis validates the model's performance, showing strong predictive accuracy with an AUC of 96.1%, indicating high sensitivity and specificity in distinguishing samples from different states. Furthermore, the model's high fit and robustness suggest good generalization ability and clinical application potential.
[0098] The specific embodiments described herein are merely illustrative of the spirit of the invention. Those skilled in the art to which this invention pertains may make various modifications or additions to the described specific embodiments or use similar methods to substitute them, without departing from the spirit of the invention or exceeding the scope defined by the appended claims.
Claims
1. A genetic marker of GFPT2, LUM, TNXB, THBS4 alone or in combination, characterized in that, This includes any one or more combinations.
2. The use of a gene marker, alone or in combination, of GFPT2, LUM, TNXB, and THBS4 in the preparation of diagnostic products for risk warning and prognostic assessment of the risk of heart failure caused by hypertensive heart disease, coronary heart disease, cardiomyopathy, myocarditis, or valvular heart disease.
3. A genetic marker composition for diagnosing the risk of heart failure in a patient with cardiovascular disease, characterized by, The composition comprises a reagent for detecting one or more genes, including GFPT2, LUM, TNXB, and THBS4, in a biological sample, wherein the reagent is used to detect gene expression levels or protein levels.
4. A test kit for early warning and prognosis assessment of heart failure in cardiovascular disease patients, characterized in that, include: Primer pairs and / or probes for the specific detection of GFPT2 mRNA, LUM mRNA, TNXB mRNA and THBS4 mRNA; Alternatively, antibodies for the specific detection of GFPT2, LUM, TNXB, and THBS4 proteins; And necessary auxiliary reagents, including buffers, enzymes, and substrates.
5. A cardiovascular disease heart failure risk warning and prognostic assessment system, characterized in that, include: The detection kit is used to measure the expression level or protein concentration of four genes, GFPT2, LUM, TNXB, and THBS4, in biological samples. The data processing unit has a pre-stored risk assessment model, which is configured to receive measurement data of four gene markers and calculate a comprehensive risk score based on their expression levels (GFPT2 upregulated, LUM, TNXB, and THBS4 downregulated). The output unit is used to output the risk score and the corresponding heart failure risk level conclusion (such as low risk, medium risk, high risk).