Cardiovascular disease
Patent Information
- Application Number
- JP2023554826
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-03-08
- Filing Date
- 2022-03-08
- Publication Date
- 2025-08-15
AI Technical Summary
The complexity of cardiovascular disease (CVD) limits the predictive power of single biomarkers, necessitating a combination of genetic and omics data to identify molecular signatures for stratifying populations and optimizing clinical trial design and drug efficacy.
A method involving the detection of multiple biomarkers, including TNF-α, GSTA1, NT-proBNP, RORA, TNC, GHR, A2M, IGFBP2, APOB, SEPP1, TFF3, IL6, CHI3L1, MET, GDF15, CCL22, TNFRSF11, ANGPT2, and ReIA NF-KB, to determine their expression levels, amounts, and activities, allowing for the identification of individuals at risk of CVD through comparison with healthy controls.
Enables early identification of individuals at risk for CVD events, facilitating timely interventions to prevent or reduce the risk of the disease through personalized strategies.
Smart Images

Figure 00000061_0000 
Figure 00000061_0001 
Figure 00000062_0000
Abstract
Description
[Technical field]
[0001] The present invention relates to cardiovascular disease and in particular, but not exclusively, to cardiovascular disease biomarkers and their use in diagnostic and prognostic methods. The invention also extends to diagnostic and prognostic kits utilising the biomarkers of the invention for diagnosing or prognosticating cardiovascular disease. [Background technology]
[0002] Due to the complexity of cardiovascular disease (CVD), the ability of single biomarkers to predict cardiovascular outcome (CVO) risk is limited (De Lemos et al. 2017), but many gene loci and protein biomarkers (BMKs) have been discovered. Combining genetic information with other types of omics data should have a higher predictive ability to detect population subgroups and corresponding molecular signatures associated with clinical outcomes (Vilne et al. 2018). This should point to molecular signatures for CV risk and their potential to stratify populations with respect to disease progression. A better understanding of the drivers of T2D-related cardiovascular disease will help define subpopulations that may differ with respect to disease progression. Summary of the Invention [Problem to be solved by the invention]
[0003] We developed and tested a workflow to identify sub-structures of study populations by molecular signatures of protein and gene biomarkers. We discovered signatures that define sub-populations that progressed differently to CVD, suggesting that these signatures reflect different stages of disease progression. Combinations of biomarkers that reflect differences in CVD progression will be identified, providing strategies for optimizing clinical trial design, drug efficacy, and thus treatment. [Means for solving the problem]
[0004] Thus, in a first aspect of the present invention there is provided a method of determining, diagnosing and / or predicting the risk of suffering from cardiovascular disease in an individual, the method comprising: a. In a sample obtained from an individual: Tumor necrosis factor (TNF)-α, β glutathione S-transferase alpha 1 (GSTA1); N-terminal-prohormone BNP (NT-proBNP); retinoic acid receptor-related orphan receptor alpha (RORA); tenascin-C (TNC); growth hormone receptor (GHR); alpha-2-macroglobulin (A2M); insulin-like growth factor binding protein 2 (IGFBP2); apolipoprotein B (APOB); selenoprotein P (SEPP1); trefoil factor (TFF3); interleukin 6 (IL6); chitinase 3-like 1 (CHI3L1); hepatocyte growth factor receptor (MET); growth differentiation factor 15 (GDF15); chemokine (CC motif) ligand 22 (CCL22); tumor necrosis factor receptor superfamily, member 11 (TNFRSF11); angiopoietin 2 (ANGPT2); and v-Rel avian reticuloendotheliosis viral oncogene homolog A nuclear factor-kappa B (ReLA detecting the expression level, amount and / or activity of two or more biomarkers selected from the group consisting of: NF-KB; b. comparing the expression level, amount and / or activity of the biomarker with a reference from a healthy control population; and c. Determining, diagnosing, and / or predicting an individual's risk of suffering from cardiovascular disease when the expression level, amount, and / or activity of the biomarker deviates from a reference from a healthy control population. Includes.
[0005] Advantageously, the methods of the invention allow for the identification of individuals at risk of suffering from a CVD event, and therefore allow for early intervention to prevent or reduce the risk of the individual suffering from a CVD event. In particular, the separate detection of each biomarker allows for the identification of individuals at risk of suffering from a CVD event, and the detection of multiple biomarkers, i.e., a biomarker signature, provides a particularly effective means of allowing early intervention to prevent or reduce the risk of the individual suffering from a CVD event.
[0006] It will be appreciated by those skilled in the art that the term prediction may relate to predicting the rate and / or duration of progression or improvement of cardiovascular disease in an individual suffering from cardiovascular disease.
[0007] The method may be performed in vivo, in vitro or ex vivo. Preferably, the method is performed in vitro or ex vivo. Most preferably, the method is performed in vitro.
[0008] The expression level may relate to the level or concentration of a biomarker polynucleotide sequence. The polynucleotide sequence may be DNA or RNA. The DNA may be genomic DNA. The RNA may be mRNA.
[0009] The amount of a biomarker can be related to the concentration of the biomarker polypeptide sequence.
[0010] Biomarker activity may relate to an activity of the biomarker protein, preferably an activity associated with an interaction with the biomarker.
[0011] Classification biomarkers: LLQQ=2 → inactivation(-1) LLQQ=3~4 → No activation (0) LLQQ=5 and higher levels → Activation (+1)
[0012] For quantitative biomarkers, we used tertiles of the distribution. Thus, the expression level, amount, and / or activity of biomarkers can be detected by: sequencing methods (e.g., Sanger, next generation sequencing, RNA-SEQ), hybridization-based methods, including those used in biochip arrays, mass spectrometry (e.g., laser desorption / ionization mass spectrometry), fluorescence (e.g., sandwich immunoassays), surface plasmon resonance, ellipsometry, and atomic force microscopy. Expression levels of markers (e.g., polynucleotides, polypeptides, or other analytes) can be compared by procedures well known in the art, such as RT-PCR, Northern blotting, Western blotting, flow cytometry, immunohistochemistry, magnetic and / or antibody-coated beads, in situ hybridization, fluorescent in situ hybridization (FISH), flow chamber adhesion assays, ELISA, microarray analysis, or colorimetric assays. The methods may further include one or more of electrospray ionization mass spectrometry (ESI-MS), ESI-MS / MS, ESI-MS / (MS)n, matrix-assisted laser desorption ionization time-of-flight mass spectrometry (MALDI-TOF-MS), surface-enhanced laser desorption / ionization time-of-flight mass spectrometry (SELDI-TOF-MS), desorption / ionization on silicon (DIOS), secondary ion mass spectrometry (SFMS), quadrupole time-of-flight (Q-TOF), atmospheric pressure chemical ionization mass spectrometry (APCI-MS), APCI-MS / MS, APCI-(MS)n, atmospheric pressure photoionization mass spectrometry (APPI-MS), APPI-MS / MS, and APPI-(MS)n, quadrupole mass spectrometry, Fourier transform mass spectrometry (FTMS), and ion trap mass spectrometry, where n is an integer greater than 0.
[0013] Thus, the method of detection may involve a probe that can hybridize to the biomarker DNA or RNA sequence. The term "probe" may be defined as an oligonucleotide. A probe may be single-stranded upon hybridization to a target. Probes include, but are not limited to, primers, i.e., oligonucleotides that can be used to prime a reaction, for example, at least in a PCR reaction.
[0014] Preferably, a decrease in the expression, amount and / or activity of TNF-α, GSTA1, NT-proBNP, RORA and / or TNC compared to the reference indicates that the individual has a higher risk of suffering from or having a poor prognosis due to cardiovascular disease.
[0015] Preferably, an increase in the expression, amount and / or activity of GHR, A2M, IGFBP2, APOB, SEPP1, TFF3, IL6 and / or CHI3L1 compared to a reference indicates that an individual has a higher risk of suffering from or having a poor prognosis due to cardiovascular disease.
[0016] Preferably, a decrease in the expression, amount and / or activity of MET, GDF15, CCL22, TNFRSF11, ANGPT2 and / or ReIA NF-KB compared to the reference indicates that the individual has a lower risk of suffering from or having a good prognosis for cardiovascular disease.
[0017] Preferably, step a) comprises detecting the expression level, amount and / or activity of at least three biomarkers selected from the group consisting of TNF-□; GSTA1; NT-proBNP; RORA; TNC; GHR; A2M; IGFBP2; APOB; SEPP1; TFF3; IL6; CHI3L1; MET; GDF15; CCL22; TNFRSF11; ANGPT2; and ReIA NF-KB in a sample obtained from the individual. However, step a) may comprise detecting the expression level, amount and / or activity of at least four or at least five biomarkers.
[0018] Step a) may comprise detecting the expression level, amount and / or activity of at least 6 biomarkers or at least 7 biomarkers. Alternatively, step a) may comprise detecting the expression level, amount and / or activity of at least 8 biomarkers or at least 9 biomarkers. In another embodiment, step a) may comprise detecting the expression level, amount and / or activity of at least 10 biomarkers, at least 11 biomarkers, at least 12 biomarkers, at least 13 biomarkers, at least 14 biomarkers or at least 15 biomarkers. In another embodiment, step a) may comprise detecting the expression level, amount and / or activity of at least 16 biomarkers, at least 17 biomarkers or at least 18 biomarkers.
[0019] Preferably, step a) comprises detecting the expression level, amount and / or activity of the following biomarkers in a sample obtained from the individual: TNF-□; GSTA1; NT-proBNP; RORA; TNC; GHR; A2M; IGFBP2; APOB; SEPP1; TFF3; IL6; CHI3L1; MET; GDF15; CCL22; TNFRSF11; ANGPT2 and ReIA NF-KB.
[0020] The cardiovascular disease may be selected from the group consisting of cardiovascular mortality; myocardial infarction; stroke; and heart failure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0021] In one embodiment, RORA is provided by gene bank locus ID: HGNC:10258; Entrez Gene:6095; and / or Ensembl:ENSG00000069667. The protein sequence may be represented by GeneBank ID P35398-2, provided herein as SEQ ID NO:1, and is as follows: MESAPAAPDPAASEPGSSGADAAAGSRETPLNQESARKSEPPAPVRRQSYSSTSRGISVTKKTHTSQIEIIPCKICGDKSSGIHYGVITCEGCKGFFRRSQQSNATYSCPRQKNCLIDRTSRNRCQHCRL QKCLAVGMSRDAVKFGRMSKKQRDSLYAEVQKHRMQQQQRDHQQQPGEAEPLTPTYNISANGLTELHDDLSNYIDGHTPEGSKADSAVSSFYLDIQPSPDQSGLDINGIKPEPICDYTPASGFFPYCSFTN GETSPTVSMAELEHLAQNISKSHLETCQYLREELQQITWQTFLQEEIENYQNKQREVMWQLCAIKITEAIQYVVEFAKRIDGFMELCQNDQIVLLKAGSLEVVFIRMCRAFDSQNNTVYFDGKYASPDVFK SLGCEDFISFVFEFGKSLCSMHLTEDEIALFSAFVLMSADRSWLQEKVKIEKLQQKIQLALQHVLQKNHREDGILTKLICKVSTLRALCGRHTEKLMAFKAIYPDIVRLHFPPLYKELFTSEFEPAMQIDG [SEQ ID NO: 1]
[0022] Thus, preferably, RORA comprises or consists of an amino acid sequence substantially as shown in SEQ ID NO:1, or a fragment or variant thereof.
[0023] In one embodiment, RORA is encoded by the nucleotide sequence provided herein as SEQ ID NO:2, as follows: [SEQ ID NO:2]
[0024] Thus, preferably, RORA comprises or consists of a nucleotide sequence substantially as shown in SEQ ID NO:2, or a fragment or variant thereof. → In one embodiment, GHR is provided by GeneBank locus ID: HGNC:4263; Entrez Gene:2690; and / or Ensembl:ENSG00000112964. The protein sequence is represented by GeneBank ID P10912, which is provided herein as SEQ ID NO: 4, as follows: MDLWQLLLTLALAGSSDAFSGSEATAAILSRAPWSLQSVNPGLKTNSSKEPKFTKCRSPERETFSCHWTDEVHHGTKNLGPIQLFYTRRNTQEWTQEWKECPDYVSAGENSCYFNSSFTSIWIPYCIKLTSNGGTVDEKCFSVDEIVQPDPPIALNWTL LNVSLTGIHADIQVRWEAPRNADIQKGWMVLEYELQKGWMVLEYELQYKEVNETKWKMMDPILTTSVPVYSLKVDKEYEVRVRSKQRNSGNYGEFSEVLYVTLPQMSQFTCEEDFYFPWLLIIIFGIFGLTVMLFVFLFSKQQRIKMLILPPVPVPKIKGIDPDLLKEGKL EEVNTILAIHDSYKPEFHSDDSWVEFIELDIDEPDEKTEESDTDRLLSSSDHEKSHSNLGVKDGDSGRTSCCEPDILETDFNANDIHEGTSEVAQPQRLKGEADLLCLDQKNQNNSPYHDACPATQQPSVIQAEKNKPQPLPTEGAESTHQAAHIQLSNP SSLSNIDFYAQVSDITPAGSVVLSPGQKNKAGMSQCDMHPEMVSLCQENFLMDNAYFCEADAKKCIPVAPHIKVESHIQPSLNQEDIYITTESLTTAAGRPGTGEHVPGSEMPVPDYTSIHIVQSPQGLILNATALPLPDKEFLSSCGYVSTDQLNKIMP [Sequence number 4].
[0025] Thus, preferably the GHR comprises or consists of an amino acid sequence substantially as shown in SEQ ID NO: 4, or a fragment or variant thereof.
[0026] In one embodiment, the GHR is encoded by the nucleotide sequence provided herein as SEQ ID NO:5, as follows: [SEQ ID NO:5]
[0027] Thus, preferably the GHR comprises or consists of a nucleotide sequence substantially as shown in SEQ ID NO:5, or a fragment or variant thereof.
[0028] In one embodiment, TNF-□ is provided by GeneBank locus ID: HGNC:11892; Entrez Gene:7124; and / or Ensembl:ENSG00000232810. The protein sequence may be represented by GeneBank ID P01375, which is provided herein as SEQ ID NO:6, as follows: MSTESMIRDVELAEEALPKKTGGPQGSRRCLFLSLFSFLIVAGATTLFCLLHFGVIGPQREEFPRDLSLISPLAQAVRSSSRTPSDKPVAHVVANPQAEGQLQWLNRRANALLANG VELRDNQLVVPSEGLYLIYSQVLFKGQGCPSTHVLLTHTISRIAVSYQTKVNLLSAIKSPCQRETPEGAEAKPWYEPIYLGGVFQLEKGDRLSAEINRPDYLDFAESGQVYFGIIAL [SEQ ID NO:6]
[0029] Thus, preferably TNF-□ comprises or consists of an amino acid sequence substantially as shown in SEQ ID NO: 6, or a fragment or variant thereof.
[0030] In one embodiment, TNF-□ is encoded by the nucleotide sequence provided herein as SEQ ID NO:7, as follows: [SEQ ID NO: 7]
[0031] Thus, preferably TNF-□ comprises or consists of a nucleotide sequence substantially as shown in SEQ ID NO: 7, or a fragment or variant thereof.
[0032] In one embodiment, GSTA1 is provided by GeneBank locus ID: HGNC:4626; Entrez Gene:2938; Ensembl:ENSG00000243955; OMIM:138359; and / or UniProtKB:P08263. The protein sequence may be represented by GeneBank ID:ENST00000334575.6, which is provided herein as SEQ ID NO:8, as follows: MAEKPKLHYFNARGRMESTRWLLAAAGVEFEEKFIKSAEDLDKLRNDGYLMFQQVPMVEIDGMKLVQTRAILNYIASKYNLYGKDIKERALIDMYIEGIADLGEMILLLPV CPPEEKDAKLALIKEKIKNRYFPAFEKVLKSHGQDYLVGNKLSRADIHLVELLYYVEELDSSLISSFPLLKALKTRISNLPTVKKFLQPGSPRKPPMDEKSLEEARKIFRF [SEQ ID NO:8]
[0033] Thus, preferably GSTA1 comprises or consists of an amino acid sequence substantially as shown in SEQ ID NO:8, or a fragment or variant thereof.
[0034] In one embodiment, GSTA1 is encoded by the nucleotide sequence provided herein as SEQ ID NO:9, as follows: [SEQ ID NO: 9]
[0035] Thus, preferably GSTA1 comprises or consists of a nucleotide sequence substantially as shown in SEQ ID NO: 9, or a fragment or variant thereof.
[0036] In one embodiment, NT-proBNP is provided by GeneBank locus ID: HGNC:7940; Entrez Gene:4879; Ensembl:ENSG00000120937; OMIM:600295; and / or UniProtKB:P16860. The protein sequence may be represented by GeneBank ID P16860, which is provided herein as SEQ ID NO: 10, as follows: MDPQTAPSRALLLLLFLHLAFLGGRSHPLGSPGSASDLETSGLQEQRNHLQGKLSELQVEQTSLEPLQESPRPTGVWKSREVATEGIRGHRKMVLYTLRAPRSPKMVQGSGCFGRKMDRISSSSGLGCKVLRRH [SEQ ID NO: 10]
[0037] Thus, preferably, the NT-proBNP comprises or consists of an amino acid sequence substantially as shown in SEQ ID NO: 10, or a fragment or variant thereof.
[0038] In one embodiment, NT-proBNP is encoded by the nucleotide sequence provided herein as SEQ ID NO:11, as follows: AGGAGGAGCACCCCGCAGGCTGAGGGCAGGTGGGAAGCAAACCCGGACGCATCGCAGCAGCAGCAGCAGCAGCAGAAGCAGCAGCAGCAGCCTCGCAGTCCCTCCAGAGACATGGATCCCCAGACAGCACCTTCCGGGCGCTCCTGCTCCTGCTCTTCTTGCATCTGGCTTTCCT GGGAGGTCGTTCCCACCCGCTGGGCAGCCCGGTTCAGCCTCGGACTTGGAAACGTCCGGGTTACAGGAGCAGCGCAACCATTTGCAGGGCAAACTGTCGGAGCTGCAGGTGGAGCAGACATCCCTGGAGCCCCTCCAGGAGAGCCCCGTCCCACAGGTGTCTGGAAGTCCCGGGA GGTAGCCACCGAGGGCATCCGTGGGCACCGCAAAATGGTCCTCTACACCCTGCGGGCACCACGAAGCCCCAAGATGGTGCAAGGGTCTGGCTGCTTTGGGAGGAAGATGGACCGGATCAGCTCCTCCAGTGGCCTGGGCTGCAAAGTGCTGAGGCGGCATTAAGAGGAAGTCCTGGC TGCAGACACCTGCTTCTGATTCCACAAGGGGCTTTTTCCTCAACCCTGTGGCCGCCTTTGAAGTGACTCATTTTTTTAATGTATTTATGTATTTATTTGATTGTTTTATAAGATGGTTTCTTACCTTTGAGCACAAAATTTCCACGGTGAAATAAAGTCAACATTATAAGCTTTA [SEQ ID NO: 11]
[0039] Thus, preferably, the NT-proBNP comprises or consists of a nucleotide sequence substantially as shown in SEQ ID NO: 11, or a fragment or variant thereof.
[0040] In one embodiment, TNC is provided by GeneBank locus ID: HGNC:5318; Entrez Gene:3371; Ensembl:ENSG00000041982; OMIM:187380; and / or UniProtKB:P24821. The protein sequence may be represented by GeneBank ID P24821, which is provided herein as SEQ ID NO: 12, as follows: [SEQ ID NO: 12]
[0041] Thus, preferably the TNC comprises or consists essentially of the amino acid sequence shown in SEQ ID NO: 12, or a fragment or variant thereof.
[0042] In one embodiment, TNC is encoded by the nucleotide sequence provided herein as SEQ ID NO: 13, as follows: [SEQ ID NO: 13]
[0043] Thus, preferably the TNC comprises or consists essentially of the nucleotide sequence shown in SEQ ID NO: 13, or a fragment or variant thereof.
[0044] In one embodiment, A2M is provided by GeneBank locus ID: HGNC:7; Entrez Gene:2; Ensembl: ENSG00000175899; OMIM: 103950; and / or UniProtKB: P01023. The protein sequence may be represented by GeneBank ID P01023, which is provided herein as SEQ ID NO: 14, as follows: [SEQ ID NO: 14]
[0045] Thus, preferably, A2M comprises or consists essentially of the amino acid sequence shown in SEQ ID NO: 14, or a fragment or variant thereof.
[0046] In one embodiment, A2M is encoded by the nucleotide sequence provided herein as SEQ ID NO: 15, as follows: [SEQ ID NO: 15]
[0047] Thus, preferably, A2M comprises or consists of a nucleotide sequence substantially as shown in SEQ ID NO: 15, or a fragment or variant thereof.
[0048] In one embodiment, IGFBP2 is provided by GenBank locus ID HGNC:5471; Entrez Gene:3485; Ensembl:ENSG00000115457; OMIM:146731 and / or UniProtKB:P18065.
[0049] The protein sequence can be represented by GeneBank ID P18065, which is provided herein as SEQ ID NO: 16, as follows: MLPRVGCPALPLPPPPLLPLLLLLLGASGGGGGARAEVLFRCPPCTPERLAACGPPPVAPPAAVAAVAGGARMPCAELVREPGCGCCSVCARLEGEACGVYTPRCGQGLRCYPHPGSELPLQALVMGEGTCEKRRDAEYGASPEQVADNGDDHSEGGLVENH VDSTMNMLGGGGSAGRKPLKSGMKELAVFREKVTEQHRQMGKGGKHHLGLEEPKKLRPPPARTPCQQELDQVLERISTMRLPDERGPLEHLYSLHIPNCDKHGLYNLKQCKMSLNGQRGECWCVNPNTGKLIQGAPTIRGDPECHLFYNEQQEARGVHTQRMQ [SEQ ID NO: 16]
[0050] Thus, preferably IGFBP2 comprises or consists of an amino acid sequence substantially as shown in SEQ ID NO: 16, or a fragment or variant thereof.
[0051] In one embodiment, IGFBP2 is encoded by the nucleotide sequence provided herein as SEQ ID NO:17, as follows: [SEQ ID NO: 17]
[0052] Thus, preferably, IGFBP2 comprises or consists of a nucleotide sequence substantially as shown in SEQ ID NO: 17, or a fragment or variant thereof.
[0053] In one embodiment, APOB is provided by GeneBank locus ID: HGNC:603; Entrez Gene:338; Ensembl:ENSG00000084674; OMIM:107730; and / or UniProtKB:P04114. The protein sequence may be represented by GeneBank ID P04114, which is provided herein as SEQ ID NO: 18, as follows: [SEQ ID NO:18]
[0054] Thus, preferably, APOB comprises or consists of an amino acid sequence substantially as shown in SEQ ID NO: 18, or a fragment or variant thereof.
[0055] In one embodiment, APOB is encoded by the nucleotide sequence provided herein as SEQ ID NO:19, as follows: [SEQ ID NO: 19]
[0056] Thus, preferably, the APOB comprises or consists of a nucleotide sequence substantially as shown in SEQ ID NO: 19, or a fragment or variant thereof.
[0057] In one embodiment, SEPP1 is provided by GeneBank locus ID: HGNC:10751; Entrez Gene:6414; Ensembl:ENSG00000250722; OMIM:601484; and / or UniProtKB:P49908. The protein sequence may be represented by GeneBank ID P49908, which is provided herein as SEQ ID NO:20, as follows: MWRSLGLALALCLLPSGGTESQDQSSLCKQPPAWSIRDQDPMLNSNGSVTVVALLQASUY LCILQASKLEDLRVKLKKEGYSNISYIVVNHQGISSRLKYTHLKNKVSEHIPVYQQEENQ TDVWTLLNGSKDDFLIYDRCGRLVYHLGLPFSFLTFPYVEEAIKIAYCEKKCGNCSLTTL KDEDFCKRVSLATVDKTVETPSPHYHHEHHHNHGHQHLGSSELSENQQPGAPNAPTHPAP PGLHHHHKHKGQHRQGHPENRDMPASEDLQDLQKKLCRKRCINQLLCKLPTDSELAPRSU CCHCHRLIFEKTGSAITUQCKENLPSLCSUQGLRAEENITESCQURLPPAAUQISQQLIP TEASASURUKNQAKKUEUPSN [SEQ ID NO:20]
[0058] Thus, preferably, SEPP1 comprises or consists of an amino acid sequence substantially as shown in SEQ ID NO: 20, or a fragment or variant thereof.
[0059] In one embodiment, SEPP1 is encoded by the nucleotide sequence provided herein as SEQ ID NO:21, as follows: T [SEQ ID NO:21]
[0060] Thus, preferably, SEPP1 comprises or consists of a nucleotide sequence substantially as shown in SEQ ID NO: 21, or a fragment or variant thereof.
[0061] In one embodiment, TFF3 is provided by GeneBank locus ID: HGNC:11757; Entrez Gene:7033; Ensembl:ENSG00000160180; OMIM:600633; and / or UniProtKB:Q07654. The protein sequence may be represented by GeneBank ID Q07654, which is provided herein as SEQ ID NO:22, as follows: MKRVLSCVPEPTVVMAARALCMLGLVLALLSSSSAEEYVGLSANQCAVPAKDRVDCGYPH VTPKECNNRGCCFDSRIPGVPWCFKPLQEAECTF [SEQ ID NO:22]
[0062] Thus, preferably TFF3 comprises or consists of an amino acid sequence substantially as shown in SEQ ID NO: 22, or a fragment or variant thereof.
[0063] In one embodiment, TFF3 is encoded by the nucleotide sequence provided herein as SEQ ID NO:23, as follows: [SEQ ID NO:23]
[0064] Thus, preferably TFF3 comprises or consists of a nucleotide sequence substantially as shown in SEQ ID NO: 23, or a fragment or variant thereof.
[0065] In one embodiment, IL6 is provided by GeneBank locus ID; HGNC:6018; Entrez Gene:3569; Ensembl:ENSG00000136244; OMIM:147620; and / or UniProtKB:P05231. The protein sequence may be represented by GeneBank ID P05231, which is provided herein as SEQ ID NO:24, as follows: MNSFSTSAFGPVAFSLGLLLVLPAAFPAPVPPGEDSKDVAAPHRQPLTSSERIDKQIRYILDGISALRKETCNKSNMCESSKEALAENNLNLPKMAEKDGCFQSGF NEETCLVKIITGLLEFEVYLEYLQNRFESSEEQARAVQMSTKVLIQFLQKKAKNLDAITTPDPTTNASLLTKLQAQNQWLQDMTTHLILRSFKEFLQSSLRALRQM [SEQ ID NO:24]
[0066] Thus, preferably IL6 comprises or consists of an amino acid sequence substantially as shown in SEQ ID NO: 24, or a fragment or variant thereof.
[0067] In one embodiment, IL6 is encoded by the nucleotide sequence provided herein as SEQ ID NO:25, as follows: [SEQ ID NO:25]
[0068] Thus, preferably IL6 comprises or consists of a nucleotide sequence substantially as shown in SEQ ID NO: 25, or a fragment or variant thereof.
[0069] In one embodiment, CHI3L1 is provided by GeneBank locus ID; HGNC:1932; Entrez Gene:1116; Ensembl:ENSG00000133048; OMIM:601525; and / or UniProtKB:P36222. The protein sequence may be represented by GeneBank ID P36222, which is provided herein as SEQ ID NO:26, as follows: MGVKASQTGFVVLVLLQCCSAYKLVCYYTSWSQYREGDGSCFPDALDRFLCTHIIYSFANISNDHIDTWEWNDVTLYGMLNTLKNRNPNLKTLLSVGGWNFGSQRFSKIASNTQSRRTFIKSVPPFLRTHGFDGLDLAWLYPGRRDKQHFTTLIKEMKAEFIKEAQPGKKQLLLSAALSAGKVTIDSSYDI AKISQHLDFISIMTYDFHGAWRGTTGHHSPLFRGQEDASPDRFSNTDYAVGYMLRLGAPASKLVMGIPTFGRSFTLASSETGVGAPISGPGIPGRFTKEAGTLAYYEICDFLRGATVHRILGQQVPYATKGNQWVGYDDQESVKSKVQYLKDRQLAGAMVWALDLDDFQGSFCGQDLRFPLTNAIKDALAAT [SEQ ID NO:26]
[0070] Thus, preferably, CHI3L1 comprises or consists of an amino acid sequence substantially as shown in SEQ ID NO: 26, or a fragment or variant thereof.
[0071] In one embodiment, CHI3L1 is encoded by the nucleotide sequence provided herein as SEQ ID NO:27, as follows: [SEQ ID NO:27]
[0072] Thus, preferably, CHI3L1 comprises or consists of a nucleotide sequence substantially as shown in SEQ ID NO:27, or a fragment or variant thereof.
[0073] In one embodiment, MET is provided by GeneBank locus ID: HGNC:7029; Entrez Gene:4233; Ensembl:ENSG00000105976; OMIM:164860; and / or UniProtKB:P08581. The protein sequence may be represented by GeneBank ID P08581, which is provided herein as SEQ ID NO:28, as follows: [SEQ ID NO:28]
[0074] Thus, preferably MET comprises or consists of an amino acid sequence substantially as shown in SEQ ID NO: 28, or a fragment or variant thereof.
[0075] In one embodiment, MET is encoded by the nucleotide sequence provided herein as SEQ ID NO:29, as follows: [SEQ ID NO:29]
[0076] Thus, preferably MET comprises or consists of a nucleotide sequence substantially as shown in SEQ ID NO: 29, or a fragment or variant thereof.
[0077] In one embodiment, GDF15 is provided by GeneBank locus ID: HGNC:30142; Entrez Gene:9518; Ensembl:ENSG00000130513; OMIM:605312; and / or UniProtKB:Q99988. The protein sequence may be represented by GeneBank ID Q99988, which is provided herein as SEQ ID NO:30, as follows: MPGQELRTVNGSQMLLVLLVLSWLPHGALSLAEASRASFPGPSELHSEDSRFRELRKRYEDLLTRLRANQSWEDSNTDLVPAPAVRILTPEVRLGSGGHLHLRISRAALPEGLPEASRLHRALFRLSPTASRSWDVTRPLRRQLSLARPQAPA LHLRLSPPSQSDQLLAESSSARPQLELHLRPQAARGRRRARARNGDHCPLGPGRCCRLHTVRASLEDLGWADWVLSPREVQVTMCIGACPSQFRAANMHAQIKTSLHRLKPDTVPAPCCVPASYNPMVLIQKTDTGVSLQTYDDLLAKDCHCI [SEQ ID NO:30]
[0078] Thus, preferably GDF15 comprises or consists of an amino acid sequence substantially as shown in SEQ ID NO: 30, or a fragment or variant thereof.
[0079] In one embodiment, GDF15 is encoded by the nucleotide sequence provided herein as SEQ ID NO:31, as follows: [SEQ ID NO:31]
[0080] Thus, preferably GDF15 comprises or consists of a nucleotide sequence substantially as shown in SEQ ID NO: 31, or a fragment or variant thereof.
[0081] In one embodiment, CCL22 is provided by GeneBank locus ID: HGNC:10621; Entrez Gene:6367; Ensembl:ENSG00000102962; OMIM:602957; and / or UniProtKB:O00626. The protein sequence may be represented by GeneBank ID O00626, which is provided herein as SEQ ID NO:32, as follows: MDRLQTALLVVLVLLAVALQATEAGPYGANMEDSVCCRDYVRYRLPLRVVKHFYWTSDSC PRPGVVLLTFRDKEICADPRVPWVKMILNKLSQ [SEQ ID NO:32]
[0082] Thus, preferably CCL22 comprises or consists of an amino acid sequence substantially as shown in SEQ ID NO: 32, or a fragment or variant thereof.
[0083] In one embodiment, CCL22 is encoded by the nucleotide sequence provided herein as SEQ ID NO:33, as follows: [SEQ ID NO:33]
[0084] Thus, preferably CCL22 comprises or consists of a nucleotide sequence substantially as shown in SEQ ID NO: 33, or a fragment or variant thereof.
[0085] In one embodiment, TNFRSF11 is provided by GeneBank locus ID: HGNC:11909; Entrez Gene:4982; Ensembl:ENSG00000164761; OMIM:602643; and / or UniProtKB:O00300. The protein sequence may be represented by GeneBank ID O00300, which is provided herein as SEQ ID NO:34, as follows: MNNLLCCALVFLDISIKWTTQETFPPKYLHYDEETSHQLLCDKCPPGTYLKQHCTAKWKTVCAPCPDHYYTDSWHTSDECLYCSPVCKELQYVKQECNRTHNRVCECCKEGRYLEIEFCLKHRSCPPGFGVVQAGTPERNTVCKRCPDGFFSNETSSKAPCRKHTNCSVFGLLLTQKGNATHDNICSGNSESTQKCGIDVT LCEEAFFRFAVPTKFTPNWLSVLVDNLPGTKVNAESVERIKRQHSSQEQTFQLLKLWKHQNKDQDIVKKIIQDIDLCENSVQRHIGHANLTFEQLRSLME SLPGKKVGAEDIEKTIKACKPSDQILKLLSLWRIKNGDQDTLKGLMHALKHSKTYHFPKTVTQSLKKTIRFLHSFTMYKLYQKLFLEMIGNQVQSVKISCL [SEQ ID NO:34]
[0086] Thus, preferably TNFRSF11 comprises or consists of an amino acid sequence substantially as shown in SEQ ID NO: 34, or a fragment or variant thereof.
[0087] In one embodiment, TNFRSF11 is encoded by the nucleotide sequence provided herein as SEQ ID NO:35, as follows: [SEQ ID NO:35]
[0088] Thus, preferably, TNFRSF11 comprises or consists essentially of the nucleotide sequence shown in SEQ ID NO: 35, or a fragment or variant thereof. In one embodiment, ANGPT2 has a genebank locus ID HGNC:485; Entrez Gene:285; Ensembl:ENSG00000091879; OMIM:601922; and / or UniProtKB:O15123. The protein sequence may be represented by Genebank ID O15123, which is provided herein as SEQ ID NO:36, as follows: MWQIVFFTLSCDLVLAAAYNNFRKSMDSIGKKQYQVQHGSCSYTFLLPEMDNCRSSSSPYVSNAVQRDAPLEYDDSVQRLQVLENIMENNTQWLMKLENYIQDNNMKKEMVEIQQNAVQNQTAVM IEIGTNLLNQTAEQTRKLTDVEAQVLNQTTRLELQLLEHSLSTNKLEKQILDQTSEINKLQDKNSFLEKKVLAMEDKHIIQLQSIKEEKDQLQVLVSKQNSIIEELEKKIVTATVNNSVLQKQQ HDLMETVNNLLTMMSTSNSAKDPTVAKEEQISFRDCAEVFKSGHTTNGIYTLTFPNSTEEIKAYCDMEAGGGGWTIIQRREDGSVDFQRTWKEYKVGFGNPSGEYWLGNEFVSQLTNQQRYVLK IHLKDWEGNEAYSLYEHFYLSSEELNYRIHLKGLTGTAGKISSISQPGNDFSTKDGDNDKCICKCSQMLTGGWWFDACGPSNLNGMYYPQRQNTNKFNGIKWYYWKGSGYSLKATTMMIRPADF [SEQ ID NO:36]
[0089] Thus, preferably, ANGPT2 comprises or consists of an amino acid sequence substantially as shown in SEQ ID NO: 36, or a fragment or variant thereof.
[0090] In one embodiment, ANGPT2 is encoded by the nucleotide sequence provided herein as SEQ ID NO:37, as follows: [SEQ ID NO:37]
[0091] Thus, preferably, ANGPT2 comprises or consists of a nucleotide sequence substantially as shown in SEQ ID NO: 37, or a fragment or variant thereof.
[0092] In one embodiment, ReIA NF-KB is provided by GeneBank locus ID: HGNC:9955; Entrez Gene:5970; Ensembl:ENSG00000173039; OMIM:164014; and / or UniProtKB:Q04206. The protein sequence may be represented by GeneBank ID Q04206, which is provided herein as SEQ ID NO:38, as follows: MDELFPLIFPAEPAQASGPYVEIIEQPKQRGMRFRYKCEGRSAGSIPGERSTDTTKTHPTIKINGYTGPGTVRISLVTKDPPHRPHPHELVGKDCRDGFYEAELCPDRCIHSFQNLGIQCVKKRDLEQAISQRIQTN NNPFQEEQRGDYDLNAVRLCFQVTVRDPSGRPLRLPPVLSHPIFDNRAPNTAELKICRVNRNSGSCLGGDEIFLLCDKVQKEDIEVYFTGPGWEARGSFSQADVHRQVAIVFRTPPYADPSLQAPVRVSMQLRRPSD RELSEPMEFQYLPDTDDRHRIEEKRKRTYETFKSIMKKSPFSGPTDPRPPPRRIAVPSRSSASVPKPAPQPYPFTSSLSTINYDEFPTMVFPSGQISQASALAPAPPQVLPQAPAPAPAPAMVSALAQAPAPVPVLA PGPPQAVAPPAPKPTQAGEGTLSEALLQLQFDDEDLGALLGNSTDPAVFTDLASVDNSEFQQLLNQGIPVAPHTTEPMLMEYPEAITRLVTGAQRPPDPAPAPLGAPGLPNGLLSGDEDFSSIADMDFSALLSQISS [SEQ ID NO:38]
[0093] Thus, preferably ReIA NF-KB comprises or consists essentially of the amino acid sequence shown in SEQ ID NO: 38, or a fragment or variant thereof.
[0094] In one embodiment, ReIA NF-KB is encoded by the nucleotide sequence provided herein as SEQ ID NO:39, as follows: [SEQ ID NO:39]
[0095] Thus, preferably, ReIA NF-KB comprises or consists of a nucleotide sequence substantially as shown in SEQ ID NO: 39, or a fragment or variant thereof.
[0096] In one embodiment, the biomarker is RORA, and a decrease in expression, amount and / or activity of RORA compared to a reference indicates that the individual has a higher risk of suffering from cardiovascular disease.
[0097] In one embodiment, the biomarker is GHR, and an increase in the expression, amount and / or activity of GHR compared to a reference indicates that the individual has a higher risk of suffering from cardiovascular disease.
[0098] Preferably, the sample comprises a biological sample. The sample can be any material obtainable from a subject from which protein, RNA and / or DNA can be obtained. Furthermore, the sample can be blood, plasma, serum, spinal fluid, urine, sweat, saliva, tears, breast aspirate, prostatic fluid, semen, vaginal fluid, stool, cervical scraping, cytes, amniotic fluid, intraocular fluid, mucous membranes, exhaled water, animal tissue, cell lysate, tumor tissue, hair, skin, buccal scraping, lymph, interstitial fluid, nail, bone marrow, cartilage, prion, bone powder, ear wax, or combinations thereof.
[0099] Preferably, however, the sample comprises blood, urine or tissue.
[0100] In one embodiment, the sample comprises a blood sample. The blood may be venous or arterial. The blood sample may be assayed immediately. Alternatively, the blood sample may be stored at low temperature, for example in a refrigerator, or frozen until the method is performed. Detection may be performed on whole blood. However, preferably, the blood sample comprises serum. Preferably, the blood sample comprises plasma.
[0101] The blood may be further processed before the method is performed. For example, an anticoagulant such as citrate (e.g., sodium citrate), hirudin, heparin, PPACK, or sodium fluoride may be added. Thus, the sample collection container may contain an anticoagulant to prevent the blood sample from clotting. Alternatively, the blood sample may be centrifuged or filtered to produce a plasma or serum fraction that can be used for analysis. Therefore, the method is preferably performed on a plasma or serum sample. The expression level, amount and / or activity of the biomarkers is preferably measured in vitro from a serum or plasma sample taken from an individual.
[0102] The present invention also provides kits for determining, diagnosing, and / or predicting CVD risk.
[0103] Thus, in a second aspect, there is provided a kit for determining, diagnosing and / or predicting the risk of an individual to suffer from cardiovascular disease, the kit comprising: a. a detection means for detecting the expression level, amount and / or activity of two or more biomarkers selected from the group consisting of TNF-a; GSTA1; NT-proBNP; RORA; TNC; GHR; A2M; IGFBP2; APOB; SEPP1; TFF3; IL6; CHI3L1; MET; GDF15; CCL22; TNFRSF11; ANGPT2 and ReIA NF-KB in a sample obtained from a test subject; and b. Reference values from a healthy control population for the expression level, amount and / or activity of two or more biomarkers selected from the group consisting of TNF-a; GSTA1; NT-proBNP; RORA; TNC; GHR; A2M; IGFBP2; APOB; SEPP1; TFF3; IL6; CHI3L1; MET; GDF15; CCL22; TNFRSF11; ANGPT2 and ReIA NF-KB; Including, wherein the kit comprises: i) a decrease in the expression, amount and / or activity of TNF-a; GSTA1; NT-proBNP; RORA and / or TNC compared to a reference; and / or an increase in the expression, amount and / or activity of GHR; A2M; IGFBP2; APOB; SEPP1; TFF3; IL6 and / or CHI3L1 compared to a reference, for determining, diagnosing and / or predicting that an individual has a higher risk of suffering from cardiovascular disease; and / or ii) A decrease in the expression, amount, and / or activity of MET; GDF15; CCL22; TNFRSF11; ANGPT2, and / or ReIA NF-KB, as compared to a reference, for determining, diagnosing, and / or predicting that an individual has a lower risk of suffering from cardiovascular disease. is used to identify
[0104] The cardiovascular disease, biomarkers, detection and sample may be as defined in the first aspect.
[0105] Preferably, a decrease in the expression, amount and / or activity of TNF-□, GSTA1, NT-proBNP, RORA and / or TNC compared to a reference indicates that an individual has a higher risk of suffering from cardiovascular disease or a poor prognosis.
[0106] Preferably, an increase in the expression, amount and / or activity of GHR, A2M, IGFBP2, APOB, SEPP1, TFF3, IL6 and / or CHI3L1 compared to a reference indicates that an individual has a higher risk of suffering from cardiovascular disease or has a poor prognosis.
[0107] Preferably, a decrease in the expression, amount and / or activity of MET, GDF15, CCL22, TNFRSF11, ANGPT2 and / or ReIA NF-KB compared to the reference indicates that the individual has a lower risk of suffering from cardiovascular disease or has a good prognosis.
[0108] The expression level, amount and / or activity of the biomarker may be as defined in the first aspect.
[0109] The kit may comprise detection means for detecting the expression level, amount and / or activity of at least three biomarkers or at least four biomarkers. The kit may comprise detection means for detecting the expression level, amount and / or activity of at least five biomarkers. The kit may comprise detection means for detecting the expression level, amount and / or activity of at least six biomarkers or at least seven biomarkers. Alternatively, the kit may comprise detection means for detecting the expression level, amount and / or activity of at least eight biomarkers or at least nine biomarkers. In another embodiment, the kit may comprise detection means for detecting the expression level, amount and / or activity of at least ten biomarkers, at least eleven biomarkers, at least twelve biomarkers, at least thirteen biomarkers, at least fourteen biomarkers or at least fifteen biomarkers. In another embodiment, the kit may comprise detection means for detecting the expression level, amount and / or activity of at least sixteen biomarkers, at least seventeen biomarkers or at least eighteen biomarkers.
[0110] Preferably, the kit of the second aspect may be a kit for determining, diagnosing and predicting the risk of an individual to suffer from cardiovascular disease, said kit comprising: a. detection means for detecting the expression level, amount and / or activity of TNF-a; GSTA1; NT-proBNP; RORA; TNC; GHR; A2M; IGFBP2; APOB; SEPP1; TFF3; IL6; CHI3L1; MET; GDF15; CCL22; TNFRSF11; ANGPT2 and ReIA NF-KB in a sample obtained from a test subject; and b. Reference values from a healthy control population for expression levels, amounts and / or activity of TNF-α; GSTA1; NT-proBNP; RORA; TNC; GHR; A2M; IGFBP2; APOB; SEPP1; TFF3; IL6; CHI3L1; MET; GDF15; CCL22; TNFRSF11; ANGPT2 and ReIA NF-KB. wherein the kit comprises: i) a decrease in the expression, amount and / or activity of TNF-a; GSTA1; NT-proBNP; RORA and TNC compared to a reference; and / or an increase in the expression, amount and / or activity of GHR; A2M; IGFBP2; APOB; SEPP1; TFF3; IL6 and CHI3L1 compared to a reference, for determining, diagnosing and / or predicting that an individual has a higher risk of suffering from cardiovascular disease; and ii) A decrease in the expression, amount, and / or activity of MET; GDF15; CCL22; TNFRSF11; ANGPT2, and ReIA NF-KB, as compared to a reference, for determining, diagnosing, and / or predicting that an individual has a lower risk of suffering from cardiovascular disease. is used to identify
[0111] The detection means may detect the expression level, e.g., the level or concentration, of a biomarker polynucleotide sequence, e.g., DNA or RNA. The DNA may be genomic DNA. The RNA may be mRNA. Alternatively, the detection means may detect the polypeptide concentration and / or activity of the biomarker.
[0112] Thus, detection means may include: sequencing methods (e.g., Sanger, next generation sequencing, RNA-SEQ), hybridization-based methods, including those used in biochip arrays, mass spectrometry (e.g., laser desorption / ionization mass spectrometry), fluorescence (e.g., sandwich immunoassays), surface plasmon resonance, ellipsometry, and atomic force microscopy. Expression levels of markers (e.g., polynucleotides, polypeptides, or other analytes) may be compared by procedures well known in the art, such as RT-PCR, Northern blotting, Western blotting, flow cytometry, immunocytochemistry, magnetic and / or antibody-coated beads, in situ hybridization, fluorescent in situ hybridization (FISH), flow chamber adhesion assays, ELISA, microarray analysis, or colorimetric assays. The methods may further include one or more of electrospray ionization mass spectrometry (ESI-MS), ESI-MS / MS, ESI-MS / (MS)n, matrix-assisted laser desorption ionization time-of-flight mass spectrometry (MALDI-TOF-MS), surface-enhanced laser desorption / ionization time-of-flight mass spectrometry (SELDI-TOF-MS), desorption / ionization on silicon (DIOS), secondary ion mass spectrometry (SFMS), quadrupole time-of-flight (Q-TOF), atmospheric pressure chemical ionization mass spectrometry (APCI-MS), APCI-MS / MS, APCI-(MS)n, atmospheric pressure photoionization mass spectrometry (APPI-MS), APPI-MS / MS, and APPI-(MS)n, quadrupole mass spectrometry, Fourier transform mass spectrometry (FTMS), and ion trap mass spectrometry, where n is an integer greater than 0.
[0113] Preferably, the kit comprises a detection means for detecting RORA present in a sample from a test subject, wherein a decrease in RORA expression, amount and / or activity indicates that the individual has a higher risk of suffering from cardiovascular disease.
[0114] Preferably, the kit comprises a detection means for detecting GHR present in a sample from a test subject, wherein increased expression, amount and / or activity of GHR indicates that the individual has a higher risk of suffering from cardiovascular disease.
[0115] Using the methods described herein, the inventors were able to identify SNPs within RORA and GHR that can be used in diagnosis and prognosis, and in particular genetic variants associated with CVD risk.
[0116] Thus, in a third aspect of the present invention, there is provided a method for determining, diagnosing and / or predicting an individual's risk of suffering from cardiovascular disease, the method comprising detecting a single nucleotide polymorphism (SNP) in the RORA gene in a sample obtained from the individual, wherein the presence of the SNP indicates that the individual has an increased risk of suffering from cardiovascular disease.
[0117] Preferably, the RORA, the sample, the detection and the cardiovascular disease are as defined in the first aspect.
[0118] The method may be performed in vivo, in vitro or ex vivo. Preferably, the method is performed in vitro or ex vivo. Most preferably, the method is performed in vitro.
[0119] Preferably, the SNP is located in a region of chromosome 15, preferably at nucleic acid position 60542728 of reference sequence NC_000015.10.
[0120] Preferably, the SNP comprises an adenine (A) to guanine (G) substitution, or an adenine (A) to cytosine (C) substitution.
[0121] Preferably, the SNP comprises an adenine (A) to guanine (G) substitution.
[0122] Thus, preferably the SNP may be referenced by the sequence variant GRCh38.p12 chr 15;NC 000015.10:g.60542728A>G.
[0123] Preferably, the SNP comprises an adenine (A) to cytosine (C) substitution.
[0124] Thus, preferably the SNP may be referenced by the sequence variant GRCh38.p12 chr 15;NC 000015.10:g.60542728A>C.
[0125] Preferably, the SNP may be the reference SNP cluster ID: rs73420079.
[0126] Thus, in one embodiment, the SNP is present in the sequence represented by reference SNP cluster ID rs73420079, referred to herein as SEQ ID NO:3, as follows: AGGCGCACCT CACACGGCAC ACAGGCACAT CTCACACATG GCACACATGC ACACCTCACA CAGATGGCAC ACATGCACAC CTCACACACA CGGCACGCAT GCACACCTCA CACACACGGC ACGCATGCAC ACCTCACACA CGGCACACAT GCACACCTCA CACACGACAC ACGGGCACAC CTCACACACA TGGCACACGG GCACACCTCC CACACACGGC ACACGGGCAC ACCTCCCACA CACGGCACAC V GGCACACCTC AAACGACACA CGGCACACC TCACACACAA GTCTATTCAG CTGCAAGTCC TGCCTCCACT TGCTGAGAAC CTGCATGACT GGGCACCAAG GATACGGCAC ACACACGCAC CCACCCCACA TACATACAGT CCACACACAC ACAACACATA TACACCACAC GCACCACAGA TGCACACCAC ACATGCCACA CACACATACA CTGCACACGC ACCCTACACA CACCCCCCAC ATGCTTACAC [SEQ ID NO:3] where "V" represents the SNP position. Thus, in one embodiment, the SNP may comprise or consist of a sequence substantially as set forth in SEQ ID NO: 3, or a fragment or variant thereof.
[0127] Preferably, the SNP comprises a substitution at nucleic acid position X in SEQ ID NO:3.
[0128] Thus, preferably, the RORA comprises a single nucleotide polymorphism (SNP), the presence of which is associated with an individual having an increased risk of suffering from CVD.
[0129] Preferably, the method for detecting the presence of a SNP comprises a probe capable of hybridizing to a biomarker sequence. Preferably, the probe is capable of hybridizing to SEQ ID NO: 3 such that the SNP is detected.
[0130] In a fourth aspect of the present invention, there is provided a method of determining, diagnosing and / or predicting an individual's risk of suffering from cardiovascular disease, the method comprising detecting a single nucleotide polymorphism (SNP) in the GHR gene in a sample obtained from the individual, wherein the presence of the SNP indicates that the individual has an increased risk of suffering from cardiovascular disease.
[0131] Preferably, the GHR, the sample, the detection and the cardiovascular disease are as defined in the first aspect.
[0132] Preferably, detecting the SNP in a subject is indicative of an increased risk of suffering from cardiovascular disease.
[0133] The method may be performed in vivo, in vitro or ex vivo. Preferably, the method is performed in vitro or ex vivo. Most preferably, the method is performed in vitro.
[0134] Preferably, the SNP is present in a region of chromosome 5, preferably at nucleic acid position 42546623 of reference sequence NC_000005.10.
[0135] Preferably, the SNP comprises a substitution of guanine (G) for adenine (A).
[0136] Preferably, the SNP may be the reference SNP cluster ID rs4314405.
[0137] Thus, in one embodiment, the SNP is present in the sequence represented by the reference SNP cluster ID: rs73420079, referred to herein as SEQ ID NO: 40, as follows: AGGCGCACCT CACACGGCAC ACAGGCACAT CTCACACATG GCACACATGC ACACCTCACA CAGATGGCAC ACATGCACAC CTCACACACA CGGCACGCAT GCACACCTCA CACACACGGC ACGCATGCAC ACCTCACACA CGGCACACAT GCACACCTCA CACACGACAC ACGGGCACAC CTCACACACA TGGCACACGG GCACACCTCC CACACACGGC ACACGGGCAC ACCTCCCACA CACGGCACAC V GGCACACCTC AAACGACACA CGGCACACC TCACACACAA GTCTATTCAG CTGCAAGTCC TGCCTCCACT TGCTGAGAAC CTGCATGACT GGGCACCAAG GATACGGCAC ACACACGCAC CCACCCCACA TACATACAGT CCACACACAC ACAACACATA TACACCACAC GCACCACAGA TGCACACCAC ACATGCCACA CACACATACA CTGCACACGC ACCCTACACA CACCCCCCAC ATGCTTACAC [SEQ ID NO: 40] where "V" represents the SNP position. Thus, in one embodiment, the SNP may comprise or consist of a sequence substantially as set forth in SEQ ID NO: 40, or a fragment or variant thereof.
[0138] Preferably, the SNP comprises a substitution at nucleic acid position X in SEQ ID NO:3.
[0139] Thus, preferably, the GHR comprises a single nucleotide polymorphism whose presence is associated with an individual having an increased risk of suffering from CVD. In a fifth aspect, there is provided GHR and / or RORA for use in diagnosis or prognosis.
[0140] Preferably, RORA and GHR may comprise SNPs as defined in the third and fourth aspects.The cardiovascular disease may be as defined in the first aspect.
[0141] In a sixth aspect, there is provided GHR and / or RORA for use in diagnosing or predicting an individual's risk of suffering from cardiovascular disease.
[0142] Preferably, RORA and GHR may comprise SNPs as defined in the third and fourth aspects.The cardiovascular disease may be as defined in the first aspect.
[0143] Preferably, the method for detecting the presence of the SNP comprises a probe capable of hybridizing to a biomarker sequence. Preferably, the probe is capable of hybridizing to SEQ ID NO: 40 such that the SNP is detected.
[0144] In a seventh aspect, there is provided a kit for determining, diagnosing and / or predicting an individual's risk of suffering from cardiovascular disease, the kit comprising detection means for detecting single nucleotide polymorphisms (SNPs) in the RORA gene and / or the GHR gene in a sample obtained from a test subject, wherein the presence of the SNP is used to determine, diagnose and / or predict that an individual has a higher risk of suffering from cardiovascular disease.
[0145] The RORA gene, the GHR gene, the sample, the detection and the cardiovascular disease may be as defined in the first aspect.
[0146] The detection means may be as defined in the second aspect. The single nucleotide polymorphisms (SNPs) in the RORA gene and / or the GHR gene may be as defined in the third and fourth aspects.
[0147] In an eighth aspect, there is provided a method of treating an individual having a higher risk of suffering from cardiovascular disease, the method comprising: (a) analyzing in a sample obtained from the subject the expression level, amount and / or activity of two or more biomarkers selected from the group consisting of: TNF-□; GSTA1; NT-proBNP; RORA; TNC; GHR; A2M; IGFBP2; APOB; SEPP1; TFF3; IL6; CHI3L1; MET; GDF15; CCL22; TNFRSF11; ANGPT2 and ReIA NF-KB; (b) comparing the expression levels, amount and / or activity of the biomarkers with a reference from a healthy control population, [wherein a decrease in the expression, amount and / or activity of TNF-□, GSTA1, NT-proBNP, RORA and / or TNC compared to the reference, and an increase in the expression, amount and / or activity of GHR, A2M, IGFBP2, APOB, SEPP1, TFF3, IL6 and / or CHI3L1 compared to the reference indicates that the individual has a higher risk of suffering from cardiovascular disease]; and (c) administering or having administered to an individual a therapeutic agent that prevents or reduces the likelihood that the individual will suffer from cardiovascular disease; Includes.
[0148] Preferably, the biomarkers, detection of the biomarkers, cardiovascular disease, expression levels, amount and / or activity of the biomarkers and the samples are as defined in the first aspect.
[0149] Preferably, the treatment method comprises analyzing and comparing the expression levels, amounts and / or activity of 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18 or more of the following biomarkers: TNF-□; GSTA1; NT-proBNP; RORA; TNC; GHR; A2M; IGFBP2; APOB; SEPP1; TFF3; IL6; CHI3L1; MET; GDF15; CCL22; TNFRSF11; ANGPT2 and ReIA NF-KB.
[0150] The treatment method may comprise analyzing and comparing the expression levels, amounts and / or activities of at least six biomarkers or at least seven biomarkers. Alternatively, the treatment method may comprise analyzing and comparing the expression levels, amounts and / or activities of at least eight biomarkers or at least nine biomarkers. In another embodiment, the treatment method may comprise analyzing and comparing the expression levels, amounts and / or activities of at least ten biomarkers, at least eleven biomarkers, at least twelve biomarkers, at least thirteen biomarkers, at least fourteen biomarkers or at least fifteen biomarkers. In another embodiment, the treatment method may comprise analyzing and comparing the expression levels, amounts and / or activities of at least sixteen biomarkers, at least seventeen biomarkers or at least eighteen biomarkers.
[0151] Preferably, the treatment method includes analysing and comparing the expression levels, amounts and / or activity of the following biomarkers: TNF-□; GSTA1; NT-proBNP; RORA; TNC; GHR; A2M; IGFBP2; APOB; SEPP1; TFF3; IL6; CHI3L1; MET; GDF15; CCL22; TNFRSF11; ANGPT2 and ReIA NF-KB.
[0152] The clinician will be able to make decisions regarding the preferred course of treatment required, for example the type and dosage of therapeutic agents according to the eighth and ninth aspects to be administered.
[0153] Suitable therapeutic agents may include: statins, including statins selected from the group consisting of atorvastatin; simvastatin; rosuvastatin; and pravastatin; beta-blockers; low-dose aspirin; antithrombotic agents, including antithrombotic agents selected from the group consisting of clopidogrel; rivaroxaban; ticagrelor and prasugrel; nitrates; angiotensin-converting enzyme (ACE) inhibitors; angiotensin II receptor antagonists; amlodipine; calcium channel blockers, including calcium channel blockers selected from the group consisting of verapamil and diltiazem; and / or diuretics.
[0154] Treatment may involve prescribing lifestyle changes.
[0155] In a ninth aspect, there is also provided a method of treating an individual having a higher risk of suffering from cardiovascular disease, the method comprising: (a) detecting single nucleotide polymorphisms (SNPs) in the RORA gene and / or the GHR gene in a sample obtained from a subject, wherein the presence of the SNP indicates that the individual has a higher risk of suffering from cardiovascular disease; and (b) administering or having administered to an individual a therapeutic agent that prevents or reduces the likelihood that the individual will suffer from cardiovascular disease; Includes.
[0156] Preferably, the biomarkers, detection of the biomarkers, cardiovascular disease, expression levels, amount and / or activity of the biomarkers and the samples are as defined in the first aspect.
[0157] Preferably, the single nucleotide polymorphisms (SNPs) RORA and / or GHR are as defined in the third and / or fourth aspect.
[0158] The clinician will be able to make decisions regarding the preferred course of treatment required, for example the type and dosage of therapeutic agents according to the eighth and ninth aspects to be administered.
[0159] Suitable therapeutic agents may include: statins, including statins selected from the group consisting of atorvastatin; simvastatin; rosuvastatin; and pravastatin; beta-blockers; low-dose aspirin; antithrombotic agents, including antithrombotic agents selected from the group consisting of clopidogrel; rivaroxaban; ticagrelor and prasugrel; nitrates; angiotensin-converting enzyme (ACE) inhibitors; angiotensin II receptor antagonists; amlodipine; calcium channel blockers, including calcium channel blockers selected from the group consisting of verapamil and diltiazem; and / or diuretics.
[0160] Treatment may involve prescribing lifestyle changes.
[0161] It will be appreciated that the present invention extends to any nucleic acid or peptide, or variant, derivative or analogue thereof, that substantially comprises the amino acid or nucleic acid sequence of any of the sequences referred to herein, including variants or fragments thereof. The terms "substantially an amino acid / nucleotide / peptide sequence", "variant" and "fragment" may refer to a sequence that has at least 40% sequence identity with the amino acid / nucleotide / peptide sequence of any one of the sequences referred to herein, such as a sequence that has 40% identity with the sequences identified as SEQ ID NOs: 1-40.
[0162] Also envisaged are amino acid / polynucleotide / polypeptide sequences having greater than 65% sequence identity, more preferably greater than 70%, even more preferably greater than 75%, and even more preferably greater than 80% sequence identity to any of the sequences referenced herein. Preferably, the amino acid / polynucleotide / polypeptide sequence has at least 85% identity, more preferably at least 90% identity, even more preferably at least 92% identity, even more preferably at least 95% identity, even more preferably at least 97% identity, even more preferably at least 98% identity, and most preferably at least 99% identity to any of the sequences referenced herein.
[0163] Those skilled in the art understand how to calculate the identity percentage between two amino acid / polynucleotide / polypeptide sequences. To calculate the identity percentage between two amino acid / polynucleotide / polypeptide sequences, the alignment of the two sequences must first be prepared, and then the sequence identity value must be calculated. The identity percentage for two sequences can take different values depending on:- (i) the method used to align sequences, such as ClustalW, BLAST, FASTA, Smith-Waterman (implemented in different programs), or structural alignment from 3D comparison; and (ii) the parameters used by the alignment method, such as local vs. global alignment, the pair-score matrix used (e.g., BLOSUM62, PAM250, Gonnet, etc.), and gap-penalty, such as function form and constants.
[0164] Once an alignment is made, there are many different ways to calculate the identity percentage between two sequences. For example, one method can divide the number of identities by: (i) the length of the shortest sequence; (ii) the length of the alignment; (iii) the average length of the sequences; (iv) the number of non-gap positions; or (v) the number of equivalenced positions excluding overhangs. Moreover, it is understood that the identity percentage is also strongly length-dependent. Thus, the shorter the sequence pair, the higher the sequence identity that can be expected to occur by chance.
[0165] Therefore, it is not surprising that accurate alignment of protein or DNA sequences is a complex process. The well-known multiple alignment program ClustalW (Thompson et al., 1994, Nucleic Acids Research, 22, 4673-4680; Thompson et al., 1997, Nucleic Acids Research, 24, 4876-4882) is the preferred method for generating protein or DNA multiple alignments according to the present invention. Suitable parameters for ClustalW may be as follows: for DNA alignments: Gap Open Penalty=15.0, Gap Extension Penalty=6.66, and Matrix=Identity. For protein alignments: Gap Open Penalty=10.0, Gap Extension Penalty=0.2, and Matrix=Gonnet. For DNA and protein alignments: ENDGAP=-1, and GAPDIST=4. It will be apparent to one skilled in the art that it may be necessary to vary these and other parameters for optimal sequence alignment.
[0166] Preferably, the calculation of the identity percentage between two amino acid / polynucleotide / polypeptide sequences can then be calculated from the above alignment as (N / T)*100, where N is the number of positions where the sequences share identical residues, and T is the total number of positions compared, including gaps, and including or excluding overhangs. Preferably, the overhangs are included in the calculation. Therefore, the most preferred method for calculating the identity percentage between two sequences includes (i) preparing a sequence alignment using the ClustalW program, for example, using an appropriate set of parameters as described above; and (ii) inserting the values of N and T into the following formula:- sequence identity=(N / T)*100.
[0167] Alternative methods for identifying similar sequences will be known to those skilled in the art. For example, a substantially similar nucleotide sequence is encoded by a sequence that hybridizes to a DNA sequence or its complement under stringent conditions. By stringent conditions, we mean hybridizing the nucleotides to filter-bound DNA or RNA in 3x sodium chloride / sodium citrate (SSC) at about 45°C, followed by at least one wash in 0.2xSSC / 0.1% SDS at about 20-65°C. Alternatively, a substantially similar polypeptide may differ from a sequence shown, for example, in SEQ ID NOs: 1-40, by at least one amino acid, but by fewer than 5, 10, 20, 50, or 100 amino acids.
[0168] Due to the degeneracy of the genetic code, it is clear that any nucleic acid sequence described herein can be varied or altered to produce a functional variant thereof without substantially affecting the sequence of the protein encoded thereby. Suitable nucleotide variants are those having a sequence altered by substitution of different codons that code for the same amino acid within the sequence, thus resulting in a silent (synonymous) change. Other suitable variants include all or a portion of a sequence having a homologous nucleotide sequence, but altered by substitution of different codons that code for amino acids with side chains with similar biophysical properties to the amino acid it replaces, resulting in a conservative change. For example, small non-polar hydrophobic amino acids include glycine, alanine, leucine, isoleucine, valine, proline, and methionine. Large non-polar hydrophobic amino acids include phenylalanine, tryptophan, and tyrosine. Polar neutral amino acids include serine, threonine, cysteine, asparagine, and glutamine. Positively charged (basic) amino acids include lysine, arginine, and histidine. The negatively charged (acidic) amino acids include aspartic acid and glutamic acid. It will therefore be understood which amino acids can be replaced with amino acids having similar biophysical properties, and the skilled artisan will know the nucleotide sequences encoding these amino acids.
[0169] All of the features described in this specification (including any of the accompanying claims, abstract and drawings), and / or all of the steps of any method or process so disclosed may be combined in any combination with any of the above aspects, except combinations in which at least some of such features and / or steps are mutually exclusive.
[0170] For a better understanding of the present invention, and to show how embodiments of the same may be carried into effect, reference will now be made, by way of example, to the accompanying drawings in which: [Brief description of the drawings]
[0171] [Figure 1] Figure 1 shows that after disease network identification (top left panel), single-level patient data are integrated based on a Bayesian probabilistic graphical model with patient-specific pathway activities inferred for each of the BMKs enrolled in the model. The calculated activities per patient and molecule were clustered by hierarchical clustering. Survival analysis was performed in an independent step to define the relevance of the identified clusters for disease progression (patient stratification). BMK = biomarker; TF = transcription factor. [Figure 2A] Figure 2a-c shows the network of molecular entities prioritized in the overconnectivity analysis. To infer patient-specific activities (workflow Fig. 1), we used patient data from: i) 13 proteins (nodes highlighted in blue), and ii) genotypes of two genes (nodes marked with stars). The activities of four nodes (red circles) were inferred by Bayesian statistics based on known molecular interactions from the literature (Supplementary Table S5) (connection tangents, red = inhibition, or green = activation; in the case of bidirectionality, right tangent = from the central node to the neighboring molecule, and left tangent = from the neighboring molecule to the central node). [Figure 2B] Figure 2a-c shows the network of molecular entities prioritized in the overconnectivity analysis. To infer patient-specific activities (workflow Fig. 1), we used patient data from: i) 13 proteins (nodes highlighted in blue), and ii) genotypes of two genes (nodes marked with stars). The activities of four nodes (red circles) were inferred by Bayesian statistics based on known molecular interactions from the literature (Supplementary Table S5) (connection tangents, red = inhibition, or green = activation; in the case of bidirectionality, right tangent = from the central node to the neighboring molecule, and left tangent = from the neighboring molecule to the central node). [Figure 2C]Figure 2a-c shows the network of molecular entities prioritized in the overconnectivity analysis. To infer patient-specific activities (workflow Fig. 1), we used patient data from: i) 13 proteins (nodes highlighted in blue), and ii) genotypes of two genes (nodes marked with stars). The activities of four nodes (red circles) were inferred by Bayesian statistics based on known molecular interactions from the literature (Supplementary Table S5) (connection tangents, red = inhibition, or green = activation; in the case of bidirectionality, right tangent = from the central node to the neighboring molecule, and left tangent = from the neighboring molecule to the central node). [Figure 3A] Figure 3 shows that molecular signatures of population subtypes are associated with CVD progression. a, Two major patient clusters (highlighted in red and blue boxes) were identified based on molecular signatures in Caucasians (left) and Hispanics (right). bc, Kaplan-Meier analysis comparing survival probability for identified clusters (a, red and blue) associated with the first and second major extended CV composite in Caucasians (b) and Hispanics (c). d, e, Proportion of CVO events or all deaths per cluster (blue = higher and blue = lower CVO risk clusters) with the calculated relative risk of events for patients in the blue highlighted cluster versus the red highlighted cluster (a), respectively. There was a significant difference in survival between clusters (log-rank P i 0.0000 in bc) for each of the outcomes adjusted for CVO risk factors (Cox regression model P i 0.0000 in bc). A, C, B represent the major clusters of pathway entities (Table 1). PA=pathway activity; MI=myocardial infarction; Hospital HF=hospitalization for heart failure. [Figure 3B]Figure 3 shows that molecular signatures of population subtypes are associated with CVD progression. a, Two major patient clusters (highlighted in red and blue boxes) were identified based on molecular signatures in Caucasians (left) and Hispanics (right). bc, Kaplan-Meier analysis comparing survival probability for identified clusters (a, red and blue) associated with the first and second major extended CV composite in Caucasians (b) and Hispanics (c). d, e, Proportion of CVO events or all deaths per cluster (blue = higher and blue = lower CVO risk clusters) with the calculated relative risk of events for patients in the blue highlighted cluster versus the red highlighted cluster (a), respectively. There was a significant difference in survival between clusters (log-rank P i 0.0000 in bc) for each of the outcomes adjusted for CVO risk factors (Cox regression model P i 0.0000 in bc). A, C, B represent the major clusters of pathway entities (Table 1). PA=pathway activity; MI=myocardial infarction; Hospital HF=hospitalization for heart failure. [Figure 3C]Figure 3 shows that molecular signatures of population subtypes are associated with CVD progression. a, Two major patient clusters (highlighted in red and blue boxes) were identified based on molecular signatures in Caucasians (left) and Hispanics (right). bc, Kaplan-Meier analysis comparing survival probability for identified clusters (a, red and blue) associated with the first and second major extended CV composite in Caucasians (b) and Hispanics (c). d, e, Proportion of CVO events or all deaths per cluster (blue = higher and blue = lower CVO risk clusters) with the calculated relative risk of events for patients in the blue highlighted cluster versus the red highlighted cluster (a), respectively. There was a significant difference in survival between clusters (log-rank P i 0.0000 in bc) for each of the outcomes adjusted for CVO risk factors (Cox regression model P i 0.0000 in bc). A, C, B represent the major clusters of pathway entities (Table 1). PA=pathway activity; MI=myocardial infarction; Hospital HF=hospitalization for heart failure. [Figure 3D]Figure 3 shows that molecular signatures of population subtypes are associated with CVD progression. a, Two major patient clusters (highlighted in red and blue boxes) were identified based on molecular signatures in Caucasians (left) and Hispanics (right). bc, Kaplan-Meier analysis comparing survival probability for identified clusters (a, red and blue) associated with the first and second major extended CV composite in Caucasians (b) and Hispanics (c). d, e, Proportion of CVO events or all deaths per cluster (blue = higher and blue = lower CVO risk clusters) with the calculated relative risk of events for patients in the blue highlighted cluster versus the red highlighted cluster (a), respectively. There was a significant difference in survival between clusters (log-rank P i 0.0000 in bc) for each of the outcomes adjusted for CVO risk factors (Cox regression model P i 0.0000 in bc). A, C, B represent the major clusters of pathway entities (Table 1). PA=pathway activity; MI=myocardial infarction; Hospital HF=hospitalization for heart failure. [Figure 3E]Figure 3 shows that molecular signatures of population subtypes are associated with CVD progression. a, Two major patient clusters (highlighted in red and blue boxes) were identified based on molecular signatures in Caucasians (left) and Hispanics (right). bc, Kaplan-Meier analysis comparing survival probability for identified clusters (a, red and blue) associated with the first and second major extended CV composite in Caucasians (b) and Hispanics (c). d, e, Proportion of CVO events or all deaths per cluster (blue = higher and blue = lower CVO risk clusters) with the calculated relative risk of events for patients in the blue highlighted cluster versus the red highlighted cluster (a), respectively. There was a significant difference in survival between clusters (log-rank P i 0.0000 in bc) for each of the outcomes adjusted for CVO risk factors (Cox regression model P i 0.0000 in bc). A, C, B represent the major clusters of pathway entities (Table 1). PA=pathway activity; MI=myocardial infarction; Hospital HF=hospitalization for heart failure. [Figure 4A] Figure 4 shows the Cox regression results expressed as hazard ratios (HRs) with confidence intervals. ab, HR for the second multi-primary composite of CVO in Caucasians (a) and Hispanics (b) for known CVO risk factors and cluster identities (defined in Figure 3, panels ab). c, HR for the first multi-primary composite of CVO for the same factors above considering the entire population (Caucasians and Hispanics). Cox regression models were significant P < 0.0000). [Figure 4B]Figure 4 shows the Cox regression results expressed as hazard ratios (HRs) with confidence intervals. ab, HR for the second multi-primary composite of CVO in Caucasians (a) and Hispanics (b) for known CVO risk factors and cluster identities (defined in Figure 3, panels ab). c, HR for the first multi-primary composite of CVO for the same factors above considering the entire population (Caucasians and Hispanics). Cox regression models were significant P < 0.0000). [Figure 4C] Figure 4 shows the Cox regression results expressed as hazard ratios (HRs) with confidence intervals. ab, HR for the second multi-primary composite of CVO in Caucasians (a) and Hispanics (b) for known CVO risk factors and cluster identities (defined in Figure 3, panels ab). c, HR for the first multi-primary composite of CVO for the same factors above considering the entire population (Caucasians and Hispanics). Cox regression models were significant P < 0.0000). [Diagram 5]Figure 5 shows the list of gene names from the identified disease networks. Three major clusters of genes (A, B, C in Figure 3) were identified (highlighted in red = lower activity, and shaded in blue = higher activity of molecules in patients at lower versus higher risk for CVO events). Individual patient data (G = genotype and P = protein level) for these genes were used for the patient stratification workflow (Figure 1). AI = molecular activity was inferred by Bayesian statistics. RORA and GHR variants have not been previously reported to be associated with CVD, and the effect allele was rare in European populations, but common in African populations (MAF = 0.41), and GHR variants in Asian populations (MAF = 0.13). Results from standard BMK analysis for CVO in the ORIGIN cohort are reported for (a) the first and (b) the second multimajor cardiovascular composites, as previously published (citing publications). (First coprimary endpoint = composite of CV death or nonfatal MI or nonfatal stroke, and second coprimary endpoint = composite of first coprimary or revascularization procedure or hospitalization for heart failure). ns = biomarkers that were not significantly associated with CVO but were associated with death in the ORIGIN CVO trial. [Figure 6A] Figures 6a and 6b (s2) show locus zoom plots of GWAs results from the ORIGIN study. SNPs are plotted according to their location on each chromosome on the x-axis against association with CAD on the y-axis (shown as -log10 P). Loci used for patient stratification (rs73420079, rs4314405; P<5 x 10-8) are shown in locus zoom plots b and c, respectively. With the exception of MIR3169, which had a MAF=18%, the MAF for these loci was approximately 1% in Europeans. [Figure 6B]Figures 6a and 6b (s2) show locus zoom plots of GWAs results from the ORIGIN study. SNPs are plotted according to their location on each chromosome on the x-axis against association with CAD on the y-axis (shown as -log10 P). Loci used for patient stratification (rs73420079, rs4314405; P<5 x 10-8) are shown in locus zoom plots b and c, respectively. With the exception of MIR3169, which had a MAF=18%, the MAF for these loci was approximately 1% in Europeans. [Figure 7] Figure 7 shows the network analysis design. Statistical summaries from the ORIGIN CAD GWAs were used to identify the most relevant genetic associations with CVO in this cohort. In a post-hoc analysis, these genes and a list of biomarker (BMK) coding genes were used to create a network of overconnected genes. This procedure was repeated using GWAs statistical summaries for CAD outcomes from the CARDIOGRAM consortium (CARDIoGRAMplusC4D, Nikpey et al. 2015) and the same list of BMKs associated with CAD outcomes. [Figure 8A] Figure 8 shows networks of molecular entities prioritized in hyperconnectivity analysis. a, b, c, The first three top networks in discovery network analysis using GWAs from ORIGIN dataset and candidates from BMK analysis. c, d, e Recursive network analysis showing the top sub-networks (ranked 1st, 4th and 8th) identified when using CARDIOGRAM GWAs results from ORIGIN and BMK analysis results. These were most similar to the network identified in a-c (blue recursive BMK). Asterisks = genes associated with CVOs in ORIGIN GWAs. [Figure 8B]Figure 8 shows networks of molecular entities prioritized in hyperconnectivity analysis. a, b, c, The first three top networks in discovery network analysis using GWAs from ORIGIN dataset and candidates from BMK analysis. c, d, e Recursive network analysis showing the top sub-networks (ranked 1st, 4th and 8th) identified when using CARDIOGRAM GWAs results from ORIGIN and BMK analysis results. These were most similar to the network identified in a-c (blue recursive BMK). Asterisks = genes associated with CVOs in ORIGIN GWAs. [Figure 8C] Figure 8 shows networks of molecular entities prioritized in hyperconnectivity analysis. a, b, c, The first three top networks in discovery network analysis using GWAs from ORIGIN dataset and candidates from BMK analysis. c, d, e Recursive network analysis showing the top sub-networks (ranked 1st, 4th and 8th) identified when using CARDIOGRAM GWAs results from ORIGIN and BMK analysis results. These were most similar to the network identified in a-c (blue recursive BMK). Asterisks = genes associated with CVOs in ORIGIN GWAs. [Figure 8D] Figure 8 shows networks of molecular entities prioritized in hyperconnectivity analysis. a, b, c, The first three top networks in discovery network analysis using GWAs from ORIGIN dataset and candidates from BMK analysis. c, d, e Recursive network analysis showing the top sub-networks (ranked 1st, 4th and 8th) identified when using CARDIOGRAM GWAs results from ORIGIN and BMK analysis results. These were most similar to the network identified in a-c (blue recursive BMK). Asterisks = genes associated with CVOs in ORIGIN GWAs. [Figure 8E]Figure 8 shows networks of molecular entities prioritized in hyperconnectivity analysis. a, b, c, The first three top networks in discovery network analysis using GWAs from ORIGIN dataset and candidates from BMK analysis. c, d, e Recursive network analysis showing the top sub-networks (ranked 1st, 4th and 8th) identified when using CARDIOGRAM GWAs results from ORIGIN and BMK analysis results. These were most similar to the network identified in a-c (blue recursive BMK). Asterisks = genes associated with CVOs in ORIGIN GWAs. [Figure 8F] Figure 8 shows networks of molecular entities prioritized in hyperconnectivity analysis. a, b, c, The first three top networks in discovery network analysis using GWAs from ORIGIN dataset and candidates from BMK analysis. c, d, e Recursive network analysis showing the top sub-networks (ranked 1st, 4th and 8th) identified when using CARDIOGRAM GWAs results from ORIGIN and BMK analysis results. These were most similar to the network identified in a-c (blue recursive BMK). Asterisks = genes associated with CVOs in ORIGIN GWAs. [Figure 9A] Figure 9 shows Kaplan-Meier survival estimates by patient cluster for CV outcomes measured in Caucasians and Hispanics. (a) Identified clusters and, bc, survival curves (in the top panel, red = clusters highlighted in red, blue = clusters highlighted in blue). (b) Caucasians and Hispanics (c). There were significant differences in survival between clusters in all cases (log-rank P<0.000 was less significant). [Figure 9B]Figure 9 shows Kaplan-Meier survival estimates by patient cluster for CV outcomes measured in Caucasians and Hispanics. (a) Identified clusters and, bc, survival curves (in the top panel, red = clusters highlighted in red, blue = clusters highlighted in blue). (b) Caucasians and Hispanics (c). There were significant differences in survival between clusters in all cases (log-rank P<0.000 was less significant). [Figure 9C] Figure 9 shows Kaplan-Meier survival estimates by patient cluster for CV outcomes measured in Caucasians and Hispanics. (a) Identified clusters and, bc, survival curves (in the top panel, red = clusters highlighted in red, blue = clusters highlighted in blue). (b) Caucasians and Hispanics (c). There were significant differences in survival between clusters in all cases (log-rank P<0.000 was less significant). [Figure 10] Figure 10 shows box plots of biomarker levels including clusters of genes (A, B, C) identified when clustering patient-specific BMK activity (Figure 2). Patients in clusters with higher risk for CVO (left box plots in each figure) had higher BMK levels, except for GSTalpha, which was lower in the higher risk group. [Figure 11]FIG. 11 shows visual inspection of cardiovascular risk factors in clusters of high (1) CVO risk versus clusters of low (2) CVO risk (Caucasian and Latino subpopulations). Top panel, box plots for age, BMI and BMK levels measured routinely in clinic. Patients in clusters with higher risk for CVO (left box plots in each figure) were, on average, similar to patients in clusters with lower CVO risk. Bottom panel, categorical CVO risk variables such as sex (lower left panel), smoking (lower center panel), and albuminuria (lower right panel) or reported albuminuria differed slightly between clusters. Further details on these and other risk factors are shown in Table 1. Statistical analyses of CVO risk (logistic regression) and survival (cox regression) included these as covariates to calculate risk ratios and hazard ratios per cluster as shown by MS body. NormalTC=normalized total cholesterol; normalSBP / DBP=normalized systolic blood pressure / diastolic blood pressure. [Figure 12A] Figure 12 shows the subcluster analysis: (a) Subclusters within clusters in both Caucasians and Hispanics (cluster 2 is now split into 2 and 3); (b) Proportion of CVO events per cluster 1-3 and the corresponding number of individuals, N. [Figure 12B] Figure 12 shows the subcluster analysis: (a) Subclusters within clusters in both Caucasians and Hispanics (cluster 2 is now split into 2 and 3); (b) Proportion of CVO events per cluster 1-3 and the corresponding number of individuals, N. [Figure 13] Figure 13 shows Kaplan-Meier survival estimates for outcomes measured in Caucasians (left panel) and Hispanics (right panel) for the three identified subclusters (as shown in Figure 13). In all cases there were significant differences in survival between clusters (log-rank P<0.0000). [Figure 14]Figure 14 (S10) shows NTproBNP biomarker levels in the three clusters of patients obtained using (left panel) or not (right panel) NT-proBNT levels as input for the network calculation that generated these clusters in Latinos. There were no significant differences in NT-proBNT levels when using the results of the initial or subsequent analyses as classifiers for the patients (clusters 1-3). EXAMPLES
[0172] The inventors set out to identify biomarkers and biomarker combinations associated with CVD progression with the goal of developing strategies that would allow the identification of individuals at risk of suffering from a CVD event and therefore allow early intervention to prevent or reduce the risk of individuals suffering from a CVD event. Identification of suitable biomarkers and biomarker networks will also optimize clinical trial design, drug efficacy, and optimize treatment.
[0173] Example 1 – Identification of biomarkers and biomarker networks Materials and Methods ORIGIN cohort Participants in the ORIGIN biomarker sub-study (N=8,401, Supplementary Table S1) were randomly selected from patients treated with Lantus or placebo. A smaller subset of this sample was genotyped (5,078 samples). After quality control (see below), the sample size was 4,390 individuals. Only these genotyped subjects were included in the analyses described in this protocol for patient stratification (demographic characteristics in Supplementary Table S1).
[0174] Biomarker quality control and normalization Quality controls for protein biomarkers measured in the ORIGIN cohort have been described previously (Gerstein et al., 2015). Normalization = non-normally distributed biomarkers were log-transformed. Standard biomarker analysis for CVO prediction in this cohort has been described elsewhere (Gerstein et al., 2015).
[0175] Quantitative BMK measures were also transformed into categorical variables (-1, 0, +1) based on the percentage of distribution for each BMK, separately for subjects of Caucasian, Latino and African origin. The algorithm PARADIGM cannot handle continuous traits for clustering.
[0176] Genotyping quality control and population stratification analysis Genotyping of the ORIGIN cohort, N=5,078 samples, was performed using the Illumina HumanCore Exome DNA Analysis BeadChip (Illumina Omni2.5). Over 540,000 genetic variants were called, including a wide range of coding variants, both common and rare. Single nucleotide polymorphisms (SNPs) were excluded with call rates <0.99, minor allele frequencies <0.01, or deviations from Hardy-Weinberg equilibrium (P<1x10-6). Individuals were excluded if their self-reported sex, ethnicity, and relatedness were not consistent with their genetic information. After quality control, the sample size was 4,390 individuals and 284,024 SNPs (Note: SNPs excluded due to low allele frequencies were genotyped correctly for the most part and will be included in future analyses). Genome-wide genotype imputation was performed using Impute v2.3.0. We used NCBI build 37 of Phase I Integrated Variants 1000 Genomic Haplotypes (SHAPEIT2) as the reference panel. Imputed SNPs were excluded with an imputation certainty score <0.3. The final number of imputed SNPs was 10,501,330.
[0177] Principal components were generated based on whole-genome genotyping separately for Caucasians and Hispanics, and these were used as covariates in genetic analyses for genome-wide association studies (GWAs) from the ORIGIN study.
[0178] Due to population ethnicity substructure, subpopulations were defined by ethnic (Caucasian and Latino) groups and analyzed separately. HWE (P>0.001) for the SNPs participating in the analysis was tested as part of quality control to define whether subgroups of the population met the expected locus distribution. Subpopulations were meta-analyzed and association results were visualized with Manhattan and QQ plots (Figure 6).
[0179] Genome-wide association study GWA analyses were performed separately for Caucasians (N = 1,931) and Hispanics (N = 2,216). Directly genotyped SNPs and imputed SNPs (N = 4.9-9 Genotypes consisting of both phenotypes (P < 0.01, P < 0.01, Mio) were entered into the GWA analysis. To avoid overinflation of test statistics due to population structure or association, we applied genomic control methods for the GWA analysis. Principal components were generated based on whole genome genotyping separately for Caucasians and Hispanics. These were used as covariates in genome-wide association studies from the ORIGIN study. Linear regression for association with normalization (PLINK) was performed under an additive model with SNP allele dosage and age and sex as predictors.
[0180] A meta-analysis was performed. Corresponding to the Bonferroni adjustment for 1 million independent tests, we specified a P threshold for genome-wide significance. The CARDIOGRAM Consortium dataset used was the CARDioGRAMplusC4D 1000 Genome-Based GWAS Meta-Analysis Statistical Summary. It mainly includes GWAS studies from Europe, South Asia, and East Asia, and the lineage was complemented using the 1000 Genomes phase 1 v3 training set, which contains 38 million variants. This study examined 9.4 million variants and included 60,801 CVD cases and 123,504 controls (Nikpey et al., 2015). To assess the number of independent loci associated with CVD, correlated SNPs were grouped using a LD-based outcome clumping procedure (PLINK, Purcell et al., 2007). This procedure was used for genetic mapping of loci to be included in the overconnectivity network analysis (Supplementary Table S2). Variants associated with CVD with p-values less than 10-6, substituted with p-values equal to or less than 10-5 in the ORIGIN cohort, and p-values less than 10-7, substituted with p-values equal to or less than 10-6 in the CARDIOGRAM cohort, were mapped to genes that could be considered in the network analysis. We excluded alleles with less than 1 percent and poor imputation quality (Info less than 0.4) from the clumping procedure.
[0181] Workflow The complete workflow steps (Figure 1) are as follows: (1) Candidate molecule selection The list of entities included in the discovery and replication studies is shown in Table S3. (2) Gene mapping SNPs associated with CAD in ORIGIN and CARDIOGRAM were mapped to genes based on empirical assessment of linkage disequilibrium (LD) between single nucleotide polymorphisms (SNPs) using a clustering procedure available in PLINK (Purcell et al., 2007). We estimated LD between variants using 1000 Genomes (phase 1 release v3) as a reference dataset; clumping analysis was performed using LD r2>0.8 (--clump-r2 0.8) and clumping variants to the index SNP within a range of 250 kb (--clump-kb 250). To identify clumped genomic regions corresponding to genes. We used the --clump-range function with a gene list (hg19) (Supplementary Table S2); (3) Network analysis to identify disease networks We used hyperconnectivity analysis to construct networks representative of cardiovascular disease, as described in more detail below. (4) Single-level patient data curation To begin the patient stratification workflow (Figure 1), we first generated a disease network and used patient-specific data to prioritize molecules (Table 1). Patient-specific data quality control is described above. (5) Probabilistic graphical model analysis for identification of patient-specific pathway activity The methods are described in more detail in the manuscript and below. (6) Clustering We used the R package as described below. (7) Linking identified patient clusters to CVOs Further details about the analyses, outcomes and covariates are provided below. (8) Single biomarker comparison These were performed by visual inspection of box plots and median comparisons. (9) Characterization of clustered populations.
[0182] Disease Network Identification Disease networks were identified using an overconnectivity algorithm as implemented in the R-based Computational Biology for Drug discovery (CDDD) package developed by Clarivate Analytics. The specificity of the network to a disease depends on the molecules associated with the disease (Figure 7) selected to generate the network and the underlying library of protein-protein interactions in humans. To this end, we extracted a manually curated systems biology knowledge base of high-confidence interactions from Ingenuity (IPA from QIAGEN Inc.) and Metabase (Metacore from Clarivate Analytics) and used it as a library for the CBDD package. To create the CVD disease network, we: i) exploited morphological features of the human protein interaction network to identify direct regulators in the immediate vicinity (one-step away) of the dataset that were statistically overconnected with objects from the dataset (hypergeometric distribution), and ii) used names of genes with evidence of association to CVD as the input dataset (see Bayesian network analysis below and in Figure 7). As input data for the hyperconnectivity analysis, we used the names of genes corresponding to the 16 proteins BMK in the discovery analysis and the names of genes from the 8 loci associated with CVO in the Origin cohort GWAs (Table S3). To confirm the initially identified networks, we reran these analyses by replacing the genetically associated loci from the ORIGIN study with 90 other genes from an independent large-scale GWAs meta-analysis from the CARDIOGRAM consortium (CARDIoGRAMplusC4D, cited Nikpey et al. 20115) (Figure 7, Table S3).
[0183] Bayesian Network Analysis for Data Integration PARADIGM is a data integration approach based on probabilistic graphical models. It draws pathways or networks as probabilistic graphical models (PGMs) and learns its parameters from a fed omics dataset. The model allows the estimation of true activity scores for each node in a pathway, taking into account different omics measurements for the node. PARADIGM allows predictions at the level of individual patients and can accommodate data types such as gene / protein expression, copy number changes, metabolomics, direct protein activity assays such as kinase activity measurements. PARADIGM combines a large number of genome-scale measurements at the sample level to infer the activities of genes, products and abstract processes in a pathway or sub-network. Edges in the original network connect hidden variables of different nodes (e.g., the hidden variable of activity of node A affects the hidden variable of protein or DNA of node B depending on the mechanism of AB links).
[0184] Each node is assigned a conditional probability distribution when the model is created. The distribution indicates how likely it is to observe the node in a particular state given the states of its parents in the model. Three states are allowed for each node (activated, inhibited, unchanged). The distributions for the hidden variables are defined in the first stage. The distributions for each molecular-level observed variable are learned by the EM algorithm using the input data. After the model is completed, it is possible to infer the probability of observing a hidden node in a particular state without the observed data (prior probability) or taking the data into account (posterior probability).
[0185] The primary output is an integrated pathway activity (IPA) matrix A, where A_ij represents the inferred activity of entity i in patient sample j. The value of A is signed and non-zero if the patient data makes activation or inhibition of hidden nodes more likely compared to the prior. A is expected to be used in place of the original dataset for purposes such as patient stratification or association analysis to reveal biological entities whose activities are associated with clinical traits.
[0186] The output is a matrix of activation scores for each node in the network and each sample. The activation scores represent the signed log-likelihood ratio (positive if the node is predicted to be active, negative if the node is predicted to be inhibited). The pathway is converted into a probabilistic graphical model that includes both hidden states for each node and observed states for the nodes that may correspond to the input dataset. There are two possible modes for assessing IPA significance. Both involve permutation-computation of IPA scores over many randomized samples.
[0187] For "within" permutations, permuted data samples are created by creating a new set of evidence (i.e., states for observed variables in gene expression and gene copy number) by assigning random nodes in the pathway / sub-network and values of the random samples to each observed node.
[0188] For "arbitrary" permutations, the procedure was the same, but the random node selection step was allowed to select a node from anywhere in the input data (regardless of whether a particular pathway / sub-network contained such a node or not).
[0189] For both permutation types, we generated iteratively permuted samples and calculated an IPA score for each permuted sample. We used the distribution of scores from the permuted data as a null distribution to estimate the significance of the IPA scores in the real dataset. We used the CBDD package developed by CLARIVATES in SCRIPT:Paradigm in R (Vaske et al., 2010).
[0190] Clusters were generated for discovery (Caucasian; N=1908) and replication (Hispanic; N=2146 subpopulations).
[0191] 1. Network: Interactome (Sanofi Network from IPA & Metabase) Human High quality.
[0192] 2. Matrix: Paradigm was run using patient-specific information from 13 protein biomarkers and genotypes from two GWAs genes. These molecular entities appeared in the first three sub-networks obtained from the network analysis using the overconnectivity algorithm. Proteins were coded as -1, 0, 1 for each patient (according to their distribution). The matrix containing the three sub-networks obtained using the OVERCONNECTIVITY algorithm was also used as input for the paradigm (although this included molecular entities for which we did not provide BMK measurements as input).
[0193] 3. Level: DNA and protein List of proteins / genes included in the analysis: Phenotype matrix (proteins) 1) Alpha 2 Macroglobulin (First Multiple Major and Fatal) 2) Angiopoietin 2 (1st, 2nd CVO and death) 3) Apolipoprotein B (1st, 2nd CVO and death) 4) YLK-40 (CHI3L1) (dead) 5) Glutathione S-transferase alpha (1st, 2nd CVO and death) 6) Tenascin-c (deceased) 7) IGF-binding protein 2 (death) 8) Hepatocyte growth factor receptor (MET) (1st, 2nd CVO and death) 9) Osteoprotegerin (TNFRSF11) (1st, 2nd CVO and death) 10) Macrophage-derived chemokine (CCL22) (death) 11) Selenoprotein P (death) 12) Trefoil factor 3 (first CVO and death) 13) GDF15 (1st, 2nd CVO and death) Genotype from gene: 1) RORA (rs73420079 CAD influence factor G -AA=-1, AG=0, GG=1), 2) GHR (rs4314405 CAD Influence Factor A - GG=-1, AG=0, AA=1).
[0194] Input network key members and predicted biomarkers: TNF-alpha, IL6, RelA NF-KB subunit, NT-proBNP (NPPB).
[0195] Patient stratification We started patient stratification analysis based on the disease sub-networks identified in the ORIGIN cohort (Figure 2) and using protein or genotype data specific to each patient (Workflow Figure 1; Figure 2 and Table 1). Due to population ethnicity background. We considered two main subpopulations for which protein biomarkers and genotypes were available (Caucasian and Latino), with similar sample sizes (respectively; N=1,908 and N=2,146) and sample characteristics for CVO risk factors (Supplementary Table S1). These are then considered as discovery and validation samples for i) Bayesian network analysis, ii) hierarchical clustering, and iii) CVO risk and survival analysis. Bayesian network analysis was performed using the PARADIGM algorithm (Vaske et al.) implemented in R as part of the CDDD package developed by Clarivate Analytics. The data integration approach is based on probabilistic graphical models. It plots pathways or networks as probabilistic graphical models (PGMs) and learns their parameters from the fed omics dataset (Figure 2). The model allows the estimation of the true activity score for each node (molecule) in the pathway, taking into account the different omics measurements (per patient) for the node (Figure 2; Table 1). Each node is assigned a conditional probability distribution when the model is created. The distribution indicates how likely it is to observe the node in a particular state given the state of its parents in the model. Three states are allowed for each node (activated, repressed, unchanged). The distributions for the hidden variables are defined in the first step (e.g., a transcription factor regulating a network of genes, see Figure 1). The distributions of the observed variables at each molecule level are learned by the EM algorithm using the input data. After the model is finalized, it is possible to infer the probability of observing the hidden node in a particular state without the observed data (prior probability) or taking the data into account (posterior probability). The output is a matrix of activity scores for each node in the network and each sample.The activity score represents a signed log-likelihood ratio (positive if the node is predicted to be active, negative if the node is predicted to be repressed). Pathways are converted into probabilistic graphical models that include both hidden states for each node and observed states for the nodes that may correspond to the input dataset. The function receives a matrix of activity scores and calculates a p-value for each value using a permutation approach. Hierarchical clustering of the calculated pathway activities per patient and molecule was performed using the R hclust function. Dissimilarity between clusters was calculated using squared Euclidean distance.
[0196] Linking identified patient clusters to CVO To validate the association of the identified patient clusters in both Caucasian and Latino populations (Figure 3) with CV outcomes measured in the ORIGIN cohort, we used logistic regression, Cox proportional hazards models, and survival analysis. To confirm the consistency of the results, these analyses were repeated for subclusters within the initially identified clusters (Figures 12 and 13). Logistic regression was used to estimate whether known CVO risk factors alone were able to distinguish patients from clusters of higher versus lower CVO risk identified by molecular signatures. Because variations in cumulative disease incidence can be evaluated in prospective studies, Cox proportional hazards models were used to estimate the relationship between clusters and time-to-event. Kaplan-Meier survival estimates were used to visualize the association of the identified clusters with disease progression, and the log-rank test was used to test the equality of survivor functions. The CV outcomes used in the analysis were: myocardial infarction (MI), stroke, cardiovascular death and heart failure (HF) with hospitalization, besides death from any cause (Figures 9 and 12). Two additional MACEs were used: (i) the first multiprimary endpoint (=CV composite): a composite of CV death, or non-fatal MI or non-fatal stroke, and (ii) the second multiprimary endpoint (=extended CV composite): a composite of first multiprimary or revascularization or hospitalization for heart failure (Figure 3). Covariates used in the model were: age, sex, BMI (kg / m2), HbA1c (%), c-peptide, HDL-C mmol / L, LDL-C (mmol / L), TG (mmol / L), TC (mmol / L), SBP (mmHg), DBP (mmHg), smoking status, albuminuria or reported albuminuria. All continuous variants were normalized by inverse normal transformation. See Supplementary Table S1 and Figure 10 for sample characteristics for clusters identified with higher CVO risk versus lower CVO risk in Caucasians and Latinos. Box plots of original biomarker levels were generated for visual inspection of differences in BMK levels between clusters (Figure 9). Analyses were performed using STATA version 15.
[0197] result ORIGIN GWAs results In the GWAs we performed in the ORIGIN cohort, we identified a small number of variants that were associated or likely to be associated with CVD (Figure 11). These were mostly rare variants (MAF=0.01 in EUR) or were not located in genic regions (Supplementary Table S2). These variants were not present in the CARDIOGRAM dataset, and we were unable to confirm these findings using the Cardiovascular Disease Knowledge Portal (http: / / broadcvdi.org / home / portalHome). We mapped the GWAs results to nearby loci (Supplementary Table S2) to inform network analyses and identify biological interactions of these loci with other CVD biomarkers (Supplementary Table S3). Results from the GWA analysis are shown in Table 1 for loci prioritized in the network analysis, and all other findings are reported in Figure 6 and Supplementary Table S2.
[0198] CVD Disease Network Prioritization We constructed a CVD network (Figure 7) based on proteins reported to be associated with CVO or death and loci associated with CVD in the ORIGIN CVO study (Supplementary Table S3) with the aim of identifying statistically overconnected direct regulators of disease. From the top eight loci (GHR, RORA, CLINT1, GRM8, LOC101928784, MYH16, SPOCK1, and TMTC2) from the CVD GWAs in the ORIGIN cohort (Supplementary Table S2), only two (RORA and GHR; Figure 6) were prioritized in the overconnectivity analysis (Figure 2). To validate the identified major network (Figure 2), we reran these analyses using an independent GWAs dataset from the CARDIOGRAM consortium (Supplementary Table S2 and Figure 7 for study design). Hyperconnectivity analysis showed that 14 of the 16 protein BMKs (Supplementary Table S3) previously identified as associated with CVOs in the ORIGIN cohort (Table 1) were associated with CVOs (12 in the replication study; (Figure 8). These were part of the top significant networks (Supplementary Table S4) in the discovery and replication analyses (Figures 7 and 8). It should be noted that although 90 other candidate genes were fed as input to the network calculation in the replication study (Supplementary Table S3), only two protein BMKs (NPPB and TFF3) and two genetically associated genes (RORA and GHR) were specific to the subnetwork identified in the discovery analysis (Figure 8).
[0199] Patient stratification associated with CVO progression The prioritized subnetworks and their molecular directional interactions (Figure 2; Supplementary Table S5 for interactions with respective supporting references) were used to inform subsequent patient stratification analyses as described in the workflow (Figure 1; see Methods). Clustering of calculated molecular activities was agnostic to CVO: three major gene clusters (A, B, C in Figure 3, and panel a) and two major patient clusters were identified in Caucasians (1,059 and 849 patients) (Figure 3, panel a) and replicated in Latinos (1,078 and 1,068 patients). Pathway activity for patients at higher CVO risk was clearly suppressed in gene cluster A and activated in gene cluster BMK, as well as for the majority of patients in cluster C (Figure 3 and Table 1). Protein BMK was previously reported to be associated with CVO or death for all causes in the ORIGIN cohort (Gerstein et al., 2015).
[0200] Associations of clusters with the first and second primary composites of CVO were found in the second stage during survival analysis (log-rank analysis to test for the quality of Kaplan-Meier survival estimates and survival functions) and in Cox regression models adjusted for CVO risk factors (Figure 3, panels bc; and Figure 4, panels ac). Similarly, patient clusters were associated with myocardial infarction, stroke, cardiovascular death, and heart failure with hospitalization, as well as all-cause mortality (Figure 9 for Kaplan-Meier survival estimates), with higher-risk clusters having on average about two-fold higher odds of event occurrence than lower-risk clusters (Figure 3, panels de).
[0201] The distribution of CVO risk factors (Supplementary Table 1 and Figure 11) was largely similar between higher and lower CVO risk patient clusters in both Caucasians and Latinos (Supplementary Table 1). In a logistic regression model including age, sex, BMI, HbA1c, c-peptide, HDL-C, LDL-C, TG, TC, SBP, DBP, smoking status, albuminuria or reported albuminuria; only albuminuria, total cholesterol, and age were associated with the lower CVO risk cluster versus the separate higher CVO risk cluster in both Caucasians and Latinos (Supplementary Table 6). Smoking habits were quite different for the clusters identified in Caucasians, while only gender and BMI contributed to the separate clusters in Latinos (Supplementary Table 6). To test how well the clusters were associated and predicted CVO after adjusting for these known CVO risk factors, we used Cox regression models (Figure 3 and more specifically Figure 4). Hazard ratios (HRs) for the second multi-major composite of CVO were largest for gender (women had reduced HR compared to men) and cluster (clusters with lower CVO risk were protective) in both Caucasians and Hispanics (Figure 4, ab). Smoking had a large effect in Caucasians but a moderate effect in Hispanics. Consistently across populations, HDL, TC, SBP, and albuminuria were also associated, but to a lesser extent than gender and cluster. More than 80 percent of the population was diabetic (Supplementary Table 1); HbA1C was more associated with HR in Hispanics, but not C-peptide in Caucasians (Figure 4, ab). For the first multi-major composite of CVO, as shown in Figure 4 (Panel C) - the results were similar to the analysis with the second multi-major composite of CVO. We repeated this analysis for all outcomes for the combined Caucasian and Hispanic populations (Figure 3, de).To test whether the clustering was not randomly related to CVO survival, we repeated the analysis shown in Figure 3 for subclusters of the initially identified cluster, with concordant results (Supplementary Figures 12 and 13). The mean levels of protein BMK in the clusters of higher versus lower CVO risk (Figure 10) were as expected and corresponded to the trend of association with CVO in the ORIGIN cohort (Table 1) for these proteins. Also, the measured levels of protein BMK in the clusters of higher versus lower CVO risk largely corresponded to the higher versus lower calculated activity for each molecule, except for NT-proBNP (Figure 3, panel a; Table 1). Thereby, the calculated activity does not need to reflect the level of BMK, since the activity is inferred based on molecular interactions (e.g., what is the probability that the activity of protein x is inactivated?). Although NT-proBNP levels were higher in the cluster of patients with higher CVO risk compared to those with lower CVO risk (Figure 10), the inferred activity for this gene in the network was lower in the higher CVO risk group than in the lower risk group (Table 1). If the network activity is recalculated using NT-proBNP levels as input for the subset of patients whose levels were measured, the level of the protein per cluster is equivalent to its level in the similar cluster when the network activity is calculated without using NT-proBNP levels (Figure 14). Similarly, in the absence of TNF-alpha measurements for the ORIGIN study, the network activity was calculated taking into account only neighboring molecules, which indicated that TNF-alpha was less active in patients with higher CVO risk, although the actual protein level may have been higher. The molecular activities of IL6 and RelA NF-KB, the upstream regulators of the network, were also estimated to be higher in the higher CVO risk group (Figure 3, Table 1).These networks were less informative than TNF alpha, which already contained all interacting molecules (Figure 2), but IL6 and RelA NF-KB were central to their activity and could therefore be calculated.
[0202] Consideration Here we developed a workflow that goes from identifying CVD networks to proof of concept that the identified molecular interactions can be used to stratify patients with respect to disease progression. Our computational biology approach identified a group of biomarkers associated with CVD outcomes and for the first time linked SNPs associated with CVD risk. Our work identifies molecular interactions and shows the interdependence of molecular activities associated with various stages of CVD progression.
[0203] These results may allow for the identification of individuals at risk of suffering from a CVD event, and therefore allow for early intervention to prevent or reduce the risk of the individual suffering from a CVD event. In particular, detection of each biomarker alone allows for the identification of individuals at risk of suffering from a CVD event, and detection of multiple biomarkers, i.e., a biomarker signature, provides a particularly effective means of allowing for early intervention to prevent or reduce the risk of the individual suffering from a CVD event.
[0204] This is the first evidence to link reported GHR and RORA variants with CVD, and Example 2 shows how to determine a subject's risk of cardiovascular disease based on the detection of SNPs in the GHR and RORA genes. Although the genetic evidence of association with CVD is not strong, RORA is known to regulate a number of genes involved in lipid metabolism, such as apolipoprotein AI, APOA5, CIII, CYP71 and PPAR gamma, possibly acting as a receptor for cholesterol or one of its derivatives (cited Uniprot). Furthermore, overexpression of RORA isoforms suppresses TNF alpha-induced expression of adhesion molecules in human umbilical vein endothelial cells and regulates inflammatory responses (Migita et al., 2004). Our network analysis shows that RORA and GHR interact with key regulatory hubs of the identified CVD network of BMK. Consistent with that, their genotypes contribute to the calculation of network activity and to the clustering of patients, which in turn are associated with CVD progression.
[0205] RORA and GHR connected to the main network through upstream regulators (TNF-alpha and IL6) common to the prioritized molecules. Although TNF-alpha and IL6 were not part of the input information to create the network, we were able to reveal them as hidden master regulators of the input molecules. Although the link of TNF-alpha and IL6 to CVD is known (Lopez-Candales et al., 2017), the network allows visualization of the direction of molecular interactions (Figure 2). These were used to inform Bayesian statistical methods to calculate patient-specific pathway activities, thereby elucidating meaningful patient subgroups and providing summaries of the different omics layers as a more robust signature to understand mechanistic interactions.
[0206] One advantage of Bayesian network analysis over single BMK analysis is that it can estimate the activity of BMKs that were not measured or not measured with sufficient quality in patient samples (e.g., TNF-alpha and IL6 in the ORIGIN study), especially when they are central to disease networks. This can lead to the discovery of new BMKs in disease progression. TNF-alpha and IL6 have been intensively studied in CVD models, but in humans their protein detection is somewhat cumbersome and their levels are also affected by physical activity (Vijayaraghava et al., 2017).
[0207] Here, we could infer their activity based on downstream interacting proteins, from which measured levels were fed into the Bayesian model (Figure 2). The calculated activity can be interpreted as the activation of the molecule rather than the actual measurement of protein levels, i.e., transcription factors are calculated to be active or inactive based on downstream targets, not just effective measured protein levels. For example, the TNF-alpha pathway was less active in patients with a higher risk for CV events. This may seem contradictory at first, since in functional studies, inhibition of TNF-alpha had beneficial effects on cardiac function and outcomes. Nevertheless, most predictive studies using TNF-alpha alpha inhibitor binding directly to TNF-alpha reported disappointing and contradictory results on CVO risk. For example, in rheumatoid arthritis (RA) patients, it is the control of the disease itself (e.g., inflammatory pathways) rather than the reduction in TNF-alpha levels itself that is necessary for the prevention of CVO (Coblyn et al., 2016). Thereby, specific protein interactions must be related to the inflammatory process, and TNF-alpha is just one piece of the puzzle. In this context, IL6, which is also an inflammation marker, was predicted to be more active in the higher CVO risk cluster in our study. The transcription factor nuclear factor kappa-B, which was predicted to be more active in the higher CVO risk cluster in our study, regulates the expression of many pro-inflammatory cytokines and is cardioprotective during acute hypoxia and reperfusion injury. As for TNF-alpha, the activity of NTproBNT, a cardiac hormone that may function as a paracrine antifibrotic factor in the heart, was predicted to be lower in patients with higher CVO risk. Here, we could confirm that the majority of these patients actually had higher levels of the protein than patients with a lower propensity for CVO. BNP and NTproBNP peptides are released into the blood circulation in response to ventricular pressure and volume overload.Cleavage of the pre-proBNP precursor in cardiomyocytes results in the formation of proBNP, which is then cleaved into N-terminal (NTproBNP) and C-terminal (BNP) fragments. Most biological effects of BNP are the result of its binding to natriuretic peptide receptors (Dhingra 2002). Thus, the calculated activity of NTproBNP reflects the molecular interactions that indicate its activity rather than the actual protein level. Although the inventors do not wish to be bound by any particular theory, they hypothesize that NTproBNP levels are higher in patients at higher CVO risk, but are not as effective as they should be.
[0208] NPPB and TFF3 were top associated BMKs by conventional statistics in the ORIGIN cohort, but were not found in the network constructed using the CARDIOGRAM gene set. Because the hyperconnectivity algorithm identifies the shortest path connecting genes, we do not wish to be bound by any particular theory, but we can conclude that: i) there is an active biological network connecting GWAs-associated genes to protein biomarkers that were associated with CVD, for which TNF-alpha plays a central role; ii) NPPB and TFF3 are clearly relevant biomarkers for CVD, but they appear to be more downstream in the cascade of events related to TNF-alpha than the gene set identified in the CARDIOGRAM study. Nevertheless, for the purpose of identifying disease-associated networks for BMK discovery, protein BMKs should be more associated than CVD-associated loci, since GWA results often map to genes that do not code for circulating proteins. In patient stratification analysis, we mainly used circulating proteins (N=13 plus 2 genetic markers) to make it suitable for translational application, since tissue-specific samples are often difficult to obtain or not suitable in standard clinical practice. Although the complexity of this computational approach does not make it a process that can be easily used in the clinic, the resulting patient strata can be further examined to identify ideal BMK (including known clinical parameters) combinations and ratios that represent different stages of disease progression.
[0209] Identifying clusters of subpopulations that progress differently to CVO will inform: iii) how the molecular signature of each cluster can be translated into a combination of biomarkers with prognostic value; and iv) how these markers respond to treatment in further cohorts. The molecular signature of disease progression should thereby inform clinical trial design and strategies for optimizing drug efficacy and thus treatment.
[0210] Supplementary Table S1. Study sample characteristics for the total sample and high versus low CVO risk clusters for Caucasian and Latino subpopulations.
[0211] Supplementary Table S2. List of molecular entities used in the hyperconnectivity analysis. Results from the GWA analyses performed in the ORIGIN and CARDIOGRAM consortia were assigned to gene loci (see Methods) and protein biomarkers were assigned to each of those genes.
[0212] Supplementary Table S3. List of associated loci included in the hyperconnectivity network analysis. Loci identified with P<10-6 in the ORIGIN cohort and with P<10-8 in the CARDIOGRAM consortium were enrolled in the analysis. Attached excel file.
[0213] Supplementary Table S4. List of subnetworks identified in the overconnectivity analysis along with their respective p-values.
[0214] Supplementary Table S5. Connectivity network results table. Nodes and relationships of the entities comprising the networks used for the Bayesian statistical-based network analysis.
[0215] Supplementary Table S6. Logistic regression models to identify standard CVO risk factors (age, sex, BMI, HbA1c, c-peptide, HDL-C, LDL-C, TG, TC, SBP, DBP, smoking status, albuminuria or reported albuminuria) contributing to cluster separation in Caucasians and Latinos.
[0216] Example 2 – Detection of SNPs associated with CVD risk Having identified two SNPs associated with individuals at increased risk for cardiovascular disease (CVD), namely reference SNP cluster ID: rs73420079 (SEQ ID NO: 3) and reference SNP cluster ID: rs4314405 (SEQ ID NO: 40), our work enables the identification of individuals at increased risk of CVD using oligonucleotide probes designed to detect the SNPs.
[0217] Oligonucleotide probes for detecting the presence of SNPs can be produced and synthesized based on the SNP rs number by any available oligonucleotide probe design tool. For example, probes for SNPs can be sent to Illumina®'s Illumina assay design tool for scoring based on the rs number format to produce assay-ready probes.
[0218] A sample may be isolated from a patient, and the sample may be a blood sample. The individual's nucleic acid is isolated from the sample. Isolation may occur by any means convenient to the practitioner. For example, isolation may occur by first lysing the cells using detergents, enzymatic digestion, or physical disruption. Contaminants are then removed from the nucleic acid using, for example, enzymatic digestion, organic solvent extraction, or chromatographic methods.
[0219] The individual's nucleic acid may be purified and / or concentrated by any means, including precipitation with alcohol, centrifugation, and / or dialysis. The individual's nucleic acid is then assayed for the presence or absence of one or more of the SNPs using oligonucleotide probes capable of hybridizing to a nucleic acid sequence that contains one or more of the SNPs.
[0220] Detection of reference SNP cluster ID: rs73420079 (SEQ ID NO: 3) and / or reference SNP cluster ID: rs4314405 (SEQ ID NO: 40) in a sample indicates that the individual has an increased risk of CVD.
Claims
1. 1. A method for assisting in determining, diagnosing, and / or predicting an individual's risk of suffering from cardiovascular disease, the method comprising: a. In a sample obtained from an individual: the biomarkers tumor necrosis factor (TNF)-α and retinoic acid receptor-related orphan receptor alpha (RORA), and optionally glutathione S-transferase alpha 1 (GSTA1); N-terminal-prohormone BNP (NT-proBNP); tenascin C (TNC); growth hormone receptor (GHR); alpha-2-macroglobulin (A2M); insulin-like growth factor binding protein 2 (IGFBP2); apolipoprotein B (APO B); selenoprotein P (SEPP1); trefoil factor (TFF3); interleukin 6 (IL6); chitinase 3-like 1 (CHI3L1); hepatocyte growth factor receptor (MET); growth differentiation factor 15 (GDF15); chemokine (C-C motif) ligand 22 (CCL22); tumor necrosis factor receptor superfamily, member 11 (TNFRSF11); angiopoietin 2 (ANGPT2); and v-Rel avian reticuloendotheliosis viral oncogene homolog A nuclear factor-kappa B (ReLA NF-KB); b. Comparing the expression level, amount and / or protein activity of the biomarker with a reference from a healthy control population; and c. Assisting in determining, diagnosing, and / or predicting an individual's risk of suffering from cardiovascular disease when the expression level, amount, and / or protein activity of the biomarker deviates from a reference from a healthy control population. The above method, comprising:
2. 2. The method of claim 1, wherein a decrease in the expression, amount and / or protein activity of TNF-α, GSTA1, NT-proBNP, RORA and / or TNC compared to a reference indicates that the individual has a higher risk of suffering from cardiovascular disease or has a poor prognosis.
3. An increase in the expression, amount and / or protein activity of GHR, A2M, IGFBP2, APOB, SEPP1, TFF3, IL6 and / or CHI3L1 as compared to a reference indicates that the individual has a higher risk of suffering from cardiovascular disease or a poor prognosis.
3. The method according to claim 1 or 2.
4. 4. The method of any one of claims 1 to 3, wherein a decrease in the expression, amount and / or protein activity of MET, GDF15, CCL22, TNFRSF11, ANGPT2 and / or ReIA NF-KB compared to a reference indicates that the individual has a lower risk of suffering from cardiovascular disease or has a good prognosis.
5. 5. The method of any one of claims 1 to 4, wherein step a) comprises detecting the expression levels, amounts and / or protein activities of at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17 or at least 18 biomarkers selected from the group consisting of TNF-α; GSTA1; NT-proBNP; RORA; TNC; GHR; A2M; IGFBP2; APOB; SEPP1; TFF3; IL6; CHI3L1; MET; GDF15; CCL22; TNFRSF11; ANGPT2; and ReIA NF-KB.
6. The method of any one of claims 1 to 5, wherein the cardiovascular disease is selected from the group consisting of: cardiovascular death; myocardial infarction; stroke; and heart failure.
7. The method of any one of claims 1 to 6, wherein the sample comprises blood.
8. 1. A kit for determining, diagnosing, and / or predicting an individual's risk of suffering from cardiovascular disease, the kit comprising: a. detection means for detecting the polypeptide amount and / or protein activity of the biomarkers TNF-α and RORA, and optionally one or more biomarkers selected from the group consisting of GSTA1; NT-proBNP; TNC; GHR; A2M; IGFBP2; APOB; SEPP1; TFF3; IL6; CHI3L1; MET; GDF15; CCL22; TNFRSF11; ANGPT2 and ReIA NF-KB in a sample obtained from the test subject; and b. Reference values from a healthy control population for the polypeptide amount and / or protein activity of the biomarkers TNF-α and RORA, and optionally one or more biomarkers selected from the group consisting of GSTA1; NT-proBNP; TNC; GHR; A2M; IGFBP2; APOB; SEPP1; TFF3; IL6; CHI3L1; MET; GDF15; CCL22; TNFRSF11; ANGPT2 and ReIA NF-KB, Including, wherein the kit comprises: i) a decrease in the polypeptide amount and / or protein activity of TNF-α; GSTA1; NT-proBNP; RORA and / or TNC compared to a reference; and / or an increase in the polypeptide amount and / or protein activity of GHR; A2M; IGFBP2; APOB; SEPP1; TFF3; IL6 and / or CHI3L1 compared to a reference, for determining, diagnosing and / or predicting that an individual has a higher risk of suffering from cardiovascular disease; and / or ii) A decrease in polypeptide amount and / or protein activity of MET; GDF15; CCL22; TNFRSF11; ANGPT2 and / or ReIA NF-KB compared to a reference to determine, diagnose and / or predict that an individual has a lower risk of suffering from cardiovascular disease. The above kit is used to identify