Chromosome conformation markers in prostate cancer and lymphoma
By detecting the chromosomal status of the subpopulation in the cancer, and using the chromosome interaction typing system, the problem that the existing technology cannot effectively clarify cancer regulation and pathogenic factors is solved, and the accurate prediction of cancer prognosis and personalized treatment is achieved.
Patent Information
- Application Number
- CN202080044081.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-04-29
- Filing Date
- 2020-05-06
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2040-05-06
AI Technical Summary
Existing DNA and protein typing methods cannot effectively elucidate regulatory and pathogenic factors in cancer disease processes, especially in diffuse large B-cell lymphoma and prostate cancer.
By detecting the chromosomal status of subpopulations representing the population, a chromosomal interaction typing system is used to determine the subpopulations related to cancer prognosis, and then personalized treatment is carried out.
Accurate prediction of cancer prognosis and the provision of personalized treatment plans are achieved, and the effectiveness of treatment and patient survival rate are improved.
Smart Images

Figure CN114008218B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to disease processes. Background Art
[0002] The regulatory and pathogenic factors in the cancer disease process are complex and cannot be easily elucidated using existing DNA and protein profiling methods.
[0003] Diffuse large B-cell lymphoma (DLBCL) is a cancer of B cells, a type of white blood cell responsible for producing antibodies. Diffuse large B-cell lymphoma is the most common type of non-Hodgkin lymphoma in adults, with an average annual incidence of 7-8 cases per 100,000 people per year in the United States and the United Kingdom. However, little is known about the outcomes of the disease process.
[0004] Prostate cancer is caused by the abnormal and uncontrolled growth of cells in the prostate gland. While prostate cancer survival rates have continued to improve over the past few decades, the disease is still largely considered incurable. According to the American Cancer Society, the one-year relative survival rate for all stages of prostate cancer combined is 20%, while the five-year relative survival rate is 7%. Summary of the invention
[0005] The present inventors have identified subtypes of prostate cancer, diffuse large B-cell lymphoma (DLBCL), and lymphoma patients defined by chromosome conformational features.
[0006] According to the present invention, there is provided a method (process) for detecting the chromosome state representing a subpopulation in a population, comprising determining whether there is a chromosome interaction associated with the chromosome state in a defined region of the genome; and
[0007] - wherein said chromosomal interactions have optionally been identified by a method of determining which chromosomal interactions are associated with the chromosomal state of said subpopulation corresponding to said population comprising the steps of contacting a first set of nucleic acids from subpopulations of chromosomes having different states with a second set of index nucleic acids and allowing complementary sequences to hybridize, wherein nucleic acids in said first set of nucleic acids and second set of nucleic acids represent ligation products comprising sequences from two chromosomal regions that have been brought together in a chromosomal interaction, and wherein the hybridization pattern between said first set of nucleic acids and second set of nucleic acids allows determination of which chromosomal interactions are specific to said subpopulation; and
[0008] - wherein the subpopulation is associated with prognosis of prostate cancer, and the chromosomes interact:
[0009] (i) present in any region or gene listed in Table 6; and / or
[0010] (ii) corresponds to any chromosomal interaction represented by any of the probes shown in Table 6, and / or
[0011] (iii) is present in a 4000 base region including or flanking (i) or (ii);
[0012] or
[0013] - wherein the subpopulation is associated with the prognosis of DLBCL and the chromosomes interact:
[0014] a) present in any region or gene listed in Table 5; and / or
[0015] b) corresponds to any of the chromosomal interactions represented by any of the probes shown in Table 5, and / or
[0016] c) is present in a 4000 base region including or flanking (a) or (b);
[0017] or
[0018] - wherein said subpopulation is associated with the prognosis of lymphoma and said chromosomes interact:
[0019] (iv) present in any of the regions or genes listed in Table 8; and / or
[0020] (v) corresponds to any of the chromosome interactions shown in Table 8, and / or
[0021] (vi) is present in a 4000 base region including or flanking (iv) or (v). BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 A principal component analysis (PCA) for a prostate cancer study is shown.
[0023] Figure 2 A VENN comparison of two PCA prognostic classifiers is shown.
[0024] Figure 3 PCA analysis of DLBCL is shown.
[0025] Figure 4 PCA of seven BTK markers in DLBCL (OBD RD051) is shown.
[0026] Figure 5 An example of how chromosome interaction typing can be performed is shown.
[0027] Figure 6Markers from canine lymphoma research that can be used for the method of the present invention are shown. The figure shows the reduction of markers. 70% of the 38 samples were used as training sets (28) and for marker selection. The remaining 10 samples were used as test sets. Multiple training sets and test sets were used. Univariate analysis, Fisher's exact test (D and E columns results) and multivariate analysis penalized logistic modeling (GLMNET, B and C columns results). Markers 2 to 18 are lymphoma markers, and markers 19 to 23 are controls. The first 11 rings present in lymphomas were selected for classification.
[0028] Figure 7 Canine markers for human genes are shown. The table shows the top 11 canine markers with the closest mapped genomic regions that map to the human genome (Hg38). The adjacent network was constructed using the 11 markers (dark colors), lighter colored nodes, and connexins using the NCI database.
[0029] Figure 8 Canine markers for human genes are shown. As before, the network was enriched for pathways. Only the 11 canine-mapped loci were used for enrichment, and connectivity patterns were omitted during enrichment. Nodes with lighter colors belong to KEGGCML pathways.
[0030] Fig. 9 Shows the XGBoost 11-logo model for training set 1 and test set 1
[0031] Fig.10 Shows the XGBoost 11-logo model for training set 2 and test set 2
[0032] Fig.11 Shows the training set 3 and test set 3 XGBoost 11 logo model
[0033] Fig.12 The logistic PCA of training set 1 is shown
[0034] Fig.13 The logistic PCA of training set 1 and test set 1 is shown. The logistic PCA model was used to predict the test set 1 (triangles). The dark triangles are lymphomas from the test set (labeled) and the light triangles are controls from the test set. The training lymphoma samples are dark and the controls are light.
[0035] Fig.14 Shows the ROC & AUC of training set 1 and test set 1
[0036] Fig.15 Shows PFS in patients with NFKB1 EpiSwitch TM Call and ring dynamics. 118 people use EpiSwitch TMPatients in the 10-marker human model were called ABC or GCB, and PFS was modeled using this calling and the dynamics of the loops. GCB with loops did not die, also indicating that the human model is effective for disease prognosis.
[0037] Fig.16 PFS EpiSwitch in 118 patients with NFATC1 was shown TM Same as before, but for NFATc1, which again shows that the human prognostic model using this marker as one of the 10 human markers is very good at classification.
[0038] Fig.17 A three-step approach to identify, evaluate, and validate diagnostic and prognostic biomarkers for prostate cancer (PCa) is shown.
[0039] Fig.18 Shown is a PCA of five markers applied to 78 samples comprising two groups. In the first group, 49 known samples (24 PCa and 25 healthy controls (Cntrl)) were combined with a second group of 29 samples (including 24 PCa samples and 5 healthy Cntrl samples).
[0040] Fig.19 The workflow for developing a classifier is shown.
[0041] Fig. 20 The relevant gene sets used for the classifiers are shown.
[0042] Fig.21 Overlap of EpiSwitchDLBCL-CCS and Fluidigm subtype calls and ROC curves when applied to the discovery cohort are shown. A. Subtype calls made by EpiSwitch DLBCL-CCS and Fluidigm analyses on samples of known subtype. 60 of 60 samples were called equally by both analyses. B. DLBCL-CCS receiver operating curve (ROC) when applied to the discovery cohort. C. Kaplan-Meier survival analysis (by progression-free survival) of samples called as ABC or GCB by DLBCL-CCS. Samples called as ABC showed significantly worse long-term survival than samples called as GCB.
[0043] Fig. 22 Assignment of DLBCL subtypes in type III samples by EpiSwitch and Fluidigm analysis is shown.
[0044] Fig.23Comparison of long-term survival using EpiSwitch and Fluidigm for baseline DLBCL subtype calling in type III samples is shown. Kaplan-Meier survival curves for 58 DLBCL patients classified as ABC, GCB, or unclassified by Fluidigm analysis (A) or EpiSwitch DLBCL-CCS (B). Fluidigm classified 15 samples as ABC, 22 as GCB, and 21 as UNC. EpiSwitch classified 34 as ABC and 24 as GCB.
[0045] Fig.24 Shown are the mean survival times by EpiSwitch and Fluidigm classification in the validation cohort.
[0046] Fig.25 A preliminary assessment of possible DLBCL subtypes is shown.
[0047] Fig.26 Shown is the PCA of DLBCL patients with baseline ABC / GCB subtype calling by EpiSwitch in the discovery cohort. DETAILED DESCRIPTION
[0048] Aspects of the Invention
[0049] The present invention relates to the determination of prostate cancer prognosis, particularly whether the cancer is aggressive or inert. The determination is by typing any relevant marker disclosed herein (e.g., in Table 6), or a combination of preferred markers, or markers in a defined specific region disclosed herein. Therefore, the present invention relates to a method of typing a prostate cancer patient to determine whether the cancer is aggressive or inert.
[0050] The present invention also relates to the determination of DLBCL prognosis, in particular, the prognosis in terms of survival. The determination is by typing any relevant marker disclosed herein (e.g., in Table 5), or a combination of preferred markers, or markers in a defined specific region disclosed herein. Therefore, the present invention relates to a method for typing a patient with DLBCL to determine the prognosis of the patient in terms of survival, such as determining the expected rate of progression of the disease and / or time of death.
[0051] In the method of the present invention, the subpopulation of prostate cancer or DLBCL is basically identified by typing of markers. Therefore, for example, the present invention relates to a group of epigenetic markers associated with the prognosis of these diseases. Therefore, the present invention allows to provide personalized treatment that accurately reflects the needs of patients to patients.
[0052] The present invention also relates to determining the prognosis of lymphoma based on the typing chromosome interactions defined in Table 8 or Table 9.
[0053] Preferably, Tables 5 to 7 are relevant to determining the prognosis in humans. Preferably, Tables 8 and 9 are relevant to determining the prognosis in dogs.
[0054] Based on the results of the methods, any of the therapies mentioned herein, such as drugs, may be administered to the individual.
[0055] Marker sets are disclosed in the figures and tables. In one embodiment, the present invention uses at least 10 markers from any disclosed marker set. In another embodiment, the present invention uses at least 20% of the markers from any disclosed marker set.
[0056] Invention Method
[0057] The method of the present invention includes a typing system for detecting chromosome interactions associated with prognosis. The typing can be performed using the EpiSwitch TM The system is based on the chromosome cross-linked regions that are gathered together in the chromosome interaction, the chromosome DNA is cut, and then the nucleic acids present in the cross-linked entities are connected to obtain the connected nucleic acids having sequences from the two regions forming the chromosome interaction. Detection of the connected nucleic acids can be used to determine whether a specific chromosome interaction exists.
[0058] Chromosomal interactions can be identified by the methods described above using the first and second nucleic acid populations. These nucleic acids can also be identified using EpiSwitch TM Technology generation.
[0059] Epigenetic interactions relevant to the present invention
[0060] As used herein, the terms "epigenetic" and "chromosomal" interactions generally refer to interactions between chromosome terminal regions that are dynamic and change, form, or break depending on the state of the chromosome regions.
[0061] In ad hoc methods of the present invention, chromosome interaction is usually first detected by generating connection nucleic acid, and the connection nucleic acid comprises the sequence from two regions of the chromosome as a part for interaction. In such method, these regions can be cross-linked by any suitable means. In a preferred aspect, the interaction uses formaldehyde to be cross-linked, but any aldehyde or D-biotin-e-aminocaproic acid-N-hydroxysuccinimide ester or digoxin-3-O-methylcarbonyl-e-aminocaproic acid-N-hydroxysuccinimide ester can also be cross-linked. Formaldehyde can make the DNA chain at a distance of 4 angstroms cross-linked. Preferably, the chromosome interaction is on the same chromosome and optionally according to 2 angstroms to 10 angstroms.
[0062] Chromosomal interactions can reflect the state of a chromosomal region, for example, whether it is transcribed or repressed in response to changes in physiological conditions. The subpopulation-specific chromosomal interactions defined herein have been found to be stable, providing a reliable way to measure differences between two subpopulations.
[0063] In addition, chromosomal interactions that are specific to features (e.g., prognosis) often occur early in the biological process, for example compared to other epigenetic markers (e.g., changes in methylation or histone binding). Therefore, the method of the present invention is able to detect early stages of biological processes. This makes early intervention (e.g., treatment) potentially more effective. Chromosome interactions also reflect the current state of an individual and can therefore be used to assess changes in prognosis. In addition, there is little variation in the relevant chromosomal interactions between individuals within the same subgroup. Since each gene has up to 50 different possible interactions, detecting chromosomal interactions provides very useful information, so the method of the present invention can query 500,000 different interactions.
[0064] Selected marker set
[0065] The term "marker" or "biomarker" herein refers to a specific chromosome interaction that can be detected (typed) in the present invention. Any of the specific markers disclosed herein can be used in the present invention. More marker sets can be used, for example, in combinations disclosed herein or digitally. The specific markers disclosed in this article are preferred, and the markers present in the genes and regions mentioned in this article are preferred. These markers can be typed by any suitable method, such as the PCR or probe-based methods disclosed herein, including qPCR methods. Markers are defined herein by position or by probe and / or primer sequences.
[0066] The where and why of epigenetic interactions
[0067] Epigenetic chromosome interactions can overlap and include chromosome regions that show coding related genes or undescribed genes, but may also be located in intergenic regions. It should also be noted that the inventors have found that epigenetic interactions in all regions are equally important in determining the state of the chromosome locus. These interactions are not necessarily in the coding region of the specific gene located at the locus, but may be located in the intergenic region.
[0068] The chromosome interactions detected in the present invention can be caused by changes in the underlying DNA sequence, environmental factors, DNA methylation, non-coding antisense RNA transcripts, non-mutagenic carcinogens, histone modifications, chromatin remodeling and specific local DNA interactions. The changes that cause chromosome interactions may be caused by changes in the underlying nucleic acid sequence, and these changes themselves do not directly affect gene products or gene expression patterns. Such changes may be, for example, gene fusions and / or gene deletions of SNPs, intergenic DNA, microRNAs and non-coding RNAs within and / or outside genes. For example, it is known that about 20% of SNPs are located in non-coding regions, so the method can also provide useful information in non-coding situations. On the one hand, the chromosome regions that gather together to form interactions are less than 5kb, 3kb, 1kb, 500 base pairs or 200 base pairs apart on the same chromosome.
[0069] The detected chromosomal interaction is preferably within any of the genes mentioned in Table 5. However, the chromosomal interaction may also be located upstream or downstream of the gene, for example up to 50,000, up to 30,000, up to 20,000, up to 10,000 or up to 5000 bases upstream or downstream of the gene or the coding sequence.
[0070] The detected chromosomal interaction is preferably within any of the genes mentioned in Table 6. However, the chromosomal interaction may also be located upstream or downstream of the gene, for example up to 50,000, up to 30,000, up to 20,000, up to 10,000 or up to 5000 bases upstream or downstream of the gene or coding sequence.
[0071] The detected chromosomal interactions are preferably within any of the genes mentioned in Table 9. However, the chromosomal interactions may also be located upstream or downstream of the gene, for example up to 50,000, up to 30,000, up to 20,000, up to 10,000 or up to 5000 bases upstream or downstream of the gene or coding sequence.
[0072] Subpopulations, time points, and personalized treatment
[0073] The object of the present invention is to determine the prognosis. This can be at one or more defined time points, for example at least 1, 2, 5, 8 or 10 different time points. The duration between at least 1, 2, 5 or 8 time points can be at least 5 days, 10 days, 20 days, 50 days, 80 days or 100 days.
[0074] As used herein, "subpopulation" preferably refers to a subpopulation of a population (a subpopulation within a population), more preferably refers to a subpopulation within a population of a specific animal, such as a specific eukaryotic organism or mammal (e.g., human, non-human, non-human primate, or rodent, such as mouse or rat). Most preferably, "subpopulation" refers to a subpopulation within a human population. The subpopulation may be a canine subpopulation, such as dog.
[0075] The present invention includes specific subgroups in detection and treatment colonies. The inventors have found that the chromosome interactions between the subsets (e.g., at least two subsets) in a given colony are different. Identifying these differences will allow doctors to classify their patients as a part of a subset of a colony as described in the method. Therefore, the present invention provides a method for providing personalized medicine to patients based on the epigenetic chromosome interactions of the patient for doctors.
[0076] In one aspect, the invention relates to testing an individual for:
[0077] - are fast or slow "progressors", and / or
[0078] - Have aggressive or indolent disease.
[0079] The present invention also allows for the determination of an individual's expected survival time.
[0080] Such tests can be used to select how to subsequently treat the patient, such as the type of drug and / or the dosage of the drug and / or the frequency of drug administration.
[0081] Generating ligated nucleic acids
[0082] Certain aspects of the invention utilize linker nucleic acids, particularly linker DNA. These linker nucleic acids include sequences from two regions that come together in a chromosome interaction and thus provide information about the interaction. TM The methods use the generation of such linked nucleic acids to detect chromosome interactions.
[0083] Thus, the methods of the invention may include the step of generating a linked nucleic acid (eg, DNA) by the following steps (including methods comprising these steps):
[0084] (i) cross-linking epigenetic-chromosomal interactions present at chromosomal loci, preferably in vitro;
[0085] (ii) optionally isolating the cross-linked DNA from the chromosomal locus;
[0086] (iii) cleaving the cross-linked DNA, for example, by restriction digestion with an enzyme that cleaves the cross-linked DNA at least once, particularly an enzyme that cleaves at least once within the chromosomal locus;
[0087] (iv) ligating the cross-linked cleaved DNA ends (particularly forming a DNA loop); and
[0088] (v) optionally identifying the presence of said connecting DNA and / or said DNA circle, in particular using techniques such as PCR (polymerase chain reaction) to identify the presence of specific chromosomal interactions.
[0089] These steps can be implemented to detect any aspect of chromosome interaction mentioned herein.These steps can also be implemented to generate the first group of nucleic acids and / or the second group of nucleic acids mentioned herein.
[0090] PCR (polymerase chain reaction) can be used for detecting or identifying connection nucleic acid, for example, the size of the PCR product produced can indicate the specific chromosome interaction of existence, and therefore can be used for identifying the locus state. In preferred aspects, at least 1,2 or 3 primers or primer pairs as shown in Table 5 are used for the PCR reaction. In other aspects, at least 1,10,20,30,50 or 80 primers or primer pairs as shown in Table 6 are used for the PCR reaction. The technician will know the multiple restriction endonucleases that can be used for cutting the DNA in the target chromosome locus. Obviously, concrete enzyme will be used according to the locus studied and the dna sequence dna therein. As described in the present invention, the limiting examples of the restriction endonuclease that can be used for cutting DNA is TaqI.
[0091] EpiSwitch TM technology
[0092] EpiSwitch TM EpiSwitch TM The epiSwitch® marker data can be used to detect phenotype-specific epigenetic chromosome conformation features. TM EpiSwitch has several advantages over conventional methods. They have low levels of random noise, for example because nucleic acid sequences from the first set of nucleic acids of the invention hybridize to the second set of nucleic acids, or fail to hybridize to the second set of nucleic acids. This provides a binary result, allowing complex mechanisms to be measured at the epigenetic level in a relatively simple way. TM The technology has fast processing time and low cost. In one aspect, the processing time is 3 hours to 6 hours.
[0093] Samples and sample handling
[0094] The methods of the invention are typically performed on a sample. The sample may be obtained at a defined time point, such as any time point defined herein. The sample typically comprises DNA from an individual. The sample typically comprises cells. In one aspect, the sample is obtained by minimally invasive means and may be, for example, a blood sample. The DNA may be extracted and cleaved with standard restriction endonucleases. This may predetermine which chromosome conformations are preserved and use EpiSwitch TM Due to the synchronicity of chromosome interactions between tissues and blood, including horizontal transfer, blood samples can be used to detect chromosome interactions in tissues, such as those associated with disease. For some conditions, such as cancer, the use of blood is advantageous because genetic noise caused by mutations can affect the chromosome interaction "signal" in the relevant tissue.
[0095] Properties of the Nucleic Acids of the Invention
[0096] The present invention relates to certain nucleic acids, such as the connection nucleic acids used or generated in the methods of the present invention as described herein. These nucleic acids may be identical to the first nucleic acid and the second nucleic acid mentioned herein, or have any of their properties. The nucleic acid of the present invention generally includes two parts, each of which includes a sequence from one of the two regions of the chromosome that are gathered together in the chromosome interaction. Generally, the length of each part is at least 8, 10, 15, 20, 30 or 40 nucleotides, for example, a length of 10 to 40 nucleotides. Preferred nucleic acids include sequences from any gene mentioned in any of the tables. Generally preferred nucleic acids include the specific probe sequences mentioned in Table 5; or fragments and / or homologues of these sequences. Preferred nucleic acids may include the specific probe sequences mentioned in Table 6; or fragments and / or homologues of these sequences.
[0097] Preferably, the nucleic acid is DNA. It should be understood that, in the case of providing a specific sequence, the present invention can use a complementary sequence as required by specific aspects. Preferably, the nucleic acid is DNA. It should be understood that, in the case of providing a specific sequence, the present invention can use a complementary sequence as required by specific aspects.
[0098] The primers shown in Table 5 can also be used in the present invention described herein. In one aspect, primers comprising any of the following are used: sequences shown in Table 5; or fragments and / or homologs of any of the sequences shown in Table 5. The primers shown in Table 6 can also be used in the present invention described herein. In one aspect, primers comprising any of the following are used: sequences shown in Table 6; or fragments and / or homologs of any of the sequences shown in Table 6. The primers shown in Table 8 can also be used in the present invention described herein. In one aspect, primers comprising any of the following are used: sequences shown in Table 8; or fragments and / or homologs of any of the sequences shown in Table 8.
[0099] The second set of nucleic acids - the "index" sequence
[0100] The second set of nucleic acid sequences functions as a set of index sequences and is essentially a set of nucleic acid sequences suitable for identifying subpopulation-specific sequences. They may represent "background" chromosome interactions and may or may not be selected in some way. They are typically a subset of all possible chromosome interactions.
[0101] The second group of nucleic acids can be obtained by any suitable method. They can be calculated or based on individual chromosome interactions. They usually represent a larger colony than the first group of nucleic acids. In a specific aspect, the second group of nucleic acids represents all possible epigenetic chromosome interactions in a specific genome. In another specific aspect, the second group of nucleic acids represents the major part of all possible epigenetic chromosome interactions present in the colony described herein. In a specific aspect, the second group of nucleic acids represents at least 50% or at least 80% epigenetic chromosome interactions in at least 20, 50, 100 or 500 genes (for example, 20 to 100 or 50 to 500 genes).
[0102] The second group of nucleic acids generally represents at least 100 possible epigenetic chromosome interactions, which modify, regulate or mediate the phenotype in the colony in any way. The second group of nucleic acids can represent the chromosome interactions of the disease state (generally relevant to diagnosis or prognosis) affecting the species. The second group of nucleic acids generally includes sequences representing epigenetic interactions relevant to and unrelated to the prognosis subgroup.
[0103] In a specific aspect, the second group of nucleic acids is at least partially derived from naturally occurring sequences in the population, and is typically obtained by in silico methods. The nucleic acid may further include single or multiple mutations compared to the corresponding portion of the nucleic acid present in the naturally occurring nucleic acid. Mutations include deletions, substitutions and / or additions of one or more nucleotide base pairs. In a specific aspect, the second group of nucleic acids may include sequences representing homologous genes and / or orthologous genes (orthologues) having at least 70% sequence identity with the corresponding portion of the nucleic acid present in the naturally occurring species. In another specific aspect, a sequence identity of at least 80% or at least 90% with the corresponding portion of the nucleic acid present in the naturally occurring species is provided.
[0104] Properties of the second group of nucleic acids
[0105] In a specific aspect, there are at least 100 different nucleic acid sequences in the second set of nucleic acids, preferably at least 1000, 2000 or 5000 different nucleic acid sequences, up to 100,000, 1,000,000 or 10,000,000 different nucleic acid sequences. Usually the number is 100 to 1,000,000, for example 1,000 to 100,000 different nucleic acid sequences. All or at least 90% or at least 50% or these nucleic acid sequences correspond to different chromosome interactions.
[0106] In a specific aspect, the second group of nucleic acids represents the chromosome interaction in at least 20 different loci or genes, preferably at least 40 different loci or genes, more preferably at least 100, at least 500, at least 1000 or at least 5000 different loci or genes, for example 100 to 10,000 different loci or genes. The length of the second group of nucleic acids makes them suitable for specific hybridization with the first group of nucleic acids according to Watson-Crick base pairing, so that the chromosome interaction of subgroup specificity can be identified. Usually, the second group of nucleic acids includes two parts corresponding to two chromosome regions gathered together in the chromosome interaction. The second group of nucleic acids usually includes a nucleic acid sequence of at least 10, preferably 20, and more preferably 30 bases (nucleotides) in length. On the other hand, the length of the nucleic acid sequence can be no more than 500, preferably no more than 100, and more preferably no more than 50 base pairs. In a preferred aspect, the second group of nucleic acids includes a nucleic acid sequence of 17 to 25 base pairs. On the one hand, at least 100%, 80% or 50% of the second group of nucleic acid sequences have above-mentioned length. Preferably, the different nucleic acids do not have any overlapping sequences, for example at least 100%, 90%, 80% or 50% of the nucleic acids do not have identical sequences over at least 5 consecutive nucleotides.
[0107] Assuming that the second set of nucleic acids acts as an "index", the same second set of nucleic acids can be used with different first sets of nucleic acids representing subsets of different traits, i.e., the second set of nucleic acids can represent a "universal" set of nucleic acids that can be used to identify chromosomal interactions associated with different traits.
[0108] The first group of nucleic acids
[0109] The first set of nucleic acids is usually from a subpopulation associated with prognosis. The first set of nucleic acids may have any of the characteristics and properties of the second set of nucleic acids mentioned herein. The first set of nucleic acids is usually from an individual sample, the individual having undergone the treatment and processing described herein, in particular EpiSwitch TM Cross-linking and cleavage steps. Typically, the first set of nucleic acids represents all or at least 80% or 50% of the chromosomal interactions present in a sample taken from the individual.
[0110] Typically, the first set of nucleic acids represents a smaller population of chromosomal interactions of the loci or genes represented by the second set of nucleic acids than the chromosomal interactions represented by the second set of nucleic acids, i.e., the second set of nucleic acids is a background or index set representing interactions in a defined locus or gene.
[0111] Nucleic acid library
[0112] Any type of nucleic acid population mentioned herein can be present in the form of a library, the library comprising at least 200, at least 500, at least 1000, at least 5000 or at least 10000 different nucleic acids of the type, such as a "first" or "second" nucleic acid. Such a library can be in an array-bound form. The library can include some or all of the probes or primer pairs shown in Table 5 or Table 6. The library can include all probe sequences from any table disclosed herein.
[0113] Hybridization
[0114] The present invention requires a means of hybridizing nucleic acid sequences that are complementary to all or part of a first group of nucleic acids and a second group of nucleic acids. In one aspect, all first groups of nucleic acids are contacted with all second groups of nucleic acids in a single analysis (i.e., in a single hybridization step). However, any suitable analysis may be used.
[0115] Labeled nucleic acids and hybridization patterns
[0116] The nucleic acid mentioned herein can be marked, preferably using an independent mark that helps to detect successful hybridization, such as a fluorophore (fluorescent molecule) or a radioactive label. Some marks can be detected under ultraviolet light. Hybridization pattern, such as on an array described herein, represents the difference of epigenetic chromosome interactions between two subgroups, and therefore provides a method for comparing epigenetic chromosome interactions and determining which epigenetic chromosome interactions are specific to the subgroup in the colony of the present invention.
[0117] The term "hybridization pattern" broadly encompasses the presence and absence of hybridization between a first group and a second group of nucleic acids, i.e., which specific nucleic acids in the first group hybridize to which specific nucleic acids in the second group, and therefore it is not limited to any particular assay or technique, or requires having a surface or array in which the "pattern" can be detected.
[0118] Selecting subpopulations with specific characteristics
[0119] The invention provides a kind of method, the method comprises detecting whether there is chromosome interaction, generally 5 to 20 or 5 to 500 such interactions, preferably 20 to 300 or 50 to 100 interactions, to determine whether there is the feature relevant to prognosis in individual.Preferably, chromosome interaction is those chromosome interactions in any gene mentioned herein.On the one hand, the chromosome interaction of typing is those represented by the nucleic acid in Table 5.On the other hand, the chromosome interaction is those represented in Table 6.On the other hand, the chromosome interaction of typing is those represented by the nucleic acid in Table 8.The column titled "Detection Ring" in the table shows the subgroup detected by each probe.Detection can be the presence of chromosome interaction in the subgroup or the absence of chromosome interaction, which is represented by "1" and "-1".
[0120] Tested individuals
[0121] Examples of species to which the individual tested is referred to herein. In addition, the individual tested in the method of the present invention may have been selected in some way. The individual may be susceptible to any of the conditions mentioned herein and / or may need any of the therapies mentioned herein. The individual may be receiving any of the therapies mentioned herein. In particular, the individual may suffer from or be suspected of suffering from prostate cancer or DLBCL. The individual may suffer from or be suspected of suffering from lymphoma.
[0122] Preferred prostate cancer gene regions, loci, genes and chromosome interactions
[0123] For all aspects of the invention, preferred gene regions, loci, genes and chromosome interactions are mentioned in the table, for example in Table 6. Typically in the methods of the invention, chromosome interactions are detected from at least 1, 2, 3, 4 or 5 relevant genes listed in Table 6. Preferably, the presence or absence of at least 1, 2, 3, 4 or 5 relevant specific chromosome interactions represented by the probe sequences in Table 6 is detected. The chromosome interaction may be upstream or downstream of any gene mentioned herein, for example 50 kb upstream or 20 kb downstream (e.g., from the coding sequence).
[0124] For all aspects of the invention, preferred gene regions, loci, genes and chromosome interactions are mentioned in Table 25. Typically in the methods of the invention, chromosome interactions are detected from at least 2, 4, 8, 10, 14 or all relevant genes listed in Table 25. Preferably, the presence or absence of at least 2, 4, 8, 10, 14 or all relevant specific chromosome interactions shown in Table 25 are detected. The chromosome interaction may be upstream or downstream of any gene mentioned herein, for example 50 kb upstream or 20 kb downstream (e.g., from the coding sequence).
[0125] In one embodiment, a combination of specific markers disclosed herein and represented (identified) by the following gene combination is typed: ETS1, MAP3K14, SLC22A3 and CASP2. This can determine the diagnosis. Preferably, at least 2 or 3 of these markers are typed.
[0126] In another embodiment, a combination of specific markers disclosed herein and represented (identified) by the following gene combination is typed: BMP6, ERG, MSR1, MUC1, ACAT1 and DAPK1. This can determine the prognosis (comparing high risk class 3 with low risk class 1 by nested PCR markers). Preferably, at least 2 or 3 of these markers are typed.
[0127] In a further embodiment, a combination of specific markers disclosed herein and represented (identified) by the following gene combination is typed: HSD3B2, VEGFC, APAF1, MUC1, ACAT1 and DAPK1. This can determine the prognosis (high risk class 3 vs. moderate risk class 2). Preferably, at least 2 or 3 of these markers are typed.
[0128] Preferred DLBCL gene regions, loci, genes and chromosome interactions
[0129] Typically, at least 10, 20, 30, 50 or 80 chromosome interactions are typed from any gene or region of the table disclosed herein or a portion of a table disclosed herein. Preferably, at least 10, 20, 30, 50 or 80 chromosome interactions are typed from any gene or region disclosed in Table 5.
[0130] Preferably, at least 2, 3, 5, or 8 markers in Table 7 are typed.
[0131] Preferably, the presence or absence of at least 10, 20, 30, 50 or 80 chromosome interactions represented by the probe sequences in Table 5 are detected. The chromosome interaction may be upstream or downstream of any gene mentioned herein, for example 50 kb upstream or 20 kb downstream (e.g., from the coding sequence).
[0132] Preferably, at least 1, 2, 5, 8 or all of the top 10 markers shown in Table 5 are typed. In one embodiment, at least 1, 2, 3 or 6 markers in Table 5 are typed, each of which corresponds to a different gene selected from STAT3, TNFRSF13B, ANXA11, MAP3K7, MEF2B and IFNAR1.
[0133] Preferred Lymphoma Gene Regions, Loci, Genes and Chromosome Interactions
[0134] Typically, at least 10, 20, 30 or 50 chromosome interactions are typed from any gene or region of a table disclosed herein or a portion of a table disclosed herein. Preferably, at least 10, 20, 30 or 50 chromosome interactions are typed from any gene or region disclosed in Table 8.
[0135] Preferably, at least 5, 10 or 15 markers in Table 9 are typed.
[0136] The chromosomal interaction may be upstream or downstream of any gene mentioned herein, for example 50 kb upstream or 20 kb downstream (eg, from the coding sequence).
[0137] In one embodiment, Figure 6 In another embodiment, at least one, two, three or six markers in Table 8 are typed, each of which corresponds to a different gene selected from STAT3, TNFRSF 13B, ANXA11, MAP3K7, MEF2B and IFNAR1.
[0138] Types of Chromosome Interactions
[0139] On the one hand, locus (including gene and / or position detected chromosome interaction) can include CTCF binding site.This is any sequence that can bind transcriptional repressor CTCF.The sequence can be composed of sequence CCCTC or include sequence CCCTC, wherein, sequence CCCTC can exist with 1, 2 or 3 copies in locus.CTCF binding site sequence can include sequence CCGCGNGGNGGCAG (IUPAC notation).CTCF binding site can be in at least 100, 500, 1000 or 4000 bases of chromosome interaction or in any chromosome region shown in table 5 or table 6.CTCF binding site can be in at least 100, 500, 1000 or 4000 bases of chromosome interaction or in any chromosome region shown in table 5 or table 6.
[0140] In one aspect, the chromosomal interaction detected is present in any of the gene regions shown in Table 5 or Table 6. If a junction nucleic acid is detected in the method, a sequence shown in any of the probe sequences in Table 5 or Table 6 may be detected.
[0141] Therefore, usually can detect the sequence from two regions of probe (i.e. from two sites of chromosome interaction). In preferred aspects, the probe used in this method comprises a sequence identical or complementary to the probe shown in any table or is composed of a sequence identical or complementary to the probe shown in any table. In some aspects, the probe used comprises a sequence homologous to any probe sequence shown in the table.
[0142] Tables provided in this article
[0143] Tables 5 and 6 show the probes (Episwitch TM Marker) data and gene data representing the chromosome interaction relevant to prognosis. The probe sequence shows the sequence of the connection product generated by two sites of the gene region gathered together in the chromosome interaction, that is, the probe will include a sequence complementary to the sequence in the connection product. The start-stop position of the first two groups shows the probe position, and the start-stop position of the latter two groups shows the relevant 4kb region. The following information is provided in the probe data table:
[0144] -HyperG_Stats: Find significant episwitches in loci based on hypergeometric enrichment parameters TM p-value of the probability of the number of markers
[0145] - Total probe count: EpiSwitch tested in the locus TM Total number of conformations
[0146] - Probe Count Sig: Discovery of statistically significant EpiSwitch in the locus TM The number of conformations
[0147] -FDR HyperG: Hypergeometric p-value corrected for multiple testing (immunoreactivity discovery rate)
[0148] -Percent Sig: Significant EpiSwitch TM Percentage of markers relative to the number of markers tested in the locus
[0149] -logFC: logarithm of epigenetic rate (FC) with base 2
[0150] -AveExpr: average log2 expression of the probe across all arrays and channels
[0151] -T: moderated t-statistic
[0152] -p-value: original p-value
[0153] -adj.p value: adjusted p value or q value
[0154] -BB statistic (lods or B) is the log-odds of differential gene expression.
[0155] -FC- non-log fold change
[0156] - FC_1 - non-log fold change centered around zero
[0157] -LS - binary value, related to the FC_1 value. FC_1 values below -1.1 are set to -1, and if the FC_1 value is above 1.1, it is set to 1. Between these values, the value is 0
[0158] Tables 5 and 6 show the genes found to have relevant chromosomal interactions. The p-values in the locus table are the same as those in HyperG Stats (based on the parameters of hypergeometric enrichment to find significant EpiSwitch in the locus). TM The LS column shows the presence or absence of an interaction associated with that particular subgroup (prognostic status).
[0159] In Table 5, DLBCL is a prognostic marker, represented by 1, and healthy is a healthy control, represented by -1.
[0160] The probe was designed to be 30 bp away from the Taq1 site. In the case of PCR, PCR primers are usually designed to detect ligation products, but their positions from the Taq1 site are different.
[0161] Probe location:
[0162] Start 1 - 30 bases upstream of the TaqI site on fragment 1
[0163] Termination 1 - TaqI restriction site on fragment 1
[0164] Start 2 - TaqI restriction site on fragment 2
[0165] Stop 2 - 30 bases downstream of the TaqI site on fragment 2
[0166] 4kb sequence position:
[0167] Start 1 - 4000 bases upstream of the TaqI site on fragment 1
[0168] Termination 1 - TaqI restriction site on fragment 1
[0169] Start 2 - TaqI restriction site on fragment 2
[0170] Stop 2 - 4000 bases downstream of the TaqI site on fragment 2
[0171] GLMNET values associated with the procedure used to fit the entire lasso or elastic net regularization with lambda set to 0.5 (elastic net).
[0172] In the tables in this article, aggressive prostate cancer subgroups refer to 3 types of patients with the following descriptions:
[0173] -PSA level is over 20 ng / ml, and
[0174] - Gleason score of 8 to 10, and
[0175] -T stage is T2c, T3, or T4
[0176] In the tables in this article, the indolent prostate cancer subgroup refers to patients in Category 1 who have the following description:
[0177] -PSA level less than 10 ng / mL, and
[0178] - Gleason score is 6 or less, and
[0179] -T stage is between T1 and T2a.
[0180] Table 7 shows preferred markers for DLBCL. Tables 8 and 9 show preferred markers for lymphoma.
[0181] Tables 5 to 7 are preferably used to type humans. Tables 8 and 9 are preferably used to type canines (e.g., dogs).
[0182] Methods for identifying markers and marker panels
[0183] The invention described herein relates to chromosome conformation maps and 3D structures that are tightly linked to phenotype as a means of regulation in their own right. Biomarker discovery is based on annotation of pattern identification and screening of representative cohorts of clinical samples representing phenotypic differences. We annotated and screened significant portions of the genome, spanning large swings of non-coding 5′ and 3′ of known genes as well as coding and non-coding portions, for identification of statistically consistent conditional chromosome conformations, such as non-coding sites anchored within open reading frames (introns) or outside of open reading frames.
[0184] In selecting the best markers, we were driven by the statistics and p-values used for marker bootstrapping. Reference to specific genes is for ease of positional reference - typically the closest gene is used as reference. The possibility cannot be excluded that the cis position of a gene and the chromosomal conformation of the associated neighborhood may have a specific regulatory component on the expression of a specific gene. Genes whose chromosomal conformation names are used as positional coordinate references during marker selection or validation do not require expression parameters. Regardless of the expression profile of the genes used in the reference, the chromosomal conformations selected and validated within the signature are themselves disseminating stratifying entities. Related regulatory patterns, such as SNPs at anchor sites, changes in gene transcriptional profiles, changes in H3K27ac levels, etc., need further investigation.
[0185] We are investigating the problem of clinical phenotypic differences and their classification from the basis of basic biology and epigenetic control of phenotypes - including, for example, from the framework of regulatory networks. Therefore, to assist in classification, changes in the network can be captured, and preferably done through the characteristics of several biomarkers, for example by following a machine learning algorithm to reduce the markers, which includes evaluating the optimal number of markers to classify the test cohort with minimal noise. This usually ends up with 3 to 17 markers, depending on the specific situation. Markers can be selected for the panel by cross-validation statistical performance (rather than, for example, by functional correlation of neighboring genes for reference names).
[0186] A panel of markers (with names of neighboring genes) was the product of cluster selection from a screening of important parts of the genome, analyzing in an unbiased manner the statistical dissemination power of more than 14,000 to 60,000 annotated EpiSwitch loci in important parts of the genome. For classification problems, it should not be considered as a tailored capture of chromosomal conformations of genes of known functional value. The total number of chromosomal interacting loci is 1.2 million, so the number of potential combinations is 1.2 million to the power of 1.2 million. Nevertheless, the approach we followed allows the identification of relevant chromosomal interactions.
[0187] The specific markers provided in this application have been selected to be statistically (significantly) associated with the disease. This is what the p-values in the relevant tables show. Each marker can be considered a representative of an epigenetic event in an organism, which manifests itself under the relevant disease as part of a network deregulation. In practice, this means that these markers are prevalent in each group of patients compared to the control group. On average, for example, individual markers can typically be present in 80% of the tested patients and 10% of the tested controls.
[0188] Simply adding all the markers does not represent some disordered network relationships. This is where the standard multivariate biomarker analysis GLMNET (R software package) is introduced. The GLMNET software package helps to identify the interdependence between certain markers, which reflects their joint role in achieving the disorder that causes the disease phenotype. Modeling and then testing the markers with the highest GLMNET scores can not only identify the minimum number of markers for accurately identifying the patient cohort, but also identify the minimum number of markers that provide the least false positive results in the control group patients, which should be attributed to the background statistical noise of the low incidence of the control group. Typically, a group (combination) of selected markers (e.g., 3 to 10) provides the best balance between detection sensitivity and specificity, appearing in the context of multivariate analysis of the individual attributes of all statistically significant markers selected for the disease.
[0189] The table in this article shows the reference names of the array probes (60-mer) used for array analysis, which overlap with the long-range interaction sites, chromosome numbers, and the junction between the start and end of two juxtaposed chromosome segments. The table also shows the standard array readout for each marker in competitive hybridization of disease samples with control samples (labeled with two different fluorescent colors). As a standard readout, it shows for each marker probe:
[0190] -Average expression signal
[0191] - Student's t-test for significant differences between fluorescent color detection of control and disease samples
[0192] - Significance p-value of the marker readout
[0193] – adjusted p-values (using Bonferroni correction for large datasets, B – background signal, FC – fold change of color detected in control samples
[0194] -FC_1-fold change of the second color assay in the case (disease or disease type) sample, LS (Loop Status)-prevalent fluorescent signal between the two color thresholds in competitive hybridization, where -1 means that the corresponding fluorescent color signal in the patient sample is blocked, when the probe on the CGH array is tested
[0195] - Directly inherited loci
[0196] - Total Probe Count - how many different position probes on the array were tested at this locus
[0197] - Probe Count Sig - how many probes are significant in distinguishing case and control samples
[0198] -Hypergeometric Stat is a statistic that shows the enrichment of loci with significant probes for disease detection
[0199] -FDR HyperG is the same statistic adjusted for large data sets based on FDR (standard procedure)
[0200] - The percentage of probes that became significant in that locus
[0201] -logFC is the logarithm of the fold change of the probe array reads. Focusing on loci with highly enriched significant probes helps select top probes representing regulatory hubs with multiple inputs relevant to the disease, thus providing markers with the best coverage (e.g., network dysregulation).
[0202] Optimal aspects of sample preparation and chromosome interaction detection
[0203] Methods for preparing samples and detecting chromosome conformation are described herein. Optimized (unconventional) versions of these methods may be used, such as described in this section.
[0204] Typically, the sample contains at least 2 × 10 5 cells. The sample can contain up to 5×10 5 In one aspect, the sample contains 2×10 5 Up to 5.5×10 5 Cells
[0205] The crosslinking of epigenetic chromosome interactions present in the chromosome locus is described herein. This can be performed before cell lysis occurs. Cell lysis can be performed for 3 minutes to 7 minutes, for example 4 minutes to 6 minutes or about 5 minutes. In some aspects, cell lysis is performed for at least 5 minutes and less than 10 minutes.
[0206] Digestion of DNA with restriction endonucleases is described herein. Typically, DNA restriction digestion is performed at about 55°C to about 70°C, such as at about 65°C, for about 10 minutes to 30 minutes, such as about 20 minutes.
[0207] Preferably, a frequently cutting restriction endonuclease is used that produces a ligated DNA fragment with an average fragment size of up to 4000 base pairs. Optionally, the restriction endonuclease produces a ligated DNA fragment with an average fragment size of about 200 to 300 base pairs (e.g., about 256 base pairs). In one aspect, the typical fragment size is 200 base pairs to 4,000 base pairs, such as 400 to 2,000 or 500 to 1,000 base pairs.
[0208] In one aspect of the EpiSwitch method, no DNA precipitation step is performed between the DNA restriction digestion step and the DNA ligation step.
[0209] DNA ligation is described herein. Typically DNA ligation is performed for 5 minutes to 30 minutes, such as about 10 minutes.
[0210] Proteins in the sample can be enzymatically digested, for example using a protease, optionally proteinase K. The proteins can be enzymatically digested for about 30 minutes to 1 hour, for example about 45 minutes. In one aspect, after protein digestion (e.g., proteinase K digestion), there is no cross-link reversal DNA extraction or phenol DNA extraction step.
[0211] In one aspect, the PCR assay is capable of detecting a single copy of the linked nucleic acid, preferably by a binary readout of the presence or absence of the linked nucleic acid.
[0212] Figure 5 A preferred method for detecting chromosome interactions is shown.
[0213] Methods and uses of the present invention
[0214] The method of the present invention can be described in different ways. It can be described as a method for preparing a linked nucleic acid, comprising (i) cross-linking chromosome regions that are brought together in chromosome interactions in vitro; (ii) cutting or restriction digesting the cross-linked DNA; and (iii) connecting the cross-linked cut DNA ends to form a linked nucleic acid, wherein the detection of the linked nucleic acid can be used to identify the chromosomal state of the locus, and wherein preferably:
[0215] - the locus may be any locus, region or gene mentioned in Table 5, and / or
[0216] - wherein the chromosome interaction may be any chromosome interaction mentioned herein or corresponds to any probe disclosed in Table 5, and / or
[0217] - wherein the ligation product may have or include (i) a sequence that is identical or homologous to any of the probe sequences disclosed in Table 5; or (ii) a sequence that is complementary to (ii).
[0218] The method of the present invention can be described as a method for detecting the status of chromosomes representing different subpopulations in a population, comprising determining whether there are chromosome interactions within a genomic defined epigenetic active region, wherein preferably:
[0219] - Subgroups defined by the presence or absence of a prognosis, and / or
[0220] - Chromosome status can be at any locus, region or gene mentioned in Table 5; and / or
[0221] - Chromosomal interactions may be any of those mentioned in Table 5 or correspond to any probe disclosed in that table.
[0222] The method of the present invention can be described as a method for preparing a linked nucleic acid, comprising (i) cross-linking in vitro chromosomal regions that are brought together in chromosomal interactions; (ii) cleaving or restriction digesting the cross-linked DNA; (iii) ligating the cross-linked cleaved DNA ends to form a linked nucleic acid, wherein detection of the linked nucleic acid can be used to determine the chromosomal status of the locus, and wherein preferably:
[0223] - the locus may be any locus, region or gene mentioned in Table 6, and / or
[0224] - wherein the chromosome interaction may be any chromosome interaction mentioned herein or corresponds to any probe disclosed in Table 6, and / or
[0225] - wherein the ligation product may have or include (i) a sequence identical or homologous to any probe sequence disclosed in Table 6; or (ii) a sequence complementary to (ii).
[0226] The method of the present invention can be described as a method for detecting the status of chromosomes representing different subpopulations in a population, comprising determining whether there are chromosome interactions within a genomic defined epigenetic active region, wherein preferably:
[0227] - Subgroups defined by the presence or absence of a prognosis, and / or
[0228] - Chromosome status can be at any locus, region or gene mentioned in Table 6; and / or
[0229] - Chromosomal interactions may be any one mentioned in Table 6 or correspond to any probe disclosed in that table.
[0230] The present invention includes detecting the chromosome interaction of any locus, gene or region mentioned in Table 5. The present invention includes using the nucleic acids and probes mentioned herein to detect chromosome interaction, for example using at least 1, 5, 10, 20 or 50 such nucleic acids or probes to detect chromosome interaction. Preferably, the nucleic acid or probe detects chromosome interaction in at least 1, 5, 10, 20 or 50 different loci or genes. The present invention includes using any primer or primer pair listed in Table 5 or using variants of these primers described herein (including primer sequences or sequences including fragments of primer sequences and / or homologous sequences) to detect chromosome interaction.
[0231] The present invention includes detecting chromosome interactions of any locus, gene or region mentioned in Table 6. The present invention includes using the nucleic acids and probes mentioned herein to detect chromosome interactions. The present invention includes using any primer or primer pair listed in Table 6 or using variants of these primers described herein (sequences comprising primer sequences or fragments of primer sequences and / or homologous sequences) to detect chromosome interactions.
[0232] When analyzing whether a chromosomal interaction occurs "within" a defined gene, region, or location, either both chromosome portions that come together in the interaction are within the defined gene, region, or location, or in some aspect only a portion of the chromosome is within the defined gene, region, or location.
[0233] Similarly, the chromosome interactions of Table 8 and Table 9 can be used in the processes and methods of the present invention.
[0234] Use of the method of the invention in identifying new therapeutic methods
[0235] Knowledge of chromosome interactions can be used to identify new treatments for disease.The invention provides methods and uses of chromosome interactions as defined herein to identify or design new therapeutic agents, for example therapeutic agents relevant to the treatment of prostate cancer or DLBCL.
[0236] Homologs
[0237] This article relates to the homologue of polynucleotide / nucleic acid (such as DNA) sequence.Such homologue usually has at least 70% homology, preferably at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98% or at least 99% homology, for example, in the region of at least 10, 15, 20, 30, 100 or more continuous nucleotides, or spans over the nucleic acid part from the chromosome region participating in chromosome interaction.Homology can be calculated based on nucleotide consistency (sometimes referred to as "hard homology").
[0238] Therefore, in specific aspects, the homologue of polynucleotide / nucleic acid (e.g., DNA) sequence refers to the sequence identity percentage in this article. Usually, such homologues have at least 70% sequence identity, preferably at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98% or at least 99% sequence identity, for example, in the region of at least 10, 15, 20, 30, 100 or more continuous nucleotides, or across the nucleic acid portion from the chromosome region participating in the chromosome interaction.
[0239] For example, the UWGCG software package provides the BESTFIT program, which can be used to calculate homology and / or sequence identity % (e.g., used under its default settings) (Devereux et al (1984) Nucleic Acids Research 12, p387-395). The PILEUP and BLAST algorithms can be used to calculate homology and / or sequence identity % and / or align sequences (e.g., identify equivalent or corresponding sequences (usually under their default settings)), such as Altschul SF (1993) J MoI Evol 36: 290-300; Altschul, S, F et al (1990) J MoI Biol 215: 403-10.
[0240] Software for performing BLAST analysis is publicly available through the National Center for Biotechnology Information. The algorithm involves first identifying high scoring sequence pairs (HSPs) by identifying short words of length W in the query sequence that match or satisfy a certain positive threshold score T when aligned with a fragment of the same length in the database sequence. T is called the neighborhood fragment score threshold (Altschul et al, supra). These initial neighborhood fragment matches (hits) serve as seeds to initiate searches to find HSPs containing them. Matched fragments (word hits) are extended in both directions along each sequence until the cumulative alignment score can be increased. Extension of the matched fragments in each direction will stop when: the cumulative alignment score drops by an amount X from its maximum achieved value; the cumulative score becomes zero or lower due to the accumulation of one or more negative score residue alignments; or the end of either sequence is reached. The BLAST algorithm parameters W5 T and X determine the sensitivity and speed of the alignment. The BLAST program uses as defaults a wordlength (W) of 11, the BLOSUM62 scoring matrix (see Henikoff and Henikoff (1992) Proc. Natl. Acad. Sci. USA 89: 10915-10919), an alignment (B) of 50, an expectation (E) of 10, M=5, N=4, and a comparison of both strands.
[0241] The BLAST algorithm performs a statistical analysis of the similarity between two sequences; see, e.g., Karlin and Altschul (1993) Proc. Natl. Acad. Sci. USA 90: 5873-5787. One measure of similarity provided by the BLAST algorithm is the smallest sum probability (P(N)), which provides an indication of the probability that a match between two polynucleotide sequences would occur by chance. For example, a sequence is considered similar to another sequence if the smallest sum probability of the first sequence compared to the second sequence is less than about 1, preferably less than about 0.1, more preferably less than about 0.01, and most preferably less than about 0.001.
[0242] Homologous sequences usually differ by 1, 2, 3, 4 or more bases, for example less than 10, 15 or 20 bases (which may be substitutions, deletions or insertions of nucleotides). These changes can be measured in any of the above regions relevant to calculating homology and / or sequence identity %.
[0243] The homology of a "primer pair" can be calculated, for example, by treating the two sequences as a single sequence (as if the two sequences were joined together) and then comparing to another primer pair that is also treated as a single sequence.
[0244] Array
[0245] The second group of nucleic acid can be attached to the array, and in one aspect, at least 15,000, 45,000, 100,000 or 250,000 different second group of nucleic acids are attached to the array, preferably, it represents at least 300, 900, 2000 or 5000 loci. On the one hand, one or more or all different populations of the second group of nucleic acid are attached to more than one different region of the array, in fact repeating on the array to allow error detection. The array can be based on AgilentSurePrint G3 customized CGH microarray platform. The combination of the first group of nucleic acid and the array can be detected by a two-color system.
[0246] Therapeutic agent (e.g., a therapeutic agent selected based on individual profiling or a therapeutic agent selected based on a test according to the present invention)
[0247] Therapeutic agents are referred to herein. The invention provides such agents, such as those identified by the methods of the invention, for use in preventing or treating disease conditions in certain individuals. This may include administering a therapeutically effective amount of the agent to an individual in need thereof. The invention provides the use of the agent in the preparation of a medicament for preventing or treating a condition in certain individuals.
[0248] The formulation of the agent will depend on the properties of the agent. The agent will be provided in the form of a pharmaceutical composition comprising the agent and a pharmaceutically acceptable carrier or diluent. Suitable carriers and diluents include isotonic saline solutions, such as phosphate buffered saline. Typical oral dosage compositions include tablets, capsules, liquid solutions and liquid suspensions. The agent can be formulated for parenteral, intravenous, intramuscular, subcutaneous, transdermal or oral administration.
[0249] The dosage of the agent can be determined according to a variety of parameters, in particular according to the substance used; the age, weight and condition of the individual being treated; the route of administration; and the desired regimen. A physician will be able to determine the route of administration and dosage required for any particular agent. However, a suitable dosage may be 0.1 mg / kg body weight to 100 mg / kg body weight, for example 1 mg / kg body weight to 40 mg / kg body weight, for example taken 1 to 3 times a day.
[0250] The therapeutic agent can be any such agent disclosed herein, or can target any "target" disclosed herein, including any protein or gene disclosed in any table herein, including Table 5 or Table 6. It should be understood that any agent disclosed in combination should also be considered to be disclosed for separate administration.
[0251] Prostate cancer treatment
[0252] Prostate cancer treatment is recommended based on the stage of disease progression. Radiation therapy, hormone therapy, and chemotherapy are three common options used in prostate cancer treatment. Single treatments or combination treatments may be used.
[0253] Chemotherapy
[0254] Chemotherapy is often used to treat prostate cancer that has invaded other organs in the body (metastatic prostate cancer). Chemotherapy destroys cancer cells by interfering with the way they reproduce. Chemotherapy cannot cure prostate cancer, but it can control it and reduce symptoms so they have less impact on your daily life.
[0255] Radiation therapy
[0256] This type of therapy can be used to treat localized prostate cancer and locally advanced prostate cancer. Radiation therapy may also be used to slow the progression of metastatic prostate cancer and relieve symptoms. Men may receive hormone therapy before chemotherapy to increase the chance of successful treatment. Hormonal therapy may also be recommended after radiation therapy to reduce the chance of recurrence.
[0257] Hormone therapy
[0258] Hormone therapy is often used in conjunction with radiation therapy. Hormone therapy alone should not usually be used to treat localized prostate cancer in men who are in good health and willing to undergo surgery or radiation therapy. Hormone therapy may be used to slow the progression of advanced prostate cancer and relieve symptoms. Hormones control the growth of prostate cells. In particular, prostate cancer needs the hormone testosterone to grow. The goal of hormone therapy is to block the effects of testosterone by either stopping its production or preventing the patient's body from using it.
[0259] Other treatments available for prostate cancer
[0260] Radical prostatectomy
[0261] High-intensity focused ultrasound therapy
[0262] Cryotherapy
[0263] Brachytherapy
[0264] ·Monitoring waiting
[0265] Transurethral prostatectomy
[0266] Treatment of advanced prostate cancer
[0267] Steroids
[0268] DLBCL treatment
[0269] The following four therapies are used to treat DLBCL:
[0270] -Chemotherapy
[0271] -Radiotherapy
[0272] -Monoclonal Antibody Therapy
[0273] -Steroid therapy
[0274] Any of the above therapies may also be used to treat lymphoma.
[0275] The form of the substances mentioned in this article
[0276] Any substance mentioned herein, such as nucleic acid or therapeutic agent, can be in purified or isolated form. They can exist in a form different from that found in nature, for example, they can exist in combination with other substances not present in nature. Nucleic acids (including parts of sequences defined herein) can have sequences different from those found in nature, for example, there are at least 1, 2, 3, 4 or more nucleotide changes in the sequence, as described in the homology section. Nucleic acids can have heterologous sequences at the 5′ end or the 3′ end. Nucleic acids can be different from those found in nature chemically, for example, they can be modified in some way, but preferably still capable of Watson-Crick base pairing. Where appropriate, nucleic acids will be provided in double-stranded or single-stranded form. The present invention provides all specific nucleic acid sequences mentioned herein in single-stranded or double-stranded form, and therefore includes the complementary strands of any sequence disclosed.
[0277] The invention provides kits for implementing any of the methods of the invention, including detecting chromosome interactions associated with prognosis. Such kits may include specific binding agents capable of detecting relevant chromosome interactions, such as agents capable of detecting the connection nucleic acids generated by the methods of the invention. Preferred agents in the kit include probes capable of hybridizing to connection nucleic acids or primer pairs, such as probes capable of amplifying connection nucleic acids in PCR reactions as described herein.
[0278] The present invention provides a device capable of detecting relevant chromosome interactions. The device preferably includes any specific binding agent, probe or primer pair capable of detecting chromosome interactions, such as any such agent, probe or primer pair described herein.
[0279] Detection Methods
[0280] In one aspect, a probe that is detectable when activated during a PCR reaction is used to quantitatively detect junction sequences associated with chromosome interactions, wherein the junction sequences include sequences from two chromosome regions that are brought together in an epigenetic chromosome interaction, wherein the method includes contacting the junction sequences with the probe during a PCR reaction and detecting the degree of activation of the probe, and wherein the probe binds to the junction site. The method can generally use a dual-labeled fluorescent hydrolysis probe to detect specific interactions in a manner consistent with MIQE.
[0281] The probe is typically labeled with a detectable tag that has an inactive and an active state, so it can only be detected when activated. The degree of activation will be related to the degree of template (ligation product) present in the PCR reaction. Detection can be performed during all or part of the PCR reaction, such as at least 50% or 80% of the PCR cycle.
[0282] The probe can include a fluorophore covalently attached to one end of the oligonucleotide and a quencher attached to the other end of the nucleotide, so that the fluorescence of the fluorophore is quenched by the quencher. On the one hand, the fluorophore is attached to the 5' end of the oligonucleotide, and the quencher is covalently attached to the 3' end of the oligonucleotide. Fluorophores that can be used for the method of the present invention include FAM, TET, JOE, Yakima yellow, HEX, Cyanine 3, ATTO 550, TAMRA, ROX, Texas Red, Cyanine 3.5, LC610, LC 640, ATTO 647N, Cyanine 5, Cyanine 5.5 and ATTO 680. Quenchers that can be used with suitable fluorophores include TAM, BHQ1, DAB, Eclip, BHQ2 and BBQ650, optionally, wherein the fluorophore is selected from HEX, Texas Red and FAM. Preferred combinations of fluorophore and quencher include FAM with BHQ1, and Texas Red with BHQ2.
[0283] Use of probes in qPCR assays
[0284] Hydrolysis probe of the present invention is usually optimized with the negative control of concentration matching to temperature gradient.Preferably, single-step PCR reaction is optimized.More preferably, calculate standard curve.An advantage of using the specific probe that passes through the junction of the connection sequence is that the specificity to the connection sequence can be achieved without using the nested PCR method.The method described herein can accurately and precisely quantify low copy number targets.Before the temperature gradient is optimized, the target connection sequence can be purified, such as gel purified.The target connection sequence can be sequenced.Preferably, about 10ng, or 5ng to 15ng, or 10ng to 20ng, or 10ng to 50ng, or 10ng to 200ng template DNA is used to carry out PCR reaction.Forward and reverse primers are designed so that a primer is combined with the sequence of one of the chromosome regions represented in the connection DNA sequence, and another primer is combined with other chromosome regions represented in the connection DNA sequence, for example, by being complementary to the sequence.
[0285] Choice of ligated DNA targets
[0286] The invention includes selecting primers and probes for use in the PCR methods defined herein, including selecting primers based on their ability to bind and amplify junction sequences, and selecting probe sequences based on the properties of the target sequence to which the probe will bind, particularly the curvature of the target sequence.
[0287] Usually design / select probe and connection sequence are combined, and described connection sequence is the restriction fragment of parallel crossing restriction site.In one aspect of the present invention, for example, use the concrete algorithm quoted in this paper, calculate the predicted curvature of the possible connection sequence relevant to specific chromosome interaction.Curvature can be expressed as the degree of each spiral angle, and for example, each spiral angle is 10.5 °.Select the connection sequence for targeting, wherein the curvature tendency peak value score of connection sequence is at least 5 ° for each spiral angle, and usually each spiral angle is at least 10 °, 15 ° or 20 °, for example, each spiral angle is 5 ° to 20 °.Preferably, for at least 20, 50, 100, 200 or 400 bases in connection site upstream and / or downstream, for example, 20 to 400 bases, calculate the curvature tendency score of each spiral angle.Therefore, on the one hand, the target sequence in connection product has any one of these curvature levels.Also can select target sequence based on minimum thermodynamic structure free energy.
[0288] Specific aspects
[0289] In one aspect, only intrachromosomal interactions are typed / detected, while extrachromosomal interactions (between different chromosomes) are not typed / detected.
[0290] In specific aspects, certain chromosomal interactions are not typed, such as any specific interaction mentioned herein (eg, defined by any probe or primer pair mentioned herein). In certain aspects, chromosomal interactions are not typed in any gene mentioned herein.
[0291] The data provided herein indicate that these markers are "spreading" markers that can distinguish cases from non-cases of the relevant disease condition. Thus, when practicing the present invention, a skilled artisan will be able to determine to which subpopulation an individual belongs by detecting the interaction. In one embodiment, a detection threshold of at least 70% of the tested markers in a form associated with the relevant disease condition (either by absence or presence) can be used to determine whether an individual belongs to the relevant subpopulation.
[0292] Screening methods
[0293] The present invention provides a method for determining which chromosome interactions are associated with the chromosome state corresponding to a prognostic subgroup in a population, comprising contacting a first group of nucleic acids from a subgroup with different chromosome states with a second group of index nucleic acids, and allowing complementary sequences to hybridize, wherein the nucleic acids in the first group of nucleic acids and the second group of nucleic acids represent connection products, including sequences from two chromosome regions that are clustered together in the chromosome interaction, and the hybridization pattern between the first group of nucleic acids and the second group of nucleic acids allows determination of which chromosome interactions are specific to the prognostic subgroup. The subgroup can be any specific subgroup defined herein, such as a subgroup associated with a specific condition or treatment.
[0294] publication
[0295] The contents of all publications mentioned herein are incorporated into the present specification by reference and may be used to further define features relevant to the present invention.
[0296] Specific aspects
[0297] EpiSwitch TM The platform technology detects epigenetic regulatory signatures that characterize regulatory changes between normal and abnormal conditions at a locus. TM The platform identifies and monitors fundamental epigenetic levels of gene regulation associated with the regulatory higher-order structures of human chromosomes, also known as chromosome conformational features. Chromosome features are unique, primary steps in the gene deregulation cascade. They are advanced biomarkers that offer a range of unique advantages over biomarker platforms that utilize late-stage epigenetic and gene expression biomarkers, such as DNA methylation and RNA analysis.
[0298] EpiSwitch TM Array analysis
[0299] Custom EpiSwitch TM The array screening platform has four densities (15K, 45K, 100K and 250K) of unique chromosome conformations, and each chimeric fragment is repeated 4 times on the array, making the effective density 60K, 180K, 400K and 1 million, respectively.
[0300] Custom Designed EpiSwitch TM Array
[0301] EpiSwitchr at 15K M Arrays can screen the entire genome, including through EpiSwitch TM EpiSwitch is a biomarker discovery technology that interrogates approximately 300 loci. TMThe arrays are built on the Agilent SurePrint G3 custom CGH microarray platform; this technology offers four probe densities (60K, 180K, 400K, and 1 million). Densities for each array are reduced to 15K, 45K, 100K, and 250K, as each EpiSwitch TM Probes were present in quadruplicate, thus allowing statistical assessment of reproducibility. Querying each locus for potential EpiSwitch TM The average number of markers was 50, so the number of loci that could be investigated was 300, 900, 2000, and 5000.
[0302] EpiSwitch TM Custom array pipeline
[0303] EpiSwitch TM The array is a two-color system, in EpiSwitch TM After library generation, one set of samples was labeled with Cy5, and the other sample to be compared / analyzed (control) was labeled with Cy3. The array was scanned using an Agilent SureScan scanner, and the resulting features were extracted using Agilent Feature Extraction software. EpiSwitch was then used in R TM Data were processed using array processing scripts. Arrays were processed using the standard two-color package in Bioconductor in R: Limma*. Normalization of arrays was done using the normalisedWithinArrays function in Limma*, relative to the Agilent positive control and EpiSwitch on-chip. TM Positive controls were completed. Data were screened based on Agilent marker calls, Agilent control probes were removed, and technical replicate probes were averaged for analysis using Limma*. Probes were modeled based on the difference between the 2 comparison scenarios and then corrected using the false discovery rate. Probes with a coefficient of variation (CV) <= 30% and a p-value of <= -1.1 or => 1.1 and passing p <= 0.1 FDR were used for further screening. To further reduce the probe set, multifactorial analysis was performed using the FactorMineR package in R.
[0304] *Note: LIMMA is a linear model and empirical Bayes method for evaluating differential expression in microarray experiments. Limma is an R package for analyzing gene expression data from microarray or RNA-Seq.
[0305] The probe pool was initially selected for final selection based on the parameters of adjusted p-value, FC and CV<30% (arbitrary cutoff point). Further analysis and the final list were drawn based on the first two parameters only (adj. p-value; FC).
[0306] Winding line
[0307] Using EpiSwitch in R TM Analysis Kit for EpiSwitch TM Screening arrays are processed to select high-value EpiSwitch TM Translate markers to EpiSwitch TM On the PCR platform.
[0308] Step 1
[0309] Probes were selected based on the corrected p-value (false discovery rate, FDR), which is the product of the modified linear regression model. Probes with p-values <= 0.1 were selected, and then probes were further narrowed down based on their epigenetic ratio (ER), which had to be <= -1.1 or => 1.1 to be selected for further analysis. The last filter was the coefficient of variation (CV), which had to be <= 0.3.
[0310] Step 2
[0311] The top 40 markers in the statistical list were selected based on their ER for selection as markers for PCR conversion. The top 20 markers with the highest negative ER amounts and the top 20 markers with the highest positive ER amounts made up the list.
[0312] Step 3
[0313] The markers obtained in step 1, i.e., the statistically significant probes, formed the basis for enrichment analysis using hypergeometric enrichment (HE). This analysis allowed the markers to be narrowed down to the list of significant probes, which together with the markers from step 2 formed the basis for transformation to EpiSwitch. TM List of probes on the PCR platform.
[0314] Statistical probes were processed by HE to determine which genetic positions were enriched with statistically significant probes, indicating which genetic positions were central to epigenetic differences.
[0315] The most significantly enriched loci based on the corrected p-values were selected for generating probe lists. Genetic positions with p-values below 0.3 or 0.2 were selected. The statistical probes mapped to these genetic positions were combined with the markers from step 2 to form the EpiSwitch TM High-value markers for PCR.
[0316] Array design and processing
[0317] Array design
[0318] 1. Process the loci using SII software (currently v3.2) to:
[0319] a. Extract the genomic sequences of these specific loci (50 kb upstream and 20 kb downstream gene sequences)
[0320] b. Define the probability that the sequence in this region participates in CC
[0321] c. Use specific RE cutting sequence
[0322] d. Determine which restriction fragments are likely to interact in a certain orientation
[0323] e. Rank the likelihood of different CCs interacting together.
[0324] 2. Determine the array size, and thus the number of available probe positions (x).
[0325] 3. Extract x / 4 interactions.
[0326] 4. For each interaction, define a 30 bp sequence from part 1 to the restriction site and a 30 bp sequence from part 2 to the restriction site. Check if those regions are duplicates, if so, exclude and record the next interaction in the list. Connect the two 30 bp to define the probe.
[0327] 5. Create a list of x / 4 probes plus the defined control probes and replicate 4 times to create the list to be built on the array.
[0328] 6. Upload the probe list to the Agilent Sure design website for custom CGH arrays.
[0329] 7. Design of Agilent custom CGH arrays using probe sets.
[0330] Array processing
[0331] 1. Use EpiSwitch TM Standard operating procedures (SOPs) for processing samples for template production.
[0332] 2. Clean up by ethanol precipitation at the array processing laboratory.
[0333] 3. Samples were processed according to the Agilent SureTag Complete DNA Labeling Kit - Agilent oligonucleotide array-based CGH is used for enzymatic labeling of genomic DNA analysis of blood, cells or tissues.
[0334] 4. Scanning was performed using an Agilent C scanner using Agilent feature extraction software.
[0335] EpiSwitch TM The biomarker signatures demonstrate high robustness, sensitivity, and specificity in classifying complex disease phenotypes. The technology leverages recent breakthroughs in epigenetics, monitoring and assessing chromosome conformation signatures as a class of highly informative epigenetic biomarkers. Current research methods used in academic settings require biochemical processing of cellular material for CCS detection over a 3-day to 7-day period. These procedures have limited sensitivity and reproducibility; furthermore, these procedures do not have the same TM The benefits of analytical packages for targeted insights provided during the design phase.
[0336] EpiSwitch in Computer Biomarker Identification TM Array
[0337] CCS sites throughout the genome are mapped by EpiSwitch TM The array was evaluated directly on clinical samples from the test cohort to identify all relevant classification-guiding biomarkers. TM Array platforms were used for marker identification due to their high throughput and ability to rapidly screen large numbers of loci. The arrays used were Agilent custom CGH arrays, which allow interrogation of markers identified by computer software.
[0338] EpiSwitch TM PCR
[0339] Via EpiSwitch TM PCR or DNA sequencer (Roche 454, Nanopore MinION, etc.) TM The best PCR markers that were statistically significant and showed the best reproducibility were selected for further reduction to the final EpiSwitch TM EpiSwitch TM PCR can be performed by trained technicians following established SOP protocols. All protocols and reagent preparation are performed under ISO 13485 and 9001 certification to ensure quality of work and ability to transfer protocols. EpiSwitch TM PCR and EpiSwitch TMThe array biomarker platform is compatible with both whole blood and cell line assays. These tests are sensitive enough to detect very low copy number abnormalities using small amounts of blood.
[0340] Paragraphs showing embodiments of the invention
[0341] 1. A method for detecting a chromosome state representative of a subpopulation in a population, comprising determining whether there is a chromosome interaction associated with the chromosome state within a defined region of a genome; and
[0342] - wherein said chromosomal interactions have optionally been identified by a method of determining which chromosomal interactions are associated with a chromosomal state corresponding to said subpopulation of said population comprising the steps of contacting a first set of nucleic acids from a subpopulation of different chromosomes having a state with a second set of index nucleic acids and allowing complementary sequences to hybridize, wherein nucleic acids in said first set of nucleic acids and second set of nucleic acids represent ligation products comprising sequences from two chromosomal regions that have been brought together in a chromosomal interaction, and wherein the hybridization pattern between said first set of nucleic acids and second set of nucleic acids allows determination of which chromosomal interactions are specific to said subpopulation; and
[0343] - wherein the subpopulation is associated with prognosis of prostate cancer, and the chromosomes interact:
[0344] (i) present in any region or gene listed in Table 6; and / or
[0345] (ii) corresponds to any chromosomal interaction represented by any of the probes shown in Table 6, and / or
[0346] (iii) is present in a 4000 base region including or flanking (i) or (ii);
[0347] or
[0348] - wherein the subpopulation is associated with the prognosis of DLBCL and the chromosomes interact:
[0349] a) present in any region or gene listed in Table 5; and / or
[0350] b) corresponds to any of the chromosomal interactions represented by any of the probes shown in Table 5, and / or
[0351] c) is present in a 4000 base region including or flanking (a) or (b).
[0352] 2. A method according to paragraph 1, wherein:
[0353] - the prognosis of the prostate cancer is related to whether the cancer is aggressive or indolent;
[0354] and / or
[0355] - The prognosis of DLBCL is related to survival.
[0356] 3. The method of paragraph 1 or 2, wherein the subpopulation is associated with prostate cancer and a specific combination of chromosomal interactions is profiled:
[0357] (i) including all of the chromosome interactions represented by the probes in Table 6; and / or
[0358] (ii) comprising at least 1, 2, 3 or 4 of the chromosomal interactions represented by the probes in Table 6; and / or
[0359] (iii) the chromosomal interactions are present in at least one, two, three or four of the regions or genes listed in Table 6; and / or
[0360] (iv) wherein at least 1, 2, 3 or 4 of the chromosomal interactions are typed, and the typed chromosomal interactions are present in a 4,000 base region that includes or flanks the chromosomal interactions represented by the probes in Table 6.
[0361] 4. The method of paragraph 1 or 2, wherein the subpopulation is associated with DLBCL and is characterized by a combination of chromosomal interaction features:
[0362] (i) including all of the chromosome interactions represented by the probes described in Table 5; and / or
[0363] (ii) comprising at least 10, 20, 30, 50 or 80 of the chromosome interactions represented by the probes described in Table 5; and / or
[0364] (iii) the chromosomal interactions are present in at least 10, 20, 30 or 50 of the regions or genes listed in Table 5; and / or
[0365] (iv) wherein at least 10, 20, 30, 50 or 80 of said chromosomal interactions are typed, and said typed chromosomal interactions are present in a 4,000 base region, said 4,000 base region including or flanking said chromosomal interactions represented by said probes in Table 5.
[0366] 5. The method of paragraph 1 or 2, wherein the subpopulation is associated with DLBCL and is characterized by a combination of chromosomal interaction features:
[0367] (i) includes all of the chromosome interactions shown in Table 7; and / or
[0368] (ii) comprising at least 1, 2, 5 or 8 of the chromosome interactions shown in Table 7.
[0369] 6. A method according to any of the preceding paragraphs, wherein at least 10, 20, 30, 40 or 50 chromosome interactions are typed, and preferably at least 10 chromosome interactions are typed.
[0370] 7. The method according to any of the preceding paragraphs, wherein the chromosome interaction is typed:
[0371] - in samples from individuals, and / or
[0372] - by detecting the presence of a DNA loop at the chromosomal interaction site, and / or
[0373] - detecting the presence of chromosome terminal regions that are clustered together in chromosome conformation, and / or
[0374] - by detecting the presence of a junction nucleic acid generated during said typing process and the sequence of said junction nucleic acid comprises two regions, each region corresponding to a region of said chromosome that is brought together in said chromosome interaction, wherein said junction nucleic acid is preferably detected by the following method:
[0375] (i) in the case of prostate cancer prognosis, by a probe having at least 70% identity to any of the specific probe sequences mentioned in Table 6, and / or (ii) by a primer pair having at least 70% identity to any of the primer pairs in Table 6; or
[0376] (ii) in the case of DLBCL prognosis, by a probe having at least 70% identity to any of the specific probe sequences mentioned in Table 5, and / or (b) by a primer pair having at least 70% identity to any of the primer pairs in Table 5.
[0377] 8. A method according to any preceding paragraph, wherein:
[0378] - the second set of nucleic acids is derived from a larger population of individuals than the first set of nucleic acids; and / or
[0379] - the first set of nucleic acids is from at least 8 individuals; and / or
[0380] - the first set of nucleic acids is derived from at least 4 individuals from a first subpopulation and at least 4 individuals from a second subpopulation, the second subpopulation preferably not overlapping with the first subpopulation; and / or
[0381] - The method is performed in order to select an individual for medical treatment.
[0382] 9. A method according to any preceding paragraph, wherein:
[0383] - said second set of nucleic acids represents an unselected population; and / or
[0384] wherein said second set of nucleic acids is bound to the array at defined locations; and / or
[0385] - wherein the second set of nucleic acids represents chromosomal interactions in at least 100 different genes; and / or
[0386] - wherein the second set of nucleic acids comprises at least 1,000 different nucleic acids representing at least 1,000 different chromosome interactions; and / or
[0387] - wherein said first set of nucleic acids and said second set of nucleic acids comprise at least 100 nucleic acids having a length of 10 to 100 nucleotide bases.
[0388] 10. A method according to any of the preceding paragraphs, wherein the first set of nucleic acids can be obtained by a method comprising the following steps:-
[0389] (i) cross-linking of chromosomal regions that are brought together in chromosomal interactions;
[0390] (ii) cleaving the cross-linked regions, optionally by restriction digestion with an enzyme; and
[0391] (iii) ligating the cross-linked cleaved DNA ends to form the first set of nucleic acids (particularly comprising ligated DNA).
[0392] 11. The method of any preceding paragraph, wherein the defined region of the genome:
[0393] (i) including single nucleotide polymorphisms (SNPs); and / or
[0394] (ii) expression of microRNA (miRNA); and / or
[0395] (iii) expression of non-coding RNA (ncRNA); and / or
[0396] (iv) expressing a nucleic acid sequence encoding at least 10 consecutive amino acid residues; and / or
[0397] (v) expression control elements; and / or
[0398] (vii) Includes CTCF binding site.
[0399] 12. A method according to any of the preceding paragraphs, implemented to determine whether prostate cancer is aggressive or indolent, the method comprising typing at least 5 chromosome interactions defined in Table 6.
[0400] 13. The method according to any of the preceding paragraphs, implemented to determine the prognosis of DLBLC, comprising typing at least 5 chromosome interactions defined in Table 5.
[0401] 14. The method of any preceding paragraph, implemented to identify or design a therapeutic agent for prostate cancer;
[0402] - wherein preferably, the method is used to detect whether a candidate agent can cause a change in chromosome state associated with different prognostic levels;
[0403] - wherein said chromosomal interaction is represented by any probe in Table 6; and / or
[0404] - the chromosomal interaction is present in any of the regions or genes listed in Table 6;
[0405] and wherein optionally:
[0406] - identifying said chromosomal interaction by the method of determining which chromosomal interactions are associated with chromosomal states as defined in paragraph 1, and / or
[0407] - using (i) a probe having at least 70% identity to any of the probe sequences mentioned in Table 6, and / or (ii) a primer pair having at least 70% identity to any of the primer pairs in Table 6 to monitor changes in said chromosomal interactions.
[0408] 15. The method according to any of the preceding paragraphs 1 to 13, implemented to identify or design a therapeutic agent for DLBCL;
[0409] - wherein preferably, the method is used to detect whether a candidate agent can cause a change in chromosome state associated with different prognostic levels;
[0410] - wherein said chromosomal interaction is represented by any probe in Table 5; and / or
[0411] - the chromosomal interaction is present in any of the regions or genes listed in Table 5;
[0412] and wherein optionally:
[0413] - identifying said chromosomal interaction by the method of determining which chromosomal interactions are associated with chromosomal states as defined in paragraph 1, and / or
[0414] - using (i) a probe having at least 70% identity to any of the probe sequences mentioned in Table 5, and / or (ii) a primer pair having at least 70% identity to any of the primer pairs in Table 5 to monitor changes in said chromosomal interactions.
[0415] 16. A method according to paragraph 14 or 15, comprising selecting a target based on detection of the chromosomal interaction, and preferably screening modulators of the target to identify therapeutic agents for immunotherapy, wherein the target is optionally a protein.
[0416] 17. The method according to any one of paragraphs 1 to 16, wherein the typing or detection comprises specific detection of the ligation product by quantitative PCR (qPCR), wherein the quantitative PCR uses primers capable of amplifying the ligation product and a probe that binds to the ligation site during the PCR reaction, wherein the probe comprises a sequence complementary to a sequence of each of the chromosomal regions that are brought together in the chromosomal interaction, wherein preferably, the probe comprises:
[0417] An oligonucleotide that specifically binds to the ligation product, and / or
[0418] a fluorophore covalently linked to the 5' end of the oligonucleotide, and / or
[0419] a quencher covalently linked to the 3' end of the oligonucleotide, and
[0420] Optionally,
[0421] The fluorophore is selected from HEX, Texas Red and FAM; and / or
[0422] The probe comprises a nucleic acid sequence of 10 to 40 nucleotide bases in length, preferably a nucleic acid sequence of 20 to 30 nucleotide bases in length.
[0423] 18. A method according to any one of paragraphs 1 to 17, wherein:
[0424] - present the results of the described method in a report, and / or
[0425] - Using the results of the method to select a patient treatment plan and preferably a specific therapy for an individual.
[0426] 19. A therapeutic agent for use in a method of treating prostate cancer or DLBCL in an individual identified as being in need of the therapeutic agent by the method according to any one of paragraphs 1 to 13 and 17.
[0427] The present invention is illustrated by the following examples:
[0428] Example 1
[0429] Using EpiSwitch TM (Chromosome conformation characteristics) marker
[0430] We have been observing highly disseminated EpiSwitch TM The markers were highly consistent with primary and secondary affected tissues and obtained strong validation results. TM Biomarker signatures have shown high robustness, sensitivity, and specificity in the classification of complex disease phenotypes.
[0431] EpiSwitch TM The technology provides an efficient method for screening, early detection, companion diagnosis, monitoring and prognostic analysis of major diseases associated with abnormal gene and response gene expression. The main advantages of the OBD method are that it is non-invasive, rapid, and relies on targets based on highly stable DNA as part of chromosomal features rather than unstable protein / RNA molecules.
[0432] CCSs form a stable regulatory framework for epigenetic control and access to genetic information across the entire genome of a cell. Changes in CCSs reflect early changes in regulatory patterns and gene expression long before the consequences manifest as overt abnormalities. Simply put, CCSs are topological arrangements in which different remote regulatory parts of DNA are brought into close proximity to influence each other's function. These connections are not made randomly; they are highly regulated and are recognized as high-level regulatory mechanisms with significant biomarker classification capabilities.
[0433] Prognostic classification of prostate cancer
[0434] The markers were developed based on retrospective annotation of class I (low risk, indolent), class II (intermediate) and class III (severe high risk). These markers showed reliable classification of patients compared with healthy controls and also differentiated between the classes. The samples were obtained from the United Kingdom.
[0435] Identification of an EpiSwitch capable of distinguishing blood from prostate cancer patients and healthy controls TM Biomarkers
[0436] Custom EpiSwitch TMThe microarray study was initially used to identify and screen approximately 15,000 potential CCSs over 425 loci to distinguish between 8 prostate cancer (PCa) individuals and 8 control individuals. The most statistically significant markers were converted to a nested PCR analysis and screened in a larger sample set including 24 PCa samples and 25 healthy control samples. A classifier was developed using the first 5 CCSs converted from the microarray, which classified PCa samples and control samples with a sensitivity and specificity of 100% (95% CI -86.2% to 100%) and 100% (95% CI -86.7% to 100%), respectively.
[0437] Figure 1 A principal component analysis of the top 5 markers for the 49 samples of the development sample cohort is shown.
[0438] The diagnostic classifier was used to classify another blinded independent cohort of 24 PCa samples and 5 healthy control samples (n = 29) with an accuracy of 83%. EpiSwitch was evaluated using an additional sample cohort of 95 PCa samples and 97 control samples (n = 192). TM The prostate cancer assay was further developed. This in turn was validated in a blinded cohort of 20 samples (10 PCa, 10 controls). The validation results are shown in Table 1.
[0439] Table 1. Classification results of the blind sample cohort (n = 20)
[0440]
[0441] The latest project in the PCA program has developed an alternative PCR format for PCa diagnosis using real-time quantitative PCR (qPCR) based on hydrolysis probes. The performance of the 6-marker model is shown in Table 2.
[0442] Table 2. Performance of the 6-marker qPCR model
[0443]
[0444] Summarize
[0445] EpiSwitch developed in PCa diagnostic program TM Three independent blind validations of the PCa diagnostic signature, using samples from the United States and the United Kingdom at different disease stages, achieved a sensitivity and specificity of >80% for the diagnosis of prostate cancer. The prostate-specific antigen (PSA) blood test is the gold standard clinical assay for detecting PCa, which itself depends on a variety of other variables, with a typical sensitivity and specificity ranging from 32% to 68%. In addition, a parallel research direction led to the EpiSwitchTM Development of assays to assess prostate cancer prognosis to aid in clinical management and treatment selection for individual patients diagnosed with PCa.
[0446] Another custom EpiSwitch was made TM Microarray studies were performed to identify and screen approximately 15,000 potential CCSs over 426 loci to distinguish between 8 patients with aggressive prostate cancer (category 3) and 8 patients with indolent PCa (category 1), the classifications of which are described in the Appendix. The most statistically significant markers were converted to a nested PCR assay and screened in a larger cohort of 42 category 1, 25 category 2, and 19 category 3 PCa samples.
[0447] The top 6 statistically significant markers were used to develop a prognostic classifier to classify class 1 (low risk) and class 3 (high risk) PCa. The performance of the classifier in an independent sample cohort of 42 class 1 samples and 25 class 3 samples (n=27) is shown in Table 3.
[0448] Table 3. Performance of the 6-marker prognostic classifier (1 class vs 3 classes)
[0449]
[0450]
[0451] Another analysis found an additional six markers that classified between PCa class 2 and class 3. The two classifiers shared two markers, and each classifier had four unique markers.
[0452] Figure 2 A VENN comparison of two PCA prognostic classifiers is shown.
[0453] The performance of the 2-class and 3-class PCa classifiers is shown in Table 4.
[0454] Table 4. Performance of the 6-marker prognostic classifier (2 categories vs 3 categories) n = 44
[0455]
[0456] in conclusion
[0457] The development of diagnostic and prognostic biomarkers was achieved in multiple cohorts of clinical samples. All markers screened and selected were based on systematic, blood-based epigenetic changes monitored by chromosome conformational signatures in patients with different stages of prostate cancer (stages 1 to 3) versus healthy controls (diagnostic applications), and in patients with aggressive, high-risk class 3 prostate cancer versus patients with indolent, low-risk class 1 prostate cancer (prognostic applications), or patients with intermediate-risk class 2 prostate cancer.
[0458] Results of the development of the classification of PCa versus healthy controls in the test cohort and in a series of blinded validations showed up to >80% sensitivity and specificity. The classification of high-risk class 3 versus low-risk class 1 PCa showed up to 80% sensitivity and up to 92% specificity in cohorts of up to 67 samples, while the classification of high-risk class 3 versus intermediate-risk class 2 showed up to 84% sensitivity and up to 88% specificity in cohorts of up to 44 samples.
[0459] appendix
[0460] Low Risk - Category 1
[0461] Localized prostate cancer is classified as low risk if
[0462] PSA Level Less than 10 ng / ml, and
[0463] Gleason score No more than 6, and
[0464] T stage between T1 and T2a
[0465] Moderate Risk - Category 2
[0466] Localized prostate cancer is classified as intermediate risk if you have at least one of the following:
[0467] PSA level is 10ng / ml to 20ng / ml
[0468] Gleason score of 7
[0469] T Stage is T2b
[0470] High Risk - Category 3
[0471] Localized prostate cancer is classified as high risk if you have at least one of the following:
[0472] PSA level exceeds 20 ng / ml
[0473] Gleason score of 8 to 10
[0474] T stage is T2c, T3, or T4
[0475] If the cancer is stage T3 or T4, it means it has broken through the outer fibrous covering of the prostate (capsule), so it is classed as locally advanced prostate cancer.
[0476] Example 2. Identification of markers for DLBCL
[0477] Summarize
[0478] This is relevant to the identification of major groups of patients with poor and good prognosis for subsequent selection of treatment (i.e., R-CHOP). Biomarkers were developed on the basis of retrospective overall survival. Typically, patients are classified according to disease subtype such as ABC (poor prognosis) or GCB (better prognosis) using biopsy-based gene expression criteria (such as Nanostring or Fluidigm). However, not all patients can be classified as ABC or GCB (so-called type III, or unclassified patients). We identified biomarkers to classify survival prognosis at baseline, before treatment, regardless of ABC or GCB criteria classification.
[0479] Identification of markers
[0480] DLBCL shows clear differences in patient survival (poor vs good prognosis) and is divided into subtypes characterized by multiple molecular readouts. In current clinical practice, the various subtypes also have different treatments. For example, this includes combination chemotherapy with rituximab and CHOP combination. There are various approaches.
[0481] The molecular readout in current practice is based on gene expression profiling by arrays, performed in biological material obtained directly from biopsies. These include Nanostring and Fluidigm array-based tests for the extreme ABC and GCB types. ANC subtypes are generally associated with a poor prognosis. Not every patient can be classified as ABC or GCB, and many patients remain unclassified (or type III) in terms of the established gene expression profiles and any association with a poor survival prognosis. We have established systemic biomarkers that can directly classify patients into poor or good prognosis, regardless of other forms of transcriptional gene expression profiles.
[0482] First step: We used the Episwitch screening array to compare the epigenetic profiles of a panel of cell lines representing poor and good prognosis for survival of DLBCL. This allowed identification of array-based markers and design of nested PCR primers for the same target in a PCR format.
[0483] Step 2: We used the first 10 nested PCR-based markers to read baseline blood samples from 57–58 unclassified DLBCL patients with known retrospective survival annotations. Table 6 provides detailed information on the markers, final features, and claimed performance of the classifier model.
[0484] Our work shows how the baseline is judged as poor prognosis / good prognosis for these patients compared to the clinical survival data. This is a Cox estimate of the hazard ratio, i.e. our baseline classification as poor prognosis shows a higher probability of belonging to the poor prognosis survival group than the clinical post hoc annotated good prognosis group, with a specific value > 1. The latter is of particular value and interest to the clinical team in trial design.
[0485] Detailed description
[0486] Diffuse large B-cell lymphoma (DLBCL) is the most common type of non-Hodgkin lymphoma in adults. It can occur anytime between adolescence and old age and affects 7-8 people per 100,000 per year in the United States, although its incidence increases with age. Gene expression profiling has revealed two major types of DLBCL - germinal center B-cell-like (GCB) and activated B-cell-like (ABC). GCB DLBCL originates in secondary lymphoid organs such as lymph nodes, where naive B cells do not stop dividing after an infection has been cleared. ABC DLBCL is thought to begin in a subset of B cells, plasmablasts, that are ready to leave the germinal center and become plasma cells, but the reality is more complicated because different forms of DLBCL can occur throughout the B cell life cycle.
[0487] The prognosis of different subtypes varies, with a 5-year survival rate of 60% for GCB DLBCL and only 35% for ABC DLBCL. Each subtype is characterized by different gene expression. In GCB DLBCL, the transcriptional repressor BCL6 is often overexpressed, while in ABC DLBCL, the NF-KB pathway is often found to be constitutively activated. There is also a third type of DLBCL, called type III, which is not well understood, but its gene expression profile is thought to be between the two main types.
[0488] Current diagnostic approaches involve excisional biopsy of affected lymph nodes followed by immunohistochemistry (IHC). Currently, the treatment of DLBCL is the same regardless of the subtype. Because the pathogenesis, treatment response, and outcomes of the various subtypes vary greatly, there remains a need to develop a robust, non-invasive assay to distinguish subtypes to help develop differentiated treatment strategies. Although a large number of studies have been conducted to search for predictive and prognostic biomarkers for DLBCL, there is no consensus on a single test that can be used to distinguish subtypes.
[0489] Identification of an EpiSwitch that can distinguish different DLBCL subtypes in the blood of DLBCL patients TM Biomarkers
[0490] We use EpiSwitch TM The array platform looked at DLBCL cell lines and blood samples and identified biomarkers that were not present in healthy control patients, and then confirmed those biomarkers in a 70-patient cohort consisting of 30 ABC, 30 GCB, and 10 healthy control samples.
[0491] EpiSwitch TM Array
[0492] EpiSwitch TM Custom arrays allow screening of thousands of possible CCSs using probes designed using pattern recognition software. TM The different long-range chromosomal interactions captured by the technology reflect the epigenetic regulatory framework imposed on the loci of interest and correspond to individual different inputs from co-regulatory signaling pathways that contribute to these loci. In summary, the combination of different inputs regulates gene expression. Before integrating all input signals into a gene expression profile, the identification of abnormal or different chromosome conformational features under specific physiological conditions provides important evidence for the specific contribution to the dysregulation.
[0493] Using data from multiple sources, 98 loci were selected and analyzed using proprietary software, and 13,332 probes for potential chromosome conformations were tested. Looking at one locus is not the same as looking at one marker, as there may be one, multiple, or no markers for higher-order epigenetic chromosome conformation at a particular locus. Once prepared, cell lines and blood samples from DLBCL patients and healthy controls were processed, labeled, and hybridized to the array using the EpiSwitch approach.
[0494] Samples for diagnostic development
[0495] We used 16 cell lines corresponding to different subtypes with varying confidence in the subtyping. The most well-defined ABC and GCB subtype cell lines were used for the analysis. In addition, blood samples from 4 DLBCL patients and 11 healthy controls were used. After the first part of the biomarker identification, 60 further samples were provided to the OBD, consisting of 30 ABC and 30 GCB blood samples that were well characterized by the Fluidigm test and supplemented by 10 healthy control samples provided by the OBD.
[0496] result
[0497] Array analysis
[0498] 72 chromosomal signature loci from the microarray were selected for screening based on two criteria:
[0499] Their ability to sort between ABC and GCB cells (high ABC high GCB)
[0500] and / or
[0501] . Low CV values (median of 5 arrays analyzed, High ABC v High GCB, DLBCL1 v Healthy Control, DLBCL2 v Healthy Control, DLBCL3 v Healthy Control, and DLBCL4 v Healthy Control)
[0502] Array to EpiSwitch TM Transformation of PCR Platform
[0503] After analyzing the sequences surrounding the probes of interest on the array, 69 primer sets were designed to interrogate the chromosomal signature loci. Pooled DLBCL blood samples were then tested, and 49 of these sets met the OBD criteria for PCR products used for detection.
[0504] Each of these 49 potential markers was then tested on 6 DLBCL cell lines, 3 of which were ABC and 3 of which were GCB. Since the same classification was found using multiple different identification methods, the cell lines used were those that were most definitely ABC or GCB. This allowed the selection of markers that were most useful in distinguishing between ABC and GCB cell subtypes. 28 EpiSwitch TM Markers are used in PCR platforms. These markers are compatible with EpiSwitch TM The microarray results were consistent. In addition, potential markers were tested in four DLBCL patients and pooled healthy controls to identify markers that were present in DLBCL patients but not in healthy controls. TM Twenty-one of the markers were not present in healthy control samples but were present in DLBCL samples, so it can be used as a marker for DLBCL and can be used for subtyping.
[0505] Sample testing
[0506] The ability to translate well into EpiSwitch was then tested in a cohort of 70 patient blood samples. TM21 markers for the PCR platform. Initially, each marker was tested in 6 new ABC samples and 6 new GCB samples, and the 21 marker set was narrowed down to 10 markers that showed the greatest difference. These 10 markers were then tested in the remaining 24 ABC samples, 24 GCB samples, and 10 healthy control samples.
[0507] Each marker was then analyzed for its ability to distinguish subpopulations, colinearity with other markers, and ability to distinguish healthy from DLBCL. A subset of six markers was identified to provide the most information possible, located at the ANXA11 IFNAR, MAP3K7, MEF2B, NFATcl, and TNFRS13C loci. Figure 3 The ability of these markers to discriminate between different sample groups is shown in the PCA plot. This 6-marker panel is able to clearly distinguish between healthy controls and DLBCL patients, a key feature of any blood-based DLBCL test.
[0508] Figure 3 Shows 6 EpiSwitch based TM PCA plot of marker binary data for 60 DLBCL and 10 healthy patients. Samples were characterized as ABC subtype or GCB subtype by Fluidigm data, and healthy controls are also shown.
[0509] Classification: Identification of ABC and GCB subtypes in a DLBCL patient cohort (60 samples)
[0510] The classification was performed using a logistic regression classifier with 5-fold cross validation and the following results were obtained. The following results were obtained in the cross validation:
[0511] ABC subtype 83.3% (95% CI -65.3% to 94.3%)
[0512] GCB subtype 83.3% (95% CI -65.3% to 94.3%)
[0513] In addition, the resulting 6-marker logistic classifier model was tested on 50 permutations of the 60 patient data sets. The data was randomized each time, and the accuracy statistics were calculated using the ROC curve. The area under the curve (AUC) was 0.802, and the p value was 0.0000037 (H0 = "AUC equals 0.5"), indicating that the model is accurate and efficient.
[0514] in conclusion
[0515] In this study, we demonstrated that their EpiSwitch TMThe ability of the technology to provide answers to difficult clinical questions, specifically the differentiation of ABC and GCB subtypes of DLBCL. Using a high-throughput array approach and translating it to a simple and cost-effective PCR platform, over 13,000 potential CCSs have been tested and refined into a 6-marker panel for DLBCL subtype differentiation. The panel was able to differentiate DLBCL patients from healthy controls and accurately predicted the subtype 83.3%. For EpiSwitch TM The test also had over 80% agreement in class assignment between (whole blood based), LPS (cell origin, tissue), and Fluidigm (cell origin, tissue).
[0516] EpiSwitch TM The technology detects changes in long-range gene-gene interactions—characteristics of chromosome conformation—that lead to changes in epigenetic states and modulation of expression patterns of key genes in disease pathogenesis. TM The diagnostic procedure of the technology is a simple and rapid technique that can be transferred to other laboratories. The test consists of several molecular biology reactions followed by nested PCR detection. The test does not require complex procedures and can be performed in any laboratory capable of running PCR-based analyses.
[0517] Example 3
[0518] Further work was performed in dogs. One of the aims was to investigate markers that would aid in the initial diagnosis of suspected lymphoma to inform the clinical veterinarian of the need for a follow-up biopsy. In this study, the first 75 EpiSwitch microarray DLBCL markers (identified previously) were translated from the human genome construct (Grch37) to the current canine genome. A total of 38 canine samples (consisting of 19 patients with possible lymphoma and 19 matched control samples) were screened using all 75 DLBCL markers. To carry out this work, the following was performed:
[0519] - Based on 75 human DLBCL markers (associated with specific genes), homologous gene sequences in the dog genome (CanFam3.1) were identified from Biomart and loci were extracted.
[0520] -Run EpiSwitch TM Software to identify potential interactions among these loci
[0521] - Primer design software and other filters were added to reduce the list to 75 markers for investigation.
[0522] Work and results such as Figures 6 to 16and as shown in Tables 8 and 9.
[0523] Example 4. Further studies on prostate cancer
[0524] Current prostate cancer (PCa) diagnostic blood tests are unreliable for early disease diagnosis, resulting in a high number of unnecessary prostate biopsies in men with benign disease and false reassurance of negative biopsies in men with PCa. Predicting PCa risk is critical for making informed decisions on treatment options, as the five-year survival rate for the low-risk group exceeds 95%, and most men will benefit from less invasive treatments. Both three-dimensional genome structure and chromosome architecture change early during tumorigenesis in both tumor and circulating cells and can serve as disease biomarkers.
[0525] In this prospective study, we performed chromosome conformation screening for 14,241 chromosomal rings in the loci of 425 cancer-related genes in whole blood of newly diagnosed, untreated PCa patients (n=140) and non-cancer controls (n=96).
[0526] Our data showed that peripheral blood mononuclear cells (PBMCs) from PCa patients acquired specific chromosomal conformational changes in the loci of ETS1, MAP3K14, SLC22A3, and CASP2 genes. Blind testing on an independent validation cohort yielded a PCa test with 80% sensitivity and 80% specificity. Further analysis between PCa risk groups yielded a prognostic validation set consisting of BMP6, ERG, MSR1, MUC1, ACAT1, and DAPK1 genes for high-risk category 3 vs low-risk category 1, and HSD3B2, VEGFC, APAF1, MUC1, ACAT1, and DAPK1 genes for high-risk category 3 vs intermediate-risk category 2, which were highly similar to the conformation of primary prostate tumors. These sets achieved 80% sensitivity and 92% specificity for classifying high-risk category 3 vs low-risk category 1, and 84% sensitivity and 88% specificity for classifying high-risk category 3 vs intermediate-risk category 2.
[0527] Our results suggest that specific chromosomal conformations in the blood of PCa patients allow for high-sensitivity and high-specificity PCa diagnosis and prognosis. These conformations are shared between PBMCs and primary tumors. These epigenetic signatures may lead to the development of blood-based PCa diagnostic and prognostic tests.
[0528] introduction
[0529] In the Western world, prostate cancer (PCa) is currently the most commonly diagnosed non-skin cancer in men and the second leading cause of cancer-related death. Many men as young as 30 years old show evidence of histological PCa, most of which is microscopic and may never show clinical manifestations. For diagnosis and prognosis, prostate-specific antigen (PSA), invasive needle biopsy, Gleason score, and disease stage are used. In a large multicenter study involving 2,299 patients, a 12-site biopsy protocol outperformed all other protocols with an overall PCa detection rate of only 44.4%.
[0530] The only available PCa blood test that is widely used clinically involves measuring circulating levels of PSA (sensitivity of 21%, specificity of 91%), however, the size of the prostate, benign prostatic hyperplasia, and prostatitis may also increase PSA levels. At the current cutoff of 4.0 ng / ml, only 20% of PCa patients are detected. In early PCa, the specificity of the PSA test is insufficient to distinguish between early invasive cancer and latent, non-lethal tumors that may remain asymptomatic throughout a person's lifetime. In advanced PCa, PSA dynamics are used as a clinical surrogate endpoint for efficacy. However, while they do give an approximate prognosis, they lack specificity for individuals. A number of more specific blood tests for PCa detection have emerged, including the 4K blood test (AUC 0.8) and the PHI blood test (90% sensitivity, 17% specificity). PSA levels, disease stage, and Gleason score are used to determine the severity of PCa and to stratify patients into risk groups. To date, there is no prognostic blood test that can be used to distinguish between low-risk and high-risk PCa.
[0531] There are multiple genetic changes associated with PCa, including mutations in p53 (up to 64% of tumors), p21 (up to 55%), p73, and the MMAC1 / PTEN tumor suppressor genes, but these mutations do not explain all observed effects on gene regulation. Epigenetic mechanisms involving dynamic and multi-layered chromosome looping interactions are powerful regulators of gene expression. Chromosome conformation capture (3C) technology allows the documentation of these features. In this study, we used EpiSwitch TM The assay was used to screen, define and evaluate specific chromosomal conformations in the blood of PCa patients and to identify loci with potential as diagnostic and prognostic markers.
[0532] method
[0533] A total of 140 PCa patients and 96 controls were recruited in two cohorts. Cohort 1: Men attending urology clinics with a diagnosis of PCa (n=105) or without a diagnosis of PCa (n=77) were prospectively recruited from October 2010 to September 2013. Cohort 2: Patient samples obtained from the United States (19 controls and 35 PCa). After recruitment, a single blood sample (5 ml) was collected from PCa patients using current needle and blood collection methods and placed in a BD Plastic EDTA tubes. Blood samples were passively frozen and stored at -80°C until processing. Prostate tumor samples were obtained from previously recruited patients (n=5) who subsequently underwent radical prostatectomy. The clinical characteristics of the patients are shown in Table 17.
[0534] The primary endpoint of this study was to detect changes in chromosome conformation in PBMCs of PCa patients compared with controls. Therefore, all untreated PCa patients were eligible for this study, regardless of grade, stage, and PSA level. Patients who had previously received chemotherapy or had other cancers were excluded from this study. PCa diagnosis was determined according to clinical routine, and appropriate treatment was assigned to the patients. For prognostic studies (secondary endpoints), patients were stratified according to the relevant NCCN risk group (Table 10). No follow-up studies were performed.
[0535] Based on the preliminary findings in melanoma, an a priori power analysis was performed using the pwr.t.test function in the R package pwd. The test indicated that 15 patients per group should be sufficient to detect associations between variables (β = 5% probability of type II error, significance level; 95% power; 50% confidence interval and 40% standard deviation).
[0536] EpiSwitch TM The technology platform pairs high-resolution 3C results with regression analysis and machine learning algorithms to develop disease classifications. To select epigenetic biomarkers that can be diagnostic of cancer, samples from cancer patients were screened to identify statistically significant differences in conditional and stable genomic structural profiles compared to healthy (control) samples. The analysis is performed on whole blood samples by first fixing the chromatin with formaldehyde to capture binding within the chromatin. The fixed chromatin is then digested into fragments using the TaqI restriction endonuclease, and the DNA strands are then ligated, favoring cross-linked fragments. The cross-links are reversed using a PCR product previously developed by EpiSwitch. TM Polymerase chain reaction (PCR) was performed using primers created by EpiSwitch software. TM Used in blood samples in a three-step approach to identify, evaluate, and validate statistically significant differences in chromosome conformation between PCa patients and healthy controls ( Fig.17). For the first step, sequences from 425 manually screened PCa-associated genes (obtained from a public database (www.ensembl.org)) were used as templates for the computational probabilistic identification of regulatory signals involving chromatin interactions (Table 18). The customized CGH Agilent microarray (8x60k) platform was designed to test technical and biological replicates of 14,241 potential chromosome conformations in 425 loci. Eight PCa and eight control samples were competitively hybridized to the array, and the presence or absence of differences in each locus was defined by LIMMA linear modeling, followed by binary filtering and cluster analysis. This initially revealed 53 chromosomal interactions that were able to best distinguish PCa patients from controls ( Fig.17 ).
[0537] In the second evaluation phase, 53 biomarkers selected from the array analysis were converted to an EpiSwitch-based TM PCR detection probes were used for multiple rounds of biomarker evaluation. PCR primers were selected based on their ability to distinguish PCa from healthy controls (n=6 per group). The consistency of PCR products generated using nested primers was confirmed by direct sequencing. Therefore, after preliminary statistical analysis, the 53 biomarkers screened were narrowed down to 15, ultimately forming a 5-marker signature (Table 11). This selected chromosome conformation signature-biomarker set was then tested on a known cohort (n=49). In addition, based on the EpiSwitch array marker lead TM The 5-marker signature developed by PCR assessment was tested in an independent blinded validation cohort of 29 samples that were combined with 49 previously tested samples of known abundance (78 samples total). Principal component analysis was also used to determine abundance levels and identify potential outliers ( Fig.18 ).
[0538] In the final step, to further validate the chromosome conformation signature for informing PCa diagnosis, the 5-marker set was tested in a blinded, independent (n=20) blood sample cohort. The results were analyzed using Bayesian logistic models, p-value null hypothesis (Pr(N|z|) analysis, Fisher's exact P test, and Glmnet (Table 12). The sample cohort size in the 5-marker signature study was gradually increased to enable the selection of the best markers for distinguishing PCa samples from healthy controls. The cohort size was expanded to 95 PCa and 96 healthy control samples. Data analysis and presentation were performed according to CONSORT recommendations. All measurements were performed in a blinded manner. The STARD criteria were used to validate the analytical procedure. A similar three-step approach was used to identify prognostic markers (Table 13).
[0539] Sequence-specific oligonucleotides were designed around the selected loci to screen potential markers by nested PCR using Primer3. All PCR amplified samples were visualized by electrophoresis in LabChipGX using the LabChip DNA 1K Version2 kit (Perkin Elmer, Beaconsfield, UK), and internal DNA markers were loaded onto the DNA chip using fluorescent dyes according to the manufacturer's protocol. Fluorescence was detected by laser, and the readout of the electropherogram was converted into simulated bands on the gel image using the instrument software. The threshold we set for bands to be considered positive was 30 fluorescence units and above.
[0540] Primary tumor samples were taken from biopsies of selected patients (n=5). The crushed tissue samples were incubated in 0.125% collagenase at 37°C with gentle stirring for 30 minutes. The resuspended cells (250ul) were then centrifuged at 800g for 5 minutes in a fixed arm centrifuge at room temperature, the supernatant was removed, and the pellet was resuspended in phosphate buffered saline (PBS). Under a fixed analytical sensitivity range (dilution factor 1:2), the primary tumor and matched blood samples were analyzed for the presence of a 6-marker set for Class 3 vs Class 1 and Class 3 vs Class 2. When a matching PCR band of the correct size was detected, 1 point was assigned, and 0 points were assigned if no band was detected (Table 14).
[0541] As described in the methods, we used EpiSwitch TM A stepwise diagnostic biomarker discovery approach based on the technology. The custom CGH Agilent microarray (8x60k) platform was designed to test technical and biological replicates of 14,241 potential chromosomal conformations at 425 loci (Table 18) in 8 PCa and 8 control samples. Fig.17 The presence or absence of each locus was defined by LIMMA linear modeling, followed by binary filtering and cluster analysis. In the second evaluation phase, nested PCR was used on the 53 selected biomarkers, which were further reduced to 15 markers and finally to a 5-marker signature ( Fig.17This unique chromosome conformation disease classification signature of PCa includes chromosomal interactions of five genomic loci: ETS proto-oncogene 1, transcription factor (ETS1), mitogen-activated protein kinase 14 (MAP3K14), solute carrier family 22 member 3 (SLC22A3), and caspase 2 (CASP2) (Table 11). The genomic locations of specific chromosomal loops in the ETS1, MAP3K14, SLC22A3, and CASP2 genes in the chromosome conformation signature (Table 11) were mapped to their associated chromosomes. Two genomic loci corresponding to the junction points of each chromosome conformation signature locus of ETS1, MAP3K14, SLC22A3, and CASP2 genes were mapped to chromosome 11, from 128,260,682 to 128,537,926; chromosome 17, from 43,303,603 to 43,432,282; chromosome 6, from 160,744,233 to 160,944,757, and chromosome 7, from 142,935,233 to 143,008,163. Circos plots showing chromosome rings for ETS1, MAP3K14, SLC22A3, and CASP2 chromosome conformation signature markers were generated.
[0542] A 5-marker principal component analysis was used to determine abundance levels and identify potential outliers. This analysis was applied to 78 samples comprising two groups. In the first group, 49 known samples (24 PCa and 25 healthy controls) were combined with a second group of 29 samples (including 24 PCa samples and 5 healthy control samples) ( Fig.18 ). The final training set was constructed using 95 PCa and 96 control samples and then tested in an independent blinded validation cohort of 20 samples (10 controls and 10 PCa). The sensitivity and specificity of detecting PCa using chromosomal interactions in the five genomic loci were 80% (CI 44.39% to 97.48%) and 80% (CI 44.39% to 97.48%), respectively (Table 12).
[0543] To select epigenetic biomarkers that could classify PCa, samples from PCa patients classified into risk groups 1-3 (low, intermediate, and high, respectively, Table 10) were screened to identify statistically significant differences in conditional and stable genomic profiles. TM Used on blood samples in a three-step approach to identify, evaluate and validate statistically significant differences in chromosome conformation between PCa patients at different stages of disease ( Fig.17). In the first step, the array used covered 425 loci and had test probes for a total of 14,241 potential chromosome conformations. Patients in the high-risk PCa3 category were compared with patients in the low-risk 1 category or the intermediate-risk 2 category. Through enrichment statistics, a total of 181 potential classification marker leads for PCR evaluation were identified (Table 19). The top 70 best markers were then taken to the next stage of PCR testing to further evaluate the classification of high-risk category 3 vs low-risk category 1 patient samples, and finally a 6-marker set of high-risk category 3 vs low-risk category 1 was established (Table 13). The chi-square test was used to identify the best markers, and then a classifier was established on the test sets of category 1 (n=21) and category 3 (n=19). Independent cohorts of category 1 (n=21) and category 3 (n=6) that were not used for any marker reduction were then used for the first round of blind validation. Similarly, the 6-marker set was evaluated for high risk class 3 vs moderate risk class 2 on the test sets of class 3 and class 2, which included 25 and 19 samples, respectively. Independent cohorts of class 2 and class 3 (n=6 each) that were not used for any marker reduction were then used for the first round of blinded validation.
[0544] In the final step, to further validate the chromosome conformation signature for informing PCa prognosis, the 6-marker set for high-risk category 3 vs low-risk category 1 was tested in a larger, more representative cohort. The initial blind cohort was expanded to 67 samples, including 40 samples for marker reduction (Table 15). Similarly, the 6-marker set for high-risk category 3 vs moderate-risk category 2 was tested in a larger, more representative cohort. The initial blind cohort was expanded to 43 samples (Table 16).
[0545] A 6-marker set for class 3 vs class 1 was established. The set contained bone morphogenetic protein 6 (BMP6), ETS transcription factor ERG (ERG), macrophage scavenger receptor 1 (MSR1), mucin 1 (MUC1), acetyl-CoA acetyltransferase 1 (ACAT1), and death-associated protein kinase 1 (DAPK1) genes (Table 13). A 6-biomarker for high-risk class 3 vs moderate-risk class 2 was identified, including hydroxy-delta-5-steroid dehydrogenase, 3β- and steroid delta-isomerase 2 (HSD3B2), vascular endothelial growth factor C (VEGFC), apoptotic peptidase activating factor 1 (APAF1), MUC1, ACAT1, and DAPK1. It is noteworthy that the last three biomarkers (MUC1, ACAT1, and DAPK1) were common between class 1 vs class 3 and class 3 vs class 2 (Table 13). The classification of high-risk class 3 vs low-risk class 1 PCa using chromosomal interactions in 6 genomic loci showed 80% sensitivity (CI 59.30% to 93.17%) and 92% specificity (CI 80.52% to 98.50%) in a blind cohort of 67 samples (Table 15). Similarly, a 6-marker set of high-risk class 3 vs moderate-risk class 2 was tested in a larger, more representative cohort of 43 samples, showing a sensitivity of 84% (CI 63.92% to 95.46%) and a specificity of 88% (CI 65.29% to 98.62%) (Table 16).
[0546] Using five matched peripheral blood and primary tumor samples, we compared epigenetic markers identified in peripheral circulation (Table 13) with tumor tissue. Our results showed that many of the dysregulated markers detected in blood as part of the Class 1 vs Class 3 and Class 2 vs Class 3 classification signatures could be detected in tumor tissue (Table 14). This suggests that chromosomal interactions that can be systematically detected can be detected at the primary site of tumorigenesis under the same conditions.
[0547] Timely diagnosis of prostate cancer is crucial to reduce mortality. The European Randomized Study of Screening for PCa showed a significant reduction in PCa mortality in men who underwent routine PSA screening. However, universal screening leads to overdiagnosis of clinically insignificant disease, so new less invasive tests that can distinguish low-risk from high-risk disease are urgently needed.
[0548] Our epigenetic profiling approach offers a potentially powerful means to address this need. The binary nature of the test (chromosome ring presence or absence) and the enormous combinatorial power (potentially >10 10The combination of approximately 50,000 rings (screening ~50,000 rings) may allow the creation of signatures that accurately meet clinically well-defined criteria. In PCa, this would distinguish low-risk vs high-risk disease, or identify small but aggressive tumors and determine the most appropriate treatment. In addition, epigenetic changes are known to manifest early in tumorigenesis, making them useful for diagnosis and prognosis.
[0549] In this study, we identified and validated chromosome conformation as a unique biomarker for a non-invasive, blood-based epigenetic signature of PCa. Our data indicate the presence of stable chromatin loops in the loci of the ETS1, MAP3K14, SLC22A3, and CASP2 genes present only in PCa patients (Table 11). Validation of these markers in an independent set of 20 blind samples showed 80% sensitivity and 80% specificity (Table 12), which is significant for a PCa blood test. Interestingly, the expression of some of these genes has been associated with cancer pathophysiology. ETS1 is a member of the ETS family of transcription factors. Prostate tumors in which ETS1 is overexpressed are associated with increased cell migration, invasion, and induction of epithelial-mesenchymal transition. MAP3K14 (also known as nuclear factor-κ-β (NF-κβ) induced kinase (NIK)) is a member of the MAP3K group (or MEKK). Physiologically, MAP3K14 / NIK can activate non-canonical NF-κβ signaling and induce canonical NF-κβ signaling, especially when MAP3K14 / NIK is overexpressed. A novel role for MAP3K14 / NIK in regulating mitochondrial dynamics to promote tumor cell invasion has been described. SLC22A3 (also known as organic cation transporter 3 (OCT3)) is a member of the SLC group of membrane transporters. The expression of SLC22A3 is associated with PCa progression. CASP2 is a member of the caspase activation and recruitment domain group. Physiologically, CASP2 can act as an endogenous repressor of autophagy. Two identified genes (SLC22A3 and CASP2) have been previously shown to be negatively correlated with cancer progression. Importantly, the presence of chromatin loops may have uncertain effects on gene expression.
[0550] To screen for PCa prognostic markers, we used EpiSwitch TMCustom arrays were used to analyze competitive hybridization of peripheral blood DNA from low-risk PCa (class 1) and high-risk PCa (class 3) patients. A 6-marker set of high-risk class 3 vs low-risk class 1 was identified, including BMP6, ERG, MSR1, MUC1, ACAT1, and DAPK1. Six biomarkers of high-risk class 3 vs moderate-risk class 2 were identified, including HSD3B2, VEGFC, APAF1, MUC1, ACAT1, and DAPK1. Three of these biomarkers (MUC1, ACAT1, and DAPK1) are shared between these sets. Our data show that the chromosome conformations in the blood of primary tumors and matched PCa patients in stages 1 and 3 are highly consistent (Table 14). The prognostic significance and diagnostic value of some of these genes have been proposed previously. BMP6 plays an important role in PCa bone metastasis. In addition to ETS1, ERG is another member of the ETS transcription factor family. Overwhelming evidence indicates that ERGs are involved in several processes associated with PCa progression, including metastasis, epithelial-mesenchymal transition, epigenetic reprogramming, and inflammation. MSR1 may confer an intermediate risk for PCa. MUC1 is a membrane-bound glycoprotein that belongs to the mucin family. High expression of MUC1 in advanced PCa is associated with adverse clinicopathological tumor features and poor outcome. ACAT1 expression is elevated in high-grade and advanced PCa and serves as an indicator of reduced biochemical recurrence-free survival. DAPK1 can act as both a tumor suppressor and an oncogenic molecule in different cellular contexts. HSD3B2 plays a crucial role in steroid hormone biosynthesis and is upregulated in relevant parts of PCa characterized by adverse tumor phenotypes, increased androgen receptor signaling, and early biochemical recurrence. VEGFC is a member of the VEGF family and its increased expression is associated with lymph node metastasis in PCa specimens. In a comprehensive biochemical approach, APAF1 has been described as the core of the apoptosome.
[0551] Despite the identification of these loci, the mechanisms of cancer-associated epigenetic changes in PBMCs have not yet been determined. However, interactions can be systematically tested and can be detected at the primary site of tumorigenesis under the same conditions (Table 14). Therefore, in order for us to measure these changes, the chromatin conformation in PBMCs must be guided by external factors; presumably substances produced by PCa tumor cells. It is well known that a large part of chromosome conformation is controlled by non-coding RNAs, which regulate tumor-specific conformations. Tumor cells have been shown to secrete non-coding RNAs that are endocytosed by neighboring cells or circulating cells and may change their chromosome conformation, in which case the RNA may be a regulator. Although RNA detection as a biomarker remains extremely challenging (low stability, background drift, continuous basis for statistical classification analysis), chromosome conformation features provide a recognized stable binary advantage for biomarker targeted use, especially when tested in the nucleus, because circulating DNA present in plasma does not retain the 3D conformational topology in intact cell nuclei. It is important to note that looking at one genetic locus is not equivalent to looking at one marker, as there may be multiple chromosomal conformations representing parallel pathways of epigenetic regulation of the locus of interest.
[0552] One of the main challenges of PCa diagnosis in current clinical practice is the time required to make a definitive diagnosis. Until now, there is no single, definitive test for PCa. A high level of PSA will send the patient on a long journey of uncertainty, where he will undergo an MRI scan and, if necessary, a biopsy. Although a biopsy is more reliable than a PSA test, it is a major surgery and missing cancerous lesions remains a problem. The five-biomarker panel described in this article is based on a relatively inexpensive and well-established molecular biology technique (PCR). The samples are based on biological fluids, which are easy to collect and provide clinicians with a rapidly available clinical readout within a few hours. This in turn saves a lot of time and cost and facilitates informed diagnostic decisions, thus filling a gap in the current definitive diagnostic protocol for PCa.
[0553] Predicting the risk of PCa is critical to making informed decisions about treatment options. With a five-year survival rate of over 95% in the low-risk group, most men will benefit from less invasive treatments. Currently, PCa risk classification is based on a combined assessment of circulating PSA, tumor grade (from biopsy), and tumor stage (from imaging findings). The ability to obtain similar information using a simple blood test would significantly reduce costs and speed up the diagnostic process. Of particular importance in PCa treatment is the identification of the minority of tumors that initially appear low risk but subsequently progress to high risk. Therefore, this subset of individuals would benefit from faster and more aggressive interventions.
[0554] In conclusion, here we have identified a subset of chromosome conformations in patient PBMCs that are strongly indicative of the presence and prognosis of PCa. These features have great potential for the development of rapid diagnostic and prognostic blood tests for PCa and significantly exceed the specificity of the currently used PSA test. Preferred markers and combinations include
[0555] -ETS1, MAP3K14, SLC22A3 and CASP2. This is diagnostic, by nested PCR markers
[0556] - BMP6, ERG, MSR1, MUC1, ACAT1, and DAPK1. This is a prognostic signature (high risk category 3 vs low risk category 1 by nested PCR markers)
[0557] -HSD3B2, VEGFC, APAF1, MUCl, ACAT1, and DAPK1. This is a prognostic signature (high risk category 3 vs intermediate risk category 2)
[0558] Example 5. Further work on DLBLC
[0559] Diffuse large B-cell lymphoma (DLBCL) is a heterogeneous blood cancer, but can be broadly divided into two major subtypes, germinal center B-cell-like (GCB) and activated B-cell-like (ABC). GCB and ABC subtypes have very different clinical courses, with ABC having a much worse survival prognosis. Patients with different subtypes have been observed to respond differently to therapeutic interventions, and in fact, some have suggested that ABC and GCB can be considered entirely different diseases. Due to this variability in response to treatment, assays to determine DLBCL subtypes are important in guiding clinical approaches with existing therapies as well as the development of new drugs. The current gold standard assay for DLBCL typing uses gene expression profiling of formalin-fixed paraffin-embedded (FFPE) tissue to determine the “cell of origin” and thus the disease subtype. However, this approach has some significant clinical limitations as it 1) requires a biopsy 2) requires complex, expensive, and time-consuming analytical methods, and 3) does not classify all patients with DLBCL.
[0560] Here, we took an epigenomic approach and developed a blood-based chromosome conformation signature (CCS) to identify DLBCL subtypes. Using clinical samples from 118 DLBCL patients, an iterative approach was used to define a panel of six markers (DLBCL-CCS) to subtype the disease. The performance of DLBCL-CCS was then compared with conventional gene expression profiling (GEX) from FFPE tissues.
[0561] DLBCL-CCS was accurate in classifying ABC and GCB in samples of known status, providing the same call in 100% (60 / 60) of samples in the discovery cohort used to develop the classifier. Additionally, in the evaluation cohort, DLBCL-CCS was able to make DLBCL subtype calls in 100% (58 / 58) of samples of the intermediate subtype (type III) defined by GEX analysis. Most importantly, when these patients were followed longitudinally throughout the course of their disease, EpiSwitch TM The relevant calls could be better tracked using the known survival patterns of the ABC and GCB subtypes.
[0562] This study provides an initial indication that a simple, accurate, cost-effective, and clinically feasible blood-based diagnostic method to identify DLBCL subtypes is possible.
[0563] background
[0564] Diffuse large B-cell lymphoma (DLBCL) is the most common type of blood cancer and is genetically and biologically heterogeneous, as evidenced by numerous studies using different approaches. The two major molecular subtypes of DLBCL are germinal center B-cell-like (GCB) and activated B-cell-like (ABC), although more refined molecular subtype definitions have also been proposed. These two major subtypes are of high clinical relevance as they have been observed to have significantly different disease courses, with the ABC subtype having a much worse survival prognosis. Perhaps more importantly, as novel investigational agents are evaluated in the clinical setting for the treatment of both GCB and ABC (or non-GCB) subtypes and historical observations of low overall response rates in unselected patients, there is a pressing need to define patient subtype prior to initiation of treatment. Historically, DLBCL subtypes have been defined by identification of the “cell of origin” (COO). The original COO classification was based on the observed similarity of DLBCL gene expression to activated peripheral blood B cells or normal germinal center B cells as determined by hierarchical cluster analysis (3). This COO classification by genome-wide expression profiling (GEP) classifies DLBCL into activated B-cell-like (ABC), germinal center B-cell-like (GCB), and type III (unclassified) subtypes, with ABC-DLBCL characterized by poor prognosis and constitutive NF-κB activation. In their seminal work, Wright et al. identified 27 genes that were most discriminatory in expression between ABC-DLBCL and GCB-DLBCL and developed a linear predictor score (LPS) algorithm for the COO classification. These original studies were based entirely on retrospective studies of fresh frozen (FF) lymphoma tissues. A major challenge in applying this COO classification in clinical practice is to establish a robust clinical analysis method that is applicable to routine formalin-fixed paraffin-embedded (FFPE) diagnostic biopsies. Several studies have also investigated the potential for COO classification of DLBCL using FFPE tissues by quantitatively measuring mRNA expression, including quantitative nuclease protection assays, GEP using the Affymetrix HGU133 Plus 2.0 platform or Illumina whole-genome DASLassay, and NanoString Lymphoma Subtyping Test (LST) technology. Several immunohistochemistry (IHC)-based algorithms have also been investigated to recapitulate COO classification by GEP. Overall, these studies have demonstrated high confidence in COO classification of DLBCL using FFPE tissues and robust separation of overall survival between ABC and GCB subtypes, but there are issues with reproducibility, especially lack of consistency between analyses. In addition, any IHC-based measurement requires baseline tissue, which is not always available, and the current long turnaround time from sample collection to analytical readout makes implementation in clinical practice challenging.
[0565] Among the methods historically used for DLBCL subtyping, one COO assessment method uses an assay that measures the expression of 27 genes in FFPE tissue by quantitative reverse transcription PCR (qRT-PCR) using the Fluidigm BioMark HD system. Although this approach has several advantages over existing technologies, it still faces several major barriers that limit its clinical application as it 1) requires a tissue biopsy and 2) relies on expensive, nonstandard, and time-consuming laboratory procedures. Therefore, blood-based assays could advance the field by providing a simple, reliable, and cost-effective method for DCBCL subtyping and enhancing clinical applicability.
[0566] In this study, we used a novel blood-based assay to identify COO classifications in patients with DLBCL by focusing on detecting changes in genomic structure. As part of the epigenetic regulatory framework, genomic regions can change their three-dimensional structure as a way to functionally regulate gene expression. The consequence of this regulatory mechanism is the formation of chromatin loops at different genomic loci. The presence or absence of these loops can be measured empirically using chromosome conformation capture (3C). Multiple genomic regions contribute to epistatic regulation by forming stable, conditional, long-range chromosomal interactions. Collective measurements of chromosome conformation at multiple genomic loci generate chromosome conformation signatures (CCS), or molecular barcodes that reflect the response of the genome to its external environment. For the detection, screening, and monitoring of CCS, we used the EpiSwitch platform, a well-established, high-resolution, and high-throughput method for detecting CCS. Based on 3C, the EpiSwitch platform has been developed to assess changes in chromatin structure at defined genetic loci as well as long-range noncoding cis- and trans-regulatory interactions. Advantages of using EpiSwitch for patient triage include its binary nature, reproducibility, relatively low cost, rapid turnaround time (samples can be processed within 24 hours), the need for only a small amount of blood (~50 mL), and compliance with FDA standards for PCR-based assays. Thus, chromosome conformation provides a stable, binary readout of cell state and represents an emerging class of biomarkers.
[0567] Here, we used an approach based on the assessment of chromosomal structural changes to develop a blood-based diagnostic test for COO subtyping of DLBCL. We hypothesized that interrogation of genomic structural changes in blood samples from patients with DLBCL could provide an alternative to tissue-based COO classification methods and offer a novel, noninvasive, and more clinically applicable approach to guide clinical decision making and trial design.
[0568] A total of 118 DLBCL patients with known COO subtype and 10 healthy controls (HC) were used in this study. These samples were a subset of samples collected in a phase III, randomized, placebo-controlled trial of rituximab combined with bevacizumab for the treatment of aggressive non-Hodgkin lymphoma. Briefly, adult patients aged ≥18 years with newly diagnosed CD20-positive DLBCL were randomized to receive R-CHOP or R-CHOP combined with bevacizumab (RA-CHOP). Blood samples collected from 60 DLBCL patients were used as a development cohort to identify, evaluate, and improve CCS biomarker leads. Patients in this cohort were all subtyped as high / strong GCB (30) or ABC (30) with high subtype-specific LPS (linear predictor score). The remaining 58 DLBCL samples had an intermediate LPS and were determined by Fluidigm testing to be ABC, GCB, or unclassified ( Fig.25 ). These patient samples were not used for CCS biomarker discovery and development; but were used at a later stage to evaluate the resulting classifier. The Fluidigm test was performed using tissue obtained from lymph nodes (either as a needle biopsy or removed during surgery), and the EpiSwitch analysis was performed using matched peripheral whole blood drawn before the patient received any treatment.
[0569] In addition to patient samples, 12 cell lines (6 ABC and 6 GCB) were used in the initial stage of biomarker screening to identify a set of chromosome conformations that could best distinguish between ABC and GCB disease subtypes (Table 20). Cell lines were obtained from the American Type Culture Collection (ATCC), the German Collection of Microorganisms and Cell Cultures (DSMZ), and the Japan Health Science Foundation (JHSF).
[0570] RNA was isolated and purified from preprocessed FFPE biopsies. DLBCL subtypes were determined by adapting the algorithm of Wright et al. to expression data from a customized Fluidigm gene expression panel (27 genes containing DLBCL subtype predictors). Validation of the COO analysis by comparing Fludigm qRT-PCR with Affymetrix data in a cohort of 15 non-trial subjects revealed a high correlation between qRT-PCR measurements from matched fresh-frozen (FF) and FFPE samples for the 19 classifier genes used. We also found a high correlation between Affymetrix microarray and Fluidigm qRT-PCR measurements from the same FF samples. Classifier gene weights calculated from qRT-PCR data from the Fluidigm COO analysis were highly consistent with weights obtained from previous microarray data in an independent patient cohort. We observed a high correlation (76% concordance) between LPS from Fluidigm analysis, data in FFPE tumors, and LPS from Affymetrix microarray data in matched FF tissues in the Technical Registry cohort.
[0571] Pattern recognition algorithms are used to annotate sites that may form long-range chromosome conformations in the human genome. Pattern recognition software is run based on Bayesian modeling and provides probability scores for regions to participate in long-range chromatin interactions. Sequences from 97 loci (Table 21) are processed by pattern recognition software to generate a list of 13,322 chromosome interactions that are most likely to distinguish DLBCL subtypes. For initial screening, array-based comparisons were performed. 60-mer oligonucleotide probes are designed to query these potential interactions and are uploaded to the Agilent SureDesign website as custom arrays. Each probe is present in quadruplicate on the EpiSwitch microarray. In order to subsequently evaluate potential CCS, nested PCR (EpiSwitch PCR) is performed using sequence-specific oligonucleotides designed using Primer3. The specificity of the oligonucleotides is tested using oligonucleotide-specific BLAST.
[0572] The top ten genomic loci identified as dysregulated in DLBCL were uploaded as a protein list to the Reactome Functional Interaction Network plugin in Cytoscape to generate a network of epigenetic dysregulation in DLBCL. These ten loci were also uploaded to STRING (a search tool for searching the interacting gene / protein database) (https: / / string-db.org / ), which contains more than 9 million known and predicted protein-protein interactions. Restricted to human interactions, a main network was generated (i.e., unconnected nodes were excluded). The highest false discovery rate (FDR) corrected functional enrichment was identified through the Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) databases. The top ten genomic loci were also uploaded to the KEGG pathway database (http: / / www.genome.jp / kegg / pathway.html) to identify specific biological pathways that are dysregulated in DLBCL.
[0573] Exact tests and Fisher's exact tests (for categorical variables) were used to identify discriminative markers. The statistical significance level was set at p ≤ 0.05, and all tests were two-sided. A random forest classifier was used to assess the ability of EpiSwitch markers to identify DLBCL subtypes. Long-term survival analysis was performed by Kaplan-Meier analysis using the survival and survminer packages in R (38). Mean survival times were calculated using a two-tailed t-test.
[0574] We used a stepwise approach to discover and validate a CCS biomarker panel that can distinguish DLBCL subtypes ( Fig.19). As the first step in discovering the EpiSwitch classifier, 97 loci (Table 21) were selected and annotated with the predicted presence of chromosome conformation interaction sites, and their empirical presence was screened using the EpiSwitchCGH Agilent array. The annotated array design represents 13,322 candidate chromosome interactions, and an average of 99 different cis interactions (99 ± 64; mean ± SD) were tested on each locus. The discovery array is used to screen and identify smaller chromosome conformation libraries that can distinguish two major DLBCL subtypes. The samples used for this step are from GCB and ABC cell lines (Table 20) and whole blood from four typing DLBCL patients (two GCB and two ABC) and four HCs. Cell lines are divided into high ABC and high GCB and low ABC and low GCB based on gene expression analysis. The comparisons used on the array were: 1) individual comparisons of DLBCL patients versus pooled HC 2) pooled DLBCL samples versus pooled HC samples 3) pooled high ABC versus pooled high GCB cell lines, and 4) pooled low ABC versus pooled low GCB cell lines.
[0575] From the array analysis, we identified 1,095 statistically significant chromosomal interactions that were discriminatory between high ABC and GCB cell lines and were present in blood samples from DLBCL patients but absent in HC. A set of statistical filters were used to further reduce these interactions to the top 293, of which 151 were associated with the ABC subtype and 143 of which were associated with the GCB subtype. The top 72 interactions from either subtype (36 interactions for ABC and 36 interactions for GCB) were selected for further refinement in 60 typed DLBCL patient samples using the EpiSwitch PCR platform. For all 118 DLBCL samples, an initial subtype classification was assigned according to the Wright algorithm, which calculates a linear predictor score (LPS) from the expression of a set of 27 genes. 60 samples were classified as ABC or GBC and were used to develop the EpiSwitch classifier (the “discovery cohort”), and 58 samples had intermediate LPS scores and were used to evaluate the performance of the EpiSwitch classifier (the “evaluation cohort”). Fig.19 ).
[0576] The 72 interactions identified in the initial screen were narrowed down to a smaller pool using DLBCL patient samples during the discovery step and a second cohort consisting of 60 DLBCL subtyped (30 ABC and 30 GCB) patient samples and 12 HC ( Fig.19). DLBCL subtype calls made by the EpiSwitch assay were confirmed using the Fluidigm platform. Fluidigm gene expression analysis was performed on tissue biopsy samples, while whole blood from the same patients was used for the EpiSwitch PCR assay. The initial step in refinement was to confirm by PCR that the 72 chromosomal interactions identified in the initial screen were specific to DLBCL and absent in HC samples. This was first tested on six untyped DLBCL samples and two HCs, resulting in the identification of 21 interactions that were specific to DLBCL. Next, we tested 24 blood samples from typed DLBCL patient samples (12 ABC and 12 GCB) using EpiSwitch PCR to identify DLBCL-specific chromosomal interactions using Fisher's test. This yielded a set of 10 discriminative chromosome conformation interactions that could accurately distinguish between ABC and GCB subtypes, and was further evaluated in blood samples from an additional set of 36 DLBCL samples (18 ABC and 18 GCB) ( Fig.19 ).
[0577] To test the accuracy, performance, and robustness of the 10-marker panel, we used exact tests for feature selection in 80% of the full sample cohort (48 samples in total: 24 ABC and 24 GCB), and the remaining 20% (12 samples, 6 ABC and 6 GCB) were used for later testing of the final selected CCS markers. The data was split 10 times, and the exact test was run on each split using 80% of the training set for each split. The markers were then ranked using the combined p-values of the 10 markers in the 10 splits. The analysis identified six chromosomal conformations in the IFNAR1, MAP3K7, STAT3, TNFRSF13B, MEF2B, and ANXA11 loci. Collectively, these six interactions formed the DLBCL chromosomal conformation signature (DLBCL-CCS) ( Fig. 20 ).
[0578] The six markers in DLBCL-CCS were used to generate a random forest classifier model and applied to classify the test set of each data split (12 samples, 6 ABC and 6 GCB) in the discovery cohort with known disease subtypes. Using principal component analysis (PCA), the DLBCL-CCS classifier was able to separate ABC and GCB patients from healthy controls ( Fig.26). The combined predictive probability of DLBCL-CCS is shown in Table 22 together with the odds ratio of each marker and the odds ratio of the model generated using logistic regression. The model provided predictive probability scores for ABC and GCB ranging from 0.186 to 0.81 (0 = ABC, 1 = GCB). The probability cutoff for correct classification of ABC was set to ≤0.30, and the probability cutoff for correct classification of GCB was set to ≥0.70. The true positive rate (sensitivity) for a score of ≤0.30 was 100% (95% confidence interval [95% CI] 88.4-100%), while the true negative response rate (specificity) for a score of ≥0.70 was 96.7% (95% CI 82.8-99.9%). Compared with Fluidigm's call for subtyping, 60 of 60 patients (100%) were correctly classified as ABC or GCB using the DLBCL-CCS classifier ( Fig.21 A, Table 22). The AUC under the receiver operating characteristic (ROC) curve of the DLBCL-CCS classifier in this sample cohort was 1 ( Fig.21 B). Finally, we compared the DLBCL subtype calls generated by DLBCL-CCS with the long-term survival curves of patients with known disease subtypes. Patients called ABC had significantly worse survival than those called GBC ( Fig.21 C).
[0579] Next, we evaluated the performance of DLBCL-CCS on a cohort of 58 DLBCL patients with more intermediate LPS values. We applied DLBCL-CCS to assign these patients to DLBCL subtypes and compared the readouts to those from Fluidigm. DLBCL-CCS made subtype calls for all 58 samples, while the Fluidigm analysis made subtype calls for 37 samples, leaving 21 samples as “unclassified” ( Fig. 22 Of the 37 samples that could be subtyped by both assays, 15 samples (40%) were called identically by both assays (8 ABC and 7 GCB) ( Fig. 22 Next, we evaluated the performance of DLBCL subtype calls made by DLBCL-CCS and Fluidigm by comparing subtype calls made at diagnosis with long-term survival curves of patients with type III. Fig.23As shown in the Kaplan-Meier survival curves in , the ABC / GBC calls made by DLBCL-CCS were able to separate the two populations based on known survival trends in DLBCL, with the ABC subtype having a worse prognosis. In contrast, the ABC and GCB populations defined by Fluidigm showed the opposite of clinical observations, with samples classified as ABC having longer survival times than those classified as GCB. Although not statistically significant, the subtype calls made by DLBCL-CCS matched historical clinical observations of survival differences between subtypes through hazard ratio analysis. We did find a significant difference in mean survival time between the two methods. The mean survival of patients classified as ABC and GCB by Fluidigm was 651 days and 626 days, respectively (p=0.854), while the mean survival of patients classified as ABC and GCB by DLBCL-CCS analysis was 550 days and 801 days, respectively (p=0.017) ( Fig.24 ).
[0580] To explore the relationship between the epigenetically dysregulated loci observed in this study and previously reported biological mechanisms associated with DLBCL, we performed a series of network and pathway analyses using the top 10 dysregulated loci as input. First, we explored how these loci were biologically related by constructing the Reactome Functional Interaction Network in Cytoscape, which revealed a network centered on NFKB1, STAT3, and NFATC1. A similar picture emerged when the 10 loci were used to construct a network through STRING DB, with the most connected hubs centered on NFKB1, STAT3, and MAP3K7 and CD40. The most enriched GO terms for biological process were “positive regulation of transcription, DNA templated”, the most enriched GO terms for molecular function were “transcription activator activity, RNA polymerase II transcriptional regulatory region sequence-specific binding”, and “Toll-like receptor signaling pathway” was the most enriched KEGG pathway (Table 22). When we mapped the top ten loci to the KEGG Toll-like receptor signaling pathway, we found that the specific cascade was associated with the production of proinflammatory cytokines and co-stimulatory molecules through the NF-κB and interferon-mediated JAK-STAT signaling cascades.
[0581] Due to the observed differences in disease progression between different DLBCL subtypes, there is an urgent clinical need for a simple and reliable test that can distinguish between ABC and GBC disease subtypes. Given the aggressive nature of the disease, DLBCL requires immediate treatment. These two major subtypes have different clinical management paradigms, and several subtype-specific treatment modalities are under development. When clinical management depends on understanding the disease subtype, rapid and accurate disease diagnosis is essential. The field of COO classification in DLBCL has expanded from IHC-based methods to DNA microarrays, parallel quantitative reverse transcription PCR (qRT-PCR), and digital gene expression. The current popular method is based on the identification of COO on FFPE tissues by GEP, but it is subject to several technical and logistical limitations that limit its widespread application in clinical settings. In addition, there are many factors that affect the performance and reliability of GEP for COO classification on FFPE tissues; including the nature / quality of lymphoma specimens, experimental methods for data collection; data normalization and transformation, type of classifier used, and probability cutoffs used for subtype assignment. Finally, using the Fluidigm approach is a complex and time-consuming process from sample collection to final readout, with many steps in between that have the potential to introduce performance variability. All of these factors impact the overall turnaround time of the assay and limit how it can be used clinically to diagnose and inform disease treatment with existing drugs, as well as to select patients for late-stage trials of novel DLBCL therapies. Therefore, a simple, minimally invasive, and reliable assay is needed to differentiate DLBCL subtypes.
[0582] Using a stepwise discovery approach, we identified a 6-marker epigenetic biomarker panel that the DLBCL-CCS could accurately distinguish DLBCL subtypes. There was complete agreement with the subtype results derived from the gene expression signature; this was expected, as these were the samples used to develop the classifier. When applied to samples with intermediate LPS, agreement between the two analyses was lower (just over 40%). This may be expected, as a lack of overall agreement in DLBCL subtype calls using different classification methods has been noted, and type III samples may be a more heterogeneous population reflecting a more intermediate biological start. However, when we evaluated the predictive classification ability of the EpiSwitch assay in longitudinal disease progression in patients with type III DLBCL, baseline predictions of disease subtype using the EpiSwitch assay based on observed survival curves of patients with unclassified disease were better at predicting actual disease subtype. The regulatory 3D genomics-based epigenetic readout used herein is more consistent with actual clinical outcomes than the gold standard transcription-based molecular approach, representing an actionable advance in the management of DLBCL. This is also consistent with the systems biology assessment of regulatory 3D genomics as molecular patterns that are closely associated with phenotypic differences in tumor conditions. We do note that DLBCL operates on a biological continuum with significant heterogeneity in disease biology between subtypes. By design, the DLBCL-CCS was built to stratify type III samples into either ABC or GCB subtypes. Based on the GEX analysis, type III samples were identified as having intermediate subtype biology and therefore may represent a more heterogeneous patient population. However, the overall observations suggest that the DLBCL-CCS is a better predictor of disease subtype as measured by clinical progression than using a GEX-based approach, and the fact that the EpiSwitch analysis was able to make subtype calls in all samples provides preliminary evidence that this approach can be applied in the clinical setting to inform prognostic landscape, potentially guide treatment decisions, and provide predictions of response to novel therapeutic agents currently in development.
[0583] In the network analysis, NF-κB and STAT3 signaling cascades emerged as putative mediators that distinguish DLBCL subtypes. The role of NF-κB signaling in DLBCL has been studied previously, and indeed, one of the distinguishing features of the ABC subtype is the constitutive expression of NF-κB target genes, a hypothesized mechanism for the poor prognosis of these patients. Moreover, mutations leading to activation of constitutive signaling were predominantly observed in multiple NF-κB pathway genes in the ABC subtype, including TNFAIP3 and MYD88.
[0584] In addition to validating known mechanisms of DLBCL, the network analysis here also identified new potential targets for therapeutic intervention in DLBCL. For example, ANXA11 is a calcium-regulated phospholipid-binding protein that has been implicated in other neoplastic diseases such as colorectal cancer, gastric cancer, and ovarian cancer, and may be a new therapeutic intervention point for DLBCL.
[0585] One of the major clinical advantages of the DLBCL subtyping approach described here lies in the simplified laboratory methods and workflow. Conventional gold standard subtyping of GEP can be performed using a variety of commercial platforms, but generally follows (and requires) a four-step approach: 1) obtaining a tissue biopsy, 2) preparing FFPE tissue sections, 3) gene expression analysis, and 4) algorithmic classification of subtypes. Obtaining a fine needle tissue biopsy of an enlarged peripheral lymph node is an invasive medical procedure that requires an inpatient visit to a clinical site and anesthesia. Once obtained, the fresh biopsy needs to be prepared for paraffin embedding. This is a multistep process, but generally involves soaking in a liquid fixative (such as formalin) long enough to allow it to penetrate the entire specimen, sequential dehydration through an ethanol gradient, and then washing in xylene (a toxic chemical). Finally, the biospecimen needs to be infiltrated with paraffin and cooled to solidify, after which it can be cut into micron-scale sections using a microtome and mounted onto laboratory slides. The entire process from fresh tissue to FFPE sections on slides can take several days. Next, for gene expression analysis, intrinsically unstable RNA is extracted from slide-mounted tissue sections and prepared for hybridization to microarrays according to the array manufacturer's instructions, a process that can take up to a day. After microarray hybridization, a digital readout of relative gene expression levels is obtained and input into a classification algorithm to determine DLBCL subtype. In summary, the process from a patient suspected of DLBCL to a subtype readout can take up to a week or more and involves many different experimental steps using expensive technology, each of which has the potential to introduce experimental variability in the process. In the methods described herein, the time and number of steps from biofluid collection to subtype readout are significantly reduced. Patients suspected of DLBCL can go to the outpatient clinic for a routine small volume (approximately 1 mL) blood draw. Fresh frozen blood can then be shipped to a central, accredited reference laboratory to analyze the absence / presence of the chromosome conformations identified in this study; the process uses a smaller volume (approximately 50 mL) of whole blood as input and specific PCR primer sets and reaction conditions to detect chromosome conformations using simple and routine PCR instruments in less than 24 hours after receiving the sample. The DLBCL subtyping approach described here offers the added advantage that there is the potential for further refinement using the proposed methodology. In this study, the final readout of DLBCL-CCS was to detect the chromosome conformation that constitutes the classifier using a set of nested PCR reactions. This PCR-based output can be further refined to utilize quantitative PCR as a readout and operate under the published Minimum Information for Real-Time Quantitative PCR Experiments (MIQE) guidelines, which aim to improve experimental reproducibility and reliability across reference laboratories and testing sites. Finally, the approach described here is applicable to the evolving understanding of the disease itself, such as the different physiologically heterogeneous forms of the disease.
[0586] In summary, here we developed a robust, complementary approach for non-invasive COO assignment from whole blood samples using the EpiSwitch CCS readout. We demonstrated the clinical validity of this classification approach on a large cohort of DLBCL patients. The EpiSwitch platform has several attractive features as a biomarker modality with clinical utility. CCS has very high biochemical stability, can be detected using very small amounts of blood (typically ~50 μl), and the assay is performed using established laboratory methods and standard PCR readouts, including MIQE-compliant qPCR. Finally, the rapid turnaround time of the EpiSwitch assay (~8 h-16 h) compares favorably to the >48 h turnaround time of the Fluidigm platform.
[0587] Example 6. Further studies of canine DLBCL
[0588] Here, we use EpiSwitch TM We investigated the potential of a platform technology to evaluate chromosome conformation signature (CCS) as a biomarker for the detection of canine diffuse large B-cell lymphoma (DLBCL). TM Whether established systemic liquid biopsy biomarkers characterized in human DLBCL patients translate to dogs with the homologous disease. Orthologous sequence conversion of CCS from human to dog was first verified and validated in control and lymphoma canine cohorts.
[0589] Blood samples were obtained from dogs with DLBCL and from apparently healthy dogs. All dogs diagnosed with DLBCL were part of the LICKing lymphoma trial. Blood samples were obtained from each dog before starting treatment and on day +5 after the experimental intervention, but before starting doxorubicin chemotherapy. TM Technology for monitoring systemic epigenetic biomarkers of CCS.
[0590] An 11-marker classifier was generated from a library of 75 EpiSwitch CCSs identified in human DLBCL using whole blood from 28 dogs (14 diagnosed with DLBCL and 14 controls without overt disease). Validation of the developed diagnostic markers was performed on a second cohort of 10 dogs: 5 DLBCL and 5 controls. In the second cohort, the classifier classified DLBCL from non-DLBCL with 80% accuracy, 80% sensitivity, 80% specificity, 80% positive predictive value (PPV), and 80% negative predictive value (NPV).
[0591] EpiSwitch EstablishedTM The classifiers contain robust systematic binary markers of epigenetic dysregulation, characteristics of which are usually attributed to genetic markers: the binary states of these classifier markers are statistically significant for diagnosis.
[0592]
[0593]
[0594] Table 5a
[0595]
[0596]
[0597] Table 5b
[0598]
[0599]
[0600] Table 5c
[0601]
[0602]
[0603] Table 5d
[0604]
[0605]
[0606] Table 5e
[0607]
[0608]
[0609] Table 5f
[0610]
[0611]
[0612] Table 5g
[0613]
[0614]
[0615] Table 5h
[0616]
[0617]
[0618] Table 5i
[0619]
[0620]
[0621] Table 5j
[0622]
[0623]
[0624] Table 5k
[0625] logFC AveExpr t P. Value adj.P.Val 53 0.146814815 0.146814815 3.078806942 0.015395162 0.044697142 54 0.147791337 0.147791337 1.633707333 0.141477876 0.232950641 55 0.148349454 0.148349454 3.08222473 0.015316201 0.044558199 56 0.153758518 0.153758518 1.99378109 0.081784292 0.154640968 57 0.156103192 0.156103192 4.16566548 0.003235665 0.015172146 58 0.161073376 0.161073376 2.527153349 0.035788708 0.08344291 59 0.171050829 0.171050829 3.890660293 0.004727539 0.019533312 60 0.174144322 0.174144322 2.584598623 0.0327468 0.078039912 61 0.18112944 0.18112944 2.685388848 0.028032588 0.069592429 62 0.19092131 0.19092131 3.73604023 0.005879148 0.022572713 63 0.194400918 0.194400918 2.492510608 0.037760627 0.086786653 64 0.195707712 0.195707712 2.71102351 0.026948639 0.067475149 65 0.204252124 0.204252124 3.975216299 0.004202345 0.018041687 66 0.210054656 0.210054656 2.338335953 0.047960033 0.103796265 67 0.210247736 0.210247736 2.234440593 0.056352274 0.11719604 68 0.213090816 0.213090816 2.087748617 0.07073665 0.138732545 69 0.226250319 0.226250319 2.25409429 0.054659526 0.114573853 70 0.246810388 0.246810388 5.953590771 0.000359019 0.0048119 71 0.247739738 0.247739738 4.372950932 0.002449301 0.012749359 72 0.248261394 0.248261394 6.173828851 0.000282141 0.004322646 73 0.251556432 0.251556432 4.318526719 0.00263344 0.013280362 74 0.253919456 0.253919456 1.95465634 0.086862876 0.161472968 75 0.256754187 0.256754187 5.535121315 0.000577231 0.005715942 76 0.257160612 0.257160612 3.449233265 0.008887517 0.030040982 77 0.259132781 0.259132781 4.366813249 0.002469352 0.012816471 78 0.287279843 0.287279843 2.539709619 0.035100109 0.082104681 79 0.31600033 031600033 3.153558526 0.013761553 0.041279885 80 0.358221647 0358221647 3.100524122 0.014900586 0.043739937 81 0.364193755 0.364193755 3.369619436 0.009987239 0.03271603 82 0.453457772 0.453457772 3.247156175 0.011968978 0.037200176 83 0.180533568 0.180533568 5.147835975 0.000914678 0.007187473 84 0.182697701 0.182697701 5.877748203 0.00039063 0.004938906 85 -0.148364769 -0.148364769 -4.986366569 0.001115061 0.008026755 86 -0.538084185 -0.538084185 -6.494881534 0.000200669 0.003807401 87 -0.545447375 -0.545447375 -6.02027801 0.000333544 0.004684915 88 -0.554745602 -0.554745602 -8.383072026 0.0000337 0.002483007 89 0.503059535 0.503059535 6.535294395 0.000192409 0.003731412 90 0.36623319 0.36623319 5.026075307 0.001061678 0.007815282 91 0.338959712 0338959712 4.957835746 0.001155226 0.008192382 92 0.127634089 0.127634089 5.070996787 0.001004634 0.007593027
[0626] Table 51
[0627] B FC FC_1 LS Detection ring 53 -3.422371897 1.107122465 1.107122465 1 DBLCL 54 -5.595015473 1.1078721 1.1078721 1 DBLCL 55 -3.417108405 1.108300771 1.108300771 1 DBLCL 56 -5.084251662 1.112463898 1.112463898 1 DBLCL 57 -1.806028568 1.114273349 1.114273349 1 DBLCL 58 -4.275957963 1.118118718 1.118118718 1 DBLCL 59 -2.201619835 1.125878252 1.125878252 1 DBLCL 60 -4.187168091 1.128295002 1.128295002 1 DBLCL 61 -4.031097147 1.133771131 1.133771131 1 DBLCL 62 -2.428494364 1.141492444 1.141492444 1 DBLCL 63 -4.329420018 1.14424891 1.14424891 1 DBLCL 64 -3.991368624 1.145285841 1.145285841 1 DBLCL 65 -2.078878648 1.152088963 1.152088963 1 DBLCL 66 -4.56625722 1.156732005 1.156732005 1 DBLCL 67 -4.724458483 1.156886824 1.156886824 1 DBLCL 68 -4.945085913 1.159168918 1.159168918 1 DBLCL 69 -4.694639254 1.169790614 1.169790614 1 DBLCL 70 0.493437892 1.186580836 1.186580836 1 DBLCL 71 -1.514951748 1.18734545 1.18734545 1 DBLCL 72 0.744313598 1.187774853 1.187774853 1 DBLCL 73 -1.590768631 1.190490767 1.190490767 1 DBLCL 74 -5.141601134 1.192442298 1.192442298 1 DBLCL 75 -0.002180366 1.194787614 1.194787614 1 DBLCL 76 -2.856981615 1.195124248 1.195124248 1 DBLCL 77 -1.523480143 1.196759104 1.196759104 1 DBLCL 78 -4.25656399 1.220337199 1.220337199 1 DBLCL 79 -3.307417374 1.24487452 1.24487452 1 DBLCL 80 -3.388938618 1.281844844 1.281844844 1 DBLCL 81 -2.977498786 1.287162103 1.287162103 1 DBLCL 82 -3.164028291 1369318233 1.369318233 1 DBLCL 83 -0.483711076 1.13330295 1.13330295 1 DBLCL 84 0.405473595 1.135004251 1.135004251 1 DBLCL 85 -0.691112407 0.902272569 -1.108312537 -1 Ctrl 86 1.098164859 0.688684835 -1.452043008 -1 Ctrl 87 0.570114643 0.685178898 -1.459472852 -1 Ctrl 88 2.923843236 0.680777092 -1.46890959 -1 Ctrl 89 1.14173325 1.417215879 1.417215879 1 DBLCL 90 -0.639742862 1.288982958 1.288982958 1 DBLCL 91 -0.728168867 1.26484422 1.26484422 1 DBLCL 92 -0.581917188 1.092500613 1.092500613 1 DBLCL
[0628] Table 5m
[0629]
[0630]
[0631] Table 5n
[0632]
[0633]
[0634] Table 5o
[0635]
[0636]
[0637] Table 5p
[0638]
[0639]
[0640] Table 5q
[0641]
[0642] Table 5r
[0643]
[0644]
[0645] Table 6a
[0646] HyperG_Stats FDR_HyperG Percent_Sig logFC AveExpr t 1 0.064790053 0.737205743 25 0.67511652 0.67511652 13.76185645 2 0.032709022 0.548212211 19.57 0.299375751 0.299375751 7.197207444 3 0.040338404 0.548212211 25 -0.168081632 -0.168081632 -3.274998031 4 0.765503518 1 7.69 -0.425291613 -0.425291613 -11.67074071 5 0.024128503 0.483719041 33.33 0.266992266 0.266992266 4.835274287 6 n / a n / a n / a n / a 4.72222828 n / a
[0647] Table 6b
[0648] P. Value adj.P.Val B FC FC_1 LS 1 0.000000031 0.0000143 9.558686586 1.596725728 1.596725728 1 2 0.0000184 0.000805368 3.154114326 1.230611817 1.230611817 1 3 0.007481356 0.033194645 -3.020586815 0.890025372 -1.123563476 -1 4 0.000000168 0.0000357 7.913111034 0.744688192 -1.342843905 -1 5 0.000536131 0.005815136 -0.328887879 1.203296575 1.203296575 1 6 0.04505295 0.4547981 n / a n / a n / a n / a
[0649] Table 6c
[0650]
[0651] Table 6d
[0652]
[0653] Table 6e
[0654]
[0655] Table 6f
[0656] <![CDATA[ Primer name ]]> <![CDATA[ Primer sequences ]]> 1 PCa119-245 AAGAAGGGATGGGACGGGACT PCa119-247 GGTACACGAATTAACTATTCCCTGT 2 PCa119-165 ACTGGTCACAGGGAACGATGG PCa119-167 AGGTGTGAATGTTACTGAACACAAA 3 PCa119-130 ACTTGGATTCCCAAAACGCCA PCa119-132 CTCTTCCCCGGTGAGTTTCCA 4 PCa119-065 CAGCCTACCTTGCCTGACACT PCa119-067 AAAGCCCAGTGATGGCCCAT 5 PCa119-154 TCCATTTTCCTTTCCCTTTGCTCTG PCa119-155 CCACACAGGGCCCTAATGACC 6 MMP 1-42F GGGGAGTGGATGGGATAAGGTG MMP 1F TGGGCCTGGTTGAAAAGCAT
[0657] Table 6g
[0658] <![CDATA[ Probe ]]> <![CDATA[ Probe sequence ]]> Gene 1 OBD119F015 AGTGTTTAATCGATAGAAATATAACATGAAACACA MIR98 2 OBD119F06 AGGGATACTCGAAGTTAATTTGCTTCTT DAPK1 3 OBD119F09 AAGAAGCTTACAGTCGAAGGTCCCAA HSD3B2 4 OBD119F08 ATTCCTTTCAAATTATGTTTTCGAGTCTGAATAATA ERG 5 SRD5A3FAM7415RC AAATAGACTTCTGCCTCGATTAAGCA SRD5A3 6 MMP1F1b2 ATCCAGCATCGAAGAGGGAAACTGCATCA MMP1
[0659] Table 6h
[0660] Markers GLMNET 1 PCa119-245.247 -5.91743E-06 2 PCa119-165.167 -1.57185E-05 3 PCa119-130.132 4.47291E-07 4 PCa119-065.067 6.32136E-06 5 PCa119-154.155 -8.00857E-08 6 MMP1-41F.MMP 1F 0
[0661] Table 6i
[0662]
[0663]
[0664] Table 7. Preferred DLBCL markers
[0665]
[0666]
[0667] Table 8a
[0668]
[0669]
[0670]
[0671] Table 8b
[0672] N Probe Markers GLMNET 1 ORF1_1_1034282_1037357_1049484_1054771_FF OBD169_001.OBD169_003 0.150341207 2 ORF5_1_1182474_1185271_1270569_1273244_RF OBD169_009.OBD169_011 -0.065057056 3 ORF5_1_1147651_1150121_1196191_1197234_RF OBD169_021.OBD169_023 0 4 ORF5_1_1146367_1147651_1165983_1167502_FF OBD169_033.OBD169_035 0 5 ORF5_1_1196191_1197234_1230936_1232838_RR OBD169_041.OBD169_043 0.122625202 6 ORF5_1_1270569_1273244_1300933_1312034_FF OBD169_049.OBD169_051 -0.050953035 7 ORF5_1_1196191_1197234_1289361_1294150_FF OBD169_061.0BD169_063 0.127785257 8 ORF5_1_1140030_1142517_1230936_1232838_RR OBD169_065.0BD169_067 -6.18144E-06 9 ORF5_1_1230936_1232838_1273244_1276010_RR OBD169_073.0BD169_075 0 10 ORF41_2_36413514_36415342_36452868_36458269_RR OBD169_105.OBD1_69_107 0.029250039 11 ORF91_7_65032142_65033242_65065127_65067650_FF OBD169_125.OBD169_127 0.005994639 12 ORF91_7_65037215_65039217_65065127_65067650_RF OBD169_129.OBD169_131 0 13 ORF9_10_23456592_23460302_23494817_23496168_RR OBD169_133.OBD169_135 0.161924686 14 ORF16_11_40371218_40374048_40393587_40395559_RF OBD169_153.OBD169_155 0 15 ORF31_15_29619588_29621525_29646237_29648560_RR OBD169_165.OBD169_167 0 16 ORF30_15_10476260_10484217_10545581_10548270_RR OBD169_169.OBD169_171 0.063674679 17 ORF32_16_10747182_10750815_10792291_10794979_RF OBD169_185.OBD169_187 0.248790766 18 ORF70_26_27894296_27895372_27963114_27965001_RR OBD169_229.OBD169_231 -0.042293888 19 ORF70_26_27890569_27893929_27906620_27909025_RR OBD169_233.OBD169_235 0.052029568 20 ORF79_32_24013860_24017127_24028587_24030780_RR OBD169_265.OBD169_267 0.141700302 21 ORF82_32_9652472_9664654_9692674_9698030_RR OBD169_277.0BD169_279 -0.097352472 22 ORF104_X_109508063_109510622_109526507_109531763_FF OBD169_293.0BD169_295 0
[0673] Table 9a
[0674]
[0675]
[0676] Table 9b
[0677] Table 10. Prostate cancer risk group categories
[0678]
[0679] * According to the 2018 updated NCCN guidelines, T2c is considered intermediate risk.
[0680] Abbreviations. PSA: prostate-specific antigen.
[0681] Table 11. 5-marker signature for the diagnosis of prostate cancer
[0682]
[0683] Table 12.
[0684] Pathology and EpiSwitch TM Comparison of results
[0685]
[0686] Blind sample classification results (n=20).
[0687]
[0688] (*) These values depend on disease prevalence.
[0689] Abbreviations: 95% CI: 95% confidence interval.
[0690]
[0691]
[0692] Table 15.
[0693] Category 3 vs Category 1 Pathology and EDiSwitch TM Comparison of results.
[0694]
[0695] Blind classification results of 3-class vs 1-class classifier (n=67).
[0696]
[0697] (*) These values depend on disease prevalence.
[0698] Abbreviations: 95% CI: 95% confidence interval.
[0699] Table 16.
[0700] Class 3 vs Class 2 Pathology and EpiSwitch TM Comparison of results.
[0701]
[0702] Blind classification results of 3-class vs 2-class classifiers (n=43).
[0703]
[0704] (*) These values depend on disease prevalence.
[0705] Abbreviations: 95% CI: 95% confidence interval.
[0706] Table 17. Clinical characteristics of patients participating in the study.
[0707]
[0708] Abbreviations: PSA: prostate-specific antigen.
[0709] Table 18. List of 425 prostate cancer-associated loci tested in the initial array.
[0710]
[0711]
[0712]
[0713]
[0714]
[0715]
[0716]
[0717]
[0718]
[0719]
[0720]
[0721] Table 20. DLBCL cell lines used in this study. Cell lines were obtained from the American Type Culture Collection (ATCC), the German Collection of Microorganisms and Cell Cultures (DSMZ), and the Japan Health Sciences Foundation (JHSF).
[0722]
[0723] Table 21. 97 genomic loci used for initial biomarker discovery screening
[0724]
[0725] Table 22. Composite predicted probability of DLBCL-CCS in the discovery cohort.
[0726]
[0727] Table 23. DLBCL-CCS and Fluidigm subtype calls in the discovery cohort. Subtype calls for samples with known DLBCL subtypes by the EpiSwitch DLBCL-CCS and Fluidigm assays. 60 of the 60 samples were consistently called as ABC or GCB by both assays.
[0728]
[0729] Table 24. Biological function enrichment of the top 10 DLBCL-CCS loci.
[0730]
[0731]
[0732] Table 25a
[0733]
[0734]
[0735] Table 25b
[0736]
[0737] Table 25c
[0738]
[0739] Table 25d
[0740]
[0741] Table 25e
[0742]
[0743] Table 25f
[0744]
[0745]
[0746] Table 25g
[0747]
[0748] Table 25h
Claims
1. Use of a probe set for detecting whether a chromosome interaction exists in an individual in the preparation of a kit for determining the prognosis of prostate cancer in the individual, in, The presence or absence of the chromosome interaction is detected by a method comprising the following steps: (i) cross-linking chromosomal regions of the individuals that have been brought together in chromosomal interactions; (ii) cutting the cross-linked region; (iii) ligating the cross-linked cleaved DNA ends to form a ligated nucleic acid; as well as (iv) detecting the presence or absence of the linked nucleic acid corresponding to the chromosome interaction by a probe; Wherein, in the prognosis, the individuals are divided into 1 category or 3 categories, the 1 category represents low-risk, indolent prostate cancer, the 3 category represents high-risk, aggressive prostate cancer, and the probe set includes: (a) A probe having the sequence ACGTC GTTACAGTTTTAATTTTTCTACTTCGATGTTAATCTCCTAAA AAACATCCAACCA for detecting chromosome interaction in the BMP6 gene; (b) a probe having the sequence TCTTGA ATGTGCTTAGTATTATTCAGACTCGAAAACATAATTTGAAA GGAATTCATTCTG for detecting chromosome interaction in the EGR gene; (c) a probe having the sequence CACCA GTTGGTAATTCTATGTGTAAGTTTCGAGCTTATAAGATCAAT CAGGAATTATTCC for detecting chromosome interaction in the MSR1 gene; (d) a probe with the sequence GCAGG GTGGCTATAGCTCAGGAGAGTGCTCGACGGAGTCTTGCTCT TTCACCCAGGCTGG for detecting chromosome interaction in the MUC1 gene; (e) a probe having a sequence of CAAT TGGTGGATATAGAAAGGTCTAAATTCGATAAGTATAGACTCAGAATGCAAAAATGT for detecting chromosome interaction in the ACAT1 gene; and (f) The probe for detecting chromosome interaction in the DAPK1 gene has the sequence ACTA ATCCCCTGAAGAAGCAAATTAACTTCGAGTATCCCTTTAAGT TTGTTTTTAAAATA.
2. The use according to claim 1, wherein detecting whether the chromosomal interaction exists comprises performing specific detection of the connecting nucleic acid by quantitative PCR (qPCR), wherein the quantitative PCR uses primers capable of amplifying the connecting nucleic acid, and during the PCR reaction, the probe binds to the connecting site.
3. The use according to claim 2, wherein a fluorophore is covalently linked to the 5' end of the probe. The use according to claim 2 , wherein the quencher is covalently linked to the 3′ end of the probe.
Citation Information
Patent Citations
Application of epigenetic chromsomal interactions in cancer diagnostics
WO2018100381A1