Biomarker

Chromosomal conformation signatures address the limitations of binary biomarkers by providing stable, early indicators of disease subgroups, enabling precise diagnosis and prognosis for ALS and Huntington's disease through methods like the Episwitch system, facilitating individualized treatment.

JP7701419B2Active Publication Date: 2025-07-01OXFORD BIODYNAMICS LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2023145557
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2017-10-02
Filing Date
2023-09-07
Publication Date
2025-07-01
Estimated Expiration
2038-10-01

AI Technical Summary

Technical Problem

Current biomarkers for diseases like ALS and Huntington's disease, such as RNA expression patterns and protein markers, provide binary readouts without magnitude, leading to challenges in classification statistics and data analysis due to varying biomarker magnitudes across patients, making it difficult to stratify patient cohorts effectively.

Method used

The use of chromosomal conformation signatures (CCSs) as biomarkers, which are stable and reflect early biological changes, allowing for the detection of specific chromosomal interactions associated with disease subgroups through methods like the Episwitch system, involving cross-linking, ligation, and hybridization of nucleic acids to identify distinct chromosomal states.

Benefits of technology

Enables accurate classification of patient subgroups by detecting stable chromosomal interactions, facilitating early diagnosis and prognosis with high sensitivity and specificity, and allowing for individualized treatment approaches.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007701419000049
    Figure 0007701419000049
  • Figure 0007701419000050
    Figure 0007701419000050
  • Figure 0007701419000051
    Figure 0007701419000051
Patent Text Reader

Abstract

To provide a process for analysing chromosome regions and interactions relating to ALS and Huntington's disease.SOLUTION: A system based on chromosomal cross-linked regions which have come together in a chromosome interaction is used, subjecting the chromosomal DNA to cleavage and then ligating the nucleic acids present in the cross-linked entity, to detect a ligated nucleic acid with sequences from both the regions which formed the chromosomal interaction. This detection allows for determination of the presence or absence of a particular chromosome interaction.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to detecting chromosomal interactions.

Background Art

[0002] Disease characteristics can be identified by biomarkers. Currently commonly used biomarkers include RNA expression patterns and protein markers.

Summary of the Invention

[0003] There are specific chromosomal conformation signatures (CCSs) at loci, or they do not exist for regulatory epigenetic control settings related to pathology or treatment. CCSs have a gentle off-rate, and when representing a special phenotype or pathology, CCSs change only in pharmacological signaling, or change as a result of external interference. Furthermore, the measurement of these events is binary, so this readout is in stark contrast to the continuous readouts of various levels of DNA methylation, histone modification, and most non-coding RNAs. The magnitude of the change in a specific biomarker varies greatly from patient to patient, and since using it to stratify patient cohorts causes problems in classification statistics, the continuous readouts used for most molecular biomarkers so far pose challenges in data analysis. These classification statistics are suitable for using biomarkers that provide only a binary score of "yes, or no" for phenotypic differences without magnitude - chromosomal conformation (EpiSwitch (trademark)) biomarkers indicate that they are excellent resources for potential diagnostic, prognostic, and predictive biomarkers.

[0004] The inventors have identified chromosomal interactions in regions of the genome associated with amyotrophic lateral sclerosis (ALS) or Huntington's disease using an approach that enables the identification of subgroups within a population. Accordingly, the present invention provides a process for detecting chromosomal states representative of subgroups within a population, the process comprising determining whether chromosomal interactions are present or absent within a disease-associated region of the genome defined as ALS or Huntington's disease. A chromosomal interaction is a method for determining which chromosomal interactions are associated with a chromosomal state corresponding to an ALS or Huntington's disease subgroup within a population, the method comprising contacting a first set of nucleic acids from a subgroup of different states of a chromosome with a second set of indicator nucleic acids and hybridizing complementary sequences, wherein the nucleic acids in the first and second sets of nucleic acids represent ligation products containing sequences from both chromosomal regions that come together in the chromosomal interaction, and optionally identifying or being identifiable (or derivable) by a method by which the pattern of hybridization between the first and second sets of nucleic acids determines which chromosomal interactions are specific to an ALS or Huntington's disease subgroup. The ALS or Huntington's disease subgroup may be related to diagnosis (presence of ALS or Huntington's disease) or prognosis (e.g., rate of progression of ALS or Huntington's disease). Any of the specific associated chromosomal interactions (markers) described herein, including combinations of markers, can be used as the basis of the present invention.

[0005] The present invention is a process for detecting the state of a chromosome indicative of a disease subgroup within a population, the process comprising detecting whether a chromosomal interaction related to the state of the chromosome is present or absent within a defined region of the genome, wherein the disease subgroup is an amyotrophic lateral sclerosis (ALS) subgroup; and - The chromosomal interaction is optionally identified by a method that determines which chromosomal interactions are associated with the chromosomal state corresponding to an ALS subgroup of the population, the method comprising contacting a first set of nucleic acids from subgroups having various states of the chromosome with a second set of index nucleic acids and allowing complementary sequences to hybridize, wherein the nucleic acids in the first and second sets of nucleic acids represent ligation products containing sequences from both chromosomal regions that come together in the chromosomal interaction, and the hybridization pattern between the first and second sets of nucleic acids allows determination of which chromosomal interactions are specific to the ALS subgroup; and - The chromosomal interaction is (i) present in any of the regions or genes listed in Table 1 or Table 5; and / or (ii) corresponds to any of the chromosomal interactions represented by any of the probes shown in Table 1 or Table 5; and / or (iii) corresponds to any of the chromosomal interactions shown in Table 10 or Table 11; and / or (iv) is included in (i), (ii) or (iii), or is present in a 4,000 base region adjacent to (i), (ii) or (iii) A process is provided.

[0006] The present invention is a process for detecting the chromosomal state indicative of a disease subgroup in a population, the process comprising detecting whether a chromosomal interaction related to the chromosomal state is present or absent within a defined region of the genome, wherein the disease subgroup is a Huntington's disease subgroup; and - said chromosomal interactions have optionally been identified by a method for determining which chromosomal interactions are associated with chromosomal states corresponding to Huntington's disease subgroups of a population, the method comprising contacting a first set of nucleic acids from subgroups with various chromosomal states with a second set of index nucleic acids and allowing complementary sequences to hybridize, where the nucleic acids in the first and second sets of nucleic acids represent ligation products comprising sequences from both chromosomal regions converged at the chromosomal interaction, and the pattern of hybridization between the first and second sets of nucleic acids allows the determination of which chromosomal interactions are specific to a Huntington's disease subgroup; and -Chromosomal interactions (i) in any of the regions or genes listed in Table 12; and / or (ii) corresponds to any of the chromosomal interactions represented by any of the probes shown in Table 12; and / or (iii) corresponds to any of the dye-and-dye interactions shown in Table 12; and / or (iv) containing (i), (ii), or (iii) or present in a 4,000-base region adjacent to (i), (ii), or (iii); Provide a process. [Brief description of the drawings]

[0007]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Mode for Carrying Out the Invention

[0008] The process of the present invention The process of the present invention includes a classification system for detecting chromosomal interactions associated with ALS or Huntington's disease. This classification is the Episwitch™ system described herein, which uses a system based on the cross-linked regions of chromosomes that gather chromosomal interactions to cleave chromosomal DNA, and then ligates the nucleic acids present in the cross-linked entities to induce a ligation diffusion having sequences from both regions that formed the chromosomal interaction. Detection of this ligated diffusion enables detection of the presence or absence of a specific chromosomal interaction.

[0009] Chromosomal interactions can be identified using the methods described above in which a first and a second population of nucleic acids are used. These nucleic acids can also be generated using Episwitch™ technology.

[0010] Epigenetic interactions related to the present invention As used herein, the terms "epigenetic" and "chromosomal" interactions typically mean interactions between distal regions of loci on a chromosome, which are dynamic and transform, form, or break depending on the state of the regions of the chromosome.

[0011] In a particular process of the present invention, chromosomal interactions are detected by first generating a ligated nucleic acid containing sequences from both regions of the chromosome that are part of the interaction. In such a process, the two regions can be cross-linked by any suitable means. In a preferred embodiment, the interaction is cross-linked using formaldehyde, although any aldehyde, or D-biotinoyl-e-aminocaproic acid-N-hydroxysuccinimide ester or digoxigenin-3-O-methylcarbonyl-e-aminocaproic acid-N-hydroxysuccinimide ester may also be used for cross-linking. Paraformaldehyde can cross-link DNA strands that are 4 angstroms apart. Preferably, the chromosomal interactions are on the same chromosome and are optionally 2 to 10 angstroms apart.

[0012] Chromosomal interactions can reflect the state of chromosomal regions, such as whether it is being transcribed or repressed in response to changes in physiological conditions. Chromosomal interactions that are specific to the subgroups defined herein have been found to be stable and thus provide a reliable means of measuring differences between two subgroups.

[0013] In addition, chromosomal interactions specific to a certain feature (such as a disease situation) are generally thought to occur early in the biological process, for example as compared to other epigenetic markers such as changes in methylation or histone protein binding. Thus, the process of the present invention can detect the early stage of a biological process. This, in turn, enables earlier intervention (such as treatment) that may be more effective. Furthermore, there is little variation in the relevant chromosomal interactions among individuals within the same subgroup. Detecting chromosomal interactions is very beneficial as there can be up to 50 different possible interactions per gene, and thus the process of the present invention can search through 500,000 different interactions.

[0014] Set of preferred markers As used herein, the term "marker" or "biomarker" means a specific chromosomal interaction that can be detected (classified) in the present invention. Specific markers are disclosed herein. Any of them can be used in the present invention. Additional sets of markers may be used, for example, in the combinations or numbers disclosed herein. The markers disclosed in the tables herein are preferred. These can be classified by any suitable method, such as PCR or probe-based methods disclosed herein, such as qPCR. Markers are defined herein by their location or by probe and / or primer sequences.

[0015] Location and cause of epigenetic interactions Epigenetic chromosomal interactions can overlap and include regions of the chromosome that are shown to encode related or un-described genes, but can also be in intergenic regions. It should further be noted that the inventors have discovered that epigenetic interactions in all regions are equally important in determining the state of a chromosomal locus. These interactions do not necessarily lie in the coding regions of specific genes located at the locus, but may be in intergenic regions.

[0016] The chromosomal interactions detected in the present invention can be caused by changes to the underlying DNA sequence, environmental factors, DNA methylation, non-coding antisense RNA transcripts, non-mutagenic carcinogens, histone modifications, chromatin remodeling, and specific local DNA interactions. The changes leading to chromosomal interactions can be changes to the underlying nucleic acid sequence that do not themselves directly affect the gene product or the mode of gene expression. Such changes can be, for example, SNPs inside and / or outside of genes, gene fusions, and / or deletions of intergenic DNA, microRNAs, and non-coding RNAs. For example, it is known that approximately 20% of SNPs are in non-coding regions, and thus the processes described are also beneficial in non-coding scenarios. In one aspect, the regions of the chromosome that together form an interaction are less than 5 kb, 3 kb, 1 kb, 500 base pairs, or 200 base pairs apart on the same chromosome.

[0017] The detected chromosomal interactions are preferably within any of the genes mentioned in Table 1 or 5. However, it may be upstream or downstream of the gene, for example up to 50,000, up to 30,000, up to 20,000, up to 10,000, or up to 5000 base pairs upstream or downstream from the gene or from the coding sequence.

[0018] Subgroups, Diagnosis, and Individualized Therapy The object of the present invention is to enable the detection of chromosomal interactions associated with subgroups of ALS or Huntington's disease. Thus, the process may or may not be used for the diagnosis of ALS or Huntington's disease. The process may or may not be used for the prognosis of ALS or Huntington's disease.

[0019] When the process is used for the diagnosis of ALS, the classification of markers related to Table 1 is preferred (i.e., specific disclosed markers and markers in the genes, regions and adjacent regions disclosed in Table 1). In one embodiment regarding diagnosis, only the markers related to Table 1 are classified and other markers are not classified. In another embodiment regarding diagnosis, at least 1, 2, 3, 4, 5 or more markers related to Table 5 (such as represented by the listed probe sequences) are not classified.

[0020] When the process is used for prognosis, the classification of markers related to Table 5 is preferred (i.e., specific disclosed markers and markers in the genes, regions and adjacent regions disclosed in Table 5). In one embodiment regarding prognosis, only the markers related to Table 5 are classified and other markers are not classified. In another embodiment regarding prognosis, at least 1, 2, 3, 4, 5 or more markers related to Table 1 (such as represented by the listed probe sequences) are not classified.

[0021] Generally, "prognosis" relates to the progression of ALS and an individual can be classified into a subgroup of progression rate. It can be measured using a progression wave, the ALS-FRS-R score (ALS Functional Rating Scale), and an individual can be classified as greater than or less than a specific value, for example, the decrease in points per 30 days is greater than or less than 0.5. Thereby, a predictive prognosis can be made.

[0022] As used herein, "subgroup" preferably means a collective subgroup (a subgroup within a collective), more preferably a subgroup within a specific population of animals such as a specific eukaryote or mammal (e.g., human, non-human, non-human primate, or rodent such as mouse or rat). Most preferably, "subgroup" means a subgroup within the human population.

[0023] The present invention relates to detecting and treating a specific subgroup in a population. The inventors have discovered that chromosomal interactions differ between subsets (e.g., at least two subsets) in a given population. By identifying these differences, a physician can classify their patients as part of one subset of the population described in the process. Therefore, the present invention provides a physician with a process for individualizing drug therapy for a patient based on epigenetic chromosomal interactions.

[0024] In one embodiment, the subgroup (e.g., the subgroup related to prognosis) is defined by the ALS Functional Rating Scale (e.g., as described in Cedarbaum, J.M., Stambler, N., Malta, E., Fuller, C., Hilt, D., Thurmond, B. and Nakanishi, A. (1999) The ALSFRS-R: a revised ALS functional rating scale that incorporates assessments of respiratory function. BDNF ALS Study Group (Phase III), Journal of the Neurological Sciences. 169(1-2), 13-21.).

[0025] In one embodiment, a subgroup (e.g., a subgroup related to prognosis) is defined by forced vital capacity (FVC) (e.g., as described in Talakad, N.S., Pradhan, C., Nalini, A., Thennarasu, K. and Raju T.R. (2009) Assessment of Pulmonary Function in Amyotrophic Lateral Sclerosis, Indian J Chest Dis Allied Sci. 51(2):87-91.).

[0026] Generate ligated nucleic acids Certain forms of the invention utilize ligated nucleic acids, particularly ligated DNA. These contain sequences from both regions that come together in a chromosomal interaction and thus provide information about the interaction. The EpiSwitch™ method described herein uses the generation of such ligated nucleic acids to detect chromosomal interactions.

[0027] Accordingly, the process of the invention includes the following steps (such as methods including these steps): (i) Crosslinking epigenetic chromosomal interactions present at chromosomal loci, preferably in vitro; (ii) Optionally, isolating crosslinked DNA from the chromosomal locus; (iii) Cutting the crosslinked DNA by restriction digestion with an enzyme that cuts it at least once, particularly an enzyme that cuts it at least once within the chromosomal locus; (iv) Ligating the ends of the crosslinked and cut DNA, particularly to form a DNA loop; and (v) Optionally, identifying the presence of the ligated DNA and / or DNA loop, particularly using techniques such as PCR (polymerase chain reaction), to identify the presence of specific chromosomal interactions including generating ligated nucleic acids (e.g., DNA) by the above.

[0028] These steps may be performed to detect chromosomal interactions for any of the embodiments described herein, such as to determine whether an individual is a member of a subgroup of ALS or Huntington's disease. The steps may also be performed to generate the first set and / or the second set of nucleic acids described herein.

[0029] PCR (polymerase chain reaction) can be used to detect or identify ligated nucleic acids. For example, the size of the produced PCR product can indicate the presence of a specific chromosomal interaction and can therefore be used to identify the state of a locus. In preferred embodiments, at least 1, 2, 3, 4, 5, 6, 7, or 8 of the primers or primer pairs shown in Table 2 or 7 are used in the PCR reaction. In another preferred embodiment, at least 1, 2, 3, 4, 5, 6, 7, or 8 of the primers or primer pairs shown in another table are used in the PCR reaction. One of ordinary skill in the art will be aware of the numerous restriction enzymes that can be used to cut DNA within the locus of interest. It will be apparent that the particular enzyme used depends on the locus being investigated and the sequence of the DNA located therein. A non-limiting example of a restriction enzyme that can be used to cut DNA as described in the present invention is TaqI.

[0030] Embodiments such as EpiSwitch™ technology The EpiSwitch™ technology relates to the use of microarray EpiSwitch™ marker data in the detection of epigenetic chromosomal conformation signatures specific to a phenotype. Embodiments such as EpiSwitch™ that utilize ligated nucleic acids in the manner described herein have several advantages. For example, nucleic acid sequences derived from a first set of nucleic acids of the invention hybridize or do not hybridize with a second set of nucleic acids, so their probabilistic noise is at a low level. This provides a binary result that enables a relatively straightforward way to measure complex mechanisms at the epigenetic level. The EpiSwitch™ technology also has a rapid processing time and low cost. In one embodiment, the processing time is 3 to 6 hours.

[0031] Samples and Sample Treatment The processes of the invention will typically be performed on a sample. The sample is typically considered to contain DNA from an individual. It is typically considered to contain normal cells. In one embodiment, the sample can be obtained by minimally invasive means, such as a blood sample. The DNA may be extracted and may be fragmented with standard restriction enzymes. This can pre-determine which chromosomal conformations are retained and detected by the EpiSwitch™ platform. By synchronizing chromosomal interactions between tissues and blood, including horizontal transfer, chromosomal interactions in tissues such as tissues related to a disease can be detected using a blood sample. For certain medical conditions such as cancer, genetic noise due to mutations can affect the chromosomal interaction "signal" in the relevant tissue, and therefore it is advantageous to use blood.

[0032] Properties of Nucleic Acids of the Invention The present invention relates, in this specification, to specific nucleic acids such as ligated nucleic acids that are described as being used in or generated by the process of the present invention. These may be the same as the first and second nucleic acids described herein, or may have any of the properties of the first and second nucleic acids. The nucleic acids of the present invention typically comprise two portions each containing a sequence from one of two regions of a chromosome that come together in a chromosomal interaction. Typically, each portion is at least 8, 10, 15, 20, 30, or 40 nucleotides in length, for example 10 - 40 nucleotides in length. Preferred nucleic acids comprise sequences derived from any of the genes described in any of the tables, in particular. Typically, preferred nucleic acids comprise the specific probe sequences described in Table 1 or 5, or fragments and / or homologs of such sequences. Another preferred nucleic acid comprises the specific probe sequences described in Tables 10, 11 or 12, or fragments and / or homologs of such sequences. Preferably, the nucleic acid is DNA. It is understood that where specific sequences are provided, the present invention may use the complementary sequences required in certain embodiments.

[0033] The primers shown in Table 2 or 7 may also be used in the present invention as described herein. In one embodiment, a primer is used that comprises either the sequence shown in Table 2 or 7; or a fragment and / or homolog of any of the sequences shown in Table 2 or 7.

[0034] Second set of nucleic acids - "index" sequences The second set of nucleic acids has the function of being a set of index sequences and is essentially a set of nucleic acid sequences suitable for identifying subgroup-specific sequences. They can represent "background" chromosomal interactions and may or may not be selected in any way. They are generally a subset of all possible chromosomal interactions.

[0035] The second set of nucleic acids can be derived by any suitable method. They may be computer-derived, or they may be based on chromosomal interactions in an individual. They typically represent a larger population group than the first set of nucleic acids. In one particular embodiment, the second set of nucleic acids represents all possible epigenetic chromosomal interactions in a particular gene set. In another particular embodiment, the second set of nucleic acids represents a majority of all possible epigenetic chromosomal interactions present in the population described herein. In one particular embodiment, the second set of nucleic acids represents at least 50% or at least 80% of the epigenetic chromosomal interactions in at least 20, 50, 100, or 500 genes, such as 20 - 100, or 50 - 500 genes.

[0036] The second set of nucleic acids typically represents at least 100 possible epigenetic chromosomal interactions that modify, regulate, or in some way mediate a disease state / phenotype in the population. The second set of nucleic acids may represent chromosomal interactions that affect a disease state (typically related to diagnosis or prognosis) in a species. The second set of nucleic acids typically includes sequences representing both epigenetic interactions associated with and not associated with an ALS subgroup. In one particular embodiment, the second set of nucleic acids is at least partially derived from naturally occurring sequences in the population and is typically obtained by in silico methods. The nucleic acids may further include single or multiple mutations compared to the corresponding portions of the nucleic acids present in naturally occurring nucleic acids. Mutations include deletions, substitutions, and / or additions of one or more nucleotide base pairs. In one particular embodiment, the second set of nucleic acids may include sequences representing homologs and / or orthologs having at least 70% sequence identity with the corresponding portions of nucleic acids present in a naturally occurring species. In another particular embodiment, at least 80% or at least 90% sequence identity is provided with the corresponding portions of nucleic acids present in a naturally occurring species.

[0037] Properties of the second set of nucleic acids In one particular embodiment, the second set of nucleic acids comprises at least 100 different nucleic acid sequences, preferably at least 1000, 2000, or 5000 different nucleic acid sequences, and up to 100,000, 1,000,000, or 10,000,000 different nucleic acid sequences. Typical numbers would be in the range of 100 - 1,000,000 different nucleic acid sequences, such as 1,000 - 100,000 different nucleic acid sequences. All or at least 90% or at least 50% of these would correspond to different chromosomal interactions.

[0038] In one particular embodiment, the second set of nucleic acids represents chromosomal interactions at at least 20 different loci or genes, preferably at least 40 different loci or genes, and more preferably at least 100, at least 500, at least 1000, or at least 5000 different loci or genes, such as 100 - 10,000 different loci or genes. The length of the second set of nucleic acids is suitable such that they specifically hybridize to the first set of nucleic acids according to Watson - Crick base pairing and enable the identification of chromosome - interaction - specific subgroups. Typically, the second set of nucleic acids will contain two portions whose sequences correspond to two chromosomal regions that come together in a chromosomal interaction. The second set of nucleic acids typically comprises nucleic acid sequences that are at least 10, preferably 20, and more preferably 30 bases (nucleotides) in length. In another embodiment, the nucleic acid sequences can be at most 500, preferably at most 100, and more preferably at most 50 base pairs in length. In a preferred embodiment, the second set of nucleic acids comprises nucleic acid sequences that are between 17 - 25 base pairs in length. In one embodiment, at least 100, 80%, or 50% of the second set of nucleic acid sequences have the lengths described above. Preferably, the different nucleic acids do not have any overlapping sequences, e.g., at least 100%, 90%, 80%, or 50% of the nucleic acids do not have the same sequence over at least 5 consecutive nucleotides.

[0039] Considering that a second set of nucleic acids acts as an "index", the same set of second nucleic acids may be used with various sets of first nucleic acids representing various characteristics of a subgroup, i.e., the second set of nucleic acids may represent a "universal" collection of nucleic acids that can be used to identify chromosomal interactions associated with various characteristics.

[0040] The first set of nucleic acids The first set of nucleic acids is typically from a subgroup associated with the diagnosis and prognosis of ALS or Huntington's disease. The first nucleic acids may have any of the characteristics and properties of the second set of nucleic acids referred to herein. The first set of nucleic acids is typically from a sample from an individual that has undergone the treatments and processes described herein, particularly the cross-linking and cleavage steps of EpiSwitch™. Typically, the first set of nucleic acids represents all or at least 80% or 50% of the chromosomal interactions present in a sample taken from an individual.

[0041] Typically, the first set of nucleic acids represents a smaller population of chromosomal interactions across the loci or genes represented by the second set of nucleic acids as compared to the chromosomal interactions represented by the second set of nucleic acids, i.e., the second set of nucleic acids represents a background or index set of interactions in a defined set of loci or genes.

[0042] The nucleic acid library Any of the types of nucleic acid populations referred to herein may exist in the form of a library containing at least 200, at least 500, at least 1000, at least 5000, or at least 10,000 different nucleic acids of that type, such as "first" or "second" nucleic acids. Such libraries may be in a form bound to an array.

[0043] Hybridization The present invention requires means for enabling nucleic acid sequences that are wholly or partially complementary from a first set of nucleic acids and a second set of nucleic acids to hybridize. In one embodiment, all of the first set of nucleic acids and all of the second set of nucleic acids are contacted in a single assay, i.e., a single hybridization step. However, any suitable assay can be used.

[0044] Labeled Nucleic Acids and Hybridization Patterns The nucleic acids referred to herein can preferably be labeled with an independent label such as a fluorophore (fluorescent molecule) or a radioactive label that aids in the detection of achieved hybridization. Certain labels can be detected under UV light. For example, the pattern of hybridization on the arrays described herein represents the differences in epigenetic chromosomal interactions between two subgroups, and thus provides a process for comparing epigenetic chromosomes and a determination of which epigenetic chromosomal interactions are specific to subgroups of the population of the present invention.

[0045] The term "pattern of hybridization" broadly covers the presence and absence of hybridization between a first set and a second set of nucleic acids, i.e., which nucleic acids from a particular first set hybridize with which nucleic acids from a particular second set, and it is not limited to any particular assay or technique, but requires a surface or array on which the "pattern" can be detected.

[0046] Selecting Subgroups with Special Characteristics The present invention provides a process comprising detecting the presence or absence of chromosomal interactions, typically 5 to 20 or 5 to 500 such interactions, preferably 20 to 300 or 50 to 100 interactions, to determine the presence or absence of characteristics associated with ALS or Huntington's disease in an individual. Preferably, the chromosomal interactions are in any of the genes referred to herein. In one embodiment, the classified chromosomal interactions are those represented by the nucleic acids of Table 1 or Table 5. Preferably, the classified chromosomal interactions are those represented by the nucleic acids of Table 10, Table 11 or Table 12. In the tables, the column named "detected loop" indicates which subgroup is detected by each probe (ALS or control). As can be seen, the process of the present invention can detect either the ALS subgroup and / or the control group (non-ALS) as part of the test.

[0047] Individuals to be tested Examples of species from which the individuals to be tested are derived are referred to herein. In addition, the individuals to be tested in the process of the present invention may be selected in several ways. The individuals may, for example, be susceptible to ALS or Huntington's disease.

[0048] Preferred gene regions, loci, genes, and chromosomal interactions For all embodiments of the present invention, preferred gene regions, loci, genes, and chromosomal interactions are referred to in tables, such as Table 1 and Table 5. Typically, in the process of the present invention, chromosomal interactions are detected from at least 1, 2, 3, 4, 5, 6, 7, or 8 related genes listed in Table 1 or Table 5. Preferably, the presence or absence of at least 1, 2, 3, 4, 5, 6, 7, or 8 related specific chromosomal interactions represented by the probe sequences of Table 1 or Table 5 is detected. Chromosomal interactions can be upstream or downstream of any of the genes referred to herein, for example from the coding sequence, for example 50 kb upstream or 20 kb downstream.

[0049] Preferably, at least 5, 8, 10, 15 or all of the chromosomal interactions in Table 1 are classified.

[0050] In typical embodiments, at least 5, 7, 8 or all of the chromosomal interactions in Table 5 are classified.

[0051] Preferably, at least 4, 6 or all of the chromosomal interactions in Table 10 are classified.

[0052] Typically, at least 1, 2 or all of the chromosomal interactions in Table 11 are classified.

[0053] Preferably, at least 4, 6 or all of the chromosomal interactions in Table 12 are classified.

[0054] Chromosomal interactions may be classified to determine the presence of a disease. Chromosomal interactions may be classified to determine whether the individual will progress quickly or slowly.

[0055] The specific probe and primer sequences (or derivatives such as fragments or homologs) disclosed in the tables can be used for any of the classification methods disclosed herein. Their use in diagnostic or prognostic methods is provided.

[0056] In one embodiment, the locus (including the gene and / or location where the chromosomal interaction is detected) may include a CTCF binding site. This is any sequence that can bind to the transcriptional repressor CTCF. The sequence may consist of, or include, the CCCTC sequence that may be present in 1, 2 or 3 copies at the locus. The CTCF binding site sequence may include the CCGCGNGGNGGCAG sequence (in IUPAC notation). The CTCF binding site may be within at least 100, 500, 1000 or 4000 bases of the chromosomal interaction, or within any of the chromosomal regions shown in Table 1 or Table 5.

[0057] In one embodiment, the chromosomal interactions detected are present in any of the genetic regions shown in Table 1 or 5. If the ligated nucleic acid is detected in the process, then subsequently, a sequence shown in any of the probe sequences of Table 1 or 5 can be detected.

[0058] Thus, typically, sequences from both regions of the probe (i.e., from both sites of the chromosomal interaction) can be detected. In a preferred embodiment, a probe comprising or consisting of a sequence that is the same as or complementary to the probe shown in any of the tables is used in the process. In some embodiments, a probe comprising a sequence homologous to any of the probe sequences shown in the table is used.

[0059] The tables provided herein Tables 1 and 5 show probe (Episwitch™ marker) data and gene data representing chromosomal interactions associated with ALS. The probe sequences show sequences that can be used to detect ligation products produced from both sites of the gene region that converges on the chromosomal interaction. That is, the probe contains a sequence complementary to the sequence of the ligation product. The two sets of the first Start-End positions indicate the probe positions, and the two sets of the second Start-End positions indicate the associated 4 kb regions. The following information is provided in the table of probe data. -HyperG_States: p-value for the probability of finding the number of significant Episwitch™ markers at a locus based on parameters of hypergeometric enrichment -Probe_Count_Total: Total number of Episwitch™ conformations tested at that locus -Probe_Count_Sig: Number of Episwitch™ conformations found to be statistically significant at that locus -FDR HyperG: Hypergeometric p-value corrected for multiple testing (false discovery rate) -Percent_Sig: Ratio of significant Episwitch™ markers to the number of markers tested at that locus -logFC: Base 2 logarithm of the epigenetic ratio (FC) -AveExpr: Average log2-expression for all array and channel probes -T: Moderated t-statistic -p-value: Unadjusted p-value -adj.p-value: Adjusted p-value or q-value -B: The B-statistic (lod or B) is the log odds that the gene is expressed separately. -FC: Non-log Fold Change -FC_1: Non-log Fold Change centered around zero -LS: Binary values are for FC_1 values. FC_1 values less than 1.1 are set to -1, and when the FC_1 value is greater than 1.1, it is set to 1. Between these values, the value is 0.

[0060] Tables 1 and 5 show that the p-values in the table of loci indicating genes where related chromosomal interactions occur are the same as the HyperG_Stats (p-value for the probability of finding the number of significant Episwitch™ markers at the locus based on the parameters of hypergeometric enrichment).

[0061] Probes are designed at positions 30 bp away from the Taq1 site. For PCR, the PCR primers are also designed to detect the ligation product, but the position from Taq1 is different.

[0062] Probe position: Start1 - 30 bases upstream of TaqI on fragment 1 End1 - TaqI restriction site on fragment 1 Sart2 - TaqI restriction site on fragment 2 End2 - 30 bases downstream of TaqI on fragment 2

[0063] 4 kb array position: Start1 - 4000 bases upstream of TaqI on Fragment 1 End1 - TaqI restriction site on Fragment 1 Sart2 - TaqI restriction site on Fragment 2 End2 - 4000 bases downstream of TaqI on Fragment 2

[0064] GLMNET value (λ (elastic - net) set to 0.5) related to the procedure for fitting the lasso whole or elastic - net regularization.

[0065] Tables 1 - 4 relate to the diagnosis of ALS, Tables 5 - 9 relate to the prognosis of ALS, and in one embodiment, the detections related to diagnosis are performed based on Tables 1 and 4, and the detections related to prognosis are performed based on Tables 5 - 9.

[0066] The LS column in Table 1 indicates 1 or - 1. 1 means present in ALS cases, and - 1 means not present in ALS cases.

[0067] The LS column in Table 5 indicates 1 or - 1. 1 means the marker is present in fast progressors and not in slow progressors, and - 1 means the marker is present in slow progressors but not in fast progressors.

[0068] Tables 10 and 11 relate to the prognosis of ALS. Table 10 includes the markers shown in the previous tables. The "detected loop" column in Table 11 means the marker is present in fast progressors and not in slow progressors.

[0069] Markers are uniquely identified in the table by reference to the relevant probes for the products of the EpiSwitch3C method. In the case of hydrolysis probes, these are on the Taq site (TCGA) and cover both genomic sites in the EpiSwitch interaction. The same junctions as the 60-base array probes (30 bases at each end of the sequence tag) are evaluated, but over the adjusted lengths of the sequences at both ends. An example of this is provided in Table 19.

[0070] Preferred embodiments for sample preparation and chromosome interaction detection Methods for preparing samples and methods for detecting chromosome conformation are described herein. Optimized (non-conventional) versions of these methods can be used, as described in this section for example.

[0071] Typically, the sample will contain at least 2×10 5 cells. The sample can contain up to 5×10 5 cells. In one embodiment, the sample will contain 2×10 5 to 5.5×10 5 cells.

[0072] Crosslinking of epigenetic chromosome interactions present at chromosomal loci is described herein. This can be performed before cell lysis occurs. Cell lysis can be performed for 3 to 7 minutes, such as 4 to 6 or about 5 minutes. In some embodiments, cell lysis is performed for at least 5 minutes and less than 10 minutes.

[0073] Digestion of DNA with restriction enzymes is described herein. Typically, DNA restriction is performed for a period of about 10 to 30 minutes, such as about 20 minutes, at about 55°C to about 70°C, such as about 65°C.

[0074] Preferably, a high-frequency cutter restriction enzyme is used that yields fragments of ligated DNA having an average fragment size of up to 4000 base pairs. Optionally, the restriction enzyme yields fragments of ligated DNA having an average fragment size of about 200 - 300 base pairs, such as about 256 base pairs. In one embodiment, typical fragment sizes are from 200 base pairs to 4000 base pairs, such as 400 - 2,000 or 500 - 1,000 base pairs.

[0075] In one embodiment of the EpiSwitch method, the DNA precipitation step is not performed between the DNA restriction digestion step and the DNA ligation step.

[0076] DNA ligation is described herein. Typically, DNA ligation is performed for 5 - 30 minutes, such as about 10 minutes.

[0077] Proteins in the sample can be enzymatically digested, for example, using a protease, optionally proteinase K. The protein can be enzymatically digested for a period of about 30 minutes to 1 hour, such as about 45 minutes. In one embodiment, there is no step of crosslink reversal or phenol DNA extraction after digestion of the protein, such as proteinase K digestion.

[0078] In one embodiment, PCR detection preferably comprises a binary readout for the presence / absence of ligated nucleic acid and can detect a single copy of the ligated nucleic acid.

[0079] The processes and uses of the present invention The process of the present invention can be described in various ways. It is a method for producing a ligated nucleic acid comprising: (i) a step of in vitro cross-linking chromosomal regions that converge on chromosomal interactions; (ii) a step of excising the cross-linked DNA or subjecting it to restriction digestion cleavage; and (iii) a step of ligating the ends of the cross-linked and cleaved DNA to form a ligated nucleic acid, wherein the chromosomal state at a locus can be determined using the detection of the ligated nucleic acid, and preferably, - the locus can be any of the loci, regions or genes mentioned in Table 1 or Table 5, - and / or the chromosomal interaction can be any chromosomal interaction corresponding to any of the probes mentioned herein or disclosed in Table 1 or Table 5, and / or - the ligated product can have or contain a sequence that is the same as or homologous to any of the probe sequences disclosed in Table 1 or Table 5; or (ii) a sequence that is complementary to (ii). It can be described as a method.

[0080] The process of the present invention is a method for detecting chromosomal states representing various subgroups within a population, comprising a step of determining whether a chromosomal interaction is present or absent within a defined epigenetically active (disease-related) region of the genome, preferably, - the subgroup is defined by the presence or absence of ALS or a feature (such as prognosis or progression) related to ALS, and / or - the chromosomal state can be at any of the loci, regions or genes mentioned in Table 1 or Table 5; and / or - the chromosomal interaction can be any of those mentioned in Table 1 or Table 5 or any corresponding to any of the probes disclosed in that table. It can be described as a method.

[0081] The present invention includes the detection of chromosomal interactions at any locus, gene or region referred to in Table 1 or 5. The present invention includes the use of the nucleic acids and probes described herein for detecting chromosomal interactions, for example, the use of at least 1, 2, 4, 6 or 8 such nucleic acids or probes for detecting chromosomal interactions at at least 1, 2, 4, 6 or 8 different loci or genes. The present invention includes the detection of chromosomal interactions using any of the primers or primer pairs listed in Table 2 or 7, or variants of these primers described herein (sequences including primer sequences or fragments and / or homologs of primer sequences).

[0082] When analyzing whether a chromosomal interaction occurs "within" a defined gene, region, or location, both portions of the chromosomes that come together in the interaction are within the defined gene, region, or location, or in some embodiments only a portion of the chromosome is within the defined gene, region, or location.

[0083] Use of the methods of the invention for identifying new therapies Knowledge of chromosomal interactions can be used to identify new therapies for disease states. The present invention provides chromosomal methods and uses as defined herein for identifying or designing novel therapeutic agents (including therapies related to prognosis) for ALS.

[0084] Homolog Homologs of polynucleotide / nucleic acid (e.g., DNA) sequences are referred to herein. Such homologs typically have at least 70% homology, preferably at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% homology, over a region of, for example, at least 10, 15, 20, 30, 100 or more contiguous nucleotides, or over a portion of a nucleic acid derived from a region of a chromosome involved in chromosomal interactions. Homology can be calculated based on nucleotide identity (sometimes referred to as "hard homology").

[0085] Thus, in certain embodiments, homologs of polynucleotide / nucleic acid (e.g., DNA) sequences are referred to herein with reference to % sequence identity. Typically, such homologs have at least 70% sequence identity, preferably at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% sequence identity, over a region of, for example, at least 10, 15, 20, 30, 100 or more contiguous nucleotides, or over a portion of a nucleic acid derived from a region of a chromosome involved in chromosomal interactions.

[0086] For example, the UWGCG package provides the BESTFIT program (e.g., as used in its default settings) that can be used to calculate homology and / or % sequence identity (Devereux et al (1984) Nucleic Acids Research 12, p387-395). Homology and / or % sequence identity can be calculated and / or sequences can be aligned using, for example, the PILEUP and BLAST algorithms as described in Altschul S. F. (1993) J Mol Evol 36:290-300; Altschul, S, F et al (1990) J Mol Biol 215:403-10 (identifying equivalent or corresponding sequences, etc., typically in their default settings).

[0087] Software for performing BLAST analysis is publicly available through the National Center for Biotechnology Information. This algorithm first identifies high-scoring sequence pairs (HSPs) by identifying short words of length W in the query sequence that match when aligned with words of the same length in the database sequence or satisfy a threshold score T having some positive value. T is referred to as the neighborhood word score threshold (Altschul et al, supra). These initial neighborhood word hits act as seeds for beginning a search to find the HSPs that contain them. The word hits are extended in both directions along each sequence as long as the cumulative alignment score can increase. The extension of the word hits in each direction stops when the cumulative alignment score drops by an amount X from its maximum reached value; when the cumulative score goes to zero or below due to the accumulation of one or more residue alignments taking negative scores; or when the end of either sequence is reached. The W, T, and X of the BLAST algorithm parameters determine the sensitivity and speed of the alignment. The BLAST programs use, by default, an 11 word length (W), a 50 BLOSUM62 score matrix (see Henikoff and Henikoff (1992) Proc. Natl. Acad. Sci. USA 89:10915-10919) alignment (B), an expectation value (E) of 10, M = 5, N = 4, and comparison of both strands.

[0088] The BLAST algorithm performs a statistical analysis of similarity between two sequences; see, e.g., Karlin and Altschul (1993) Proc. Natl. Acad. Sci. USA 90:5873-5787. One measure of similarity provided by the BLAST algorithm is the minimum sum probability (P(N)), which provides an indication of the probability that a match between two polynucleotide sequences would occur by chance. For example, if the minimum sum probability of comparing a first sequence to a second sequence is less than about 1, preferably less than about 0.1, more preferably less than about 0.01, and most preferably less than about 0.001, the sequences are considered to be similar to each other.

[0089] Identical sequences typically differ by only 1, 2, 3, 4, or more bases, such as less than 10, 15, or 20 bases (which can be nucleotide substitutions, deletions, or insertions). These changes can be measured over any of the regions mentioned above with respect to the step of calculating percent homology and / or sequence identity.

[0090] The homology of a "primer pair" can be calculated, for example, by considering the two sequences as a single sequence (as if the two sequences were concatenated), and then comparing it to another primer pair that is again considered as a single sequence.

[0091] array A second set of nucleic acids can be bound to the array, and in one embodiment, there are at least 15,000, 45,000, 100,000, or 250,000 different second nucleic acids bound to the array, preferably corresponding to at least 300, 900, 2000, or 5000 loci. In one embodiment, one or more or all of the various populations of the second nucleic acids are bound to a number of individual regions on the array that exceeds one, are in effect repeated on the array, and allow for error detection. The array is based on the Agilent SurePrint G3 Custom CGH microarray platform. Detection of the binding of the first nucleic acid to the array can be performed by a two-color system.

[0092] Therapeutic agent Therapeutic agents are mentioned herein. The present invention provides such agents for use in preventing or treating ALS in an individual, for example an individual identified by the process of the present invention. This may involve administering to the individual in need a therapeutically effective amount of the agent. The present invention provides the use of the agent in the manufacture of a medicament for preventing or treating ALS in an individual.

[0093] Preferred therapeutic agents for ALS are as follows. Rilutek: This drug is usually taken as a pill and suppresses the progression of the disease by reducing the level of the messenger (glutamate) in the brain. Glutamate is present at high levels in ALS patients. Radicava: This drug reduces the daily performance decline related to ALS. The drug is usually administered to patients by intravenous infusion once a month for 10 - 14 consecutive days. Alimocromol: This drug functions as an inducer of the heat shock response of motor neurons and prevents neuropathy and cell death. Talampanel: This drug reduces muscle strength decline and the rate of symptom progression. Beta-lactam antibiotics: These antibodies such as penicillin and cephalosporin maintain muscle stability and extend lifespan by upregulating the level of GLT1, glial glutamate transporter 1. Bromocriptine: This drug is a free radical scavenger that inhibits oxidative cell death caused by stress. Pramipexole and dextropramipexole: These drugs act as dopamine agonists and have free radical scavenging functions. These drugs are involved in mitochondrial dysfunction. Dextropramipexole is the optical isomer of pramipexole. Stem cell therapy: The growth of stem cells reduces the progression of neuronal diseases or replaces motor neurons. Stem cells can create spinal cord motor neurons, extend axons, receive synapses with muscles, and generate. Mesenchymal stem cells (MSCs) derived from adult stem cells release trophic factors, anti-inflammatory cytokines, and immunomodulatory chemokines, delaying the progression of the disease. Immunotherapy: Antibody therapies such as injecting the D3H5 antibody via the ICV route maintain body weight for a long time and extend the lifespan of ALS transgenic mouse models.

[0094] The following is a list of therapeutic agents for Huntington's disease. These can help reduce some symptoms of movement and mental disorders.

[0095] Medications for movement disorders · Tetrabenazine (Xenazine) · Antipsychotics such as haloperidol (Haldol), chlorpromazine, risperidone (Risperdal), and quetiapine (Seroquel) · Other medications include amantadine, levetiracetam (Keppra, etc.), and clonazepam (Klonopin). Medications for neuropathy · Antidepressants include citalopram (Celexa), escitalopram (Lexapro), fluoxetine (Prozac, Sarafem), and sertraline (Zoloft). · Antipsychotics include quetiapine (Seroquel), risperidone (Risperdal), and olanzapine (Zyprexa). · Mood stabilizers include valproate (Depacon), carbamazepine (Carbatrol, Epitol, Tegretol), and lamotrigine (Lamictal).

[0096] Early diagnosis of Huntington's disease helps manage the treatment of symptoms as the disease progresses.

[0097] (For ALS or Huntington's disease) The formulation of a drug is believed to depend on the nature of the drug. The drug is provided in the form of a pharmaceutical composition containing the drug and a pharmaceutically acceptable carrier or diluent. Suitable carriers and diluents include isotonic saline solutions, such as phosphate buffered saline. Typical oral dosage compositions include tablets, capsules, liquid solutions, and liquid suspensions. The drug can be formulated for parenteral, intravenous, intramuscular, subcutaneous, transdermal, or oral administration.

[0098] The dosage of a drug can be determined according to various parameters, especially the substance used; the age, weight, and medical condition of the individual being treated; the route of administration; and the regimen required. A physician can determine the required route of administration and dosage for any particular drug. However, an appropriate dosage can be, for example, 0.1 - 100 mg / kg body weight, such as 1 - 40 mg / kg body weight to be taken 1 - 3 times a day.

[0099] The forms of substances referred to in this specification Any of the substances such as nucleic acids or therapeutic agents referred to in this specification may be in a purified or isolated form. They may be in a form different from that found in nature, for example, they may exist in combination with other substances that do not occur together naturally. Nucleic acids (including a portion of the sequences defined herein) may have sequences different from those found in nature, for example, in the sequences described in the section on homology, having at least 1, 2, 3, 4, or more nucleotide changes. Nucleic acids may have heterologous sequences at the 5' or 3' ends. Nucleic acids may be chemically different from those found in nature, for example, they may be modified in some way, preferably still allowing Watson-Crick base pairing. Where appropriate, nucleic acids are provided in double-stranded or single-stranded form. The present invention provides all of the specific nucleic acid sequences referred to herein in single-stranded or double-stranded form, and thus includes the complementary strand to any of the disclosed sequences.

[0100] The present invention also provides kits for performing any of the methods of the present invention, including the detection of chromosomal interactions associated with a particular subgroup. Such kits may include specific binding agents capable of detecting the associated chromosomal interactions, such as agents capable of detecting the ligated nucleic acids generated by the methods of the present invention. Preferred agents present in the kits include probes capable of hybridizing to ligated nucleic acids or primer pairs capable of amplifying ligated nucleic acids in a PCR reaction, for example, as described herein.

[0101] The present invention also provides devices capable of detecting associated chromosomal interactions. The device preferably includes any specific binding agent, probe, or primer pair capable of detecting chromosomal interactions, such as any of the agents, probes, or primer pairs described herein.

[0102] Detection methods In one embodiment, the quantitative detection of ligated sequences related to chromosomal interactions is carried out using a detectable probe upon activation during a PCR reaction, the ligated sequences comprising sequences from two chromosomal regions that converge in an epigenetic chromosomal interaction, the method comprising contacting the ligated sequences with the probe during the PCR reaction and detecting the degree of activation of the probe, the probe binding to the ligation site. This method typically uses dual-labeled fluorescent hydrolysis probes to enable the detection of specific interactions in a MIQE-compliant manner.

[0103] Probes are generally labeled with a detectable label having an inactive state and an active state such that they are only detected when activated. The degree of activation will be related to the degree of template (ligation product) present in the PCR reaction. Detection may be carried out for all or part of the PCR, for example, during at least 50% or 80% of the PCR cycles.

[0104] The probe can comprise a fluorophore covalently attached to one end of the oligonucleotide and a quencher attached to the other end of the nucleotide, the fluorescence of the fluorophore being quenched by the quencher. In one embodiment, the fluorophore is attached to the 5' end of the oligonucleotide and the quencher is covalently attached to the 3' end of the oligonucleotide. Fluorophores that can be used in the methods of the present invention include FAM, TET, JOE, Yakima Yellow, HEX, Cy3, ATTO 550, TAMRA, ROX, Texas Red, Cy3.5, LC610, LC 640, ATTO 647N, Cy5, Cy5.5 and ATTP 680. Quenchers that can be used in conjunction with suitable fluorophores include TAM, BHQ1, DAB, Eclip, BHQ2 and BBQ650, optionally where the fluorophore is selected from HEX, Texas Red and FAM. Preferred combinations of fluorophore and quencher include FAM and BHQ1 and Texas Red and BHQ2.

[0105] Use of probes in qPCR assays The hydrolysis probe of the present invention is typically a temperature gradient optimized with a concentration-matched negative control. Preferably, a single-step PCR reaction is optimized. More preferably, a standard curve is calculated. The advantage of using a specific probe that ligates and binds in the ligation of ligated sequences is that the specificity of the ligated sequence can be achieved without using a nested PCR approach. The methods described herein enable accurate and precise quantification of low-copy number targets. The ligated sequence of the target can be purified, for example, gel purified, prior to temperature gradient optimization. The ligated sequence of the target can be sequenced. Preferably, the PCR reaction is performed using about 10 ng, or 5-15 ng, or 10-20 ng, or 10-50 ng, or 10-200 ng of template DNA. The forward primer and the reverse primer are designed, for example, to be complementary to the sequence such that one primer binds to one sequence of the chromosomal region represented by the ligated DNA sequence and the other primer binds to another chromosomal region represented by the ligated DNA sequence.

[0106] Selection of Ligated DNA Targets The present invention includes the selection of primers and probes for use in the PCR methods defined herein, which includes selecting primers based on the ability to bind and amplify ligated sequences and, in particular, selecting characteristics based on the probe sequence of the target sequence that binds at the curvature of the target sequence.

[0107] Probes are typically designed / selected to bind to a ligated array in which restriction fragments spanning a restriction site are juxtaposed. In one embodiment of the invention, the predicted curvature of a ligated array that may be associated with a particular chromosomal interaction is calculated, for example, using a particular algorithm referenced herein. Curvature can be expressed in degrees per helical turn, for example 10.5° per helical turn. The ligated array is selected for a target having a curvature trend peak score of at least 5°, usually at least 10°, 15° or 20° per helical turn, for example 5° - 20° per helical turn. Preferably, the curvature trend score per helical turn is calculated for at least 20, 50, 100, 200 or 400 bases, for example 20 - 400 bases, upstream and / or downstream of the ligation site. Thus, in one embodiment, the target sequence of the ligation product has any of these levels of curvature. The target sequence can also be selected based on the lowest thermodynamic structural free energy.

[0108] Preferred ALS embodiments 1. A process for detecting the state of a chromosome indicative of a disease subgroup in a population, the process comprising detecting whether a chromosomal interaction related to the state of the chromosome is present or absent within a defined region of the genome, wherein the disease subgroup is an amyotrophic lateral sclerosis (ALS) subgroup; and - The chromosomal interaction is optionally identified by a method for determining which chromosomal interactions are associated with the chromosomal state corresponding to the ALS subgroup of the population, the method comprising contacting a first set of nucleic acids from subgroups having various states of chromosomes with a second set of index nucleic acids and allowing complementary sequences to hybridize, wherein the nucleic acids in the first and second sets of nucleic acids represent ligation products containing sequences from both chromosomal regions that come together in the chromosomal interaction, and the hybridization pattern between the first and second sets of nucleic acids allows determination of which chromosomal interactions are specific to the ALS subgroup; and - The chromosomal interaction is (i) present in any of the regions or genes listed in Table 1 or Table 5; and / or (ii) corresponding to any of the chromosomal interactions represented by any of the probes shown in Table 1 or Table 5; and / or (iii) corresponding to any of the chromosomal interactions shown in Table 10 or Table 11; and / or (iv) present in a 4,000 base region containing (i), (ii) or (iii) or adjacent to (i), (ii) or (iii) Process.

[0109] 2. The process according to item 1, carried out for diagnosing ALS or determining the prognosis of ALS.

[0110] 3. A specific combination of chromosomal interactions is (i) including all chromosomal interactions represented by the probes in Table 1 or Table 5; or (ii) including at least 4, 5, 6 or 7 chromosomal interactions represented by the probes in Table 1 or Table 5 (iii) present together in at least 4, 5, 6 or 7 regions or genes listed in Table 1 or Table 5; or (iv) Comprising chromosomal interactions represented by the probes in Table 1 or Table 5, or comprising at least 4, 5, 6, or 7 chromosomal interactions present in a 4,000-base region adjacent to the chromosomal interactions represented by the probes in Table 1 or Table 5, or (v) Comprising all chromosomal interactions shown in Table 10 or Table 11; or (vi) Comprising at least 4, 5, 6, or 7 chromosomal interactions shown in Table 10 The process according to item 1 or 2 classified therein.

[0111] Item 4. The chromosomal interaction is - In a sample from an individual, and / or - By detecting the presence or absence of a DNA loop at the site of the chromosomal interaction, and / or - By detecting the presence or absence of distal regions of chromosomes that are brought together in the chromosomal conformation, and / or - By detecting the presence of ligated nucleic acids generated during the classification, the sequences of which comprise two regions each corresponding to a region of the chromosome that is brought together in the chromosomal interaction, wherein the detection of the ligated nucleic acids preferably uses (i) a probe having at least 70% identity to any of the specific probe sequences described in Table 1 or Table 5, and / or (ii) a primer pair having at least 70% identity to any of the primer pairs in Table 2 or Table 7 The process according to any one of items 1 to 3 classified therein.

[0112] Item 5. The second set of nucleic acids is from a larger group of individuals than the first set of nucleic acids; and / or - The first set of nucleic acids is from at least 8 individuals; and / or - The first set of nucleic acids is from at least 4 individuals from a first subgroup, and preferably from at least 4 individuals from a second subgroup that does not overlap with the first subgroup; and / or - The process is carried out to select an individual for medical treatment The process according to any one of items 1 to 4.

[0113] Item 6. The second set of nucleic acids represents a non-selected group; and / or - The second set of nucleic acids is bound to array at a defined position; and / or - The second set of nucleic acids represents chromosomal interactions in at least 100 different genes; and / or - The second set of nucleic acids contains at least 1,000 different nucleic acids representing at least 1,000 different chromosomal interactions; and / or - The first set of nucleic acids and the second set of nucleic acids contain at least 100 nucleic acids having a length of 10 to 100 nucleotide bases The process according to any one of items 1 to 5.

[0114] Item 7. The first set of nucleic acids is (i) A step of cross-linking chromosomal regions aggregated in chromosomal interactions; (ii) A step of cleaving the cross-linked region, optionally by restriction digestion with an enzyme; and (iii) A step of ligating the cross-linked and cleaved DNA ends to generate a first set of nucleic acids (especially including the ligated DNA) The process according to any one of items 1 to 6 obtained in a process including

[0115] Item 8. The process according to any one of items 1 to 7, wherein at least 5 to 9 different chromosomal interactions are preferably classified in 5 to 9 different regions or genes.

[0116] Item 9. The defined region of the genome is (i) Containing a single nucleotide polymorphism (SNP); and / or (ii) Expressing a microRNA (miRNA); and / or (iii) Expressing a non-coding RNA (ncRNA); and / or expressing a nucleic acid sequence encoding at least 10 consecutive amino acid residues; and / or (v) expressing a control element; and / or (vii) comprising a CTCF binding site The process according to any one of items 1 to 8.

[0117] 10. A method of identifying or designing a therapeutic agent for treating ALS by selecting an agent capable of causing a change in chromosomal interaction and thereby producing a therapeutic effect, - the chromosomal interaction is represented by any of the probes in Table 1 or Table 5; and / or - the chromosomal interaction is present in any of the regions or genes listed in Table 1 or Table 5; and / or - the chromosomal interaction is any one of the interactions shown in Table 10 or Table 11, optionally: - the chromosomal interaction has been identified by a method for determining which chromosomal interactions are associated with the chromosomal state defined in claim 1; and / or - a change in chromosomal interaction is monitored using (i) a probe having at least 70% identity to any of the probe sequences described in Table 1 or Table 5, and / or (ii) a primer pair having at least 70% identity to any of the primer pairs in Table 1 or Table 5; and / or - a candidate agent is contacted with a cell and the chromosomal interaction in that cell is monitored to determine whether the candidate agent can treat ALS Method.

[0118] 11. (iv) Detection of chromosomal interaction, - the chromosomal interaction is represented by a probe in Table 1 or Table 5, and / or - present in any of the regions or genes described in Table 1 or Table 5; and / or - the chromosomal interaction is any one of the interactions shown in Table 10 or Table 11 Detection of chromosomal interactions, or (v) A probe having at least 70% identity with any of the probe sequences described in Table 1 or Table 5, or (vi) A primer pair having at least 70% identity with any of the primers identified in Table 2 or Table 7 For use in identifying or designing a therapeutic agent for ALS.

[0119] Item 12. Administration of a candidate agent, and use of the detection of the chromosomal interaction, the probe or the primer pair to detect whether there is a change in the chromosomal state, thereby determining whether the candidate agent is a therapeutic agent. The use according to Item 11 for identifying a therapeutic agent, wherein the use is optionally performed in vitro, preferably intracellularly.

[0120] Item 13. A therapeutic agent for ALS for use in a method of preventing or treating ALS in an individual identified as in need of a therapeutic agent by the process according to any of Items 1 to 9.

[0121] Item 14. The classification or detection includes specific detection of a ligation product by quantitative PCR (qPCR) using primers capable of amplifying the ligation product and a probe that binds to the ligation site during the PCR reaction. The probe includes a sequence complementary to a sequence derived from each chromosomal region where chromosomal interactions accumulate. Preferably, the probe Is an oligonucleotide that specifically binds to the ligation product, and / or A fluorophore covalently attached to the 5' end of the oligonucleotide, and / or A quencher covalently attached to the 3' end of the oligonucleotide, and Optionally, The fluorophore is selected from HEX, Texas Red and FAM; and / or The probe is a nucleic acid sequence 10 to 40 nucleotide bases in length, preferably 20 to 30 nucleotide bases in length The process, method or use according to any one of claims 1 to 13, which includes

[0122] Special embodiments In one embodiment, only chromosomal interactions are classified / detected, and extrachromosomal interactions (between different chromosomes) are not classified / detected.

[0123] Publications The content of all publications mentioned in this specification is incorporated herein by reference and can be used to further define the characteristics related to the present invention.

[0124] Specific embodiments The EpiSwitch™ platform technology detects epigenetic regulatory signatures for regulatory changes between normal and abnormal conditions at a locus. The EpiSwitch™ platform identifies and monitors the fundamental epigenetic level of gene regulation associated with the regulatory higher-order structure of the human chromosome, also known as the chromosomal conformation signature. Chromosomal signatures are individual primary steps in the cascade of gene deregulation. They are higher-order biomarkers with a unique set of advantages over biomarker platforms that utilize late epigenetic and gene expression biomarkers such as DNA methylation and RNA profiling.

[0125] EpiSwitch™ array assay The custom EpiSwitch™ array screening platform has four densities of unique chromosomal conformations of 15K, 45K, 100K, and 250K, and each chimeric fragment is repeated 4 times on the array, creating effective densities of 60K, 180K, 400K, and 100,000 respectively.

[0126] Custom-designed EpiSwitch™ array The 15K EpiSwitch™ array can screen the entire genome containing approximately 300 loci interrogated by the EpiSwitch™ biomarker discovery technology. The EpiSwitch™ array is constructed on the Agilent SurePrint G3 Custom CGH microarray platform; this technology provides probes at four densities of 60K, 180K, 400K, and 100,000. Since each EpiSwitch™ probe is presented as a quadruplicate, the density per array is reduced to 15K, 45K, 100K, and 250K, thus enabling a statistical assessment of reproducibility. The average number of potential EpiSwitch™ markers interrogated per locus is 50; therefore, the number of loci that can be examined is 300, 900, 2000, and 5000.

[0127] EpiSwitch™ Custom Array Pipeline The EpiSwitch™ array is a two-color system with a set of samples. After EpiSwitch™ library generation, one sample (control), which is to be labeled with Cy5 and compared / analyzed, is labeled with Cy3. The array is scanned using an Agilent SureScan scanner, and the resulting features are extracted using Agilent Feature Extraction software. The data is then processed using the EpiSwitch™ array processing script in R. The array is processed using the standard two-color package in Bioconductor in R: Limma * is used. Array normalization is performed using the within-array normalization function in Limma * and this is done against on-chip Agilent positive controls and EpiSwitch™ positive controls. The data is filtered based on the Agilent Flag call, Agilent control probes are removed, and technical replicate probes are averaged so that they are Limma *It is analyzed using. Probes are modeled based on their differences between two scenarios that have been compared and then corrected by using the False Discovery Rate. Probes with a coefficient of variation (CV) of < 30% that are < 1 or > 1 and pass an FDR p-value of p = 0.01 are used for further screening. To reduce the probe set, further multi-factor analysis is performed using the FactorMineR package in R.

[0128] * Note: LIMMA is a linear model and empirical Bayes method for evaluating differential expression in microarray experiments. Limma is an R package for the analysis of gene expression data generated by microarray or RNA-Seq.

[0129] The pool of probes is initially selected for final selection based on the parameters of adjusted p-value, FC, and < 30% CV (any cut-off point). Further analysis and the final list are derived based on only the first two parameters (adjusted p-value; FC).

[0130] Other embodiments 1. A process for detecting the chromosomal state representing a disease subgroup of amyotrophic lateral sclerosis (ALS), comprising detecting whether chromosomal interactions related to the chromosomal state are present or absent within a defined region of the genome, the following (a) Sequences: TCTTGTACACGGTTGGTGGT and TGTCACCTATGTGCTGAGTACTGG, TGACGAAGAAGCAATCCCTGGT and GGACCTACCTCCACTGGGTTG, CCGTGCCATATCCTCTGATTTATGC and GGCTGACCTTCAACAGATTCGC, ACTTCTTCCCAAGTCACTTTTTGC and TGGCCATCTTGCTTTGCCTC, GGATATGCAGTTTTCCTGGCACTAC and CATGCTAGGGCCGAGTAATCATCT, GCAGCACACAGGGAACTCTCTT and TTGTTGAGCCCAGCAATTCCTTT, CAGCCACTGTAGAGAGCAGT and TAACCCACCAGCAGCAAGGT, and AGTAGCTTCCCTGTTAGAGGTCTTG and AGCCAGTGACTCCACAACTTCTT A group consisting of chromosomal interactions each represented by a primer pair having At least 4 chromosomal interactions selected from

[0131] 2.(b) Probe sequence: TCACCACACATCACCCCCTTGCTCCTCCTCGAGTCTTGGTGACCACAACAGGGTGCCACC, GAGGTGGGTGAATCATGAGGTCAAGGGTTCGACAATAGTTGAGAATCTCCAACCACCTGG, GGCCTTATAGTCAGCTGATCAGGTGAAATCGATTGGTCCTTAGGATCAGCTACCATTTGC, GAGGCAGGCGGATCACAAAGTCAAAAGATCGATAACTTCAATAATAGTTACAGATGCAAA, AGCACCATATCTGGGATGTAGCTATTGCTCGAGATTGCAGTGAGCTGTGATCACACCTCT, TCTTCCCTCTTTTTAAAACCACCATTCATCGACCCCACACATCCTGTGCCACTCTACTGC, TAACCATTATGCATCACTAACATAGCATTCGATATGATATGCTCAGTTTAGTTAGGGAAA, GGCTCAGGAAGAGAACTATTTGTCTCTTTCGACACGCACATGCAGGACACTCACACGTAG, GTTGGGTGGATCCCTTGAGCTCAGGAATTCGAAGAATGATTTTTCAGCCCGTGTGGAAGG、 ATCAAAAGAAAATAGATACTTGTCTTACTCGAGTTGAATAAAATCCTCAGCTTTCTGTCC、 AAAAGAAACTGTGAAAAGTTGTCACATTTCGATTAAATCCAAAAAGGTCTTCTATGAGGC、 TTAAAAGTATAGTAGTTGGCATTAACATTCGACCTTTTTCTGTTTCAGTAACCAACCCAG、 CATCAACTAATAGTTAAACATTATAATATCGACTGAAGACCTTTCATACTGTAAGATTCA、 CATCAACTAATAGTTAAACATTATAATATCGAGTCTGCAGTGAGCTGAGATCACACTGCC、 TTATTCCTTTCCAAATAGTTAAAATTATTCGAAACTTTTAAGAATCAATATAAAATTTCC、 CATAATTATAAATTAAAAAATGACACTATCGATTATGTCCAGTGTTTCTTGGTTGGTGTC, and CAGAGCACTAAGATAGACTTCTAAGGTTTCGAGGCATATAGCTCCAGCTGTATTGAGGTA at least 4 chromosomal interactions selected from the group consisting of the chromosomal interactions represented by each of the above, and / or (c) Probe sequence: TATATTTAAAAATACATACTGGTATACATCGATCTCATGACTTTGCTATTATGCATAGTG、 CCCCAGCCCAGCAACCTGGCTCACCTGATCGAGTACATCTTCAAGCCATCCTGTGTGCCC、 TATAAATAATACAGCTCTATTTGCCTACTCGATTAAAGAATCATATTATATCCTTAATTC、 CACTATGCATAATAGCAAAGTCATGAGATCGAAAATGTTTGTCAAGCAGTAGGTTTTGGG、 GATTTTAGAATCTCTAACAAGGCTGCAATCGAGGTTAGCTGCTGCAGAAAGAAGAGAAAA、 GGTGACTAAAATGAGATTGCATTTTCTTTCGACCATTTGGCCAGCATGCCAAACACTGGT、 TAACCTCTCCTTTCTTAGGTTCTCCATATCGATAGAAAATTGTCTGCAGCCCTTAATGCC、 AGCTCACGGTCAGTGCCGTTCCGTTTGCTCGAATAAAGAACAAGGACCTTAAAAAATAGA, and AAAGTCATACAACTACTATGTAAGATATTCGAATACCTGTTAGAATAGGTGAAGGTTTAT at least 4 chromosomal interactions selected from the group consisting of chromosomal interactions represented by each of the following, and / or (d) Sequence: TCTTGTACACGGTTGGTGGT and TGTCACCTATGTGCTGAGTACTGG, GGATATGCAGTTTTCCTGGCACTAC and CATGCTAGGGCCGAGTAATCATCT, and AGTAGCTTCCCTGTTAGAGGTCTTG and AGCCAGTGACTCCACAACTTCTT at least 2 chromosomal interactions selected from the group consisting of chromosomal interactions represented by each primer pair having The process according to item 1 above, further classified.

[0132] 3. The process according to item 2 above, wherein the classified chromosomal interactions include at least 5 chromosomal interactions selected from any of the groups (a) to (c).

[0133] 4. Chromosomal interactions are -classified by detecting the presence or absence of DNA loops at sites of chromosomal interactions, and / or -classified by detecting the presence or absence of distal regions of chromosomes that are brought together in a chromosomal conformation, and / or -classified by detecting the presence of a ligated nucleic acid having a sequence containing two regions each corresponding to a region of a chromosome that is brought together in a chromosomal interaction, where the detection of the ligated nucleic acid is (i) specific probe sequences: TCACCACACATCACCCCCTTGCTCCTCCTCGAGTCTTGGTGACCACAACAGGGTGCCACC, GAGGTGGGTGAATCATGAGGTCAAGGGTTCGACAATAGTTGAGAATCTCCAACCACCTGG, GGCCTTATAGTCAGCTGATCAGGTGAAATCGATTGGTCCTTAGGATCAGCTACCATTTGC, GAGGCAGGCGGATCACAAAGTCAAAAGATCGATAACTTCAATAATAGTTACAGATGCAAA, AGCACCATATCTGGGATGTAGCTATTGCTCGAGATTGCAGTGAGCTGTGATCACACCTCT, TCTTCCCTCTTTTTAAAACCACCATTCATCGACCCCACACATCCTGTGCCACTCTACTGC, TAACCATTATGCATCACTAACATAGCATTCGATATGATATGCTCAGTTTAGTTAGGGAAA, GGCTCAGGAAGAGAACTATTTGTCTCTTTCGACACGCACATGCAGGACACTCACACGTAG, GTTGGGTGGATCCCTTGAGCTCAGGAATTCGAAGAATGATTTTTCAGCCCGTGTGGAAGG, ATCAAAAGAAAATAGATACTTGTCTTACTCGAGTTGAATAAAATCCTCAGCTTTCTGTCC, AAAAGAAACTGTGAAAAGTTGTCACATTTCGATTAAATCCAAAAAGGTCTTCTATGAGGC、 TTAAAAGTATAGTAGTTGGCATTAACATTCGACCTTTTTCTGTTTCAGTAACCAACCCAG、 CATCAACTAATAGTTAAACATTATAATATCGACTGAAGACCTTTCATACTGTAAGATTCA、 CATCAACTAATAGTTAAACATTATAATATCGAGTCTGCAGTGAGCTGAGATCACACTGCC、 TTATTCCTTTCCAAATAGTTAAAATTATTCGAAACTTTTAAGAATCAATATAAAATTTCC、 CATAATTATAAATTAAAAAATGACACTATCGATTATGTCCAGTGTTTCTTGGTTGGTGTC, and CAGAGCACTAAGATAGACTTCTAAGGTTTCGAGGCATATAGCTCCAGCTGTATTGAGGTA using a probe having at least 70% identity to any of the above, and / or (ii) Specific probe sequence: TATATTTAAAAATACATACTGGTATACATCGATCTCATGACTTTGCTATTATGCATAGTG、 CCCCAGCCCAGCAACCTGGCTCACCTGATCGAGTACATCTTCAAGCCATCCTGTGTGCCC、 TATAAATAATACAGCTCTATTTGCCTACTCGATTAAAGAATCATATTATATCCTTAATTC、 CACTATGCATAATAGCAAAGTCATGAGATCGAAAATGTTTGTCAAGCAGTAGGTTTTGGG、 GATTTTAGAATCTCTAACAAGGCTGCAATCGAGGTTAGCTGCTGCAGAAAGAAGAGAAAA、 GGTGACTAAAATGAGATTGCATTTTCTTTCGACCATTTGGCCAGCATGCCAAACACTGGT、 TAACCTCTCCTTTCTTAGGTTCTCCATATCGATAGAAAATTGTCTGCAGCCCTTAATGCC、 AGCTCACGGTCAGTGCCGTTCCGTTTGCTCGAATAAAGAACAAGGACCTTAAAAAATAGA, and AAAGTCATACAACTACTATGTAAGATATTCGAATACCTGTTAGAATAGGTGAAGGTTTAT using a probe having at least 70% identity to any of the following, and / or (iii) sequences: AAGCACTTCATTCTCCCCTCACC and ATACTCCCATCCCCTAGGCCC, ACTGAGCAATGATGGCAACAAC and CCCTACGACTGGCAAACCCA, CTAGGCCTGCGTTTCTCCGT and CCCTGGCATTCACATCACCGA, AGTTCTCTCTCTAAGAACTCAAGGA and GTCAAGCAACTGTGTCTGGGG, CCCCAGGCACTCACACCTTA and GGATCCACGATCTCCCTCCAC, GGGCACTCAACACCCTTTTGT and AGGTCAGGATGGGTACCGTTG, GATGGGAATCAAGGGCAAGGG and CTGGGATGATTCCTCTGGACTTCT, GCATTAGCCAGCAAGCATACCT and GGTGGGCCTGGGTTAGATGC, CACCGCCTGATGCAGGTCTT and CAGCTGGCCGATCCATCACC, GGCACTGTTGGTCTGAAGCAC and GAGAACGACGACCTGGCACT, GGTCTAGATGTCAGTCTTTCC and TCATCCTATCCTCTCCTAGC, GACTATAAATCTCTCCTTGTCAGC and ACTGTAGGCCAACCAGAAG, CCTTACCCGACACCAGGTAGC and CTGGTCCAGTGTCAGCGTGT, CCCGACACCAGGTAGCATTC and GTGACTCTGCACGCACTGTT, GGGCACCCCACTAGACCAC and AGGCTCAGCAGGTTTCTGCC, ATTCTCTCCCTGGTAAATCCTGGT and CTGGGCAGGTCATCCAGACAG, and CGCCAGCTCAGCAGCAATAA and TGCTAGGGCCGAGTAATCATC using a primer pair having at least 70% identity with any primer pair having, and / or (iv) Sequence: GCAGCACACAGGGAACTCTCTT and TTGTTGAGCCCAGCAATTCCTTT, GCTAGGAAGGGCCTGGGATG and CTCAGTGGGCACACACTCCA, TCCAAAATGTTTAATACTGCCTAGA and AGCCATGTGGTCTGGAATCT, CAGCCACTGTAGAGAGCAGT and TAACCCACCAGCAGCAAGGT, ACTTCTTCCCAAGTCACTTTTTGC and TGGCCATCTTGCTTTGCCTC, TGACGAAGAAGCAATCCCTGGT and GGACCTACCTCCACTGGGTTG, AATCTCTGCCCTCCTCTCATCTTG and CATGTTCCCACAGCAAGGAAGTTA, AGGTATGCAGCCAGCCTGAG and TTTCCGTGCCAGTGTCCTGT, and CCGTGCCATATCCTCTGATTTATGC and GGCTGACCTTCAACAGATTCGC using a primer pair having at least 70% identity with any primer pair having the process according to any one of 1 to 3 above.

[0134] the process according to 1 or 2 above, wherein 5.5 to 9 chromosomal interactions are classified.

[0135] 6. A process for detecting the chromosomal state representing a disease subgroup of Huntington's disease, comprising detecting whether chromosomal interactions related to the chromosomal state are present or absent within a defined region of the genome, Probe sequence: AGATCTAGTTCACAGTAGCACCAATATATCGACAGATAGCTGACATCATCCTCCCAATGT, AGAGTACTTCCTAACTCCTACTGTACACTCGACAGATAGCTGACATCATCCTCCCAATGT, TATAACCAGTGCTCCTACGAAGGCCGCTTCGAAGTCTCAAACTTCACTTCTCCTGTGCGC, AGAGTACTTCCTAACTCCTACTGTACACTCGAGGATGATCGCTCCGACAGCTCCTCCAGC, GGGTTTCGCCATGTTGGCCAGGCTGGTCTCGAAGTTGATGCATCTGTGCTCACGTTTGCA, AGATCTAGTTCACAGTAGCACCAATATATCGACTGTCTCCTGTTGGCCATCTCTCACCCT, and AGATCTAGTTCACAGTAGCACCAATATATCGAACTCCTGACCTTGTGATCCACCCACCTC the process in which at least 4 chromosomal interactions selected from the chromosomal interactions represented by are classified.

[0136] 7. The classified chromosomal interactions are Probe sequence: AGATCTAGTTCACAGTAGCACCAATATATCGACAGATAGCTGACATCATCCTCCCAATGT, AGAGTACTTCCTAACTCCTACTGTACACTCGACAGATAGCTGACATCATCCTCCCAATGT, TATAACCAGTGCTCCTACGAAGGCCGCTTCGAAGTCTCAAACTTCACTTCTCCTGTGCGC, AGAGTACTTCCTAACTCCTACTGTACACTCGAGGATGATCGCTCCGACAGCTCCTCCAGC, GGGTTTCGCCATGTTGGCCAGGCTGGTCTCGAAGTTGATGCATCTGTGCTCACGTTTGCA, AGATCTAGTTCACAGTAGCACCAATATATCGACTGTCTCCTGTTGGCCATCTCTCACCCT, and AGATCTAGTTCACAGTAGCACCAATATATCGAACTCCTGACCTTGTGATCCACCCACCTC including at least 5 chromosomal interactions represented by the process according to item 6 above.

[0137] 8. The classified chromosomal interactions are the probe sequence: AGATCTAGTTCACAGTAGCACCAATATATCGACAGATAGCTGACATCATCCTCCCAATGT, AGAGTACTTCCTAACTCCTACTGTACACTCGACAGATAGCTGACATCATCCTCCCAATGT, TATAACCAGTGCTCCTACGAAGGCCGCTTCGAAGTCTCAAACTTCACTTCTCCTGTGCGC, AGAGTACTTCCTAACTCCTACTGTACACTCGAGGATGATCGCTCCGACAGCTCCTCCAGC, GGGTTTCGCCATGTTGGCCAGGCTGGTCTCGAAGTTGATGCATCTGTGCTCACGTTTGCA、 AGATCTAGTTCACAGTAGCACCAATATATCGACTGTCTCCTGTTGGCCATCTCTCACCCT、and AGATCTAGTTCACAGTAGCACCAATATATCGAACTCCTGACCTTGTGATCCACCCACCTC including all chromosomal interactions represented by the process described in item 7 above.

[0138] 9. The chromosomal interaction is - classified by detecting the presence or absence of DNA loops at the site of the chromosomal interaction, and / or - classified by detecting the presence or absence of distal regions of chromosomes that are brought together in the chromosomal conformation, and / or - classified by detecting the presence of ligated nucleic acids having sequences containing two regions corresponding to each of the regions of the chromosomes that are brought together in the chromosomal interaction, where the detection of the ligated nucleic acids is (i) specific probe sequences: AGATCTAGTTCACAGTAGCACCAATATATCGACAGATAGCTGACATCATCCTCCCAATGT、 AGAGTACTTCCTAACTCCTACTGTACACTCGACAGATAGCTGACATCATCCTCCCAATGT、 TATAACCAGTGCTCCTACGAAGGCCGCTTCGAAGTCTCAAACTTCACTTCTCCTGTGCGC、 AGAGTACTTCCTAACTCCTACTGTACACTCGAGGATGATCGCTCCGACAGCTCCTCCAGC、 GGGTTTCGCCATGTTGGCCAGGCTGGTCTCGAAGTTGATGCATCTGTGCTCACGTTTGCA、 AGATCTAGTTCACAGTAGCACCAATATATCGACTGTCTCCTGTTGGCCATCTCTCACCCT, and AGATCTAGTTCACAGTAGCACCAATATATCGAACTCCTGACCTTGTGATCCACCCACCTC using a probe having at least 70% identity to any of the above, and / or (ii) primer pairs: GAATACCCAGGAATGCTTACTTGAGC and GATTCCAGCCACCCACCTTTCACAAG, ACCAAATGCCATCTGGGACACATCCA and TGATTCCAGCCACCCACCTTTCACAA, CCCGCTAAGTCCACCCCTCTGTA and GAACCGCACTTACCCTCAGCAGT, ACCAAATGCCATCTGGGACACATCCA and GGACAAGCAGACACACTACCTGAACT, AGTCTGCCCACTGAGGTAACTAACAA and ATCCCCTGAAACAGAAGGACCTCGTG, GAATACCCAGGAATGCTTACTTGAGC and GAGCCGTCTCATAATAACCTCAGGGT, and GAATACCCAGGAATGCTTACTTGAGC and CAGTGGTTTAGGGCAAAGAGAGGGAG using a primer pair having at least 70% identity to any of the above primer pairs. The process according to any one of 6 to 8 above.

[0139] The process according to 6 or 7 above, wherein 10.5 to 9 chromosomal interactions are classified.

[0140] 11. The detection involves specific detection of ligation products by quantitative PCR (qPCR) using primers capable of amplifying the ligation products and a probe that binds to the ligation site during the PCR reaction, wherein the probe contains a sequence complementary to a sequence derived from each chromosomal region where chromosomal interactions gather, according to any one of the above 1 to 10 processes.

[0141] 12. The probe is an oligonucleotide that specifically binds to the ligation product, and / or a fluorophore covalently attached to the 5'-end of the oligonucleotide, and / or a quencher covalently attached to the 3'-end of the oligonucleotide according to the process described in the above 11.

Example

[0142] The present invention is illustrated by the following non-limiting examples.

[0143] Example 1 Statistical pipeline To select high-value EpiSwitch™ markers for conversion to the EpiSwitch™ PCR platform, the EpiSwitch™ screening array is processed using the EpiSwitch™ analysis package in R.

[0144] Step 1 Probes are selected based on their corrected p-values (false discovery rate, FDR), which are the products of a modified linear regression model. Probes with p-values below 0.1 are selected and then further reduced by their epigenetic ratio (ER), and the probe ER must be ≤ -1.1 or ≥ 1.1 to be selected for further analysis. The final filter is the coefficient of variation (CV), and the probe must be below 0.3.

[0145] Step 2 The top 40 markers of the statistical list are selected based on their ERs for selection as markers for PCR conversion. The top 20 markers with the largest negative ER load and the top 20 markers with the largest positive ER load form the list.

[0146] Step 3 Statistically significant probes, which are markers resulting from the results of Step 1, form the basis of enrichment analysis using hypergeometric enrichment (HE). This analysis enables marker reduction from the list of significant probes and, together with the markers from Step 2, forms a list of probes to be converted to the EpiSwitch™ PCR platform.

[0147] The statistical probes are processed by HE to determine which genetic locations exhibit enrichment of statistically significant probes and to show which genetic locations are the epicenters of epigenetic differences.

[0148] The most significantly enriched loci based on the corrected p-value are selected for probe list creation. Genetic locations below a p-value of 0.3 or 0.2 are selected. Together with the markers from Step 2, the statistical probes mapping to these genetic locations form high-value markers for EpiSwitch™ PCR conversion.

[0149] Array Design and Processing Array Design 1. Process the loci using SII software (currently, v3.2) to a. Retrieve the genomic sequences (gene sequences with 50 kb upstream and 20 kb downstream) at these specific loci. b. Define the probability that the sequences within this region are involved in CCs. c. Cut the sequences using specific REs. d. Determine whether the restriction fragments may interact in a specific orientation. e. Rank the possibilities of various CCs interacting together. 2. Determine the array size and thus the number of available probe sites (x). 3. Extract x / 4 interactions. 4. For each interaction, define sequences of 30 bp from part 1 to the restriction site and 30 bp from the restriction site of part 2. Check that these regions are not repeats, and if so, exclude them and select the next interaction below in the list. Join both 30 bp together to define the probe. 5. Create a list of x / 4 probes, plus the defined control probes, and replicate four times to create the list to be created on the array. 6. Upload the list of probes to the Agilent SureDesign website for the custom CGH array. 7. Design an Agilent custom CGH array using the probe set.

[0150] Array processing 1. Process the samples using the EpiSwitch™ Standard Operating Procedure (SOP) for template generation. 2. Purify by ethanol precipitation by the array processing laboratory. 3. Process the samples as per the Agilent SureTag complete DNA Labeling Kit - Agilent Oligonucleotide Array-based CGH for Genomic DNA Analysis Enzymatic labelling for Blood, Cells or Tissues. 4. Scan using an Agilent C Scanner with Agilent feature extraction software.

[0151] Overview of EpiSwitch™ technology The EpiSwitch™ platform provides a highly effective means for screening, early detection, companion diagnostics, monitoring, and prognostic analysis of major diseases associated with abnormal and highly responsive gene expression. The main advantage of this approach is that it is non-invasive and rapid, and is based on very stable DNA-based targets as part of chromosomal signatures rather than unstable protein / RNA molecules.

[0152] EpiSwitch™ biomarker signatures exhibit high structural stability, sensitivity, and specificity in the stratification of complex disease phenotypes. This technology harnesses the latest breakthroughs in epigenetics chemistry, monitoring, and assessment of chromosomal conformation signatures as a very useful class of epigenetic biomarkers. Current research methodologies deployed in academic settings require 3 - 7 days of biochemical processing of cell material to detect CCSs. These procedures limit sensitivity and reproducibility and, furthermore, do not offer the advantage of targeted insights provided by the EpiSwitch™ analysis package at the design stage.

[0153] Computer-Aided EpiSwitch™ Array Marker Identification CCS sites in the genome are directly evaluated by the EpiSwitch™ array for clinical samples of the test cohort to identify all relevant stratified lead biomarkers. The EpiSwitch™ array platform is used for marker identification due to its high throughput capacity and ability to rapidly screen multiple loci. The array used was an Agilent custom CGH array that enables marker identification by the interrogated in silico software.

[0154] EpiSwitch™ PCR Potential markers identified by the EpiSwitch™ array are subsequently verified by EpiSwitch™ PCR or a DNA sequencer (i.e., Roche 454, Nanopore MinION, etc.). Top PCR markers that are statistically significant and show the highest reproducibility are selected to further reduce to the final EpiSwitch™ signature set and verified in an independent cohort of samples. EpiSwitch™ PCR is performed by trained technicians following established standardized operating procedure protocols. The manufacture of all protocols and reagents is done under ISO 13485 and 9001 accreditation to ensure the quality of work and transferability of the protocols. The EpiSwitch™ PCR and EpiSwitch™ array biomarker platforms are suitable for the analysis of both whole blood and cell lines. The tests are sensitive enough to detect very low copy number abnormalities using small amounts of blood.

[0155] Analysis of the ALS cohort The inventors used epigenetic chromosomal interactions as a basis for identifying biomarkers for use as a companion diagnostic for ALS. The EpiSwitch™ biomarker discovery platform was developed by the inventors to detect epigenetic regulatory signature changes that promote phenotypic changes associated with ALS. The EpiSwitch™ biomarker discovery platform identifies CCs that define the initial regulatory processes in integrating environmental cues into the epigenetic transcriptional machinery. The CCs themselves are major steps in the gene regulatory cascade. The CCs isolated by the EpiSwitch™ biomarker discovery platform have several well-documented advantages: extreme biochemical and physiological stability; their binary nature and readout; and their major positioning in the eukaryotic cascade of gene regulation.

[0156] The ability to detect perturbations in the global epigenetic control of gene expression enables both early diagnosis of ALS and the establishment of patient prognosis. By a comparative investigation of the cellular control genomic structure from blood samples of healthy and affected patients, two ALS-related epigenetic signatures (one for diagnostic potential and the second for prognosis prediction) have been revealed. In this prospective study, samples collected by the clinical research group of the Oxford University Nuffield Department of Clinical Neurosciences (NDCN)'s Oxford Movement Neuron Disorder Clinic were analyzed using EpiSwitch™, a high-throughput platform developed by Oxford Biodynamics to monitor chromosomal conformation signatures. In this study, clinical annotations of ALS-FRS-R, forced vital capacity (FVC), and other clinical findings were compared, and ALS progression subtypes were assigned by analysis of epigenetic signatures.

[0157] A total of 100 patients who attended the Oxford Movement Neuron Disorder Clinic participated in the study and were asked to return at 3 and 6 months. Controls (n = 100) were collected. During each visit, participants underwent ALS-FRS-R and FVC tests and provided blood samples. Samples were analyzed to identify ALS disease-related diagnostic signatures or prognosis disease-related signatures (at 3 and 6 months). The results of the clinical evaluations were compared with the EpiSwitch™ analyses at 0, 3, and 6 months. Using a cutoff of a 0.5-point decline per month in the ALS-FRS-R score, ALS patients were divided into progression subtypes.

[0158] Based on the rate of decline of the ALS-FRS-R score, preliminary results from 3-month and 6-month samples demonstrated that an epigenetic signature with 80% sensitivity and specificity selected fast-progressing (>0.5) and slow-progressing (<0.5) ALS patients. Previous results have shown that the prognostic signature is robust in selecting ALS subtypes over time. The results presented in the table pertain to the diagnosis and prognostic analysis of 75 ALS samples and 75 control samples (n = 150) using an independent sample cohort (n = 50).

[0159] Additional results (including Table 11) are based on 100 ALS patients observed and evaluated by standard practice of ALSFRS-R score and FVC (forced vital capacity) at baseline, 3 months, and 6 months of survival. Based on measurements at baseline and 3 months, these patients were classified as slow or fast ALS (standard classification). Reading out three markers at baseline only and classifying all patients as slow or fast (EpiSwitch classification). Next, these patients were evaluated based on survival at 6 months, and the survival rates of the slow and fast groups were compared. In the standard classification, there were a significant number of lethal outcomes in the slow group, and the difference in survival rates by the p-value between slow and fast was not statistically significant (p = 0.052). In the EpiSwitch classification, the boundary between slow and fast at 6 months was highly significant (p = 0.0097), and most deaths occurred in the fast group.

[0160] Example 2 Differences in genomic structure at the HTT locus underlie symptomatic and pre-symptomatic cases of Huntington's disease.

[0161] Huntington's disease (HD) is a progressive neurodegenerative disorder that causes widespread degeneration of neurons in the adult brain and ultimately leads to death. The underlying cause of HD is an elongated trinucleotide cytosine-adenine-guanine (CAG) repeat in the "huntingtin gene" (HTT), which results in a mutant huntingtin protein that tends to form intracellular aggregates and cause neuronal death. Importantly, the number of CAG repeats correlates with the onset and severity of the disease, and patients with more than 39 CAG repeats will develop HD at some point in their lives. The onset of the disease can vary within an individual over decades, and little is known about this pre-onset stage. With the advancement of epigenetic approaches and the emergence of new technologies for detecting epigenetic regulatory markers for DNA methylation, histone modifications, non-coding RNAs, and chromosomal conformation signatures as part of the understanding of genomic structure, we became interested in analyzing the relationship between global regulatory changes in the genomic structure near the HTT gene and its disease manifestations. Here, we examined detectable global non-invasive epigenetic differences at the HTT locus between pre-onset and symptomatic HD patients compared to non-diseased controls.

[0162] Methods: Blood samples from HD patients and healthy controls were used to evaluate chromatin structure in the immediate vicinity of the HTT gene using EpiSwitch™, a validated high-resolution industrial platform for the detection of chromosomal conformation. The presence or absence of conditionally stable chromatin structure at 20 interaction sites across 225 kb of the HTT locus was evaluated, and the resulting chromosomal conformations were compared among a group of healthy controls, validated symptomatic HD patients (CAG, n>39), and patients with genetically verified CAG expansions who did not show clinical symptoms of HD.

[0163] Results: When tested on peripheral blood samples, a consistently stable chromosomal conformation was observed across the patient group. Two structural interactions (occurring in all patient groups and controls) and seven conditional interactions that are present in HD but not in healthy controls were found. Most importantly, three conditional interactions were observed that are present only in HD patients presenting clinical symptoms (symptomatic cases) and not in premanifest cases. Eighty-five percent (6 out of 7) of the patients in the symptomatic HD cohort showed at least one of the specific chromosomal conformations associated with symptomatic HD.

[0164] Conclusions: Our results are the first evidence that the regulatory chromatin structure at the HTT locus is globally altered in HD patients. Furthermore, the chromatin structure at the HTT locus also shows conditional differences between the clinical stages of symptomatic HD patients and premanifest HD patients. Considering the high clinical utility of having molecular tools for assessing disease progression in HD, these results strongly suggest that non-invasive assessment of global chromosomal conformation signatures (CCS) could be a valuable addition to the prognostic evaluation of HD patients.

[0165] Introduction Huntington's disease (HD) is a neurodegenerative disorder characterized cellularly by the loss of neurons in the basal ganglia of the brain and clinically by uncontrollable movements, emotional problems, and loss of cognitive ability. HD is an autosomal dominant genetic disorder, and its prevalence varies widely depending on geography and ethnicity, but it is thought to affect more than 50,000 people in the United States and Europe alone. The underlying genetic cause is an expansion of the trinucleotide CAG in the huntingtin gene (HTT), discovered as a genetic marker in 1983 by James Gusella of Massachusetts General Hospital, which results in a mutant huntingtin protein (mHTT) with a toxic polyglutamine (polyQ) tract. Thus, despite decades of research and clinical trials, no successful treatment has yet been developed. The "typical" onset of HD is between the ages of 40 and 50, but up to 15% of cases have a very late onset and do not show clinical symptoms until after the age of 60. A meta-analysis of recent studies examining cases of late-onset HD (LoHD) found that over 90% of patients had a CAG repeat length of 44 or less. One of the more interesting observations in HD is that while there is a well-known correlation between the length of the polyQ repeat tract and the onset and severity of the disease, there is considerable variability among individual patients. For example, in patients with a moderate repeat length (defined here as 40 - 50), the onset of the disease can vary by 60 years among individual patients. This means that many patients who are carriers of the polyQ tract, which is a predisposing factor for the onset of the disease, can live for decades in a "pre-symptomatic" state. What controls the onset of clinical symptoms is currently unknown, complicating the prognosis assessment of HD patients.

[0166] Historically considered a single-gene disorder, extensive research on the underlying pathology of HD has suggested that the mechanisms leading to disease onset and progression are more complex than initially thought. Many different techniques have been used to examine the molecular changes underlying HD disease progression, including gene expression, proteomics, metabolomics, network analysis, genomics, single nucleotide polymorphism (SNP) profiling, etc. Given that HD is regarded as a paradigm of diseases characterized by epigenetic dysregulation, more recently, epigenetic approaches have emerged as promising new tools for assessing pathology-related changes. Most epigenetic studies in HD have focused on examining genome-wide histone modifications (acetylation, methylation) or histone modifications at specific loci associated with HD. While these approaches have provided interesting insights into the disease, they often produce conflicting results and have shown discrepancies between mouse models and human disease. Thus, a consensus picture of epigenetic deregulation in HD using histone modification readers has not yet materialized. However, not all molecular mechanisms related to epigenetic control have been evaluated in the context of HD. An important aspect of epigenetic control is at the level of three-dimensional (3D) genome structure.

[0167] The 3D organization of the genome reflects the heterogeneous influence of cues and inputs from the external environment and can be measured by the assessment of chromosomal conformations or, when multiple conformations are measured simultaneously, by the assessment of chromosomal conformation signatures (CCSs). CCSs can be considered molecular barcodes that read the epigenetic landscape of a given cell population. Given the central role of mHTT in the development of HD, we hypothesized that regulatory differences in genomic structure at the HTT locus might exist between affected individuals and healthy individuals, i.e., non - affected controls. We used EpiSwitch, an established proprietary industrial platform for monitoring CCSs, to assess chromatin structure differences between pre - symptomatic and symptomatic HD patients and healthy individuals, i.e., non - affected individuals. The EpiSwitch readout provides high - resolution, reliable, high - throughput detection of CCSs while meeting high standards of industry standards for quality control.

[0168] Methods Sample collection All samples were obtained from National BioService, LLC. In this study, a total of 20 samples were used, including 10 healthy control (HC) samples (CAG repeat, n < 35) and 10 HD samples (CAG repeat, n > 39). For the HD samples, 7 were from symptomatic patients (HD - Sym) and 3 were from pre - symptomatic patients (HD - Pre) diagnosed with HD but not yet showing clinical symptoms. One HD patient was taking tetrabenazine and one patient was taking sertraline. In all samples, human immunodeficiency virus, hepatitis B virus, hepatitis C virus, and syphilis were negative (Table 14).

[0169] Study design We sought to identify differential chromatin conformations among healthy controls (low CAG), premanifest HD patients (high CAG, no signs of disease), and symptomatic HD patients (high CAG, signs of disease). For our analysis, we focused on a region of approximately 225 kb surrounding the HTT locus (annotated as chr4:3,033,588 to 3,258,170 hg38). The CAG repeat expansion tract of exon 1 of HTT (chr4:3,054,162 to 3,095,930) was used as an anchor point (“anchor”), and five genomic zones surrounding the anchor were defined and examined for differential chromosome conformations among sample groups (Figure 2 and Table 15). These zones were selected based on the presence of potential EpiSwitch anchor sites; the presence of known disease-associated SNPs (HD and other diseases); and the enrichment of known histone modification sites (H3K4me3, H3K36me3, and H3K27ac) in HD as seen in the GWASdvV2 database (http: / / jjwanglab.org / gwasdb). (Figure 2).

[0170] Chromosome conformation identity NCBI search: The GEO database for already reported HD epigenetic data was run in February 2018. Peaks called ChIP-seq data of H3K4me3 from 12 (6 HD and 6 control samples) postmortem prefrontal cortex brain samples (bed format) were obtained (GSE68952). Additionally, Bigwig tracks of ChIP-seq data of H3K27ac and H3K36me3 from HD iPSC-derived neuronal cell lines and control cell lines were also downloaded (GSE95342). The data tracks were loaded into Integrative Genome Viewer (IGV)

[41] along with EpiSwitch and reference sequence annotations. Both visual comparison and programmatic comparison (BEDtools) were performed at the HTT locus to identify the five zones of interest.

[0171] Using appropriate software, chromatin folding interactions between the "ends" that occur in the proximal anchor zone of the CAG repeat and the other "ends" that occur in any of the five target zones were identified with high probability. A total of 61 interactions met these criteria, and for practical reasons, 20 interactions were selected to cover interactions between the anchor site and all of the target zones. Oligonucleotide pairs were designed using an automated primer design application to amplify the predicted DNA sequences that would be generated by the interactions when subjected to a chromosome conformation capture assay.

[0172] 3C and PCR Chromosome conformation capture and detection were performed by PCR. Chromatin with intact chromosomal structure was extracted from 50 μl of blood samples from each patient sample using the EpiSwitch assay according to the manufacturer's instructions (Oxford BioDynamics Plc). Quality control of all samples was performed using the detection of chromatin loops at the MMP1 locus, which is a historical internal control for 3C analysis. Pooled 3C libraries of each sample type were generated to provide a generalized population sample for each sample subgroup. Real-time PCR was performed using SYBR Green on a CFX-96 (Bio-Rad) machine to identify interactions with different PCR product detection patterns between sample types. Oligonucleotides were tested with a control template to confirm that each primer set was functioning correctly. Final nested PCR was performed three times for each sample on the follow-up data of individual HD patients according to the Royal Forensic Protocol for PCR detection. This procedure enabled more accurate detection of templates with limited copy numbers. All PCR amplification products were monitored with a PerkinElmer LabChip® GX using the PerkinElmer LabChip® DNA 1K version 2 kit, and internal DNA markers were loaded onto the DNA chip according to the manufacturer's protocol using fluorescent dyes. Fluorescence was detected with a laser, and the electrophoretic readings were converted to simulated bands on the gel image using the instrument's software. The detection threshold of the instrument was set by the manufacturer at 30 fluorescence units or more.

[0173] Statistical analysis Data analysis was performed in R (the language and environment for statistical computing). This included the stats and dplyr packages for t-tests and R-squared analysis, and the ggplot2 package for box-and-whisker plots and regression plots.

[0174] Results Clinical characteristics of the patients Samples of HC and HD were age-matched (mean age of HC = 36.9 and mean age of HD = 35.3) and gender-matched (1 / 2 male and 1 / 2 female), and the majority (70%) of HD cases were symptomatic (Table 16). All samples were of non-Hispanic or Latino white origin. The mean length of CAG repeats was 25.7 in HC and 44.2 in HD (Table 16). There was no statistical difference in age between the HC and HD samples. HC patients did not show a statistical difference in age from HD-Sym. HD-Pre patients were younger (mean age = 25.3) than HD-Sym patients (mean age = 39.6) (p = 0.02). Although neither was statistically significant, there was a weak negative relationship between disease duration and CAG repeat size, and a weak positive relationship between age at diagnosis and CAG repeat size.

[0175] There was a statistically significant increase in CAG repeat length in HD patients compared to HC (p = 1.08E -7 ). There was no statistically significant increase in CAG repeat length between HC and HD-Pre (p = 3.43E -6 ) or between HC and HD-Sym (p = 9.50E -8 ). There was no statistical difference in CAG repeat length between HD-Pre and HD-Sym (p = 0.09).

[0176] The mean age at diagnosis of the HD samples was 35.3 years, the mean disease duration was 3.8 years, and 7 out of 10 patients reported involuntary, choreic movements, or both (Table 16).

[0177] Chromosome conformation of HC, HD-Pre, and HD-Sym Of the 20 interactions evaluated (Table 17), we identified 9 beneficial interactions. We identified 2 constitutive interactions and 7 conditional interactions that are present in HD but not in healthy controls. Three of the 7 conditional interactions were present only in HD-Sym and not in HD-Pre.

[0178] Constitutive conformation All samples passed an internal QC analysis for MMP1 interaction. Two constitutive (identified in all samples) chromatin loops were identified. Both loops were between the anchor and zone 2, with the first loop spanning 28 kb and the second loop spanning 34 kb. Two constitutive interactions occurring in all patients (HD-Sym, HD-Pre and HC) were observed in this study. Both interactions (CC1&CC2) were between the anchor and zone 2, with CC1 spanning 28 kb and CC2 spanning 34 kb.

[0179] Conditional conformation We identified seven conditional chromosomal conformations that can distinguish different patient subgroups evaluated in this study. Specifically, we identified two chromosomal interactions that are present in HC but not in all HD samples. The first interaction (I1) covered 77 kb across the anchor and zone 4, and the second interaction (I2) covered 140 kb across the anchor and zone 3. We also identified two chromosomal interactions that are present in HC and HD-Pre but not in HD-Sym samples. Both interactions (I3 and I4) covered 92 kb and 104 kb, respectively, across the anchor and zone 4. Finally, we identified three chromosomal interactions that are present in HD-Sym samples but not in HD-Pre and HC samples. The first interaction (I5) covered 122 kb across the anchor and zone 3. Interestingly, this interaction contained the SNP (rs362331) known to be a factor in the predisposition to develop HD. The second and third interactions (I6 and I7) covered 185 kb and 174 kb, respectively, across the anchor and zone 1. Finally, we tested for the presence or absence of all conditional interactions in individual HD samples. In HD-Sym samples, at least one of the conditional markers (I5, I6, and I7) was found to be present in 6 out of 7 samples (Figure 3). An overview of all interactions evaluated in this study is shown in Figures 4 and 5. Odds ratios are shown in Table 18. Table 13 shows chromosomal interactions that showed no differences between subgroups.

[0180] Discussion Description of the problem and summary of results It is well known that individuals with CAG repeats exceeding 39 develop HD, but the clinical onset of the disease varies greatly among individual patients, and the factors affecting the clinical manifestation of the disease have not been well characterized. Here, we used EpiSwitch, an industrial platform for evaluating chromatin structure, to assess the epigenetic landscape of the HTT locus in HD patients and healthy (non-affected controls). Collectively, we identified a set of seven interactions that can distinguish HD from non-affected controls and, more importantly, distinguish pre-onset and symptomatic HD patients. One of these interactions specific to symptomatic HD contains a SNP (rs362331) that has been shown to be associated with the haplogroup of the predisposing disease. Collectively, these results indicate that a simple, non-invasive blood-based test for assessing CCS can serve as a surrogate biomarker for evaluating the disease progression of HD.

[0181] Biological relevance Although it is known that the expansion of the polyQ repeat tract and the generation of mHTT are the fundamental causes of HD, the molecular events leading to the onset of clinical symptoms have not been well characterized. In several studies, SNPs within the HTT locus have been investigated as potential causes of disease onset. In a recent SNP genotyping study of HD patients, at least 41 SNPs with heterozygosity were identified in at least 30% of patients, including the rs362331 C / T SNP in exon 50 of the HTT gene. Perhaps more biologically relevant is that allele-selective knockdown of the rs362331 SNP using antisense oligonucleotides, siRNA, or miRNA results in a dramatic decrease in the level of mHTT protein, achieved both in vitro and in vivo, suggesting that this SNP and its surrounding genomic context play an important role in the regulation of mHTT levels. In this study, chromosomal cytobands (I5) present only in HD-Sym patients and HCs overlapping the rs362331 SNP were observed. This suggests that the generation of neurotoxic mHTT in patients with an increased polyQ tract and the genetic predisposition to the early onset of HD due to the presence of the rs362331 SNP are regulated at the level of higher-order chromatin structure. Another unresolved issue in HD is how the disease is inherited when neither parent has received a diagnosis. Two main general hypotheses point to 1) the carrier parent died from another factor before the onset of the disease, and 2) the "unstable" CAG repeat tract expands from generation to generation. A third possibility exists, where individuals with moderate (35 - 50) repeats can be asymptomatic carriers of the disease, but their offspring may not be able to compensate for the genetic defect through an undefined mechanism and may develop the disease. All HD patients evaluated in this study have CAG repeats in this moderate range, increasing the likelihood that potential compensatory mechanisms for disease development can be mediated by differences in genomic structure.

[0182] Clinical Relevance HD is unique in that there are simple tests for definitive diagnosis of the disease, HTT gene sequencing, and measurement of the CAG repeat number. For clinical care and clinical trials, there are also several tests for measuring disease severity, such as the Unified Huntington's Disease Rating Scale (UHDRS), the Shoulson-Fahn scale, and the Mini-Mental State Test (MMSE). These assessments measure various aspects of the physical and mental health of HD patients as surrogates for disease severity, but they are all subjective and few are specific to HD. What is lacking is a specific molecular tool for monitoring disease progression.

[0183] Currently, there are 22 therapeutic agents for treating HD at various stages of preclinical and clinical development, half of which are in Phase 2 or Phase 3. If further validated, the CCS reported here can be used in clinical trials as a surrogate outcome biomarker to assess the therapeutic effect of the drug in question. In addition to monitoring the symptomatic patients' response to a specific treatment in clinical trials, another advantage of the approach described herein lies in the information available for premanifest patients. For most HD patients, the premanifest period can last for decades. Five of the seven (i3-i7) interactions identified herein clearly distinguish premanifest HD patients from symptomatic patients and, if further validated, can serve as "early warning" indicator tests for the onset of HD symptoms in premanifest carriers.

[0184] This study provides the first evidence of detectable differences in chromatin structure that are specific to the manifestation of HD and correlate with known disease haplotypes. The main strength of this study lies in its unique approach based on recent developments in understanding the regulatory role of genomic structure. There have been several historical studies of HD aimed at developing disease progression biomarkers based on clinical, imaging, and molecular measurements, but to our knowledge, this is the first application of the assessment of higher-order chromatin structure in clinically accessible biological fluids to HD.

[0185]

Table 1

[0186]

Table 2

[0187]

Table 3

[0188]

Table 4

[0189]

Table 5

[0190]

Table 6

[0191]

Table 7

[0192]

Table 8

[0193]

Table 9

[0194]

Table 10

[0195]

Table 11

[0196]

Table 12

[0197]

Table 13

[0198]

Table 14

[0199]

Table 15

[0200]

Table 16

[0201]

Table 17

[0202]

Table 18

[0203]

Table 19

Claims

A process for identifying whether a subject is undergoing rapid progression of amyotrophic lateral sclerosis (ALS), the process comprising determining the presence or absence of at least four chromosomal interactions selected from chromosomal interactions (i)-(viii), the presence being associated with rapid progression of ALS; Detecting the presence or absence of each chromosomal interaction (i)-(viii) is (a) a step of in vitro cross-linking chromosomal regions aggregated by chromosomal interaction in a sample derived from the subject; (b) a step of subjecting the cross-linked DNA to cleavage, (c) a step of ligating the cross-linked and cleaved DNA to form ligated DNA; and (d) a step of detecting the presence or absence of chromosomal interaction by detecting the ligated DNA by PCR by a method comprising; Chromosomal interaction (i) is detected by primers TCTTGTACACGGTTGGTGGT and TGTCACCTATGTGCTGAGTACTGG in PCR, Chromosomal interaction (ii) is detected by primers TGACGAAGAAGCAATCCCTGGT and GGACCTACCTCCACTGGGTTG in PCR, Chromosomal interaction (iii) is detected by primers CCGTGCCATATCCTCTGATTTATGC and GGCTGACCTTCAACAGATTCGC in PCR, Chromosomal interaction (iv) is detected by primers ACTTCTTCCCAAGTCACTTTTTGC and TGGCCATCTTGCTTTGCCTC in PCR, Chromosomal interaction (v) is detected by primers GGATATGCAGTTTTCCTGGCACTAC and CATGCTAGGGCCGAGTAATCATCT in PCR, Chromosomal interaction (vi) is detected by primers GCAGCACACAGGGAACTCTCTT and TTGTTGAGCCCAGCAATTCCTTT in PCR, Chromosomal interaction (vii) is detected by primers CAGCCACTGTAGAGAGCAGT and TAACCCACCAGCAGCAAGGT in PCR, and Chromosomal interaction (viii) is detected by primers in PCR A process detected using AGTAGCTTCCCTGTTAGAGGTCTTG and AGCCAGTGACTCCACAACTTCTT.

2. The detection includes specific detection of ligated DNA by quantitative PCR (qPCR) using the primer capable of amplifying the ligated DNA and a probe that binds to the ligation site during the PCR reaction, wherein the probe includes a sequence complementary to a sequence derived from each chromosomal region gathered by chromosomal interaction. The process according to claim 1.

3. The probe is a fluorophore covalently bound to the 5' end of the probe, and / or a quencher covalently bound to the 3' end of the probe The process according to claim 2.

Citation Information

Patent Citations

  • Chromosome conformation capture-on-chip (4c) assay

    JP2009500031A

  • Method for determining presence of disease

    JP2011092137A

  • Methods for detecting long-range chromosome interactions

    JP2011521633A

  • Chromosome Conformation Analysis

    US20120302449A1

  • Epigenetic chromosome interactions

    WO2016207647A1