Methods of producing nucleic acid fragments

The use of dsDNase and HL-dsDNase nucleases for high-resolution chromatin fragmentation addresses the limitations of existing methods, enabling detailed analysis of chromatin interactions and improving personalized medicine and diagnostics.

WO2026013416A1PCT designated stage Publication Date: 2026-01-15NUCLEOME THERAPEUTICS LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/GB2025/051544
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-12
Filing Date
2025-07-11
Publication Date
2026-01-15

AI Technical Summary

Technical Problem

Existing methods for analyzing interactions between enhancers, silencers, boundary elements, and promoters in chromatin at high resolution are limited by the resolution of nucleases used, leading to incomplete understanding of gene regulation and potential benefits for personalized medicine and diagnostics.

Method used

The use of dsDNase and HL-dsDNase nucleases, belonging to the CATH Superfamily 3.40.570.10 Extracellular Endonuclease, to fragment chromatin within permeabilized cells, achieving high-resolution nucleic acid fragments with a distinct 'ball and chain' morphology, allowing for improved analysis of chromatin interactions.

Benefits of technology

This approach enhances resolution and reduces cell number requirements, providing more accurate data representative of the whole genome, enabling the identification of novel regulatory sequences and allele-specific interactions, which can inform personalized medicine and diagnostics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure GB2025051544_15012026_PF_FP_ABST
    Figure GB2025051544_15012026_PF_FP_ABST
Patent Text Reader

Abstract

The invention provides a method of producing nucleic acid fragments, the method comprising fragmenting chromatin within a population of permeabilised cells with a CATH Superfamily 3.40.570.10 Extracellular Endonuclease, subunit A nuclease, or a nuclease variant thereof or, wherein the nuclease comprises an active site having the amino acid sequence of A / S / D / Q / T / E-K / R / Y / l / H-X-H / Q / Y / F-X(a y-N / F / G-X(7 )-E / L / M / H / Q-X(3 )-R / M / L / V / F, where X is any amino acid and a is 23 to 37, the nuclease has a ball and chain morphology or comprises an amino acid sequence having at least 40% identity to SEQ ID NO: 1 or SEQ ID NO: 2. The invention also provides products, like a pool of nucleic acid fragments obtained by using such nuclease, as well as a method of producing a 3C library, and other methods using such nuclease.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] METHODS OF PRODUCING NUCLEIC ACID FRAGMENTS

[0002] Progress in our ability to annotate regulatory elements in the genome and determine their potential function has been driven by technological advances, such as RNA-seq, ChlP-seq, DNase-seq and ATAC-seq. However, an outstanding challenge is to understand the mechanisms by which regulatory elements control specific gene promoters at a distance (10s to 1000s kb).

[0003] Using conventional Chromosome Conformation Capture (3C), it is possible to analyse in detail the interactions between enhancers, silencers, boundary elements and promoters at individual loci at high resolution.

[0004] Since the development of the original 3C method in 2002, several new 3C-based techniques have emerged such as Capture-C, Hi-C, Capture Hi-C, in situ Hi-C, Circularized Chromosome Conformation Capture (4C), 4C-seq, ChlA-PET, Carbon Copy Chromosome Conformation Capture (5C), NG Capture-C (W02017 / 068379) and most recently Micro Capture C (WO2020 / 161485).

[0005] The resolution of these methods, other than Micro Capture C, remains limited when studying mammalian genomes. In most assays, the resolution is determined by the bin size used to pool data and improve signal strength. This is a function of the depth of sequencing and the size of the organism's genome; the sequencing requirements increase with the square of the genome size or resolution, whereby a decreasing bin size requires increased sequencing depth to ensure adequate coverage is achieved. If sufficient sequence depth can be obtained then the limiting factor becomes the restriction fragment size, which equates to a theoretical limit of ~256 bp with 4 cutter restriction enzymes

[0006] Increases in resolution are potentially useful for highlighting the regulatory sequences that control genes in greater detail and to identify novel sequences that control genes. Increases in resolution should also allow single nucleotide polymorphisms identified by genome wide association studies (GWAS) to be linked to the genes or other aspects of genome function or structure that they control with greater confidence. This has potential benefits for personalised medicine, diagnostics and drug discovery. In previous methods (e.g. WO2017 / 068379), the cells have been fixed (e.g. using formaldehyde) and then homogenised in order to break open the cells and to release the chromatin.

[0007] Digestion of the chromatin has previously been carried out using a number of different enzymes, including 4 and 6 base-pair cutting restriction endonucleases, e.g. Hindlll, EcoRI, Ncol, Xbal, Bgilll, Dpnll and Nialll, and bacterial nucleases including Micrococcal nuclease and DNasel. However, digestion of chromatin with 4 and 6 base-pair cutting restriction endonucleases does not allow for accurate analysis of the interactions between enhancers, silencers, boundary elements and promoters at individual loci at high resolution. Accordingly, there is a need for improved methods for fragmenting chromatin, and for analysing the interactions between enhancers, silencers, boundary elements and promoters at individual loci at high resolution.

[0008] The present invention relates to a method of producing nucleic acid fragments, the method comprising fragmenting a nucleic acid sequence with a nuclease. In particular, the present invention relates to a method of producing nucleic acid fragments, the method comprising fragmenting chromatin with a nuclease. The invention relates to a method of producing nucleic acid fragments, the method comprising a step for fragmenting chromatin. The invention also relates to a method of producing nucleic acid fragments, the method comprising fragmenting chromatin with a means for fragmenting chromatin. The chromatin may be fragmented within a population of permeabilised cells, fragmented within isolated nuclei, or fragmented as isolated chromatin. In particular, the present invention relates to a method of producing nucleic acid fragments, the method comprising fragmenting chromatin within a population of permeabilised cells with a nuclease. The invention also relates to a method of producing nucleic acid fragments, the method comprising fragmenting chromatin with a means for fragmenting chromatin in a population of permeabilised cells. The invention also relates to a method of producing nucleic acid fragments, the method comprising a step for fragmenting chromatin within a population of permeabilised cells.

[0009] The inventors have now found a class of nucleases exemplified by the endonucleases dsDNase (also known as Northern shrimp nuclease having UniProt ID C9YSL6, and which we call NuCase) and HL-dsDNase (which we call HL-NuCase), when used for chromatin conformation assays, are able to enter fixed permeabilised cells and digest chromatin down to predominantly mononucleosome levels. These endonucleases belong to the CATH Superfamily 3.40.570.10 Extracellular Endonuclease, subunit A nuclease. In particular, the endonucleases belong to the DNA / RNA nonspecific endonucleases family and feature the family characteristic NUC domain (IPR020821 : NUC / extracellular endonuclease, subunit A domain). These endonucleases feature a remarkably high affinity for activity on dsDNA with minimal activity on ssDNA or RNA. The endonucleases dsDNase and HL-dsDNase also have a distinct “ball and chain” morphology which is clearly different from other nucleases used in chromatin conformation assays, such as Micrococcal nuclease. The resulting assay features a number of key improvements over other methods currently in use for 3D conformation analysis. The enhanced resolution previously obtained In the Micro Capture C method (WO2020 / 161485) is maintained in addition to a large reduction in cell number requirements owing to efficient digestion and a lack of over-digestion owing to the endonuclease only activity, advantageously creating the ability to explore rare cell types previously unattainable. Thus, the digest profile of NuCase distinctly lacked chromatin lengths below ~140 base pairs and had a larger average size of mononucleosomes, ~190 base pairs, relative to the MNase digests average ~160 base pairs with a large amount of chromatin content below ~140 base pairs. In addition, Micrococcal nuclease has a well described bias with a preference for cutting at AT rich regions of the genome. Advantageously, the NuCase libraries displayed much more evenly distributed cut site locations. This is advantageous as the genome features GC rich regions around genes and with better cut site preferences, the data generated with NuCase is more representative of the whole genome rather that of more AT rich regions.

[0010] In addition, the inventors have surprisingly shown that NuCase and its derivatives enter fixed permeabilised cells. Under common conditions, permeabilised cell pores are thought to be about 10 nm in size, and are believed to be spherical. Micrococcal nuclease has a molecular weight of approximately 17 kDa and is therefore small enough to enter through permeabilised cell pores without issue. However, Dpnll which has a molecular weight of approximately 37 kDa is unable to enter these pores. The inventors, nevertheless, surprisingly found that the class of nucleases exemplified by NuCase, which has a molecular weight of approximately 43 kDa, could enter permeabilised cells and was very effective for DNA and chromatin digestion in MCC methods. Whilst not wishing to be bound by any particular theory, the inventors believe that NuCase was able to work, despite its relatively large size, because of the enzyme’s shape (such as revealed using AlphaFold™).

[0011] The invention provides a method of producing nucleic acid fragments, the method comprising fragmenting chromatin within a population of permeabilised cells with a CATH Superfamily 3.40.570.10 Extracellular Endonuclease, subunit A nuclease, or a nuclease variant thereof. The invention provides a method of producing nucleic acid fragments, the method comprising a step for fragmenting chromatin within a population of permeabilised cells with a CATH Superfamily 3.40.570.10 Extracellular Endonuclease, subunit A nuclease, or a nuclease variant thereof.

[0012] The invention provides a method of producing nucleic acid fragments, the method comprising fragmenting chromatin within a population of permeabilised cells with a means for fragmenting chromatin, wherein the means is a CATH Superfamily 3.40.570.10 Extracellular Endonuclease, subunit A nuclease, or a nuclease variant thereof.

[0013] The invention also provides a method of producing nucleic acid fragments, the method comprising fragmenting chromatin within a population of permeabilised cells with a nuclease, wherein the nuclease comprises a cd00091 active site.

[0014] The invention also provides a method of producing nucleic acid fragments, the method comprising a step for fragmenting chromatin within a population of permeabilised cells with a nuclease, wherein the nuclease comprises a cd00091 active site.

[0015] The invention also provides a method of producing nucleic acid fragments, the method comprising fragmenting chromatin within a population of permeabilised cells with a means for fragmenting chromatin, wherein the means comprises a cd00091 active site.

[0016] The invention also provides a method of producing nucleic acid fragments, the method comprising fragmenting chromatin within a population of permeabilised cells with a nuclease, wherein the nuclease has a ball and chain morphology.

[0017] The invention also provides a method of producing nucleic acid fragments, the method comprising a step for fragmenting chromatin within a population of permeabilised cells with a nuclease, wherein the nuclease has a ball and chain morphology.

[0018] The invention also provides a method of producing nucleic acid fragments, the method comprising fragmenting chromatin within a population of permeabilised cells with a means for fragmenting chromatin, wherein the means has a ball and chain morphology.

[0019] The invention also provides a method of producing nucleic acid fragments, the method comprising fragmenting chromatin within a population of permeabilised cells with a nuclease comprising an amino acid sequence having at least 40% identity to SEQ ID NO: 1 or SEQ ID NO: 2.

[0020] The invention also provides a method of producing nucleic acid fragments, the method comprising a step for fragmenting chromatin within a population of permeabilised cells with a nuclease comprising an amino acid sequence having at least 40% identity to SEQ ID NO: 1 or SEQ ID NO: 2.

[0021] The invention also provides a method of producing nucleic acid fragments, the method comprising fragmenting chromatin within a population of permeabilised cells with a means for fragmenting chromatin, wherein the means comprises an amino acid sequence having at least 40% identity to SEQ ID NO: 1 or SEQ ID NO: 2.

[0022] The invention also provides a method of producing nucleic acid fragments, the method comprising fragmenting chromatin (e.g. in a population of permeabilised cells) with a nuclease (e.g., a nuclease described herein).

[0023] The invention also provides a method of producing nucleic acid fragments, the method comprising a step for fragmenting chromatin (e.g. in a population of permeabilised cells) with a nuclease (e.g., a nuclease described herein).

[0024] The invention also provides a method of producing nucleic acid fragments, the method comprising fragmenting chromatin (e.g. in a population of permeabilised cells) with a means for fragmenting chromatin (e.g., a nuclease described herein).

[0025] The means for fragmenting chromatin used in any of the methods of the invention may be a nuclease described herein. The means for fragmenting chromatin (e.g., nuclease) used in any of the methods of the invention may be a CATH Superfamily 3.40.570.10 Extracellular Endonuclease, subunit A nuclease, or a nuclease variant thereof. The means for fragmenting chromatin (e.g., nuclease) used in any of the methods of the invention may comprise a cd00091 active site. The means for fragmenting chromatin (e.g., nuclease) used in any of the methods of the invention may have a molecular weight of greater than 17 kDa. The means for fragmenting chromatin (e.g., nuclease) used in any of the methods of the invention may have a molecular weight of at least 20 kDa and up to 50 kDa. The means for fragmenting chromatin (e.g., nuclease) used in any of the methods of the invention may have a ball domain and optionally a chain domain. The means for fragmenting chromatin (e.g., nuclease) used in any of the methods of the invention may have a ball and chain morphology. The means for fragmenting chromatin (e.g., nuclease) used in any of the methods of the invention may be an endonuclease, e.g., a non-specific endonuclease, such as a DNA / RNA nonspecific endonuclease. The means for fragmenting chromatin (e.g., nuclease) used in any of the methods of the invention may be an endonuclease with minimal exonuclease activity. The means for fragmenting chromatin (e.g., nuclease) used in any of the methods of the invention may be an endonuclease which does not have exonuclease activity. The means for fragmenting chromatin (e.g., nuclease) used in any of the methods of the invention may comprise an amino acid sequence having at least 40% identity to SEQ ID NO: 1 or SEQ ID NO: 2 (e.g., may comprise the amino acid sequence of SEQ ID NO: 1 or SEQ ID NO: 2).

[0026] The invention also provides a pool of nucleic acid fragments obtainable, or obtained, by a method of the invention.

[0027] The invention also provides a pool of nucleic acid fragments obtainable, or obtained, by fragmenting chromatin with a nuclease described herein within a population of permeabilised cells.

[0028] The invention also provides a pool of nucleic acid fragments, wherein the pool of nucleic acid fragments is present in a population of cells (e.g. a population of permeabilised cells described herein), and wherein: (a) 15-35% of the nucleic acid fragments have a 3’ terminal C residue; (b) 15-35% of the nucleic acid fragments have a 3’ terminal G residue; (c) 15-35% of the nucleic acid fragments have a 3’ terminal T residue; and (d) 15-35% of the nucleic acid fragments have a 3’ terminal A residue, wherein the sum of (a), (b), (c) and (d) is 100%. In some instances, (e) 10-40% of the nucleic acid fragments have a 5’ terminal C residue; (f) 10-40% of the nucleic acid fragments have a 5’ terminal G residue; (g) 10-40% of the nucleic acid fragments have a 5’ terminal T residue; and (h) 10-40% of the nucleic acid fragments have a 5’ terminal A residue, wherein the sum of (e), (f), (g) and (h) is 100%.

[0029] The invention also provides a pool of nucleic acid fragments, wherein the pool of nucleic acid fragments is present in a population of cells (e.g. a population of permeabilised cells described herein), and wherein: (a) 10-40% of the nucleic acid fragments have a 5’ terminal C residue; (b) 10-40% of the nucleic acid fragments have a 5’ terminal G residue; (c) 10-40% of the nucleic acid fragments have a 5’ terminal T residue; and (d) 10-40% of the nucleic acid fragments have a 5’ terminal A residue, wherein the sum of (a), (b), (c) and (d) is 100%. In some instances, (e) acid fragments have a 3’ terminal G residue; (g) 15-35% of the nucleic acid fragments have a 3’ terminal T residue; and (h) 15-35% of the nucleic acid fragments have a 3’ terminal A residue, wherein the sum of (e), (f), (g) and (h) is 100%.

[0030] The invention also provides a method of producing a 3C library, the method comprising (a) ligating cross-linked nucleic acid fragments obtainable, or which have been obtained, by a method of the invention, and (b) de-crosslinking the ligated nucleic acid fragments.

[0031] The invention also provides a method of producing a 3C library, the method comprising (a) ligating a pool of cross-linked nucleic acid fragments of the invention; and (b) de-crosslinking the ligated nucleic acid fragments.

[0032] The invention also provides a method of producing a 3C library, the method comprising the steps:

[0033] (a) fragmenting chromatin by a method of the invention to produce immobilised (e.g. cross-linked) nucleic acid fragments;

[0034] (b) ligating the nucleic acid fragments to produce ligated nucleic acid fragments; and

[0035] (c) de-immobilising (e.g. de-crosslinking) the ligated nucleic acid fragments.

[0036] The invention also provides a method of producing a 3C library, the method comprising the steps:

[0037] (a) cross-linking chromatin (e.g. within a population of permeabilised cells);

[0038] (b) fragmenting the cross-linked chromatin (e.g. within the population of permeabilised cells) using a nuclease described herein to produce cross-linked nucleic acid fragments;

[0039] (c) ligating the cross-linked nucleic acid fragments to produce ligated nucleic acid fragments; and

[0040] (d) de-crosslinking the ligated nucleic acid fragments.

[0041] The invention also provides a method of producing a 3C library, the method comprising (a) a step for ligating cross-linked nucleic acid fragments obtainable, or which have been obtained, by a method of producing nucleic acid fragments of the invention, and (b) a step for decrosslinking the ligated nucleic acid fragments. The invention also provides a method of producing a 3C library, the method comprising (a) a step for ligating a pool of cross-linked nucleic acid fragments of the invention; and (b) a step for de-crosslinking the ligated nucleic acid fragments.

[0042] The invention also provides a method of producing a 3C library, the method comprising:

[0043] (a) a step for fragmenting chromatin (e.g., using a method of the invention) to produce immobilised (e.g. cross-linked) nucleic acid fragments;

[0044] (b) a step for ligating the nucleic acid fragments to produce ligated nucleic acid fragments; and

[0045] (c) a step for de-immobilising (e.g. de-crosslinking) the ligated nucleic acid fragments.

[0046] The invention also provides a method of producing a 3C library, the method comprising:

[0047] (a) a step for cross-linking chromatin (e.g. within a population of permeabilised cells);

[0048] (b) a step for fragmenting the cross-linked chromatin (e.g., within the population of permeabilised cells), (e.g., using a nuclease described herein) to produce cross-linked nucleic acid fragments;

[0049] (c) a step for ligating the cross-linked nucleic acid fragments to produce ligated nucleic acid fragments; and

[0050] (d) a step for de-crosslinking the ligated nucleic acid fragments.

[0051] The invention also provides a 3C library of nucleic acid fragments obtainable, or obtained, by a method of the invention.

[0052] The invention also provides a 3C library of nucleic acid fragments, wherein: (a) 15-35% of the nucleic acid fragments have a 3’ terminal C residue; (b) 15-35% of the 3C nucleic acid fragments have a 3’ terminal G residue; (c) 15-35% of the nucleic acid fragments have a 3’ terminal T residue; and (d) 15-35% of the nucleic acid fragments have a 3’ terminal A residue, wherein the sum of (a), (b), (c) and (d) is 100%. In some instances, (e) 10-40% of the nucleic acid fragments have a 5’ terminal C residue; (f) 10-40% of the 3C nucleic acid fragments have a 5’ terminal G residue; (g) 10-40% of the nucleic acid fragments have a 5’ terminal T residue; and (h) 10-40% of the nucleic acid fragments have a 5’ terminal A residue, wherein the sum of

[0053] (e), (f), (g) and (h) is 100%.

[0054] The invention also provides a 3C library of nucleic acid fragments, wherein: (a) 10-40% of the nucleic acid fragments have a 5’ terminal C residue; (b) 10-40% of the nucleic acid fragments have a 5’ terminal G residue; (c) 10-40% of the nucleic acid fragments have a 5’ terminal T residue; and (d) 10-40% of the nucleic acid fragments have a 5’ terminal A residue, wherein the sum of (a), (b), (c) and (d) is 100%. In some instances, (e) 15-35% of the nucleic acid fragments have a 3’ terminal C residue; (f) 15-35% of the 3C nucleic acid fragments have a 3’ terminal G residue; (g) 15-35% of the nucleic acid fragments have a 3’ terminal T residue; and (h) 15-35% of the nucleic acid fragments have a 3’ terminal A residue, wherein the sum of (e), (f), (g) and (h) is 100%.

[0055] The invention also provides a method of identifying chromatin regions within a nucleic acid sample which interact with one another, the method comprising:

[0056] (a) fragmenting a 3C library obtainable, or which has been obtained, by a method of the invention;

[0057] (b) optionally, adding sequencing adaptors to the ends of the nucleic acid fragments and / or amplifying the nucleic acid fragments;

[0058] (c) contacting the nucleic acid fragments with a targeting nucleic acid which binds to a subgroup of the nucleic acid fragments, wherein the targeting nucleic acid is labelled with the first half of a binding pair;

[0059] (d) isolating the nucleic acid fragments which have been bound by the targeting nucleic acid using the second half of the binding pair;

[0060] (e) amplifying the isolated subgroup of nucleic acid fragments;

[0061] (f) optionally repeating steps (c), (d), and (e) one or more times; and

[0062] (g) optionally sequencing the amplified isolated subgroup of nucleic acid fragments, in order to identify chromatin regions within the nucleic acid sample which interact with one another.

[0063] The invention also provides a method of identifying chromatin regions within a nucleic acid sample which interact with one another, the method comprising:

[0064] (a) fragmenting a 3C library of the invention;

[0065] (b) optionally, adding sequencing adaptors to the ends of the nucleic acid fragments and / or amplifying the nucleic acid fragments;

[0066] (c) contacting the nucleic acid fragments with a targeting nucleic acid which binds to a subgroup of the nucleic acid fragments, wherein the targeting nucleic acid is labelled with the first half of a binding pair;

[0067] (d) isolating the nucleic acid fragments which have been bound by the targeting nucleic acid using the second half of the binding pair;

[0068] (e) amplifying the isolated subgroup of nucleic acid fragments;

[0069] (f) optionally repeating steps (c), (d), and (e) one or more times; and (g) optionally sequencing the amplified isolated subgroup of nucleic acid fragments, in order to identify chromatin regions within the nucleic acid sample which interact with one another.

[0070] The invention also provides a method of identifying chromatin regions within a nucleic acid sample which interact with one another, the method comprising:

[0071] (a) producing a 3C library using a method of the invention;

[0072] (b) fragmenting the 3C library to produce nucleic acid fragments;

[0073] (c) optionally, adding sequencing adaptors to the ends of the nucleic acid fragments and / or amplifying the nucleic acid fragments;

[0074] (d) contacting the nucleic acid fragments with a targeting nucleic acid which binds to a subgroup of the nucleic acid fragments, wherein the targeting nucleic acid is labelled with the first half of a binding pair;

[0075] (e) isolating the subgroup of nucleic acid fragments which have been bound by the targeting nucleic acid using the second half of the binding pair;

[0076] (f) amplifying the isolated subgroup of nucleic acid fragments;

[0077] (g) optionally repeating Steps (d), (e) and (f) one or more times; and

[0078] (h) optionally sequencing the amplified isolated subgroup of nucleic acid fragments, in order to identify chromatin regions within the nucleic acid sample which interact with one another.

[0079] The invention also provides a method of identifying chromatin regions within a nucleic acid sample which interact with one another, the method comprising:

[0080] (a) a step for fragmenting a 3C library obtainable, or which has been obtained, by a method of the invention;

[0081] (b) optionally, a step for adding sequencing adaptors to the ends of the nucleic acid fragments and / or amplifying the nucleic acid fragments;

[0082] (c) a step for contacting the nucleic acid fragments with a targeting nucleic acid which binds to a subgroup of the nucleic acid fragments, wherein the targeting nucleic acid is labelled with the first half of a binding pair;

[0083] (d) a step for isolating the nucleic acid fragments which have been bound by the targeting nucleic acid using the second half of the binding pair;

[0084] (e) a step for amplifying the isolated subgroup of nucleic acid fragments;

[0085] (f) optionally repeating steps (c), (d), and (e) one or more times; and

[0086] (g) optionally, a step for sequencing the amplified isolated subgroup of nucleic acid fragments, in order to identify chromatin regions within the nucleic acid sample which interact with one another.

[0087] The invention also provides a method of identifying chromatin regions within a nucleic acid sample which interact with one another, the method comprising:

[0088] (a) a step for fragmenting a 3C library of the invention;

[0089] (b) optionally, a step for adding sequencing adaptors to the ends of the nucleic acid fragments and / or amplifying the nucleic acid fragments;

[0090] (c) a step for contacting the nucleic acid fragments with a targeting nucleic acid which binds to a subgroup of the nucleic acid fragments, wherein the targeting nucleic acid is labelled with the first half of a binding pair;

[0091] (d) a step for isolating the nucleic acid fragments which have been bound by the targeting nucleic acid using the second half of the binding pair;

[0092] (e) a step for amplifying the isolated subgroup of nucleic acid fragments;

[0093] (f) optionally repeating steps (c), (d), and (e) one or more times; and

[0094] (g) optionally, a step for sequencing the amplified isolated subgroup of nucleic acid fragments, in order to identify chromatin regions within the nucleic acid sample which interact with one another.

[0095] The invention also provides a method of identifying chromatin regions within a nucleic acid sample which interact with one another, the method comprising:

[0096] (a) a step for producing a 3C library (e.g., using a method of the invention);

[0097] (b) a step for fragmenting the 3C library to produce nucleic acid fragments;

[0098] (c) optionally, a step for adding sequencing adaptors to the ends of the nucleic acid fragments and / or amplifying the nucleic acid fragments;

[0099] (d) a step for contacting the nucleic acid fragments with a targeting nucleic acid which binds to a subgroup of the nucleic acid fragments, wherein the targeting nucleic acid is labelled with the first half of a binding pair;

[0100] (e) a step for isolating the subgroup of nucleic acid fragments which have been bound by the targeting nucleic acid using the second half of the binding pair;

[0101] (f) a step for amplifying the isolated subgroup of nucleic acid fragments;

[0102] (g) optionally repeating Steps (d), (e) and (f) one or more times; and

[0103] (h) optionally, a step for sequencing the amplified isolated subgroup of nucleic acid fragments, in order to identify chromatin regions within the nucleic acid sample which interact with one another.

[0104] The invention also provides a method of identifying allele-specific interaction profiles in SNP- containing regions of nucleic acids, the method comprising sequencing the amplified isolated subgroup of nucleic acid fragments which has been obtained by a method of the invention, in order to identify allele-specific interaction profiles in SNP-containing regions.

[0105] The invention also provides a method of identifying allele-specific interaction profiles in SNP- containing regions of nucleic acids, the method comprising a step for sequencing the amplified isolated subgroup of nucleic acid fragments which has been obtained by a method of the invention, in order to identify allele-specific interaction profiles in SNP-containing regions.

[0106] The invention also provides a method of identifying allele-specific interaction profiles in SNP- containing regions of nucleic acids, the method comprising carrying out a method of identifying chromatin regions within a nucleic acid sample which interact with one another of the invention, and sequencing the amplified isolated subgroup of nucleic acid fragments in order to identify allele-specific interaction profiles in SNP-containing regions.

[0107] The invention also provides a method of identifying a chromatin region that is indicative of a disease, the method comprising:

[0108] (a) quantifying a frequency of interaction between a first chromatin region and a second chromatin region within a nucleic acid sample from a subject with the disease, wherein the first chromatin region and second chromatin region have been identified as interacting with one another by a method of the invention; and

[0109] (b) comparing the frequency of the interaction in the nucleic acid sample from the subject with the disease with the frequency of interaction in a nucleic acid sample from a subject without the disease, wherein a difference in the frequency of interaction between the nucleic acid samples is indicative of the disease.

[0110] The invention also provides a method of identifying a chromatin region that is indicative of a disease, the method comprising:

[0111] (a) carrying out a method of identifying chromatin regions within a nucleic acid sample which interact with one another of the invention wherein the 3C library is from a subject with a disease; (b) quantifying a frequency of interaction between a first chromatin region and a second chromatin region in the nucleic acid sample; and

[0112] (c) comparing the frequency of the interaction between the first chromatin region and the second chromatin region from the subject with the disease with the frequency of interaction between the first chromatin region and the second chromatin region from a subject without the disease, wherein a difference in the frequency of interaction is indicative of the disease.

[0113] The invention also provides a method of identifying a chromatin region that is indicative of a disease, the method comprising:

[0114] (a) a step for quantifying a frequency of interaction between a first chromatin region and a second chromatin region within a nucleic acid sample from a subject with the disease, wherein the first chromatin region and second chromatin region have been identified as interacting with one another by a method of the invention; and

[0115] (b) a step for comparing the frequency of the interaction in the nucleic acid sample from the subject with the disease with the frequency of interaction in a nucleic acid sample from a subject without the disease, wherein a difference in the frequency of interaction between the nucleic acid samples is indicative of the disease.

[0116] The invention also provides a method of identifying a chromatin region that is indicative of a disease, the method comprising:

[0117] (a) a step for carrying out a method of identifying chromatin regions within a nucleic acid sample which interact with one another of the invention, wherein the 3C library is from a subject with a disease;

[0118] (b) a step for quantifying a frequency of interaction between a first chromatin region and a second chromatin region in the nucleic acid sample; and

[0119] (c) a step for comparing the frequency of the interaction between the first chromatin region and the second chromatin region from the subject with the disease with the frequency of interaction between the first chromatin region and the second chromatin region from a subject without the disease, wherein a difference in the frequency of interaction is indicative of the disease.

[0120] The invention also provides a method comprising producing a means for targeting, or a means which has been identified as targeting, an expression product of a nucleotide sequence, wherein the nucleotide sequence, or a nucleotide sequence having at least 90%, 95%, 97%, or 99% identity thereto, is comprised by a chromatin region which has been identified as being indicative of a disease by a method of the invention.

[0121] The invention also provides a method comprising a step for producing a means for targeting, or a means which has been identified as targeting, an expression product of a nucleotide sequence, wherein the nucleotide sequence, or a nucleotide sequence having at least 90%, 95%, 97%, or 99% identity thereto, is comprised by a chromatin region which has been identified as being indicative of a disease by a method of the invention

[0122] The invention also provides a method comprising producing an agent which targets, or has been identified as targeting, an expression product of nucleotide sequence, wherein the nucleotide sequence, or a nucleotide sequence having at least 90%, 95%, 97%, or 99% identity thereto, is comprised by a chromatin region which has been identified as being indicative of a disease by a method of the invention.

[0123] The invention also provides a method comprising:

[0124] (a) identifying a chromatin region that is indicative of a disease by carrying out a method of identifying a chromatin region that is indicative of a disease of the invention; and

[0125] (b) producing a means for targeting, or a means which has been identified as targeting, an expression product of a nucleotide sequence, wherein the nucleotide sequence, or a nucleotide sequence having at least 90%, 95%, 97%, or 99% identity thereto, is comprised by the chromatin region identified in step (a).

[0126] The invention also provides a method comprising:

[0127] (a) identifying a chromatin region that is indicative of a disease by carrying out a method of identifying a chromatin region that is indicative of a disease of the invention; and

[0128] (b) a step for producing a means for targeting, or a means which has been identified as targeting, an expression product of a nucleotide sequence, wherein the nucleotide sequence, or a nucleotide sequence having at least 90%, 95%, 97%, or 99% identity thereto, is comprised by the chromatin region identified in step (a).

[0129] The invention also provides a method comprising:

[0130] (a) identifying a chromatin region that is indicative of a disease by carrying out a method of identifying a chromatin region that is indicative of a disease of the invention; and

[0131] (b) producing an agent which targets, or which has been identified as targeting, an expression product of a nucleotide sequence, wherein the nucleotide sequence, or a nucleotide sequence having at least 90%, 95%, 97%, or 99% identity thereto, is comprised by the chromatin region identified in step (a).

[0132] The invention also provides a method comprising producing a means for targeting, or a means which has been identified as targeting, an expression product regulated by a nucleotide sequence, wherein the nucleotide sequence, or a nucleotide sequence having at least 90%, 95%, 97%, or 99% identity thereto, is comprised by a chromatin region which has been identified as being indicative of a disease by a method of the invention.

[0133] The invention also provides a method comprising a step for producing a means for targeting, or a means which has been identified as targeting, an expression product regulated by a nucleotide sequence, wherein the nucleotide sequence, or a nucleotide sequence having at least 90%, 95%, 97%, or 99% identity thereto, is comprised by a chromatin region which has been identified as being indicative of a disease by a method of the invention.

[0134] The invention also provides a method comprising producing an agent which targets an expression product regulated by a nucleotide sequence, wherein the nucleotide sequence, or a nucleotide sequence having at least 90%, 95%, 97%, or 99% identity thereto, is also comprised by a chromatin region which has been identified as being indicative of a disease by a method of the invention.

[0135] The invention also provides a method comprising:

[0136] (a) identifying a chromatin region that is indicative of a disease by carrying out a method of identifying a chromatin region that is indicative of a disease of the invention; and

[0137] (b) producing a means for targeting, or a means which has been identified as targeting, an expression product regulated by a nucleotide sequence, wherein the nucleotide sequence, or a nucleotide sequence having at least 90%, 95%, 97%, or 99% identity thereto, is comprised by the chromatin region identified in step (a).

[0138] The invention also provides a method comprising:

[0139] (a) identifying a chromatin region that is indicative of a disease by carrying out a method of identifying a chromatin region that is indicative of a disease of the invention; and

[0140] (b) a step for producing a means for targeting, or a means which has been identified as targeting, an expression product regulated by a nucleotide sequence, wherein the nucleotide sequence, or a nucleotide sequence having at least 90%, 95%, 97%, or 99% identity thereto, is comprised by the chromatin region identified in step (a). The invention also provides a method comprising:

[0141] (a) identifying a chromatin region that is indicative of a disease by carrying out a method of identifying a chromatin region that is indicative of a disease of the invention; and

[0142] (b) producing an agent which targets, or which has been identified as targeting, an expression product regulated by a nucleotide sequence, wherein the nucleotide sequence, or a nucleotide sequence having at least 90%, 95%, 97%, or 99% identity thereto, is also comprised by the chromatin region identified in step (a).

[0143] The invention also provides an agent obtainable, or which has been obtained, by a method of the invention.

[0144] The invention also provides a mean for targeting, or which has been identified as targeting, an expression product described herein.

[0145] The invention also provides a method of producing a pharmaceutical composition, the method comprising combining:

[0146] (a) an agent of the invention; or

[0147] (b) an agent obtainable, or which has been obtained, by a method of the invention, with a pharmaceutically acceptable carrier, excipient or diluent.

[0148] The invention also provides a method of producing a pharmaceutical composition, the method comprising combining:

[0149] (a) a means for targeting an expression product described herein; or

[0150] (b) a means which has been identified as targeting an expression product described herein, with a pharmaceutically acceptable carrier, excipient or diluent.

[0151] The invention also provides a pharmaceutical composition obtainable, or which has been obtained, by a method of producing a pharmaceutical composition of the invention.

[0152] The invention also provides a method comprising:

[0153] (a) carrying out a method of identifying a chromatin region that is indicative of a disease of the invention;

[0154] (b) producing an agent which targets, or which has been identified as targeting, an expression product comprising a nucleotide sequence, wherein the nucleotide sequence, or a nucleotide sequence having at least 90%, 95%, 97%, or 99% identity thereto, is also comprised by the chromatin region identified in step (a); and

[0155] (c) administering the agent to the subject to treat a disease or condition.

[0156] The invention provides a method comprising:

[0157] (a) carrying out a method of identifying a chromatin region that is indicative of a disease of the invention;

[0158] (b) producing a means for targeting, or a means which has been identified as targeting, an expression product comprising a nucleotide sequence, wherein the nucleotide sequence, or a nucleotide sequence having at least 90%, 95%, 97%, or 99% identity thereto, is also comprised by the chromatin region identified in step (a); and

[0159] (c) administering the means to the subject to treat a disease or condition.

[0160] The invention also provides a method comprising:

[0161] (a) carrying out a method of identifying a chromatin region that is indicative of a disease of the invention; and

[0162] (b) producing an agent which targets an expression product regulated by a nucleotide sequence, wherein the nucleotide sequence, or a nucleotide sequence having at least 90%, 95%, 97%, or 99% identity thereto, is also comprised by the chromatin region identified in step

[0163] (a); and

[0164] (c) administering the agent to the subject to treat a disease or condition.

[0165] The invention also provides a method comprising:

[0166] (a) carrying out a method of identifying a chromatin region that is indicative of a disease of the invention; and

[0167] (b) producing a means for targeting, or a means which has been identified as targeting, an expression product regulated by a nucleotide sequence, wherein the nucleotide sequence, or a nucleotide sequence having at least 90%, 95%, 97%, or 99% identity thereto, is also comprised by the chromatin region identified in step (a); and

[0168] (c) administering the means to the subject to treat a disease or condition.

[0169] The invention also provides a method of treating or preventing a disease or condition in a human or animal subject, the method comprising administering an agent of the invention or a pharmaceutical composition of the invention to the subject. The invention also provides an agent of the invention, or a pharmaceutical composition of the invention, for use in a method of treating or preventing a disease or condition in a human or animal subject, the method comprising administering the agent or pharmaceutical composition to the subject.

[0170] The invention also provides use of an agent of the invention, or a pharmaceutical composition of the invention in the manufacture of a medicament for treating or preventing a disease or condition in a human or animal subject.

[0171] The invention also provides a method of treating or preventing a disease or condition in a human or animal subject, the method comprising administering a means for targeting, or a means which has been identified as targeting, an expression product described herein to the subject.

[0172] The invention also provides a means for targeting, or a means which has been identified as targeting, an expression product described herein, for use in a method of treating or preventing a disease or condition in a human or animal subject, the method comprising administering the agent or pharmaceutical composition to the subject.

[0173] The invention also provides use of a means for targeting, or a means which has been identified as targeting, an expression product described herein, in the manufacture of a medicament for treating or preventing a disease or condition in a human or animal subject.

[0174] The invention also provides a kit comprising:

[0175] (a) a CATH Superfamily 3.40.570.10 Extracellular Endonuclease, subunit A nuclease, or a nuclease variant thereof;

[0176] (b) a nuclease comprising a cd00091 active site;

[0177] (c) a nuclease comprising an amino acid sequence of A / S / D / Q / T / E-K / R / Y / l / H-X- H / Q ' / F-X(a)-N / F / G-X(7)-E / L / M / H / Q-X(3)-R / M / LA / / F, where X is any amino acid and a is 23 to 37;

[0178] (d) a nuclease having a ball and chain morphology;

[0179] (e) a nuclease comprising an amino acid sequence having at least 40% identity to SEQ ID NO: 1 or SEQ ID NO: 2; or

[0180] (f) a means for fragmenting chromatin of the invention; and instructions for use of the nuclease to fragment chromatin within a population of permeabilised cells. The invention also provides an apparatus comprising:

[0181] (a) means to fragment chromatin in a population of permeabilised cells with:

[0182] (i) a CATH Superfamily 3.40.570.10 Extracellular Endonuclease, subunit A nuclease, or a nuclease variant thereof

[0183] (ii) a nuclease comprising a cd00091 active site;

[0184] (iii) a nuclease comprising an amino acid sequence of A / S / D / Q / T / E-K / R / Y / l / H- X-H / Q / Y / F-X(a)-N / F / G-X(7)-E / L / M / H / Q-X(3)-R / M / UV / F, where X is any amino acid and a is 23 to 37;

[0185] (iv) a nuclease having a ball and chain morphology;

[0186] (v) a nuclease comprising an amino acid sequence having at least 40% identity to SEQ ID NO: 1 or SEQ ID NO: 2; or

[0187] (vi) a means for fragmenting chromatin of the invention; and

[0188] (b) instructions which, when executed, cause the apparatus to carry out a method of producing nucleic acid fragments of the invention, and optionally to carry out:

[0189] (i) a method of producing a 3C library of the invention;

[0190] (ii) a method of identifying chromatin regions within a nucleic acid sample which interact with one another of the invention;

[0191] (iii) a method of identifying allele-specific interaction profiles in SNP-containing regions of nucleic acids of the invention; and / or

[0192] (iv) a method of identifying a chromatin region that is indicative of a disease of the invention.

[0193] The invention also provides a method of identifying a nuclease suitable for fragmenting chromatin within a population of permeabilised cells, the method comprising identifying:

[0194] (a) a CATH Superfamily 3.40.570.10 Extracellular Endonuclease, subunit A nuclease, or a nuclease variant thereof;

[0195] (b) a nuclease comprising a cd00091 active site;

[0196] (c) a nuclease comprising an amino acid sequence of A / S / D / Q / T / E-K / R / Y / l / H-X- H / Q / Y / F-X(a)-N / F / G-X(7)-E / L / M / H / Q-X(3)-R / M / L / V / F, where X is any amino acid and a is 23 to 37;

[0197] (d) a nuclease having a ball and chain morphology; or

[0198] (e) a nuclease comprising an amino acid sequence having at least 40% identity to SEQ ID NO: 1 or SEQ ID NO: 2; and determining whether the nuclease fragments chromatin within a population of permeabilised cells. The invention also provides a method of fragmenting chromatin within a population of permeabilised cells with a nuclease which has been identified as suitable for fragmenting chromatin within a population of permeabilised cells by a method of the invention.

[0199] Any of the nucleases described herein may be used in any of the methods of the invention. The means for fragmenting chromatin used in any of the methods of the invention may be a nuclease. Any description of a nuclease provided herein is equally applicable to a means for fragmenting chromatin.

[0200] Nucleases may be clustered into groups of related nucleases based on their sequence or structure. Clustering nucleases by their sequence or structure enables the grouping of related nucleases which share functional properties. Accordingly, the nuclease may be a nuclease which is present in a sequence- and / or structure-based cluster comprising Northern shrimp nuclease having UniProt ID C9YSL6. The nuclease may be a nuclease which is present in a structure-based cluster comprising Northern shrimp nuclease having UniProt ID C9YSL6. The nuclease may be a nuclease which is present in a sequence-based cluster comprising Northern shrimp nuclease having UniProt ID C9YSL6. The nuclease may be a nuclease which is present in a sequence- and structure- based cluster comprising Northern shrimp nuclease having UniProt ID C9YSL6. Exemplary sequence and structure clusters comprising Northern shrimp nuclease having UniProt ID C9YSL6 are described herein.

[0201] The CATH database is an exemplary hierarchical domain classification of protein structures in the Protein Data Bank. There are four major levels in the CATH hierarchy: (1) Class - structures are classified according to their secondary structure composition; (2) Architecture - structures are classified according to their overall shape as determined by the orientations of the secondary structures in 3D space but ignores the connectivity between them; (3) Topology (fold family) - structures are grouped into fold groups at this level depending on both the overall shape and connectivity of the secondary structures; and (4) Homologous superfamily - this level groups together protein domains which are through to share a common ancestor and can therefore be described as homologous. Thus, protein structures that fall within the same superfamily are likely to have diverged from a common ancestor and have similar structural and functional properties. Protein structures in the CATH database are classified using the CATH code. The first number in the CATH code represents the class; the second number represents the architecture; the third number represents the topology; and the fourth number represents the homologous superfamily. Thus, all protein structures in CATH Superfamily 3.40.570.10 Extracellular Endonuclease, subunit A nuclease are in the same class (3), architecture (40), topology (570), and homologous superfamily (10). Northern shrimp nuclease having UniProt ID C9YSL6 is part of CATH Superfamily 3.40.570.10 Extracellular Endonuclease, subunit A nuclease cluster. Accordingly, the nuclease may be a CATH Class 3 nuclease, or a nuclease variant thereof. The nuclease may be a CATH 3.40 nuclease, or a nuclease variant thereof. The nuclease may be a CATH 3.40.570 nuclease, or a nuclease variant thereof. In particular, the nuclease may be a CATH Superfamily 3.40.570.10 Extracellular Endonuclease, subunit A nuclease, or a nuclease variant thereof. The nuclease may be a CATH Superfamily 3.40.570.10 Extracellular Endonuclease, subunit A nuclease.

[0202] InterPro provides functional analysis of proteins by classifying them into families and predicting domains and important sites. Northern shrimp nuclease having UniProt ID C9YSL6 is part of the IPR001604, IPR044929, IPR020821 , IPR044925, and IPR040255 clusters. Accordingly, the nuclease may be an IPR001604, IPR044929, IPR020821 , IPR044925, and / or IPR040255 nuclease, or a nuclease variant thereof. The nuclease may be an IPR001604 nuclease, or a nuclease variant thereof. The nuclease may be an IPR044929 nuclease, or a nuclease variant thereof. The nuclease may be an IPR020821 nuclease, or a nuclease variant thereof. The nuclease may be an IPR044925 nuclease, or a nuclease variant thereof. The nuclease may be an IPR040255 nuclease, or a nuclease variant thereof. The nuclease may be an IPR001604, IPR044929, IPR020821 , IPR044925, and IPR040255 nuclease, or a nuclease variant thereof.

[0203] The PANTHER (Protein Analysis Through Evolutionary Relationships) Classification System was designed to classify proteins and their genes. Northern shrimp nuclease having UniProt ID C9YSL6 is part of the PTHR13966 and PTHR13966:SF17 clusters. Accordingly, the nuclease may be a PTHR13966 and / or PTHR13966:SF17 nuclease, or a nuclease variant thereof. The nuclease may be a PTHR13966 nuclease, or a nuclease variant thereof. The nuclease may be a PTHR13966:SF17 nuclease, or a nuclease variant thereof. The nuclease may be a PTHR13966 and PTHR13966:SF17 nuclease, or a nuclease variant thereof.

[0204] The Pfam database is a collection of protein families, each represented by multiple sequence alignments and hidden Markov models (HMM). Northern shrimp nuclease having UniProt ID C9YSL6 is part of the PF01223 cluster. Accordingly, the nuclease may be a PF01223 nuclease, or a nuclease variant thereof. SMART (a Simple Modular Architectural Research Tool) allows the identification and annotation of genetically mobile domains and the analysis of domain architectures. Northern shrimp nuclease having UniProt ID C9YSL6 is part of the SM00892 and SM00477 clusters. Accordingly, the nuclease may be a SM00892 and / or SM00477 nuclease, or a nuclease variant thereof. The nuclease may be an SM00892 nuclease, or a nuclease variant thereof. The nuclease may be an SM00477 nuclease, or a nuclease variant thereof. The nuclease may be a SM00892 and SM00477 nuclease, or a nuclease variant thereof.

[0205] The Conserved Domain Database is a resource for the annotation of functional units in proteins. Northern shrimp nuclease having UniProt ID C9YSL6 is part of the cd00091 cluster of the Conserved Domain Database. Accordingly, the nuclease may be a cd00091 nuclease, or a nuclease variant thereof.

[0206] AlphaFold™ Clusters is an exemplary clustering tool which clusters proteins based on their structure, as described in Barrio-Hernandez et al. (2023) Nature 622(7983) :637-645. In brief, AlphaFold™ Clusters uses a two-step approach to cluster proteins present in the AlphaFold™ Protein Structure Database (AFDB). First MMseqs2 is used to cluster protein sequences on the basis of 50% sequence identity and 90% sequence overlap. For each cluster, the protein with the highest confidence structure score (pLDDT) is selected as the representative. Next, using Foldseek, the representative structures are clustered without a sequence identity threshold, but still enforcing a 90% sequence overlap and an E-value of less than 0.01 for each structural alignment. Subsequently, all sequences labelled as fragments are removed from the clustering. Each resulting cluster comprises proteins predicted to have structural and functional similarity. Accordingly, the nuclease may be a nuclease which is present in an AlphaFold™ Cluster comprising Northern shrimp nuclease having UniProt ID C9YSL6, or a nuclease variant thereof. The nuclease may be a nuclease which is present in an AlphaFold™ Cluster MMseqs2 cluster comprising Northern shrimp nuclease having UniProt ID C9YSL6, or a nuclease variant thereof. The nuclease may be a nuclease which is present in an AlphaFold™ Cluster Foldseek cluster comprising Northern shrimp nuclease having UniProt ID C9YSL6, or a nuclease variant thereof. The nuclease may be a nuclease which is present in an AlphaFold™ Cluster Fragment cluster comprising Northern shrimp nuclease having UniProt ID C9YSL6, or a nuclease variant thereof. The nuclease may be a nuclease which is present in an AlphaFold™ Cluster Singleton cluster comprising Northern shrimp nuclease having UniProt ID C9YSL6, or a nuclease variant thereof. A nuclease variant of a nuclease may retain the ability to fragment chromatin (e.g. in a population of permeabilised cells). A nuclease variant may be tested for nuclease activity using any suitable assay, such as an assay set out in Example 1 . In particular, the nuclease may be incubated with chromatin (e.g. in a population of permeabilised cells) to produce nucleic acid fragments and the digestion profile of the chromatin may be assessed. The nuclease may be a suitable nuclease variant when it has the ability to produce nucleic acid fragments. In particular, the nuclease may be a suitable nuclease variant when it has the ability to produce nucleic acid fragments wherein at least 50% (e.g. at least 70%) of the nucleic acid fragments are 145 bp to 290 bp fragments. The nuclease may be a suitable nuclease variant when it has the ability to produce nucleic acid fragments wherein at least 50% (e.g. at least 70%) of the nucleic acid fragments are 145 bp to 190 bp fragments or 180 bp to 200 bp fragments. For example, the nuclease may be a suitable nuclease variant when it has the ability to produce nucleic acid fragments wherein at least 50% (e.g. at least 70%) of the nucleic acid fragments are mono-nucleosomes of 180-200 bp, e.g., with the inter-nucleosomal linkers attached. References herein to a nuclease may also encompass a nuclease variant of the nuclease.

[0207] The inventors have found that nucleases having certain morphologies are particularly suitable for entering permeabilised cells and fragmenting chromatin in situ. In particular, the inventors have found that nucleases having a domain that is substantially ball-like may be suitable for entering permeabilised cells and fragmenting chromatin in situ. Thus, the nuclease may have a ball portion. The ball portion of the nuclease may be responsible for fragmenting chromatin. The ball portion of the nuclease may have nuclease activity. The nuclease may have a ball morphology. A nuclease having a ball morphology may have a ball portion. As used herein, the term “morphology” when used in reference to a nuclease refers to the overall shape of the nuclease.

[0208] The ball portion of the nuclease may have a length of at least 5 A, 10 A, 15 A, 20 A, 25 A, 30 A, 35 A, 40 A, 45 A, 50 A, 55 A, 60 A, 65 A, or 70 A along a first axis. The ball portion of the nuclease may have a length of 125 A or less, 120 A or less, 115 A or less, 110 A or less, 105 A or less, 100 A or less, 95 A or less, 90 A or less, 85 A or less, 80 A or less, 75 A or less, 70 A or less, 65 A or less, 60 A or less, 55 A or less, or 50 A or less along a first axis. The ball portion of the nuclease may have a length of 5 A to 125 A, 50 A to 100 A or 65 A to 90 A along a first axis. The ball portion of the nuclease may have a length of 50 A to 100 A along a first axis. The ball portion of the nuclease may have a length of 65 A to 90 A along a first axis. The ball portion of the nuclease may have a length of at least 5 A, 10 A, 15 A, 20 A, 25 A, 30 A, 35 A, 40 A, 45 A, 50 A, 55 A, 60 A, 65 A, or 70 A along a second axis. The ball portion of the nuclease may have a length of 125 A or less, 120 A or less, 115 A or less, 110 A or less, 105 A or less, 100 A or less, 95 A or less, 90 A or less, 85 A or less, 80 A or less, 75 A or less, 70 A or less, 65 A or less, 60 A or less, 55 A or less, or 50 A or less along a second axis. The ball portion of the nuclease may have a length of 5 A to 125 A, 25 A to 75 A, or 30 A to 60 A along a second axis. The ball portion of the nuclease may have a length of 25 A to 75 A along a second axis. The ball portion of the nuclease may have a length of 30 A to 60 A along a second axis.

[0209] The ball portion of the nuclease may have a length of at least 5 A, 10 A, 15 A, 20 A, 25 A, 30 A, 35 A, 40 A, 45 A, 50 A, 55 A, 60 A, 65 A, or 70 A along a third axis. The ball portion of the nuclease may have a length of 125 A or less, 120 A or less, 115 A or less, 110 A or less, 105 A or less, 100 A or less, 95 A or less, 90 A or less, 85 A or less, 80 A or less, 75 A or less, 70 A or less, 65 A or less, 60 A or less, 55 A or less, or 50 A or less along a third axis. The ball portion of the nuclease may have a length of 5 A to 125 A, 25 A to 75 A or 30 A to 60 A along a third axis The ball portion of the nuclease may have a length of 25 A to 75 A along a third axis. The ball portion of the nuclease may have a length of 30 A to 60 A along a third axis.

[0210] The ball portion of the nuclease may have: (a) a length of 50 A to 100 A along a first axis; (b) a length of 25 A to 75 A along a second axis; and (c) a length of 25 A to 75 A along a third axis, wherein the first, second and third axis define the longest height, width, and depth of the ball portion. The ball portion of the nuclease may have: (a) a length of 65 A to 90 A along a first axis; (b) a length of 30 A to 60 A along a second axis; and (c) a length of 30 A to 60 A along a third axis, wherein the first, second and third axis define the longest height, width, and depth of the ball portion.

[0211] The ball portion of the nuclease may have a first axis, second axis and third axis, wherein the ratio of the lengths of the first axis : second axis : third axis is 50 A to 100 A : 25 A to 75 A : 25 A to 75 A, wherein the first, second and third axis define the longest height, width, and depth of the ball portion. The ball portion of the nuclease may have a first axis, second axis and third axis, wherein the ratio of the lengths of the first axis : second axis : third axis is 65 A to 90 A : 30 A to 60 A : 30 A to 60 A, wherein the first, second and third axis define the longest height, width, and depth of the ball portion. The first, second and third axis of the ball portion may define the longest height, width, and depth of the ball portion.

[0212] The ball portion of the nuclease may have an amino acid sequence of at least 100 amino acid residues, at least 150 amino acid residues, at least 200 amino acid residues, at least 250 amino acid residues, at least 300 amino acid residues, at least 350 amino acid residues, at least 400 amino acid residues, at least 450 amino acid residues, or at least 500 amino acid residues. The ball portion of the nuclease may have an amino acid sequence of up to 750 amino acid residues, up to 700 amino acid residues, up to 650 amino acid residues, up to 600 amino acid residues, up to 550 amino acid residues, up to 500 amino acid residues, up to 450 amino acid residues, or up to 400 amino acid residues. The ball portion of the nuclease may have an amino acid sequence of at least 100 and up to 750 amino acid residues, at least 200 and up to 600 amino acid residues, or at least 300 and up to 500 amino acid residues.

[0213] The ball portion of the nuclease may comprise alpha helices and beta sheets.

[0214] The ball portion of the nuclease may have an amino acid sequence having at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 99%, or 100% sequence identity to the amino acid sequence of a ball portion defined in Table 1 or Table 2. The ball portion of the nuclease may have an amino acid sequence having at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 99%, or 100% sequence identity to the amino acid sequence of a ball portion defined in Table 1. In particular, the ball portion of the nuclease may have an amino acid sequence having at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 99%, or 100% sequence identity to amino acids 24 to 404 of SEQ ID NO: 1. For example, the ball portion of the nuclease may have the amino acid sequence of amino acids 24 to 404 of SEQ ID NO: 1 .

[0215] The nuclease may have a chain portion. The chain portion of the nuclease may comprise a signal peptide, such as a signal peptide sequence described herein. The chain portion of the nuclease may be a signal peptide, such as a signal peptide sequence described herein. The chain portion of the nuclease may target the nuclease to a location in the cell, such as the nucleus or endoplasmic reticulum. The chain portion of the nuclease may target the nuclease to the nucleus of a cell. The chain portion of the nuclease may target the nuclease to the endoplasmic reticulum of a cell. For example, the chain portion of the nuclease may target the nuclease to the endoplasmic reticulum of a cell during synthesis of the nuclease and prior to secretion of the nuclease from the cell. The chain portion of the nuclease may be removed prior to secretion of the nuclease from a cell. The chain portion of the nuclease may not be responsible for fragmenting chromatin. The chain portion of the nuclease may not have nuclease activity.

[0216] The nuclease may comprise a chain portion and another domain or portion. For example, the nuclease may have a ball portion and a chain portion. The nuclease may have a ball and chain morphology. Nuclease having a ball and chain morphology may comprise a ball portion (e.g. a ball portion described herein) and a chain portion (e.g. a chain portion described herein). Nucleases having a ball and chain morphology described herein may be particularly suitable for entering permeabilised cells and fragmenting chromatin within the permeabilised cells.

[0217] The chain portion of the nuclease may have a length of at least 5 A, 10 A, 15 A, 20 A, 25 A, 30 A, 35 A, 40 A, 45 A, 50 A, 55 A, 60 A, 65 A, or 70 A along a first axis. The chain portion of the nuclease may have a length of 125 A or less, 120 A or less, 115 A or less, 110 A or less, 105 A or less, 100 A or less, 95 A or less, 90 A or less, 85 A or less, 80 A or less, 75 A or less, 70 A or less, 65 A or less, 60 A or less, 55 A or less, or 50 A or less along a first axis. The chain portion of the nuclease may have a length of 5 A to 125 A, 50 A to 100 A, or 60 A to 80 A along a first axis. The chain portion of the nuclease may have a length of 50 A to 100 A along a first axis. The chain portion of the nuclease may have a length of 60 A to 80 A along a first axis.

[0218] The chain portion of the nuclease may have a length of at least 0.1 A, 0.2 A, 0.3 A, 0.4 A, 0.5 A, 0.6 A, 0.7 A, 0.8 A, 0.9 A, 1 .0 A, 1 .1 A, 1 .2 A, 1 .3 A, 1 .4 A, or 1 .5 A along a second axis. The chain portion of the nuclease may have a length of 50 A or less, 45 A or less, 40 A or less, 35 A or less, 30 A or less, 25 A or less, 20 A or less, 15 A or less, 10 A or less, or 5 A or less along a second axis. The chain portion of the nuclease may have a length of 0.1 A to 50 A, 1 A to 25 A, or 1 A to 15 A along a second axis. The chain portion of the nuclease may have a length of 1 A to 25 A along a second axis. The chain portion of the nuclease may have a length of 1 A to 15 A along a second axis.

[0219] The chain portion of the nuclease may have a length of less than 50 A, 45 A, 40 A, 35 A, 30 A, 25 A, 20 A, 15 A, 10 A, 5 A, or 1 A along a third axis. The chain portion of the nuclease may have a length of less than 50 A along a third axis. The chain portion of the nuclease may have a length of less than 10 A along a third axis. The chain portion of the nuclease may have a length of less than 5 A along a third axis.

[0220] The chain portion of the nuclease may have: (a) a length of 50 A to 100 A along a first axis; (b) a length of 1 A to 25 A along a second axis; and (c) a length of less than 10 A along a third axis, wherein the first, second and third axis define the longest height, width, and depth of the chain portion. The chain portion of the nuclease may have: (a) a length of 60 A to 80 A along a first axis; (b) a length of 1 A to 15 A along a second axis; and (c) a length of less than 5 A along a third axis, wherein the first, second and third axis define the longest height, width, and depth of the chain portion.

[0221] The chain portion of the nuclease may have a first axis, second axis and third axis, wherein the ratio of the lengths of the first axis : second axis : third axis is 50 A to 100 A : 1 A to 25 A : less than 10 A, wherein the first, second and third axis define the longest height, width, and depth of the chain portion. The chain portion of the nuclease may have a first axis, second axis and third axis, wherein the ratio of the lengths of the first axis : second axis : third axis is 60 A to 80 A : 1 A to 15 A : less than 5 A, wherein the first, second and third axis define the longest height, width, and depth of the chain portion.

[0222] The first, second and third axis of the chain portion may define the longest height, width, and depth of the chain portion.

[0223] The chain portion of the nuclease may be intrinsically disordered.

[0224] The chain portion of the nuclease may have an amino acid sequence of at least 5 amino acid residues, at least 10 amino acid residues, at least 15 amino acid residues, at least 20 amino acid residues, at least 25 amino acid residues, or at least 30 amino acid residues. The chain portion of the nuclease may have an amino acid sequence of up to 50 amino acid residues, up to 45 amino acid residues, up to 40 amino acid residues, up to 35 amino acid residues, up to 30 amino acid residues, up to 25 amino acid residues, up to 20 amino acid residues, or up to 15 amino acid residues. The chain portion of the nuclease may have an amino acid sequence of at least 5 and up to 50 amino acid residues, at least 15 and up to 40 amino acid residues, or at least 20 and up to 30 amino acid residues.

[0225] The chain portion of the nuclease may comprise an amino acid sequence having at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 99%, or 100% sequence identity to the amino acid sequence of a chain portion defined in Table 1 or Table 2. The chain portion of the nuclease may comprise an amino acid sequence having at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 99%, or 100% sequence identity to the amino acid sequence of a chain portion defined in Table 1. In particular, the chain portion of the nuclease may comprise an amino acid sequence having at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 99%, or 100% sequence identity to amino acids 1 to 23 of SEQ ID NO: 1 . For example, the ball portion of the nuclease may comprise the amino acid sequence of amino acids 1 to 23 of SEQ ID NO: 1.

[0226] The angle between the first axis of the ball portion and the first axis of the chain portion may be 45 degrees, 60 degrees, 75 degrees, 90 degrees, 105 degrees, 120 degrees, 135 degrees, 150 degrees, 165 degrees, 180 degrees, 205 degrees, 220 degrees, or 235 degrees. The angle between the first axis of the ball portion and the first axis of the chain portion may be 45 degrees to 235 degrees. The angle between the first axis of the ball portion and the first axis of the chain portion may be 90 degrees to 180 degrees. The angle between the first axis of the ball portion and the first axis of the chain portion may be 105 degrees to 150 degrees. The angle between the first axis of the ball portion and the first axis of the chain portion may be rigid or flexible. In particular, the angle between the first axis of the ball portion and the first axis of the chain portion may be flexible. A flexible angle may mean that the ball portion can move relative to the chain portion.

[0227] Recent advances in the prediction of protein structure from amino acid sequence, such as AlphaFold™, have enabled the rapid identification of accurate protein structures without the need for complex biochemical analyses such as nuclear magnetic resonance, X-ray crystallography, and cryo-electron microscopy. AlphaFold™ therefore provides a rapid and accurate prediction of a protein’s dimensions and overall morphology. Accordingly, any of the morphologies and / or dimensions of the nuclease described herein may be determined using AlphaFold™. The nuclease may have a ball portion as determined using AlphaFold™. The nuclease may have a ball morphology as determined using AlphaFold™. The nuclease may have a chain portion as determined using AlphaFold™. The nuclease may have a ball portion and a chain portion as determined using AlphaFold™. The nuclease may have a ball and chain morphology as determined using AlphaFold™. Any of the references to AlphaFold™ herein may be reference to AlphaFold™ DB version 2022-11-01 , created with AlphaFold™ Monomer v2.0 pipeline.

[0228] The nuclease may have a ball portion and be a nuclease which is present in a sequence- and / or structure- based cluster comprising Northern shrimp nuclease having UniProt ID C9YSL6 described herein. The nuclease may have a ball morphology and be a nuclease which is present in a sequence- and / or structure-based cluster comprising Northern shrimp nuclease having UniProt ID C9YSL6 described herein. The nuclease may have a ball and chain morphology and be a nuclease which is present in a sequence- and / or structure-based cluster comprising Northern shrimp nuclease having UniProt ID C9YSL6 described herein. In particular, the nuclease may: (a) have a ball and chain morphology and (b) be a CATH Superfamily 3.40.570.10 Extracellular Endonuclease, subunit A nuclease, or a nuclease variant thereof.

[0229] The nuclease may comprise a signal peptide sequence. The signal peptide sequence may target the nuclease to a location in the cell, such as the nucleus or endoplasmic reticulum of the cell. The signal peptide sequence may target the nuclease to the endoplasmic reticulum of a cell. For example, the signal peptide sequence may target the nuclease to the endoplasmic reticulum of a cell during synthesis of the nuclease and prior to secretion of the nuclease from the cell. The signal peptide sequence may target the nuclease to the nucleus of a cell. Thus, the signal peptide sequence may comprise a nuclear localisation signal. The nuclease may comprise a nuclear localisation signal. The signal peptide sequence may be removed prior to secretion of the nuclease from a cell.

[0230] The inventors have shown that nucleases having a relatively high molecular weight are able to enter permeabilised cells and fragment chromatin in situ. Previously, it was believed that only nucleases which have a relatively low molecular weight were able to enter permeabilised cells and fragment chromatin in situ. For example, micrococcal nuclease has previously been used to fragment chromatin in permeabilised cells and has a molecularweight of approximately 17 kDa. In contrast, the inventors have shown that a nuclease having a molecular weight of about 45 kDa is able to enter permeabilised cells and fragment chromatin in situ. Accordingly, the nuclease may have any suitable molecular weight. The nuclease may have a molecular weight of at least about 5 kDa, at least about 10 kDa, at least about 15 kDa, at least about 17 kDa, at least about 20 kDa, at least about 25 kDa, at least about 30 kDa, at least about 35 kDa, at least about 40 kDa, or at least about 45 kDa, In particular, the nuclease may have a molecular weight of at least about 17 kDa. The nuclease may have a molecular weight of up to about 100 kDa, up to about 95 kDa, up to about 90 kDa, up to about 85 kDa, up to about 80 kDa, up to about 75 kDa, up to about 70 kDa, up to about 65 kDa, up to about 60 kDa, up to about 55 kDa, or up to about 50 kDa. The nuclease may have a molecular weight of at least about 5 kDa and up to about 100 kDa. The nuclease may have a molecular weight of at least about 10 kDa and up to about 75 kDa. The nuclease may have a molecular weight of at least about 17 kDa and up to about 50 kDa. For example, the nuclease may have a molecular weight of at least about 40 kDa up to about 50 kDa. The nuclease may have a molecular weight of about 5 kDa, about 10 kDa, about 15 kDa, about 20 kDa, about 25 kDa, about 30 kDa, about 35 kDa, about 40 kDa, about 45 kDa, about 50 kDa, about 55 kDa, about 60 kDa, about 65 kDa, about 70 kDa, about 75 kDa, about 80 kDa, about 85 kDa, about 90 kDa, about 95 kDa, or about 100 kDa. In particular, the nuclease may have a molecular weight of about 40 kDa or about 45 kDa. The nuclease may have a molecular weight of about 40 kDa. The nuclease may have a molecular weight of about 45 kDa.

[0231] The nuclease may have a molecular weight of at least 5 kDa, at least 10 kDa, at least 15 kDa, at least 17 kDa, at least 20 kDa, at least 25 kDa, at least 30 kDa, at least 35 kDa, at least 40 kDa, or at least 45 kDa, In particular, the nuclease may have a molecular weight of at least 17 kDa. The nuclease may have a molecular weight of up to 100 kDa, up to 95 kDa, up to 90 kDa, up to 85 kDa, up to 80 kDa, up to 75 kDa, up to 70 kDa, up to 65 kDa, up to 60 kDa, up to 55 kDa, or up to 50 kDa. The nuclease may have a molecular weight of at least 5 kDa and up to 100 kDa. The nuclease may have a molecular weight of at least 10 kDa and up to 75 kDa. The nuclease may have a molecular weight of at least 17 kDa and up to 50 kDa. For example, the nuclease may have a molecular weight of at least 40 kDa up to 50 kDa. The nuclease may have a molecular weight of 5 kDa, 10 kDa, 15 kDa, 20 kDa, 25 kDa, 30 kDa, 35 kDa, 40 kDa, 45 kDa, 50 kDa, 55 kDa, 60 kDa, 65 kDa, 70 kDa, 75 kDa, 80 kDa, 85 kDa, 90 kDa, 95 kDa, or 100 kDa. In particular, the nuclease may have a molecular weight of 40 kDa or 45 kDa. The nuclease may have a molecular weight of 40 kDa. The nuclease may have a molecular weight of 45 kDa.

[0232] The nuclease may be an endonuclease, an exonuclease, or an exo-endonuclease. In particular, the nuclease may be an endonuclease. An endonuclease may be a nuclease that cleaves a phosphodiester bond within a polynucleotide chain. Thus, the nuclease may be a nuclease capable of cleaving a phosphodiester bond within a polynucleotide chain. An exonuclease may be a nuclease that cleaves the end of a polynucleotide chain. The nuclease may be an exonuclease. Thus, the nuclease may be a nuclease that cleaves the end of a polynucleotide chain. An exo-endonuclease may be a nuclease that displays both endonuclease- and exonuclease-like qualities. The nuclease may be an exo-endonuclease. The nuclease may be a nuclease that displays both endonuclease- and exonuclease-like qualities.

[0233] The nuclease may have more endonuclease activity than exo-nuclease activity. A nuclease may be tested for endonuclease activity and exo-nuclease activity using any suitable assay. The nuclease may be a nuclease which has minimal ability to cleave the end of a polynucleotide chain. For example, the nuclease may be an endonuclease which has minimal ability to cleave the end of a polynucleotide chain. The nuclease may have minimal exonuclease activity. For example, the nuclease may be an endonuclease with minimal exonuclease activity. The nuclease may be a nuclease which is unable to cleave the end of a polynucleotide chain. For example, the nuclease may be an endonuclease which is unable to cleave the end of a polynucleotide chian. The nuclease may be a nuclease which has no exonuclease activity. For example, the nuclease may be an endonuclease which has no exonuclease activity.

[0234] The nuclease may fragment double stranded DNA (dsDNA), single stranded DNA (ssDNA), or RNA. In particular, the nuclease may fragment dsDNA. The nuclease may preferentially fragment dsDNA (e.g., compared to ssDNA and RNA). The nuclease may not fragment ssDNA or RNA. The nuclease may not fragment ssDNA. The nuclease may not fragment RNA. The nuclease may not fragment ssDNA and RNA. The nuclease may fragment dsDNA and may not fragment ssDNA and RNA. The nuclease may have high affinity for activity on dsDNA. The nuclease may have low affinity for activity on ssDNA or RNA. The nuclease may have low affinity for activity on ssDNA. The nuclease may have low affinity for activity on RNA. The nuclease may have more affinity for activity on dsDNA than for ssDNA or RNA. The nuclease may have more affinity for activity on dsDNA than for ssDNA. The nuclease may have more affinity for activity on dsDNA than for RNA. The nuclease may have more affinity for activity on dsDNA than for ssDNA and RNA. The nuclease may have high affinity for activity on dsDNA with minimal activity on ssDNA or RNA.

[0235] The nuclease may be a nonspecific nuclease or a specific nuclease. A nonspecific nuclease may be a nuclease that cleaves nucleic acids without regard to the sequence that is being cleaved. Meanwhile, a specific nuclease may be a nuclease that only cleaves nucleic acids at specific nucleotide sequences. The nuclease may be a nuclease that cleaves nucleic acids without regard to the sequence that is being cleaved. In particular, the nuclease may be a nonspecific nuclease. For example, the nuclease may be a nonspecific endonuclease. The nuclease may have any suitable cut site preference profile. As used herein, a “cut site preference profile” may refer to the distribution of cut sites in chromatin by the nuclease. The nuclease may have a cut site preference profile which does not exhibit sequence bias. The nuclease may have a cut site preference profile which exhibits less sequence bias than micrococcal nuclease. The nuclease may have a cut site preference profile which exhibits a reduction in A-T bias compared to micrococcal nuclease. The nuclease may have a cut site preference profile which does not exhibit A-T bias.

[0236] Sequence bias may be assessed using a sequence bias assay (such as a sequence bias assay described herein). The sequence bias assay may comprise determining the sequence content of a pool of nucleic acid fragments produced using a nuclease described herein. The sequence content of the pool of nucleic acid fragments may be compared to the sequence content of a pool of nucleic acid fragments produced using a reference nuclease such as micrococcal nuclease, e.g. to determine whether there is a reduction in sequence bias using the nuclease.

[0237] The nuclease may have a cut site preference profile such that at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, or at least 40% of nucleic acid residues immediately before a cut site are cytosine (C). The nuclease may have a cut site preference profile such that up to 50%, up to 45%, up to 40%, up to 35%, up to 30%, up to 25%, or up to 20% of nucleic acid residues immediately before a cut site are C. The nuclease may have a cut site profile such that at least 10% and up to 50%, at least 15% and up to 35%, or at least 20% and up to 30% of nucleic acid residues immediately before a cut site are C. In particular, the nuclease may have a cut site preference profile such that 15-35% of nucleic acid residues immediately before a cut site are C.

[0238] The nuclease may have a cut site preference profile such that at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, or at least 40% of nucleic acid residues immediately before a cut site are guanine (G). The nuclease may have a cut site preference profile such that up to 50%, up to 45%, up to 40%, up to 35%, up to 30%, up to 25%, or up to 20% of nucleic acid residues immediately before a cut site are G. The nuclease may have a cut site profile such that at least 10% and up to 50%, at least 15% and up to 35%, or at least 20% and up to 30% of nucleic acid residues immediately before a cut site are G. In particular, the nuclease may have a cut site preference profile such that 15-35% of nucleic acid residues immediately before a cut site are G. The nuclease may have a cut site preference profile such that at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, or at least 40% of nucleic acid residues immediately before a cut site are thymine (T). The nuclease may have a cut site preference profile such that up to 50%, up to 45%, up to 40%, up to 35%, up to 30%, up to 25%, or up to 20% of nucleic acid residues immediately before a cut site are T. The nuclease may have a cut site profile such that at least 10% and up to 50%, at least 15% and up to 35%, or at least 20% and up to 30% of nucleic acid residues immediately before a cut site are T. In particular, the nuclease may have a cut site preference profile such that 15-35% of nucleic acid residues immediately before a cut site are T.

[0239] The nuclease may have a cut site preference profile such that at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, or at least 40% of nucleic acid residues immediately before a cut site are adenine (A). The nuclease may have a cut site preference profile such that up to 50%, up to 45%, up to 40%, up to 35%, up to 30%, up to 25%, or up to 20% of nucleic acid residues immediately before a cut site are A. The nuclease may have a cut site profile such that at least 10% and up to 50%, at least 15% and up to 35%, or at least 20% and up to 30% of nucleic acid residues immediately before a cut site are A. In particular, the nuclease may have a cut site preference profile such that 15-35% of nucleic acid residues immediately before a cut site are A.

[0240] The nuclease may have a cut site preference profile such that: (a) 15-35% of nucleic acid residues immediately before a cut site are C; (b) 15-35% of nucleic acid residues immediately before a cut site are G; (c) 15-35% of nucleic acid residues immediately before a cut site are T; and (d) 15-35% of nucleic acid residues immediately before a cut site are A, wherein the sum of (a), (b), (c) and (d) is 100%.

[0241] The nuclease may have a cut site preference profile such that at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, or at least 40% of nucleic acid residues immediately after a cut site are cytosine (C). The nuclease may have a cut site preference profile such that up to 50%, up to 45%, up to 40%, up to 35%, up to 30%, up to 25%, or up to 20% of nucleic acid residues immediately after a cut site are C. The nuclease may have a cut site profile such that at least 10% and up to 50%, at least 10% and up to 40%, at least 15% and up to 35%, or at least 20% and up to 30% of nucleic acid residues immediately after a cut site are C. In particular, the nuclease may have a cut site preference profile such that 10-40% of nucleic acid residues immediately after a cut site are C. The nuclease may have a cut site preference profile such that at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, or at least 40% of nucleic acid residues immediately after a cut site are guanine (G). The nuclease may have a cut site preference profile such that up to 50%, up to 45%, up to 40%, up to 35%, up to 30%, up to 25%, or up to 20% of nucleic acid residues immediately after a cut site are G. The nuclease may have a cut site profile such that at least 10% and up to 50%, at least 10% and up to 40%, at least 15% and up to 35%, or at least 20% and up to 30% of nucleic acid residues immediately after a cut site are G. In particular, the nuclease may have a cut site preference profile such that 10-40% of nucleic acid residues immediately after a cut site are G.

[0242] The nuclease may have a cut site preference profile such that at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, or at least 40% of nucleic acid residues immediately after a cut site are thymine (T). The nuclease may have a cut site preference profile such that up to 50%, up to 45%, up to 40%, up to 35%, up to 30%, up to 25%, or up to 20% of nucleic acid residues immediately after a cut site are T. The nuclease may have a cut site profile such that at least 10% and up to 50%, at least 10% and up to 40%, at least 15% and up to 35%, or at least 20% and up to 30% of nucleic acid residues immediately after a cut site are T. In particular, the nuclease may have a cut site preference profile such that 10-40% of nucleic acid residues immediately after a cut site are T.

[0243] The nuclease may have a cut site preference profile such that at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, or at least 40% of nucleic acid residues immediately after a cut site are adenine (A). The nuclease may have a cut site preference profile such that up to 50%, up to 45%, up to 40%, up to 35%, up to 30%, up to 25%, or up to 20% of nucleic acid residues immediately after a cut site are A. The nuclease may have a cut site profile such that at least 10% and up to 50%, at least 10% and up to 40%, at least 15% and up to 35%, or at least 20% and up to 30% of nucleic acid residues immediately after a cut site are A. In particular, the nuclease may have a cut site preference profile such that 10-40% of nucleic acid residues immediately after a cut site are A.

[0244] The nuclease may have a cut site preference profile such that: (a) 10-40% of nucleic acid residues immediately after a cut site are C; (b) 10-40% of nucleic acid residues immediately after a cut site are G; (c) 10-40% of nucleic acid residues immediately after a cut site are T; and (d) 10-40% of nucleic acid residues immediately after a cut site are A, wherein the sum of (a), (b), (c) and (d) is 100%. The nuclease may be a DNA nuclease. The nuclease may be an RNA nuclease. The nuclease may be a DNA / RNA nuclease. In particular, the nuclease may be a DNA / RNA nonspecific nuclease. For example, the nuclease may be a DNA / RNA nonspecific endonuclease.

[0245] The nuclease may be suitable for entering a cell. In particular, the nuclease may be suitable for entering a permeabilised cell. The cell may have been permeabilised by a membrane- permeabilising agent described herein. For example, the cell may have been permeabilised using Digitonin. The nuclease may be suitable for entering a cell which has been permeabilised using a cell permeabilization method described herein. The nuclease may be suitable for traversing a cell pore. A cell pore may be any pore which allows movement of a molecule between the extracellular and intracellular environment of a cell. The nuclease may be suitable for traversing a permeabilised cell pore. Under common cell permeabilization conditions, permeabilised cell pores are thought to be approximately 10nm in diameter and are believed to be substantially spherical. Accordingly, the nuclease may be suitable for traversing a cell pore which has the dimensions of a cell pore produced using a permeabilization method described herein. The nuclease may be suitable fortraversing an at least about 1 nm cell pore, an at least about 2 nm cell pore, an at least about 3 nm cell pore, an at least about 4 nm cell pore, an at least about 5 nm cell pore, an at least about 6 nm cell pore, an at least about 7 nm cell pore, an at least about 8 nm cell pore, an at least about 9 nm cell pore, an at least about 10 nm cell pore, an at least about 15 nm cell pore, or an at least about 20 nm cell pore. In particular, the nuclease may be suitable for traversing an at least 10 nm cell pore. The nuclease may be suitable for traversing an up to about 150 nm cell pore, an up to about 100 nm cell pore, an up to about 90 nm cell pore, an up to about 80 nm cell pore, an up to about 70 nm cell pore, an up to about 60 nm cell pore, an up to about 50 nm cell pore, an up to about 40 nm cell pore, an up to about 30 nm cell pore, an up to about 20 nm cell pore, or an up to about 10 nm cell pore. The nuclease may be suitable for traversing an at least about 1 nm and up to about 150 nm cell pore, an at least about 5 nm and up to about 100 nm cell pore, or an at least about 10 nm and up to about 50 nm cell pore.

[0246] The nuclease may be suitable for traversing an at least 1 nm cell pore, an at least 2 nm cell pore, an at least 3 nm cell pore, an at least 4 nm cell pore, an at least 5 nm cell pore, an at least 6 nm cell pore, an at least 7 nm cell pore, an at least 8 nm cell pore, an at least 9 nm cell pore, an at least 10 nm cell pore, an at least 15 nm cell pore, or an at least 20 nm cell pore. In particular, the nuclease may be suitable for traversing an at least 10 nm cell pore. The nuclease may be suitable for traversing an up to 150 nm cell pore, an up to 100 nm cell pore, an up to 90 nm cell pore, an up to 80 nm cell pore, an up to 70 nm cell pore, an up to 60 nm cell pore, an up to 50 nm cell pore, an up to 40 nm cell pore, an up to 30 nm cell pore, an up to 20 nm cell pore, or an up to 10 nm cell pore. The nuclease may be suitable for traversing an at least 1 nm and up to 150 nm cell pore, an at least 5 nm and up to 100 nm cell pore, or an at least 10 nm and up to 50 nm cell pore.

[0247] Reference to the size of a cell pore may be reference to the diameter of a cell pore. For example, reference to a 10 nm cell pore may be reference to a cell pore having a diameter of 10 nm.

[0248] The nuclease may have dimensions suitable for traversing a permeabilised cell pore. The nuclease may have a dimension (e.g. a width or a length) which is smaller than the dimension of a cell pore, e.g., a cell pore described herein. In particular, the nuclease may have a longest width which is smaller than the dimension of a cell pore, e.g., a cell pore described herein. The nuclease may have a longest length which is smaller than the dimension of a cell pore, e.g., a cell pore described herein. The nuclease may have a longest length and a longest width which is smaller than the dimension of a cell pore, e.g., a cell pore described herein. The nuclease may have a longest width of about 125 A or less, about 120 A or less, about 115 A or less, about 110 A or less, about 105 A or less, about 100 A or less, about 95 A or less, about 90 A or less, about 85 A or less, about 80 A or less, about 75 A or less, about 70 A or less, about 65 A or less, about 60 A or less, about 55 A or less, or about 50 A or less. The nuclease may have a longest width of about 5 A to about 125 A, about 50 A to about 100 A or about 65 A to about 90 A. The nuclease may have a longest width of about 50 A to about 100 A. The nuclease may have a longest width of about 65 A to about 90 A.

[0249] The nuclease may have a longest width of 125 A or less, 120 A or less, 115 A or less, 110 A or less, 105 A or less, 100 A or less, 95 A or less, 90 A or less, 85 A or less, 80 A or less, 75 A or less, 70 A or less, 65 A or less, 60 A or less, 55 A or less, or 50 A or less. The nuclease may have a longest width of 5 A to 125 A, 50 A to 100 A or 65 A to 90 A. The nuclease may have a longest width of 50 A to 100 A. The nuclease may have a longest width of 65 A to 90 A.

[0250] The nuclease may have a longest length of about 125 A or less, about 120 A or less, about 115 A or less, about 110 A or less, about 105 A or less, about 100 A or less, about 95 A or less, about 90 A or less, about 85 A or less, about 80 A or less, about 75 A or less, about 70 A or less, about 65 A or less, about 60 A or less, about 55 A or less, or about 50 A or less. The nuclease may have a longest length of about 5 A to about 125 A, about 50 A to about 100 A or about 65 A to about 90 A. The nuclease may have a longest length of about 50 A to about 100 A. The nuclease may have a longest length of about 65 A to about 90 A.

[0251] The nuclease may have a longest length of 125 A or less, 120 A or less, 115 A or less, 110 A or less, 105 A or less, 100 A or less, 95 A or less, 90 A or less, 85 A or less, 80 A or less, 75 A or less, 70 A or less, 65 A or less, 60 A or less, 55 A or less, or 50 A or less. The nuclease may have a longest length of 5 A to 125 A, 50 A to 100 A or 65 A to 90 A. The nuclease may have a longest length of 50 A to 100 A. The nuclease may have a longest length of 65 A to 90 A.

[0252] The nuclease may have a longest length which is bigger than the dimensions of a cell pore described herein as long as the nuclease has a longest width which is smaller than the dimensions of a cell pore described herein. The nuclease may have a longest width which is bigger than the dimensions of a cell pore described herein as long as the nuclease has a longest width which is smaller than the dimensions of a cell pore described herein. A particularly suitable nuclease may have a longest length which is smaller than the dimensions of a cell pore described herein and a longest width which is smaller than the dimensions of a cell pore described herein.

[0253] The active site of a nuclease may be the region of the nuclease where a substrate molecule (e.g. chromatin) binds and undergoes a chemical reaction (e.g. fragmentation). A number of nucleases have been found to have a conserved active site amino acid sequence. In particular, Northern shrimp nuclease having UniProt ID C9YSL6 may comprise a cd00091 active site having the amino acid residues shown in Table 4. Accordingly, the nuclease may comprise a cd00091 active site. The invention provides a method of producing nucleic acid fragments, the method comprising fragmenting chromatin within a population of permeabilised cells with a nuclease, wherein the nuclease comprises a cd00091 active site.

[0254] The cd00091 active site may have an amino acid sequence of A / S / D / Q / T / E-K / R / Y / l / H-X- H / Q / Y / F-(X)a-N / F / G-(X)7-E / L / M / H / Q-(X)3-R / M / UV / F, where X is any amino acid and a is 23 to 37. The nuclease may comprise an active site comprising the amino acid sequence A / S / D / Q / T / E-K / R / Y / l / H-X-H / Q / Y / F-(X)a-N / F / G-(X)7-E / L / M / H / Q-(X)3-R / M / UV / F, where X is any amino acid and a is 23 to 37. In particular, the nuclease may comprise an active site comprising the amino acid sequence A / S / D-K / R-X-H-(X)29-N-(X)7-E-(X)3-R, where X is any amino acid. The nuclease may comprise an active site comprising the amino acid sequence: (a) A-K-X-H- (X)29-N-(X)7-E-(X)3-R, where X is any amino acid; (b) A-R-X-H-(X)29-N-(X)7-E-(X)3-R, where X is any amino acid; (c) S-R-X-H-(X)29-N-(X)7-E-(X)3-R, where X is any amino acid; (d) D-R-X-H- (X)29-N-(X)7-E-(X)3-R, where X is any amino acid; (e) D-R-X-H-(X)3I-N-(X)7-E-(X)3-R, where X is any amino acid; (f) D-R-X-H-(X)30-N-(X)7-E-(X)3-R, where X is any amino acid; (g) S-Y-X-F- (X)28-F-(X)7-Q-(X)3-V, where X is any amino acid; (h) D-R-X-H-(X)37-G-(X)7-E-(X)3-L, where X is any amino acid; (i) Q-R-X-Q-(X)26-F-(X)7-L-(X)3-M, where X is any amino acid; (j) T-H-X-F- (X)28-F-(X)7-H-(X)3-L, where X is any amino acid; (k) T-Y-X-Y-(X)23-F-(X)7-M-(X)3-L, where X is any amino acid; (I) D-K-X-H-(X)36-N-(X)7-E-(X)3-L, where X is any amino acid; (m) D-R-X-H- (X)35-G-(X)7-E-(X)3-L, where X is any amino acid; (n) E-I-X-Q-(X)26-F-(X)7-L-(X)3-M, where X is any amino acid; (o) S-R-X-H-(X)30-N-(X)7-E-(X)3-R, where X is any amino acid; (p) S-Y-X-F- (X)29-F-(X)7-H-(X)3-L, where X is any amino acid; (q) D-R-X-H-(X)33-N-(X)7-E-(X)3-F, where X is any amino acid; or (r) D-R-X-H-(X)33-N-(X)7-E-(X)3-L, where X is any amino acid. In particular, the nuclease may comprise an active site comprising the amino acid sequence: (a) A-K-X-H- (X)29-N-(X)7-E-(X)3-R, where X is any amino acid; (b) A-R-X-H-(X)29-N-(X)7-E-(X)3-R, where X is any amino acid; (c) S-R-X-H-(X)29-N-(X)7-E-(X)3-R, where X is any amino acid; or (d) D-R- X-H-(X)29-N-(X)7-E-(X)3-R, where X is any amino acid. For example, the nuclease may comprise an active site comprising the amino acid sequence A-K-X-H-(X)29-N-(X)7-E-(X)3-R, where X is any amino acid.

[0255] The nuclease may comprise an active site comprising the amino acid sequence A / S / D / Q / T / E- K / R ' / l / H-G-H / Q ' / F-(X)b-N-(X)2-P-(X)c-N / F / G-(X)3-W-(X)3-E / L / M / H / Q-(X)3-R / M / LA / / F, where X is any amino acid; b is 15 to 27 and c is 4 to 7, optionally wherein the sum of b and c is 19 to 33. In particular, the nuclease may comprise an active site comprising the amino acid sequence A-K-G-H-(X)2o-N-(X)2-P-(X)5-N-(X)3-W-(X)3-E-(X)3-R, where X is any amino acid.

[0256] The nuclease may comprise an active site having at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 99%, or 100% identity to amino acids 231 to 276 of SEQ ID NO: 1 or SEQ ID NO: 2. The nuclease may comprise an active site having at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 99%, or 100% identity to amino acids 231 to 276 of SEQ ID NO: 1 or SEQ ID NO: 2, wherein the active site comprises the amino acid sequence A-K-X-H-(X)29-N-(X)7-E-(X)3-R, where X is any amino acid. The nuclease may comprise an active site having at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 99%, or 100% identity to amino acids 231 to 276 of SEQ ID NO: 1 or SEQ ID NO: 2, wherein the active site comprises the amino acid sequence A-K-G-H-(X)2o-N-(X)2-P-(X)5-N-(X)3-W-(X)3-E-(X)3-R, where X is any amino acid.

[0257] The nuclease may comprise an active site comprising the amino acid residues of an active site as defined in Table 4. In particular, the nuclease may comprise an active site comprising the amino acid residues A231 , K232, H234, N264, E272 and R276 of SEQ ID NO: 1 . The nuclease may comprise an active site corresponding to the amino acid residues of an active site as defined in Table 4. In particular, the nuclease may comprise an active site corresponding to the amino acid residues A231 , K232, H234, N264, E272 and R276 of SEQ ID NO: 1 .

[0258] The active site of the nuclease may comprise a substrate binding site (e.g. a chromatin binding site).

[0259] The substrate binding site may correspond to A231 , K232 and R276 of SEQ ID NO: 1. Accordingly, the nuclease may comprise an active site comprising the amino acid sequence A7S7D7Q7T7E*-K7R7Y7l7H*-X-H / Q ' / F-(X)a-N / F / G-(X)7-E / L / M / H / Q-(X)3-R7M7L*A / 7F*, where X is any amino acid and a is 23 to 37, and wherein * denotes a substrate binding site. In particular, the nuclease may comprise an active site comprising the amino acid sequence A7S7D*-K7R*-X-H-(X)29-N-(X)7-E-(X)3-R*, where X is any amino acid, and wherein * denotes a substrate binding site.

[0260] The nuclease may comprise an active site comprising the amino acid sequence: (a) A*-K*-X- H-(X)29-N-(X)7-E-(X)3-R*, where X is any amino acid, and wherein * denotes a substrate binding site; (b) A*-R*-X-H-(X)29-N-(X)7-E-(X)3-R*, where X is any amino acid, and wherein * denotes a substrate binding site; (c) S*-R*-X-H-(X)29-N-(X)7-E-(X)3-R*, where X is any amino acid, and wherein * denotes a substrate binding site; (d) D*-R*-X-H-(X)29-N-(X)7-E-(X)3-R*, where X is any amino acid, and wherein * denotes a substrate binding site; (e) D*-R*-X-H-(X)3I-N-(X)7-E- (X)3-R*, where X is any amino acid, and wherein * denotes a substrate binding site; (f) D*-R*- X-H-(X)3O-N-(X)7-E-(X)3-R*, where X is any amino acid, and wherein * denotes a substrate binding site; (g) S*-Y*-X-F-(X)28-F-(X)7-Q-(X)3-V*, where X is any amino acid, and wherein * denotes a substrate binding site; (h) D*-R*-X-H-(X)37-G-(X)7-E-(X)3-L*, where X is any amino acid, and wherein * denotes a substrate binding site; (i) Q*-R*-X-Q-(X)26-F-(X)7-L-(X)3-M*, where X is any amino acid, and wherein * denotes a substrate binding site; (j) T*-H*-X-F-(X)28- F-(X)7-H-(X)3-L*, where X is any amino acid, and wherein * denotes a substrate binding site; (k) T*-Y*-X-Y-(X)23-F-(X)7-M-(X)3-L*, where X is any amino acid, and wherein * denotes a substrate binding site; (I) D*-K*-X-H-(X)36-N-(X)7-E-(X)3-L*, where X is any amino acid, and wherein * denotes a substrate binding site; (m) D*-R*-X-H-(X)35-G-(X)7-E-(X)3-L*, where X is any amino acid, and wherein * denotes a substrate binding site; (n) E*-I*-X-Q-(X)26-F-(X)7-L- (X)3-M*, where X is any amino acid, and wherein * denotes a substrate binding site; (o) S*-R*- X-H-(X)3O-N-(X)7-E-(X)3-R*, where X is any amino acid, and wherein * denotes a substrate binding site; (p) S*-Y*-X-F-(X)29-F-(X)7-H-(X)3-L*, where X is any amino acid, and wherein * denotes a substrate binding site; (q) D*-R*-X-H-(X)33-N-(X)7-E-(X)3-F*, where X is any amino acid, and wherein * denotes a substrate binding site; or (r) D*-R*-X-H-(X)33-N-(X)7-E-(X)3-L*, where X is any amino acid, and wherein * denotes a substrate binding site. In particular, the nuclease may comprise an active site comprising the amino acid sequence: (a) A*-K*-X-H- (X)29-N-(X)7-E-(X)3-R*, where X is any amino acid, and wherein * denotes a substrate binding site; (b) A*-R*-X-H-(X)29-N-(X)7-E-(X)3-R*, where X is any amino acid, and wherein * denotes a substrate binding site; (c) S*-R*-X-H-(X)29-N-(X)7-E-(X)3-R*, where X is any amino acid, and wherein * denotes a substrate binding site; or (d) D*-R*-X-H-(X)29-N-(X)7-E-(X)3-R*, where X is any amino acid, and wherein * denotes a substrate binding site. For example, the nuclease may comprise an active site comprising the amino acid sequence A*-K*-X-H-(X)29-N-(X)7-E- (X)3-R*, where X is any amino acid, and wherein * denotes a substrate binding site.

[0261] The nuclease may comprise an active site comprising the amino acid sequence A7S7D7Q7T7E*-K7R7Y7l7H*-G-H / Q ' / F-(X)b-N-(X)2-P-(X)c-N / F / G-(X)3-W-(X)3-E / L / M / H / Q- (X)3-R7M7L7V7F*, where X is any amino acid; b is 15 to 27 and c is 4 to 7, optionally wherein the sum of b and c is 19 to 33, and wherein * denotes a substrate binding site. In particular, the nuclease may comprise an active site comprising the amino acid sequence A-K*-G*-H- (X)2O-N-(X)2-P-(X)5-N-(X)3-W-(X)3-E-(X)3-R*, where X is any amino acid, and wherein * denotes a substrate binding site.

[0262] The nuclease may comprise an active site comprising the amino acid residues A231*, K232*, H234, N264, E272 and R276* of SEQ ID NO: 1 , wherein * denotes a substrate binding site. The nuclease may comprise an active site corresponding to the amino acid residues A231*, K232*, H234, N264, E272 and R276* of SEQ ID NO: 1 , wherein * denotes a substrate binding site. The active site of the nuclease may comprise a metal ion binding site. The presence of a metal ion may aid in fragmentation of chromatin. The metal ion may be any suitable metal ion. In particular, the metal ion may be a divalent metal ion, such as a Mg2+metal ion.

[0263] The metal ion (e.g. Mg2+) binding site may correspond to N264 of SEQ ID NO: 1. Accordingly, the nuclease may comprise an active site comprising the amino acid sequence A / S / D / Q / T / E- K / R ' / l / H-X-H / Q ' / F-(X)a-N+ / F+ / G+-(X)7-E / L / M / H / Q-(X)3-R / M / L / / F, where X is any amino acid and a is 23 to 37, and wherein + denotes a metal ion (e.g. Mg2+) binding site. In particular, the nuclease may comprise an active site comprising the amino acid sequence A / S / D-K / R-X-H- (X)29-N+-(X)7-E-(X)3-R, where X is any amino acid, and wherein + denotes a metal ion (e.g. Mg2+) binding site.

[0264] The nuclease may comprise an active site comprising the amino acid sequence: (a) A-K-X-H- (X)29-N+-(X)7-E-(X)3-R, where X is any amino acid, and wherein + denotes a metal ion (e.g. Mg2+) binding site; (b) A-R-X-H-(X)29-N+-(X)7-E-(X)3-R, where X is any amino acid, and wherein + denotes a metal ion (e.g. Mg2+) binding site; (c) S-R-X-H-(X)29-N+-(X)7-E-(X)3-R, where X is any amino acid, and wherein + denotes a metal ion (e.g. Mg2+) binding site; (d) D-R-X-H-(X)29- N+-(X)7-E-(X)3-R, where X is any amino acid, and wherein + denotes a metal ion (e.g. Mg2+) binding site; (e) D-R-X-H-(X)3I-N+-(X)7-E-(X)3-R, where X is any amino acid, and wherein + denotes a metal ion (e.g. Mg2+) binding site; (f) D-R-X-H-(X)30-N+-(X)7-E-(X)3-R, where X is any amino acid, and wherein + denotes a metal ion (e.g. Mg2+) binding site; (g) S-Y-X-F-(X)28-F+- (X)7-Q-(X)3-V, where X is any amino acid, and wherein + denotes a metal ion (e.g. Mg2+) binding site; (h) D-R-X-H-(X)37-G+-(X)7-E-(X)3-L, where X is any amino acid, and wherein + denotes a metal ion (e.g. Mg2+) binding site; (i) Q-R-X-Q-(X)26-F+-(X)7-L-(X)3-M, where X is any amino acid, and wherein + denotes a metal ion (e.g. Mg2+) binding site; (j) T-H-X-F-(X)28-F+- (X)7-H-(X)3-L, where X is any amino acid, and wherein + denotes a metal ion (e.g. Mg2+) binding site; (k) T-Y-X-Y-(X)23-F+-(X)7-M-(X)3-L, where X is any amino acid, and wherein + denotes a metal ion (e.g. Mg2+) binding site; (I) D-K-X-H-(X)36-N+-(X)7-E-(X)3-L, where X is any amino acid, and wherein + denotes a metal ion (e.g. Mg2+) binding site; (m) D-R-X-H-(X)35-G+-(X)7-E- (X)3-L, where X is any amino acid, and wherein + denotes a metal ion (e.g. Mg2+) binding site; (n) E-I-X-Q-(X)26-F+-(X)7-L-(X)3-M, where X is any amino acid, and wherein + denotes a metal ion (e.g. Mg2+) binding site; (o) S-R-X-H-(X)30-N+-(X)7-E-(X)3-R, where X is any amino acid, and wherein + denotes a metal ion (e.g. Mg2+) binding site; (p) S-Y-X-F-(X)29-F+-(X)7-H-(X)3-L, where X is any amino acid, and wherein + denotes a metal ion (e.g. Mg2+) binding site; (q) D- R-X-H-(X)33-N+-(X)7-E-(X)3-F, where X is any amino acid, and wherein + denotes a metal ion (e.g. Mg2+) binding site; or (r) D-R-X-H-(X)33-N+-(X)7-E-(X)3-L, where X is any amino acid, and wherein + denotes a metal ion (e.g. Mg2+) binding site. In particular, the nuclease may comprise an active site comprising the amino acid sequence: (a) A-K-X-H-(X)29-N+-(X)7-E-(X)3-R, where X is any amino acid, and wherein + denotes a metal ion (e.g. Mg2+) binding site; (b) A-R-X-H- (X)29-N+-(X)7-E-(X)3-R, where X is any amino acid, and wherein + denotes a metal ion (e.g. Mg2+) binding site; (c) S-R-X-H-(X)29-N+-(X)7-E-(X)3-R, where X is any amino acid, and wherein + denotes a metal ion (e.g. Mg2+) binding site; or (d) D-R-X-H-(X)29-N+-(X)7-E-(X)3-R, where X is any amino acid, and wherein + denotes a metal ion (e.g. Mg2+) binding site. For example, the nuclease may comprise an active site comprising the amino acid sequence A-K-X-H-(X)29- N+-(X)7-E-(X)3-R, where X is any amino acid, and wherein + denotes a metal ion (e.g. Mg2+) binding site.

[0265] The nuclease may comprise an active site comprising the amino acid sequence A / S / D / Q / T / E- K / RA' / l / H-G-H / QA' / F-(X)b-N-(X)2-P-(X)c-N+ / F+ / G+-(X)3-W-(X)3-E / L / M / H / Q-(X)3-R / M / LA / / F, where X is any amino acid; b is 15 to 27 and c is 4 to 7, optionally wherein the sum of b and c is 19 to 33, and wherein + denotes a metal ion (e.g. Mg2+) binding site. In particular, the nuclease may comprise an active site comprising the amino acid sequence A-K-G-H-(X)2o-N- (X)2-P-(X)S-N+-(X)3-W-(X)3-E-(X)3-R, where X is any amino acid, and wherein + denotes a metal ion (e.g. Mg2+) binding site.

[0266] The nuclease may comprise an active site comprising the amino acid residues A231 , K232, H234, N264+, E272 and R276 of SEQ ID NO: 1 , wherein + denotes a metal ion (e.g. Mg2+) binding site. The nuclease may comprise an active site corresponding to the amino acid residues A231 , K232, H234, N264+, E272 and R276 of SEQ ID NO: 1 , wherein + denotes a metal ion (e.g. Mg2+) binding site.

[0267] The active site of the nuclease may comprise a substrate binding site and a metal ion (e.g. Mg2+) binding site.

[0268] The substrate binding site may correspond to A231 , K232 and R276 of SEQ ID NO: 1 and the metal ion (e.g. Mg2+) binding site may correspond to N264 of SEQ ID NO: 1. Accordingly, the nuclease may comprise an active site comprising the amino acid sequence A7S7D7Q7T7E*- K7R7Y7l7H*-X-H / Q / Y / F-(X)a-N F G+-(X)7-E / L / M / H / Q-(X)3-R7M7L7V7F*, where X is any amino acid and a is 23 to 37, wherein * denotes a substrate binding site, and wherein + denotes a metal ion (e.g. Mg2+) binding site. In particular, the nuclease may comprise an active site comprising the amino acid sequence A7S7D*-K7R*-X-H-(X)29-N+-(X)7-E-(X)3-R*, where X is any amino acid, wherein * denotes a substrate binding site, and wherein + denotes a metal ion (e.g. Mg2+) binding site.

[0269] The nuclease may comprise an active site comprising the amino acid sequence: (a) A*-K*-X- H-(X)29-N+-(X)7-E-(X)3-R*, where X is any amino acid, wherein * denotes a substrate binding site, and wherein + denotes a metal ion (e.g. Mg2+) binding site; (b) A*-R*-X-H-(X)29-N+-(X)7-E- (X)3-R*, where X is any amino acid, wherein * denotes a substrate binding site, and wherein + denotes a metal ion (e.g. Mg2+) binding site; (c) S*-R*-X-H-(X)29-N+-(X)7-E-(X)3-R*, where X is any amino acid, wherein * denotes a substrate binding site, and wherein + denotes a metal ion (e.g. Mg2+) binding site; (d) D*-R*-X-H-(X)29-N+-(X)7-E-(X)3-R*, where X is any amino acid, wherein * denotes a substrate binding site, and wherein + denotes a metal ion (e.g. Mg2+) binding site; (e) D*-R*-X-H-(X)3I-N+-(X)7-E-(X)3-R*, where X is any amino acid, wherein * denotes a substrate binding site, and wherein + denotes a metal ion (e.g. Mg2+) binding site; (f) D*-R* -X-H-(X)30-N+-(X)7-E-(X)3-R*, where X is any amino acid, wherein * denotes a substrate binding site, and wherein + denotes a metal ion (e.g. Mg2+) binding site; (g) S*-Y*-X- F-(X)28-F+-(X)7-Q-(X)3-V*, where X is any amino acid, wherein * denotes a substrate binding site, and wherein + denotes a metal ion (e.g. Mg2+) binding site; (h) D*-R*-X-H-(X)37-G+-(X)7- E-(X)3-L*, where X is any amino acid, wherein * denotes a substrate binding site, and wherein + denotes a metal ion (e.g. Mg2+) binding site; (i) Q*-R*-X-Q-(X)26-F+-(X)7-L-(X)3-M*, where X is any amino acid, wherein * denotes a substrate binding site, and wherein + denotes a metal ion (e.g. Mg2+) binding site; (j) T*-H*-X-F-(X)28-F+-(X)7-H-(X)3-L*, where X is any amino acid, wherein * denotes a substrate binding site, and wherein + denotes a metal ion (e.g. Mg2+) binding site; (k) T*-Y*-X-Y-(X)23-F+-(X)7-M-(X)3-L*, where X is any amino acid, wherein * denotes a substrate binding site, and wherein + denotes a metal ion (e.g. Mg2+) binding site; (I) D*-K*-X-H-(X)36-N+-(X)7-E-(X)3-L*, where X is any amino acid, wherein * denotes a substrate binding site, and wherein + denotes a metal ion (e.g. Mg2+) binding site; (m) D*-R*-X-H-(X)35- G+-(X)7-E-(X)3-L*, where X is any amino acid, wherein * denotes a substrate binding site, and wherein + denotes a metal ion (e.g. Mg2+) binding site; (n) E*-I*-X-Q-(X)26-F+-(X)7-L-(X)3-M*, where X is any amino acid, wherein * denotes a substrate binding site, and wherein + denotes a metal ion (e.g. Mg2+) binding site; (o) S*-R*-X-H-(X)30-N+-(X)7-E-(X)3-R*, where X is any amino acid, wherein * denotes a substrate binding site, and wherein + denotes a metal ion (e.g. Mg2+) binding site; (p) S*-Y*-X-F-(X)29-F+-(X)7-H-(X)3-L*, where X is any amino acid, wherein * denotes a substrate binding site, and wherein + denotes a metal ion (e.g. Mg2+) binding site; (q) D*-R*-X-H-(X)33-N+-(X)7-E-(X)3-F*, where X is any amino acid, wherein * denotes a substrate binding site, and wherein + denotes a metal ion (e.g. Mg2+) binding site; or (r) D*-R*-X-H-(X)33-N+-(X)7-E-(X)3-L*, where X is any amino acid, wherein * denotes a substrate binding site, and wherein + denotes a metal ion (e.g. Mg2+) binding site. In particular, the nuclease may comprise an active site comprising the amino acid sequence: (a) A*-K*-X- H-(X)29-N+-(X)7-E-(X)3-R*, where X is any amino acid, wherein * denotes a substrate binding site, and wherein + denotes a metal ion (e.g. Mg2+) binding site; (b) A*-R*-X-H-(X)29-N+-(X)7-E- (X)3-R*, where X is any amino acid, wherein * denotes a substrate binding site, and wherein + denotes a metal ion (e.g. Mg2+) binding site; (c) S*-R*-X-H-(X)29-N+-(X)7-E-(X)3-R*, where X is any amino acid, wherein * denotes a substrate binding site, and wherein + denotes a metal ion (e.g. Mg2+) binding site; or (d) D*-R*-X-H-(X)29-N+-(X)7-E-(X)3-R*, where X is any amino acid, wherein * denotes a substrate binding site, and wherein + denotes a metal ion (e.g. Mg2+) binding site. For example, the nuclease may comprise an active site comprising the amino acid sequence A*-K*-X-H-(X)29-N+-(X)7-E-(X)3-R*, where X is any amino acid, wherein * denotes a substrate binding site, and wherein + denotes a metal ion (e.g. Mg2+) binding site.

[0270] The nuclease may comprise an active site comprising the amino acid sequence A7S7D7Q7T7E*-K7R7Y7l7H*-G-H / QA' / F-(X)b-N-(X)2-P-(X)c-N7F7G+-(X)3-W-(X)3- E / L / M / H / Q-(X)3-R7M7L7V7F*, where X is any amino acid; b is 15 to 27 and c is 4 to 7, optionally wherein the sum of b and c is 19 to 33, wherein * denotes a substrate binding site, and wherein + denotes a metal ion (e.g. Mg2+) binding site. In particular, the nuclease may comprise an active site comprising the amino acid sequence A-K*-G*-H-(X)2o-N-(X)2-P-(X)s- N+-(X)3-W-(X)3-E-(X)3-R*, where X is any amino acid, wherein * denotes a substrate binding site, and wherein + denotes a metal ion (e.g. Mg2+) binding site.

[0271] The nuclease may comprise an active site comprising the amino acid residues A231*, K232*, H234, N264+, E272 and R276* of SEQ ID NO: 1 , wherein * denotes a substrate binding site, and wherein + denotes a metal ion (e.g. Mg2+) binding site. The nuclease may comprise an active site corresponding to the amino acid residues A231*, K232*, H234, N264+, E272 and R276* of SEQ ID NO: 1 , wherein * denotes a substrate binding site, and wherein + denotes a metal ion (e.g. Mg2+) binding site.

[0272] The active site may be present within a nuclease-active portion of the nuclease. The nucleaseactive portion of the nuclease may be an extracellular endonuclease subunit A domain (e.g. a OATH 3.40.570.10 extracellular endonuclease subunit A domain, or an IPR020821 extracellular endonuclease subunit A domain), an endonuclease_NS domain (e.g. a PF01223 endonuclease_NS domain), a non-specific endonuclease domain (e.g. an IPR040255 nonspecific endonuclease domain), an endonuclease related domain (e.g. a PTHR13966 endonuclease related domain), a DNA / RNA non-specific endonuclease domain (e.g. an IPR001604 DNA / RNA non-specific endonuclease domain, a SM00892 DNA / RNA non-specific endonuclease domain, a PF01223 DNA / RNA non-specific endonuclease domain, or a SM00477 DNA / RNA non-specific endonuclease domain), a His-Me finger superfamily domain (e.g. an IPR00044925 His-Me finger superfamily domain or an SSF54060 His-Me finger superfamily domain), a DNA / RNA non-specific endonuclease superfamily domain (e.g. an IPR044929 DNA / RNA non-specific endonuclease superfamily domain), or a NUC domain (e.g. a cd00091 NUC domain).

[0273] The nuclease may comprise an extracellular endonuclease subunit A domain (e.g. a CATH 3.40.570.10 extracellular endonuclease subunit A domain, or an IPR020821 extracellular endonuclease subunit A domain), an endonuclease_NS domain (e.g. a PF01223 endonuclease_NS domain), a non-specific endonuclease domain (e.g. an IPR040255 nonspecific endonuclease domain), an endonuclease related domain (e.g. a PTHR13966 endonuclease related domain), a DNA / RNA non-specific endonuclease domain (e.g. an IPR001604 DNA / RNA non-specific endonuclease domain, a SM00892 DNA / RNA non-specific endonuclease domain, a PF01223 DNA / RNA non-specific endonuclease domain, or a SM00477 DNA / RNA non-specific endonuclease domain), a His-Me finger superfamily domain (e.g. an IPR00044925 His-Me finger superfamily domain or an SSF54060 His-Me finger superfamily domain), a DNA / RNA non-specific endonuclease superfamily domain (e.g. an IPR044929 DNA / RNA non-specific endonuclease superfamily domain), or a NUC domain (e.g. a cd00091 NUC domain).

[0274] The nuclease may comprise an extracellular endonuclease subunit A domain. In particular, the nuclease may comprise an amino acid sequence having at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 99%, or 100% identity to the amino acid sequence of an extracellular endonuclease subunit A domain defined in Table 1 or Table 2. In particular, the nuclease may comprise an amino acid sequence having at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 99%, or 100% identity to the amino acid sequence of an extracellular endonuclease subunit A domain defined in Table 1.

[0275] The nuclease may comprise a CATH 3.40.570.10 extracellular endonuclease subunit A domain. In particular, the nuclease may comprise an amino acid sequence having at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 99%, or 100% identity to the amino acid sequence of a CATH 3.40.570.10 extracellular endonuclease subunit A domain defined in Table 1 or Table 2. The nuclease may comprise an amino acid sequence having at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 99%, or 100% identity to the amino acid sequence of a CATH 3.40.570.10 extracellular endonuclease subunit A domain defined in Table 1. In particular, the nuclease may comprise an amino acid sequence having at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 99%, or 100% identity to amino acids 135 to 389 of SEQ ID NO: 1 or SEQ ID NO: 2. For example, the nuclease may comprise the amino acid sequence of amino acids 135 to 389 of SEQ ID NO: 1. The nuclease may comprise the amino acid sequence of amino acids 135 to 389 of SEQ ID NO: 2.

[0276] The nuclease may comprise an IPR020821 extracellular endonuclease subunit A domain. In particular, the nuclease may comprise an amino acid sequence having at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 99%, or 100% identity to the amino acid sequence of an IPR020821 extracellular endonuclease subunit A domain defined in Table 1 or Table 2. The nuclease may comprise an amino acid sequence having at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 99%, or 100% identity to the amino acid sequence of an IPR020821 extracellular endonuclease subunit A domain defined in Table 1 . In particular, the nuclease may comprise an amino acid sequence having at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 99%, or 100% identity to amino acids 192 to 376 of SEQ ID NO: 1 or SEQ ID NO: 2. For example, the nuclease may comprise the amino acid sequence of amino acids 192 to 376 of SEQ ID NO: 1. The nuclease may comprise the amino acid sequence of amino acids 192 to 376 of SEQ ID NO: 2.

[0277] The nuclease may comprise an endonuclease_NS domain. The nuclease may comprise a PF01223 endonuclease_NS domain. In particular, the nuclease may comprise an amino acid sequence having at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 99%, or 100% identity to amino acids 136 to 385 of SEQ ID NO: 1 or SEQ ID NO: 2. For example, the nuclease may comprise the amino acid sequence of amino acids 136 to 385 of SEQ ID NO: 1 . The nuclease may comprise the amino acid sequence of amino acids 136 to 385 of SEQ ID NO: 2.

[0278] The nuclease may comprise a non-specific endonuclease domain. The nuclease may comprise an IPR040255 non-specific endonuclease domain. In particular, the nuclease may comprise an amino acid sequence having at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 99%, or 100% identity to amino acids 50 to 389 of SEQ ID NO: 1 or SEQ ID NO: 2. For example, the nuclease may comprise the amino acid sequence of amino acids 50 to 389 of SEQ ID NO: 1 . The nuclease may comprise the amino acid sequence of amino acids 50 to 389 of SEQ ID NO: 2.

[0279] The nuclease may comprise an endonuclease related domain. The nuclease may comprise a PTHR13966 endonuclease related domain.

[0280] The nuclease may comprise a DNA / RNA non-specific endonuclease domain. The nuclease may comprise an IPR001604 DNA / RNA non-specific endonuclease domain. In particular, the nuclease may comprise an amino acid sequence having at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 99%, or 100% identity to amino acids 136 to 386 of SEQ ID NO: 1 or SEQ ID NO: 2. For example, the nuclease may comprise the amino acid sequence of amino acids 136 to 386 of SEQ ID NO: 1. The nuclease may comprise the amino acid sequence of amino acids 136 to 386 of SEQ ID NO: 2.

[0281] The nuclease may comprise a SM00892 DNA / RNA non-specific endonuclease domain. In particular, the nuclease may comprise an amino acid sequence having at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 99%, or 100% identity to amino acids 145 to 386 of SEQ ID NO: 1 or SEQ ID NO: 2. For example, the nuclease may comprise the amino acid sequence of amino acids 145 to 386 of SEQ ID NO: 1. The nuclease may comprise the amino acid sequence of amino acids 145 to 386 of SEQ ID NO: 2.

[0282] The nuclease may comprise a PF01223 DNA / RNA non-specific endonuclease domain. In particular, the nuclease may comprise an amino acid sequence having at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 99%, or 100% identity to amino acids 136 to 385 of SEQ ID NO: 1 or SEQ ID NO: 2. For example, the nuclease may comprise the amino acid sequence of amino acids 136 to 385 of SEQ ID NO: 1 . The nuclease may comprise the amino acid sequence of amino acids 136 to 385 of SEQ ID NO: 2.

[0283] The nuclease may comprise a SM00477 DNA / RNA non-specific endonuclease domain. In particular, the nuclease may comprise an amino acid sequence having at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 99%, or 100% identity to amino acids 192 to 376 of SEQ ID NO: 1 or SEQ ID NO: 2. For example, the nuclease may comprise the amino acid sequence of amino acids 192 to 376 of SEQ ID NO: 1 . The nuclease may comprise the amino acid sequence of amino acids 192 to 376 of SEQ ID NO: 2.

[0284] The nuclease may comprise a His-Me finger superfamily domain. The nuclease may comprise an IPR00044925 His-Me finger superfamily domain. The nuclease may comprise an SSF54060 His-Me finger superfamily domain.

[0285] The nuclease may comprise a DNA / RNA non-specific endonuclease superfamily domain. The nuclease may comprise an IPR044929 DNA / RNA non-specific endonuclease superfamily domain.

[0286] The nuclease may comprise a NUC domain. The nuclease may comprise a cd00091 NUC domain. In particular, the nuclease may comprise an amino acid sequence having at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 99%, or 100% identity to amino acids 146 to 394 of SEQ ID NO: 1 or SEQ ID NO: 2. For example, the nuclease may comprise the amino acid sequence of amino acids 146 to 394 of SEQ ID NO: 1. The nuclease may comprise the amino acid sequence of amino acids 146 to 394 of SEQ ID NO: 2.

[0287] The nuclease may have an amino acid sequence having at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 99%, or 100% sequence identity to the amino acid sequence of a nuclease shown in Table 1 , Table 2 or Table 3. The nuclease may have an amino acid sequence having at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 99%, or 100% sequence identity to the amino acid sequence of a nuclease shown in Table 1 or Table 2. The nuclease may have an amino acid sequence having at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 99%, or 100% sequence identity to the amino acid sequence of a nuclease shown in Table 1. In particular, the nuclease may have an amino acid sequence having at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 99%, or 100% sequence identity to SEQ ID NO: 1 or SEQ ID NO: 2. For example, the nuclease may have an amino acid sequence of SEQ ID NO: 1 or SEQ ID NO: 2.

[0288] The nuclease may be encoded by a nucleotide sequence encoding for a nuclease having at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 99%, or 100% sequence identity to the amino acid sequence of a nuclease shown in Table 1 , Table 2, or Table 3. The nuclease may be encoded by a nucleotide sequence encoding for a nuclease having at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 99%, or 100% sequence identity to the amino acid sequence of a nuclease shown in Table 1 or Table 2. The nuclease may be encoded by a nucleotide sequence encoding for a nuclease having an amino acid sequence having at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 99%, or 100% sequence identity to the amino acid sequence of a nuclease shown in Table 1. The nuclease may be encoded by a nucleotide sequence encoding for a nuclease having an amino acid sequence having at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 99%, or 100% sequence identity to SEQ ID NO: 1 or SEQ ID NO: 2. In particular, the nuclease may be encoded by a nucleotide sequence having an amino acid sequence of SEQ ID NO: 1 or SEQ ID NO: 2.

[0289] The nuclease may be encoded by a nucleotide sequence having at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 99%, or 100% sequence identity to SEQ ID NO: 166. In particular, the nuclease may be encoded by the nucleotide sequence of SEQ ID NO: 166. Table 1 : Table of exemplary nucleases suitable for use in the invention. For each nuclease, the Table shows the amino acid residues which correspond to particular domains of the nuclease.

[0290] SEQ ID NO: 2 is a nuclease variant of C9YSL6 and is expected to have the same domain structure and morphology as C9YSL6.

[0291] Table 2: Table of exemplary nucleases suitable for use in the invention. For a given nuclease, the Table shows the amino acid residues which correspond to particular domains of the nuclease.

[0292]

[0293]

[0294]

[0295]

[0296]

[0297] *SEQ ID NO: 140 is a nuclease variant of P13717 and is expected to have the same domain structure and morphology asP13717.

[0298] Table 3: Table of exemplary nucleases suitable for use in the invention.

[0299] Table 4: Table showing active site residues of exemplary nucleases suitable for use in the invention.

[0300] * SEQ ID NO: 2 is a nuclease variant of C9YSL6 and is expected to have the same active site residues as C9YSL6. SEQ ID NO: 140 is a nuclease variant of P13717 and is expected to have the same active site residues as P13717.

[0301] Table 5: Table showing active site residues of micrococcal nuclease.

[0302] The methods of the invention comprise fragmenting chromatin to produce nucleic acid fragments. As used herein, the term “nucleic acid” encompasses chromatin, DNA and RNA. The nucleic acid fragments may be chromatin fragments or DNA fragments. In particular, the nucleic acid fragments may be DNA fragments. The chromatin may be fragmented in order to allow the nucleic acid fragments subsequently to be ligated to other nucleic acid sequences within chromatin that were in close physical proximity in the nucleus.

[0303] In the fragmentation step, the nuclease may cleave in regions of chromatin comprising inter- nucleosomal linkers. Thus, inter-nucleosomal linkers of the chromatin may be cleaved by the nuclease. Cleavage of inter-nucleosomal linkers may produce nucleosomes. Thus, the method may comprise fragmenting chromatin to produce nucleosomes. Cleavage of inter-nucleosomal linkers may produce mono-nucleosomes. In particular, the method may comprise fragmenting chromatin to produce mono-nucleosomes. Thus, the nucleic acid fragments may comprise mono-nucleosomes.

[0304] The fragmenting of the chromatin may not be carried out to completion. In particular, the inter- nucleosomal linkers may be kept at least partially intact. The inter-nucleosomal linkers may be cleaved by the nuclease but kept at least partially intact. The inter-nucleosomal linkers may be of length 10-500, 10-200, 50-200, or 10-100 base pairs after fragmentation. In particular, the method may comprise fragmenting chromatin to produce mono-nucleosomes with internucleosome linker nucleic acids attached. For example, the method may comprise fragmenting chromatin to produce mono-nucleosomes of 180-200 bp with inter-nucleosome linker nucleic acids attached. Thus, the nucleic acid fragments may comprise mono-nucleosomes with internucleosome linker nucleic acids attached.

[0305] The fragmenting of the chromatin may be carried out to completion. Thus, the nucleic acid fragments may comprise mono-nucleosomes without inter-nucleosomal linkers attached. In particular, the nucleic acid fragments may comprise mono-nucleosomes without inter- nucleosomal linkers attached and mono-nucleosomes with inter-nucleosomal linkers attached.

[0306] The nucleic acid fragments may be any suitable size. For example, the nucleic acid fragments may have at least 145 bp, at least 150 bp, at least 155 bp, at least 160 bp, at least 165 bp, at least 170 bp, at least 175 bp, at least 180 bp, at least 185 bp, or at least 190 bp. The nucleic acid fragments have up to 290 bp, up to 275 bp, up to 250 bp, up to 225 bp, or up to 200 bp. The nucleic acid fragments may have at least 145 bp and up to 290 bp, at least 145 bp and up to 200 bp, at least 180 bp and up to 200 bp, or at least 145 bp and up to 190bp. In particular, the nucleic acid fragments may have at least 145 bp and up to 190 bp, or at least 180 bp and up to 200 bp.

[0307] The nucleic acid fragments may have a range of sizes. The nucleic acid fragments may have an average size of at least 145 bp, at least 150 bp, at least 155 bp, at least 160 bp, at least 165 bp, at least 170 bp, at least 175 bp, at least 180 bp, at least 185 bp, or at least 190 bp. The nucleic acid fragments may have an average size of 290 bp, up to 275 bp, up to 250 bp, up to 225 bp, or up to 200 bp. The nucleic acid fragments may have an average size of at least 145 bp and up to 290 bp, at least 145 bp and up to 200 bp, at least 180 bp and up to 200 bp, or at least 145 bp and up to 190bp. In particular, the nucleic acid fragments may have an average size of at least 145 bp and up to 190 bp, or at least 180 bp and up to 200 bp. The average may be the mean average, median average or mode average. In particular, the average may be the mode average.

[0308] The inventors have shown that improved resolution may be obtainable when chromatin is cleaved to mono-nucleosomes with inter-nucleosomal linkers attached, e.g., when chromatin is cleaved to produce nucleic acid fragments of at least 145 bp and up to 290 bp length, and particularly nucleic acid fragments of at least 145 bp and up to 190 bp, or at least 180 bp and up to 200 bp. Accordingly, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, or at least 90% of the nucleic acid fragments may be 145-290 bp fragments. In particular, at least 70% of the nucleic acid fragments may be 145-290bp fragments. At least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, or at least 90% of the nucleic acid fragments may be 145-190 bp fragments. In particular, at least 70% of the nucleic acid fragments may be 145-190bp fragments. At least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, or at least 90% of the nucleic acid fragments may be 180-200 bp fragments. In particular, at least 70% of the nucleic acid fragments may be 180- 200bp fragments. Up to 20%, up to 15%, up to 10%, or up to 5% of the nucleic acid fragments may be smaller than 150 bp fragments. In particular, up to 10% of the nucleic acid fragments may be smaller than 150 bp fragments. For example, at least 70% of the nucleic acid fragments may be 145-190bp fragments and up to 10% of the nucleic acid fragments may be smaller than 150 bp fragments.

[0309] The degree of chromatin fragmentation and / or extent of inter-nucleosomal linker degradation may be assayed by gel electrophoresis, e.g., by using an automated system such as the Agilent Tapestation (D1000 reagents). The nucleic acid fragments may be defined by their 5’ or 3’ terminal nucleic acid residues. At least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, or at least 40% of the nucleic acid fragments may have a 3’ cytosine (C). Up to 50%, up to 45%, up to 40%, up to 35%, up to 30%, up to 25%, or up to 20% of the nucleic acid fragments may have a 3’ C. At least 10% and up to 50%, at least 15% and up to 35%, or at least 20% and up to 30% of the nucleic acid fragments may have a 3’ C. In particular, at least 15 and up to 35% of the nucleic acid fragments may have a 3’ C.

[0310] At least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, or at least 40% of the nucleic acid fragments may have a 3’ guanine (G). Up to 50%, up to 45%, up to 40%, up to 35%, up to 30%, up to 25%, or up to 20% of the nucleic acid fragments may have a 3’ G. At least 10% and up to 50%, at least 15% and up to 35%, or at least 20% and up to 30% of the nucleic acid fragments may have a 3’ G. In particular, at least 15% and up to 35% of the nucleic acid fragments may have a 3’ G.

[0311] At least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, or at least 40% of the nucleic acid fragments may have a 3’ thymine (T). Up to 50%, up to 45%, up to 40%, up to 35%, up to 30%, up to 25%, or up to 20% of the nucleic acid fragments may have a 3’ T. At least 10% and up to 50%, at least 15% and up to 35%, or at least 20% and up to 30% of the nucleic acid fragments may have a 3’ T. In particular, at least 15 and up to 35% of the nucleic acid fragments may have a 3’ T.

[0312] At least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, or at least 40% of the nucleic acid fragments may have a 3’ adenine (A). Up to 50%, up to 45%, up to 40%, up to 35%, up to 30%, up to 25%, or up to 20% of the nucleic acid fragments may have a 3’ A. At least 10% and up to 50%, at least 15% and up to 35%, or at least 20% and up to 30% of the nucleic acid fragments may have a 3’ A. In particular, at least 15% and up to 35% of the nucleic acid fragments may have a 3’ A.

[0313] In particular: (a) at least 15 and up to 35% of the nucleic acid fragments may have a 3’ C; (b) at least 15% and up to 35% of the nucleic acid fragments may have a 3’ G; (c) at least 15 and up to 35% of the nucleic acid fragments may have a 3’ T; and (d) at least 15% and up to 35% of the nucleic acid fragments may have a 3’ A, wherein the sum of (a), (b), (c) and (d) is 100%. At least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, or at least 40% of the nucleic acid fragments may have a 5’ cytosine (C). Up to 50%, up to 45%, up to 40%, up to 35%, up to 30%, up to 25%, or up to 20% of the nucleic acid fragments may have a 5’ C. At least 10% and up to 50%, at least 10% and up to 40%, at least 15% and up to 35%, or at least 20% and up to 30% of the nucleic acid fragments may have a 5’ C. In particular, at least 10% and up to 40% of the nucleic acid fragments may have a 5’ C.

[0314] At least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, or at least 40% of the nucleic acid fragments may have a 5’ guanine (G). Up to 50%, up to 45%, up to 40%, up to 35%, up to 30%, up to 25%, or up to 20% of the nucleic acid fragments may have a 5’ G. At least 10% and up to 50%, at least 10% and up to 40%, at least 15% and up to 35%, or at least 20% and up to 30% of the nucleic acid fragments may have a 5’ G. In particular, at least 10% and up to 40% of the nucleic acid fragments may have a 5’ G.

[0315] At least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, or at least 40% of the nucleic acid fragments may have a 5’ thymine (T). Up to 50%, up to 45%, up to 40%, up to 35%, up to 30%, up to 25%, or up to 20% of the nucleic acid fragments may have a 5’ T. At least 10% and up to 50%, at least 10% and up to 40%, at least 15% and up to 35%, or at least 20% and up to 30% of the nucleic acid fragments may have a 5’ T. In particular, at least 10% and up to 40% of the nucleic acid fragments may have a 5’ T.

[0316] At least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, or at least 40% of the nucleic acid fragments may have a 5’ adenine (A). Up to 50%, up to 45%, up to 40%, up to 35%, up to 30%, up to 25%, or up to 20% of the nucleic acid fragments may have a 5’ A. At least 10% and up to 50%, at least 10% and up to 40%, at least 15% and up to 35%, or at least 20% and up to 30% of the nucleic acid fragments may have a 5’ A. In particular, at least 10% and up to 40% of the nucleic acid fragments may have a 5’ A.

[0317] In particular: (a) at least 10% and up to 40% of the nucleic acid fragments may have a 5’ C; (b) at least 10% and up to 40% of the nucleic acid fragments may have a 5’ G; (c) at least 10% and up to 40% of the nucleic acid fragments may have a 5’ T; and (d) at least 10% and up to 40% of the nucleic acid fragments may have a 5’ A, wherein the sum of (a), (b), (c) and (d) is 100%.

[0318] In previous methods, the concentration of nucleases used to fragment chromatin was carefully selected based on the amount of chromatin being fragmented. However, the inventors have shown that nucleases described herein can be used at a much wider concentration range which does not depend on the amount of chromatin being fragmented. Accordingly, the method may comprise fragmenting chromatin with a fixed concentration of the nuclease. As used herein a “fixed” concentration of a nuclease may be a concentration which does not depend on the amount of chromatin being fragmented, which in turn may depend on the number of cells in a population of cells.

[0319] The method may comprise fragmenting chromatin with at least 0.25 Units, 0.5 Units, at least 1 Unit, at least 2.5 Units, at least 5 Units, at least 10 Units, at least 15 Units, at least 20 Units, at least 25 Units, at least 30 Units, at least 35 Units, or at least 40 Units of the nuclease. The method may comprise fragmenting chromatin with up to 80 Units, up to 60 Units, up to 40 Units, up to 35 Units, up to 30 Units, up to 25 Units, or up to 20 Units of the nuclease. The method may comprise fragmenting chromatin with at least 0.25 and up to 80 Units, at least 0.5 and up to 40 Units, or at least 15 Units and up to 25 Units of the nuclease. In particular, the method may comprise fragmenting the chromatin with at least 0.5 Units and up to 40 Units of the nuclease. The method may comprise fragmenting chromatin with 0.25 Units, 0.5 Units, 1 Unit, 2.5 Units, 5 Units, 10 Units, 15 Units, 20 Units, 25 Units, 30 Units, 35 Units, or 40 Units of the nuclease. In particular, the method may comprise fragmenting chromatin with 20 Units of the nuclease. The Units may be determined using any suitable means. For example, the Units may be determined using Kunitz or a modified version of Kunitz. Kunitz is a standard spectrophotometric assay used to determine nuclease activity. In a Kunitz assay, the change in absorbance of a mixture comprising a substrate and a nuclease may be measured to determine the activity of the nuclease. The modified version of Kunitz may use a pH of 8, addition of 0.5 volumes of 12% HCI04to stop the reaction, and removal of precipitated material by centrifugation before reading the absorbance in the supernatant at 260 nm. In both Kunitz and the modified version of Kunitz, one unit of DNase activity may correspond to an increase in absorbance of 0.001 / min.

[0320] The chromatin may be fragmented within a population of cells. The population of cells may be from any suitable source. For example, the population of cells may be a population of eukaryotic cells. Examples of eukaryotic cells include cells from animals, plants and fungi. In particular, the eukaryotic cells may be higher eukaryotic cells or cells from multicellular organisms. The plants may be monocots or dicots. The eukaryotic cells may be animal cells. In particular, the eukaryotic cells may be vertebrate cells. For example, the eukaryotic cells may be mammalian cells. The population of cells may not be yeast cells. The population of cells may not comprise yeast cells.

[0321] The mammalian cells may be from a human, monkey, mouse, rat, rabbit, guinea pig, sheep, horse, pig, cow, goat, dog, or a cat. In particular, the mammalian cells may be from human cells.

[0322] The population of cells may be a population of THP-1 cells or a population of monocytes, e.g. mammalian THP-1 cells or mammalian monocytes.

[0323] The number of cells in the population of cells (e.g. population of permeabilised cells) may be 1-10,000; 10,000-100,000; 100,000-250,000; 250,000-500,000; 500,000-1 million; or 1 million to 100 million.

[0324] The population of cells may be a population of permeabilised cells. Accordingly, the method may comprise permeabilising a population of cells prior to fragmenting to produce a population of permeabilised cells. In this step, the cell membrane of the cells may be permeabilised to allow the nuclease to gain access to the chromatin in the cells. As used herein, the term “permeabilise” may mean that the membranes of the cells are rendered permeable to the nuclease but the membranes remain otherwise intact. Thus, the cells may not be lysed. The cell membranes may not be completely destroyed. The cell membranes may not be completely removed. The cell membranes may be permeabilised without removing the cell membranes. The cell membranes may be permeabilised in substantially all or all of the cells.

[0325] Cells in the population of permeabilised cells may comprise one or more cell pores. For example, cells in the population of permeabilised cells may comprise one or more permeabilised cell pores. The pores of the permeabilised cells may have any suitable size. The pores of the permeabilised cells may be at least about 1 nm cell pores, at least about 2 nm cell pores, at least about 3 nm cell pores, at least about 4 nm cell pores, at least about 5 nm cell pores, at least about 6 nm cell pores, at least about 7 nm cell pores, at least about 8 nm cell pores, at least about 9 nm cell pores, at least about 10 nm cell pores, at least about 15 nm cell pores, or at least about 20 nm cell pores. In particular, the pores of the permeabilised cells may be at least 10 nm cell pores. The pores of the permeabilised cells may be up to about 150 nm cell pores, up to about 100 nm cell pores, up to about 90 nm cell pores, up to about 80 nm cell pores, up to about 70 nm cell pores, up to about 60 nm cell pores, up to about 50 nm cell pores, up to about 40 nm cell pores, up to about 30 nm cell pores, up to about 20 nm cell pores, or up to about 10 nm cell pores. The pores of the permeabilised cells may be at least about 1 nm and up to about 150 nm cell pores, at least about 5 nm and up to about 100 nm cell pores, or at least about 5 nm and up to about 20 nm cell pores. In particular, the pores of the permeabilised cells may be at least about 5 nm and up to about 20 nm cell pores. The pores of the permeabilised cells may be about 1 nm cell pores, about 2 nm cell pores, about 3 nm cell pores, about 4 nm cell pores, about 5 nm cell pores, about 6 nm cell pores, about 7 nm cell pores, about 8 nm cell pores, about 9 nm cell pores, about 10 nm cell pores, about 15 nm cell pores, or about 20 nm cell pores. In particular, the pores of the permeabilised cells may be about 10 nm cell pores.

[0326] The pores of the permeabilised cells may be at least 1 nm cell pores, at least 2 nm cell pores, at least 3 nm cell pores, at least 4 nm cell pores, at least 5 nm cell pores, at least 6 nm cell pores, at least 7 nm cell pores, at least 8 nm cell pores, at least 9 nm cell pores, at least 10 nm cell pores, at least 15 nm cell pores, or at least 20 nm cell pores. In particular, the pores of the permeabilised cells may be at least 10 nm cell pores. The pores of the permeabilised cells may be up to 150 nm cell pores, up to 100 nm cell pores, up to 90 nm cell pores, up to 80 nm cell pores, up to 70 nm cell pores, up to 60 nm cell pores, up to 50 nm cell pores, up to 40 nm cell pores, up to 30 nm cell pores, up to 20 nm cell pores, or up to 10 nm cell pores. The pores of the permeabilised cells may be at least 1 nm and up to 150 nm cell pores, at least 5 nm and up to 100 nm cell pores, or at least 5 nm and up to 20 nm cell pores. In particular, the pores of the permeabilised cells may be at least 5 nm and up to 20 nm cell pores. The pores of the permeabilised cells may be 1 nm cell pores, 2 nm cell pores, 3 nm cell pores, 4 nm cell pores, 5 nm cell pores, 6 nm cell pores, 7 nm cell pores, 8 nm cell pores, 9 nm cell pores, 10 nm cell pores, 15 nm cell pores, or 20 nm cell pores. In particular, the pores of the permeabilised cells may be 10 nm cell pores.

[0327] The cell membranes may be permeabilised using a membrane-permeabilising agent. Examples of membrane-permeabilising agents include Digitonin, Saponin, Tergitol-type NP40, Triton X-100, sodium dodecyl sulphate, Tween 20 and glycodiosgenin. In particular, the permeabilising agent may comprise digitonin or glycodiosgenin.

[0328] The method may comprise permeabilising a population of cells and fragmenting chromatin within the population of permeabilised cells with a nuclease described herein.

[0329] The chromatin may be fragmented within nuclei. For example, the chromatin may be fragmented within nuclei of a population of permeabilised cells. The chromatin may be fragmented within nuclei isolated from a population of cells (e.g. within isolated nuclei). Nuclei may be isolated from a population of cells by removing cell membranes of the population of cells. Accordingly, the method may comprise removing the cell membranes from a population of cells, e.g., a population of cells described herein, prior to fragmenting to produce a population of nuclei. In this step, the cell membrane of the cells may be removed to allow the nuclease to gain access to the chromatin in the nuclei. Thus, the cells may be lysed. The cell membranes may be completely destroyed. The cell membranes may be completely removed. The cell membranes may be removed in substantially all or all of the cells.

[0330] The cell membranes may be removed using a membrane-removing agent. Examples of membrane-removing agents include SDS (e.g., 0.1% SDS in PBS), Saponin (e.g., 1% Saponin in PBS), Triton® X-100 (e.g., 1% Triton® X-100 in PBS) and Tween® 20 (e.g., 1% Tween® 20 in PBS). In particular, the membrane-removing agent may comprise SDS (e.g., 0.1% SDS in PBS).

[0331] The method may comprise removing cell membranes from a population of cells to produce a population of isolated nuclei and fragmenting chromatin within the population of isolated nuclei with a nuclease described herein.

[0332] The nuclei may be permeabilised nuclei. For example, the chromatin may be within permeabilised nuclei of a population of permeabilised cells. The chromatin may be within isolated permeabilised nuclei. Accordingly, the method may comprise permeabilising nuclei. The method may comprise permeabilising nuclei within a population of cells prior to fragmenting. In this step, the nuclear membranes of the nuclei may be permeabilised to allow the nuclease to gain access to the chromatin in the nuclei. Thus, the nuclear membranes may not be completely destroyed. The nuclear membranes may not be completely removed. The nuclear membranes may be permeabilised without removing the nuclear membranes. The nuclear membranes may be permeabilised in substantially all or all of the cells.

[0333] The nuclear membranes may be permeabilised using a membrane-permeabilising agent, e.g., a membrane-permeabilising agent described herein.

[0334] Thus, the method may comprise permeabilising a population of cells or removing the cell membranes from a population of cells and permeabilising nuclei prior to fragmenting. In particular, the method may comprise permeabilising a population of cells to produce a population of permeabilised cells and permeabilising nuclei within the population of permeabilised cells prior to fragmenting. The method may comprise removing the cell membranes from a population of cells to produce a population of nuclei and permeabilising the nuclei prior to fragmenting.

[0335] The chromatin may be isolated. Chromatin may be isolated by removing cell membranes and nuclear membranes of the population of cells. Accordingly, the method may comprise removing nuclear membranes from nuclei prior to fragmenting. In this step, nuclear membranes of the nuclei may be removed to allow the nuclease to gain access to the chromatin. Thus, the nuclear membranes may be completely destroyed. The nuclear membranes may be completely removed.

[0336] The nuclear membranes may be removed using a membrane-removing agent, e.g., a membrane-removing agent described herein.

[0337] Thus, the method may comprise removing the cell membranes from a population of cells to produce a population of nuclei and removing nuclear membranes from the nuclei prior to fragmenting.

[0338] Following removal of the nuclear membranes, the chromatin may then be isolated. Thus, the method may comprise isolating chromatin from a population of cells prior to fragmenting.

[0339] The population of cells may be a population of cells described herein.

[0340] The chromatin may be immobilised chromatin. In particular, the chromatin may be cross-linked chromatin, such as formaldehyde cross-linked chromatin, methanol cross-linked chromatin, disuccinimidyl glutarate (DSG) cross-linked chromatin, or ethylene glycol bis(succinimidyl succinate)) (EGS) cross-linked chromatin. For example, the chromatin may be formaldehyde cross-linked chromatin. The chromatin may be formaldehyde and methanol and DSG and EGS cross-linked chromatin.

[0341] Accordingly, the method may comprise immobilising the chromatin prior to fragmenting. For example, the method may comprise immobilising the chromatin within a population of cells prior to fragmenting the chromatin. The method may comprise immobilising the chromatin prior to permeabilising the cells. In this step, the chromatin may be immobilised such that regions within the chromatin which were interacting with one another are held or fixed in close proximity. The immobilisation may be carried out on an individual-cell basis. The immobilisation may be carried out in situ, e.g. within a cell nucleus. In particular, the immobilisation may be carried out within substantially all or all of the cells in a population of cells.

[0342] The chromatin may be immobilised using an immobilising agent. The immobilising agent may be a substance which is capable of entering cells and immobilising chromatin within the cells such that regions within the chromatin which were interacting with one another are held or fixed in close proximity. The cells may be immobilised within plugs of the immobilising agent. Examples of immobilising agents include gels, for example hydrogels, formed from crosslinked polymers, formaldehyde, methanol, DSG and EGS.

[0343] A hydrogel may be a network of polymer chains that are hydrophilic, sometimes found as a colloidal gel, in which water is the dispersion medium. The structure of the hydrogel may be changed by varying the concentration of the hydrogel-forming polymer in the hydrogel. Examples of hydrogel polymers include polyvinyl alcohol, acrylate polymers (e.g. sodium acrylate) and polymers with an abundance of hydrophilic groups. Other hydrogel polymers include agarose, alginate, methylcellulose, hyaluronan, elastin-like polypeptides and other naturally-derived polymers. In particular, the immobilising agent may be an agarose gel.

[0344] The chromatin may be immobilised by cross-linking the chromatin. The method may comprise cross-linking the chromatin. In particular, the method may comprise cross-linking the chromatin prior to fragmenting the chromatin. For example, the method may comprise cross-linking the chromatin within a population of cells prior to fragmenting the chromatin. The method may comprise cross-linking the chromatin prior to permeabilising the cells. In particular, the method may comprise cross-linking chromatin within a population of cells; permeabilising the population of cells; and fragmenting chromatin within the population of permeabilised cells with a nuclease described herein.

[0345] In the cross-linking step, the chromatin may be cross-linked such that regions within the chromatin which were interacting with one another are held or fixed in close proximity. The chromatin regions which interact with one another may be DNA elements which affect or control the expression of an associated gene or other aspects of genome function or structure. The DNA elements may be cis-regulatory regions. For example, cis-regulatory regions described herein may comprise a promoter, enhancer, repressor, intron, insulator and / or silencer. The regions of chromatin which were interacting with one another may be cross-linked directly (e.g. nucleic acid to nucleic acid) or indirectly (e.g. by cross-linking of the chromatin to moieties (e.g. proteins) which are bound to the chromatin or between proteins bound to chromatin directly or indirectly). In particular, the chromatin may be cross-linked using a cross-linking agent. The cross-linking agent may be one which is capable of entering cells, e.g. unpermeabilised cells. The cross-linking agent may be an agent which is capable of entering into cells and cross-linking the chromatin within those cells such that regions within the chromatin which were interacting with one another are held or fixed in close proximity. The cross-linking agent may comprise formaldehyde, methanol, DSG and / or EGS. The crosslinking agent may comprise formaldehyde, methanol, DSG and EGS. In particular, the crosslinking agent may comprise formaldehyde.

[0346] Thus, in the methods of the invention the nucleic acid fragments may be immobilised nucleic acid fragments. In particular, the nucleic acid fragments may be cross-linked nucleic acid fragments, such as formaldehyde cross-linked nucleic acid fragments, methanol cross-linked nucleic acid fragments, disuccinimidyl glutarate (DSG) cross-linked nucleic acid fragments, and / or ethylene glycol bis(succinimidyl succinate)) (EGS) cross-linked nucleic acid fragments. For example, the nucleic acid fragments may be formaldehyde cross-linked nucleic acid fragments. The nucleic acid fragments may be formaldehyde and methanol and DSG and EGS cross-linked nucleic acid fragments.

[0347] The present invention also relates to a pool of nucleic acid fragments obtainable, or obtained, by a method of the invention. Accordingly, the invention provides a pool of nucleic acid fragments obtainable, or obtained, by a method of the invention.

[0348] The invention also provides a pool of nucleic acid fragments obtainable, or obtained, by fragmenting chromatin with a nuclease described herein within a population of permeabilised cells.

[0349] The invention also provides a pool of nucleic acid fragments, wherein the pool of nucleic acid fragments is present in a population of cells (e.g. a population of permeabilised cells described herein), and wherein: (a) 15-35% of the nucleic acid fragments have a 3’ terminal C residue; (b) 15-35% of the nucleic acid fragments have a 3’ terminal G residue; (c) 15-35% of the nucleic acid fragments have a 3’ terminal T residue; and (d) 15-35% of the nucleic acid fragments have a 3’ terminal A residue, wherein the sum of (a), (b), (c) and (d) is 100%. In some instances, (e) 10-40% of the nucleic acid fragments have a 5’ terminal C residue; (f) 10-40% of the nucleic acid fragments have a 5’ terminal G residue; (g) 10-40% of the nucleic acid fragments have a 5’ terminal T residue; and (h) 10-40% of the nucleic acid fragments have a 5’ terminal A residue, wherein the sum of (e), (f), (g) and (h) is 100%.

[0350] The invention also provides a pool of nucleic acid fragments, wherein the pool of nucleic acid fragments is present in a population of cells (e.g. a population of permeabilised cells described herein), and wherein: (a) 10-40% of the nucleic acid fragments have a 5’ terminal C residue; (b) 10-40% of the nucleic acid fragments have a 5’ terminal G residue; (c) 10-40% of the nucleic acid fragments have a 5’ terminal T residue; and (d) 10-40% of the nucleic acid fragments have a 5’ terminal A residue, wherein the sum of (a), (b), (c) and (d) is 100%. In some instances, (e) 15-35% of the nucleic acid fragments have a 3’ terminal C residue; (f) 15-35% of the nucleic acid fragments have a 3’ terminal G residue; (g) 15-35% of the nucleic acid fragments have a 3’ terminal T residue; and (h) 15-35% of the nucleic acid fragments have a 3’ terminal A residue, wherein the sum of (e), (f), (g) and (h) is 100%.

[0351] The nucleic acid fragments in the pool of nucleic acid fragments may be immobilised nucleic acid fragments. In particular, the nucleic acid fragments in the pool of nucleic acid fragments may be cross-linked nucleic acid fragments, such as formaldehyde cross-linked nucleic acid fragments, methanol cross-linked nucleic acid fragments, disuccinimidyl glutarate (DSG) cross-linked nucleic acid fragments, or ethylene glycol bis(succinimidyl succinate)) (EGS) cross-linked nucleic acid fragments. For example, the nucleic acid fragments in the pool of nucleic acid fragments may be formaldehyde cross-linked nucleic acid fragments. The nucleic acid fragments in the pool of nucleic acid fragments may be formaldehyde and methanol and DSG and EGS cross-linked nucleic acid fragments.

[0352] The invention also relates to a method of producing a 3C library. Accordingly, the invention provides a method of producing a 3C library, the method comprising (a) ligating cross-linked nucleic acid fragments obtainable, or which have been obtained, by a method of the invention, and (b) de-crosslinking the ligated nucleic acid fragments.

[0353] The invention also provides a method of producing a 3C library, the method comprising (a) ligating a pool of cross-linked nucleic acid fragments of the invention; and (b) de-crosslinking the ligated nucleic acid fragments.

[0354] The invention also provides a method of producing a 3C library, the method comprising the steps: (a) fragmenting chromatin by a method of the invention to produce immobilised (e.g. cross-linked) nucleic acid fragments;

[0355] (b) ligating the nucleic acid fragments to produce ligated nucleic acid fragments; and

[0356] (c) de-immobilising (e.g. de-crosslinking) the ligated nucleic acid fragments.

[0357] The invention also provides a method of producing a 3C library, the method comprising the steps:

[0358] (a) cross-linking chromatin (e.g. within a population of permeabilised cells);

[0359] (b) fragmenting the cross-linked chromatin (e.g., within the population of permeabilised cells) using a nuclease described herein to produce cross-linked nucleic acid fragments;

[0360] (c) ligating the cross-linked nucleic acid fragments to produce ligated nucleic acid fragments; and

[0361] (d) de-crosslinking the ligated nucleic acid fragments.

[0362] The invention also provides a method of producing a 3C library, the method comprising (a) a step for ligating cross-linked nucleic acid fragments obtainable, or which have been obtained, by a method of producing nucleic acid fragments of the invention, and (b) a step for decrosslinking the ligated nucleic acid fragments.

[0363] The invention also provides a method of producing a 3C library, the method comprising (a) a step for ligating a pool of cross-linked nucleic acid fragments of the invention; and (b) a step for de-crosslinking the ligated nucleic acid fragments.

[0364] The invention also provides a method of producing a 3C library, the method comprising:

[0365] (a) a step for fragmenting chromatin (e.g., using a method of the invention) to produce immobilised (e.g. cross-linked) nucleic acid fragments;

[0366] (b) a step for ligating the nucleic acid fragments to produce ligated nucleic acid fragments; and

[0367] (c) a step for de-immobilising (e.g. de-crosslinking) the ligated nucleic acid fragments.

[0368] The invention also provides a method of producing a 3C library, the method comprising:

[0369] (a) a step for cross-linking chromatin (e.g. within a population of permeabilised cells);

[0370] (b) a step for fragmenting the cross-linked chromatin (e.g., within the population of permeabilised cells) (e.g., using a nuclease described herein) to produce cross-linked nucleic acid fragments; (c) a step for ligating the cross-linked nucleic acid fragments to produce ligated nucleic acid fragments; and

[0371] (d) a step for de-crosslinking the ligated nucleic acid fragments.

[0372] As used herein, the term “3C library” refers to a library of DNA fragments, wherein the DNA fragments comprise contiguously-joined DNA elements wherein the DNA elements are ones which are capable of interacting with one another (for example within a cell).

[0373] The method of producing a 3C library may comprise a step of ligating the nucleic fragments to produce ligated nucleic acid fragments. The ligated nucleic acid fragments may be ligated chromatin fragments or ligated DNA fragments.

[0374] In the ligating step, the free ends of the nucleic acid fragments which were produced by fragmenting chromatin may be ligated together in order to produce ligated nucleic acid fragments.

[0375] Ligation may occur in a random manner between the free ends of the nucleic acid fragments. However, ligation may occur most preferably between adjacent free nucleic acid ends which are held in close proximity to one another, for example by the immobilisation (e.g. cross-linking) process. In this way, regions of nucleic acid within the nucleic acid sample which previously interacted with one another may now become chemically joined (ligated) to one another.

[0376] The length of the ligated nucleic acid fragments may be greater than 200bp. Thus, there may be an increase in fragment size in profile of the nucleic acid fragment lengths. For example, there may be an increase in fragment size such that there is very little nucleic acid fragments of the size of fragments in the main mono-nucleosomal peak following the fragmentation reaction.

[0377] Prior to ligation, the ends of the nucleic acid fragments may be blunted and phosphorylated, e.g. using T4 polynucleotide kinase (PNK) and DNA Polymerase I, Large (Klenow) Fragment.

[0378] Ligation may be carried out using any suitable ligating agent, e.g. a ligase. In particular, the ligase may be a DNA ligase. Examples of suitable ligases include T4 DNA ligase. The method of producing a 3C library may comprise a step of de-immobilising or decrosslinking the ligated nucleic acid fragments. If the cells have not already been lysed, then they may lysed at this time.

[0379] If the cell membranes have not previously been removed, then the cell membranes may be removed at this time, e.g. with a lysis buffer, proteinase K or heat treatment or suitable detergent. Alternatively, sufficient amounts of a permeabilising agent (e.g. as disclosed herein) may be used. In particular, the nuclear and / or cell membranes may not be removed until this step.

[0380] In this step, the ligated nucleic acid fragments (e.g. ligated chromatin fragments) may be deimmobilised (e.g. decrosslinked) in order to produce linear nucleic acid fragments (e.g. linear chromatin fragments). For example, the immobilising agent may be removed / dissolved or the crosslinking moieties may be cleaved or removed.

[0381] The crosslinks may be removed by heating the ligated nucleic acid fragments to a high temperature, such as to 50°C, 60°C, 70°C, 80°C or greater. In particular, the de-crosslinking may be carried out using Proteinase K. Optionally, non-nucleic acid material (e.g. proteins, cross-linking agents, etc.) may also be removed at this time. RNA may also be removed from the sample at this point, for example using RNase. In particular, the ligated nucleic acid fragments may be extracted with phenol / chloroform or solid phase extraction methods (such as Qiagen spin columns).

[0382] The invention also provides a 3C library of nucleic acid fragments obtainable, or obtained, by a method of the invention.

[0383] The invention also provides a 3C library of nucleic acid fragments, wherein: (a) 15-35% of the nucleic acid fragments have a 3’ terminal C residue; (b) 15-35% of the 3C nucleic acid fragments have a 3’ terminal G residue; (c) 15-35% of the nucleic acid fragments have a 3’ terminal T residue; and (d) 15-35% of the nucleic acid fragments have a 3’ terminal A residue, wherein the sum of (a), (b), (c) and (d) is 100%. In some instances, (e) 10-40% of the nucleic acid fragments have a 5’ terminal C residue; (f) 10-40% of the 3C nucleic acid fragments have a 5’ terminal G residue; (g) 10-40% of the nucleic acid fragments have a 5’ terminal T residue; and (h) 10-40% of the nucleic acid fragments have a 5’ terminal A residue, wherein the sum of (e), (f), (g) and (h) is 100%. The invention also provides a 3C library of nucleic acid fragments, wherein: (a) 10-40% of the nucleic acid fragments have a 5’ terminal C residue; (b) 10-40% of the nucleic acid fragments have a 5’ terminal G residue; (c) 10-40% of the nucleic acid fragments have a 5’ terminal T residue; and (d) 10-40% of the nucleic acid fragments have a 5’ terminal A residue, wherein the sum of (a), (b), (c) and (d) is 100%. In some instances, (e) 15-35% of the nucleic acid fragments have a 3’ terminal C residue; (f) 15-35% of the 3C nucleic acid fragments have a 3’ terminal G residue; (g) 15-35% of the nucleic acid fragments have a 3’ terminal T residue; and (h) 15-35% of the nucleic acid fragments have a 3’ terminal A residue, wherein the sum of (e), (f), (g) and (h) is 100%.

[0384] The invention also relates to a method of identifying chromatin regions within a nucleic acid sample which interact with one another. Accordingly, the invention provides a method of identifying chromatin regions within a nucleic acid sample which interact with one another, the method comprising:

[0385] (a) fragmenting a 3C library obtainable, or which has been obtained, by a method of the invention;

[0386] (b) optionally, adding sequencing adaptors to the ends of the nucleic acid fragments and / or amplifying the nucleic acid fragments;

[0387] (c) contacting the nucleic acid fragments with a targeting nucleic acid which binds to a subgroup of the nucleic acid fragments, wherein the targeting nucleic acid is labelled with the first half of a binding pair;

[0388] (d) isolating the nucleic acid fragments which have been bound by the targeting nucleic acid using the second half of the binding pair;

[0389] (e) amplifying the isolated subgroup of nucleic acid fragments;

[0390] (f) optionally repeating steps (c), (d), and (e) one or more times; and

[0391] (g) optionally sequencing the amplified isolated subgroup of nucleic acid fragments, in order to identify chromatin regions within the nucleic acid sample which interact with one another.

[0392] The invention also provides a method of identifying chromatin regions within a nucleic acid sample which interact with one another, the method comprising:

[0393] (a) fragmenting a 3C library of the invention;

[0394] (b) optionally, adding sequencing adaptors to the ends of the nucleic acid fragments and / or amplifying the nucleic acid fragments; (c) contacting the nucleic acid fragments with a targeting nucleic acid which binds to a subgroup of the nucleic acid fragments, wherein the targeting nucleic acid is labelled with the first half of a binding pair;

[0395] (d) isolating the nucleic acid fragments which have been bound by the targeting nucleic acid using the second half of the binding pair;

[0396] (e) amplifying the isolated subgroup of nucleic acid fragments;

[0397] (f) optionally repeating steps (c), (d), and (e) one or more times; and

[0398] (g) optionally sequencing the amplified isolated subgroup of nucleic acid fragments, in order to identify chromatin regions within the nucleic acid sample which interact with one another.

[0399] The invention also provides a method of identifying chromatin regions within a nucleic acid sample which interact with one another, the method comprising:

[0400] (a) producing a 3C library using a method of the invention;

[0401] (b) fragmenting the 3C library to produce nucleic acid fragments;

[0402] (c) optionally, adding sequencing adaptors to the ends of the nucleic acid fragments and / or amplifying the nucleic acid fragments;

[0403] (d) contacting the nucleic acid fragments with a targeting nucleic acid which binds to a subgroup of the nucleic acid fragments, wherein the targeting nucleic acid is labelled with the first half of a binding pair;

[0404] (e) isolating the subgroup of nucleic acid fragments which have been bound by the targeting nucleic acid using the second half of the binding pair;

[0405] (f) amplifying the isolated subgroup of nucleic acid fragments;

[0406] (g) optionally repeating Steps (d), (e) and (f) one or more times; and

[0407] (h) optionally sequencing the amplified isolated subgroup of nucleic acid fragments, in order to identify chromatin regions within the nucleic acid sample which interact with one another.

[0408] The invention provides a method of identifying chromatin regions within a nucleic acid sample which interact with one another, the method comprising:

[0409] (a) a step for fragmenting a 3C library obtainable, or which has been obtained, by a method of the invention;

[0410] (b) optionally, a step for adding sequencing adaptors to the ends of the nucleic acid fragments and / or amplifying the nucleic acid fragments; (c) a step for contacting the nucleic acid fragments with a targeting nucleic acid which binds to a subgroup of the nucleic acid fragments, wherein the targeting nucleic acid is labelled with the first half of a binding pair;

[0411] (d) a step for isolating the nucleic acid fragments which have been bound by the targeting nucleic acid using the second half of the binding pair;

[0412] (e) a step for amplifying the isolated subgroup of nucleic acid fragments;

[0413] (f) optionally repeating steps (c), (d), and (e) one or more times; and

[0414] (g) optionally, a step for sequencing the amplified isolated subgroup of nucleic acid fragments, in order to identify chromatin regions within the nucleic acid sample which interact with one another.

[0415] The invention also provides a method of identifying chromatin regions within a nucleic acid sample which interact with one another, the method comprising:

[0416] (a) a step for fragmenting a 3C library of the invention;

[0417] (b) optionally, a step for adding sequencing adaptors to the ends of the nucleic acid fragments and / or amplifying the nucleic acid fragments;

[0418] (c) a step for contacting the nucleic acid fragments with a targeting nucleic acid which binds to a subgroup of the nucleic acid fragments, wherein the targeting nucleic acid is labelled with the first half of a binding pair;

[0419] (d) a step for isolating the nucleic acid fragments which have been bound by the targeting nucleic acid using the second half of the binding pair;

[0420] (e) a step for amplifying the isolated subgroup of nucleic acid fragments;

[0421] (f) optionally repeating steps (c), (d), and (e) one or more times; and

[0422] (g) optionally, a step for sequencing the amplified isolated subgroup of nucleic acid fragments, in order to identify chromatin regions within the nucleic acid sample which interact with one another.

[0423] The invention also provides a method of identifying chromatin regions within a nucleic acid sample which interact with one another, the method comprising:

[0424] (a) a step for producing a 3C library (e.g., using a method of the invention);

[0425] (b) a step for fragmenting the 3C library to produce nucleic acid fragments;

[0426] (c) optionally, a step for adding sequencing adaptors to the ends of the nucleic acid fragments and / or amplifying the nucleic acid fragments; (d) a step for contacting the nucleic acid fragments with a targeting nucleic acid which binds to a subgroup of the nucleic acid fragments, wherein the targeting nucleic acid is labelled with the first half of a binding pair;

[0427] (e) a step for isolating the subgroup of nucleic acid fragments which have been bound by the targeting nucleic acid using the second half of the binding pair;

[0428] (f) a step for amplifying the isolated subgroup of nucleic acid fragments;

[0429] (g) optionally repeating Steps (d), (e) and (f) one or more times; and

[0430] (h) optionally, a step for sequencing the amplified isolated subgroup of nucleic acid fragments, in order to identify chromatin regions within the nucleic acid sample which interact with one another.

[0431] In particular, the targeting nucleic acid may be a DNA oligonucleotide. In particular, the nucleic acid sample may be a sample of eukaryotic cells, for example mammalian cells.

[0432] The methods of identifying chromatin regions within a nucleic acid sample which interact with one another comprise a step of fragmenting a 3C library to produce nucleic acid fragments. In this step, the nucleic acid fragments may be DNA fragments. The lengths of the nucleic acid fragments may be reduced in size. In particular, the lengths of the nucleic acid fragments may be reduced to a size which is suitable for high throughput sequencing, capture and / or amplification.

[0433] In particular, the lengths of the nucleic acid fragments may be reduced to 100-500 base pairs, such as 100-300 or 150-250 base pairs, for example to about 250 base pairs, or to 250 base pairs.

[0434] Fragmentation may be performed by any suitable process. Examples of suitable fragmentation processes include using nucleases (e.g. restriction endonucleases) and sonication. In particular, the 3C library may be fragmented by sonication.

[0435] The methods of identifying chromatin regions within a nucleic acid sample which interact with one another comprise an optional step of adding sequencing adaptors to the ends of the nucleic acid fragments. Furthermore, the nucleic acid fragments may be amplified at this time. In this optional step, sequencing adaptors and / or amplification primers (e.g. short doublestranded nucleic acids) may be added to both ends of the nucleic acid fragments in order to facilitate the amplification and later sequencing of the nucleic acid fragments. Each sequencing adaptor may comprise a unique indexing barcode, e.g. a short nucleic acid motif which acts as a unique identifier for that nucleic acid fragment. In particular, the sequencing adaptors may be Next Generation sequencing adaptors. The sequencing adaptors may comprise P5 or P7 sequences. P5 or P7 sequences may mediate binding to the flow cell and bridge amplification. Thus, the sequencing adaptors may comprise P5 or P7 sequences which mediate binding to the flow cell and bridge amplification. Internal binding sites for sequencing primers and barcodes may also be added to allow indexing of samples. The sequencing adaptors may be added to the nucleic acid fragments by ligation-mediated PCR.

[0436] The 3C nucleic acid fragments may also be amplified (e.g. by PCR). For example, 1-20 rounds of PCR may be performed, in particular 3-10 rounds for example about 6 rounds of PCR or 6 rounds of PCR.

[0437] The indexed samples may optionally be pooled for multiplex sequence analysis.

[0438] The methods of identifying chromatin regions within a nucleic acid sample which interact with one another comprise a step of contacting the nucleic acid fragments with a targeting nucleic acid which binds to a subgroup of the nucleic acid fragments, wherein the targeting nucleic acid is labelled with the first half of a binding pair. In this step, the desired nucleic acid fragments (e.g. DNA fragments) may be prepared for isolation from the background of contaminating nucleic acid fragments.

[0439] A targeting nucleic acid may be used which has a nucleotide sequence which is complementary or substantially complementary to that of a desired region of the nucleic acids within the nucleic acid sample. The targeting nucleic acid may therefore hybridise, under appropriate conditions, to the desired region of the nucleic acid within the nucleic acid sample.

[0440] For example, the desired region of the nucleic acid may be that of a promoter from a particular gene (e.g., wherein it is desired to determine which chromatin regions interact with that promoter) or it may be that of an enhancer element (e.g., wherein it is desired to determine which genes are enhanced by that element).

[0441] The targeting nucleic acid may be single- or double-stranded. In particular, the targeting nucleic acid may be single-stranded. The targeting nucleic acid may be DNA or RNA. In particular, the targeting nucleic may be DNA (e.g. a DNA oligonucleotide). If a restriction endonuclease is used in the production of the 3C library, the targeting nucleic acid may contain the ends of the restriction fragment containing the desired region and may include the restriction endonuclease site. In this way, the targeting nucleic acid may bind to informative ligation junctions.

[0442] The concentration of the targeting nucleic acid (e.g. a DNA oligonucleotide) may be 5pM to 1 pM. In particular, the concentration of the targeting nucleic acid (e.g. a DNA oligonucleotide) may be 2.9pM to 29pM. For example, the concentration of the targeting nucleic acid (e.g. a DNA oligonucleotide) may be 1 pM to 30pM, or 300nM to 30pM.

[0443] The concentration of the targeting nucleic acid (e.g. a DNA oligonucleotide) may be 30nM to 0.3nM. The concentration of the targeting nucleic acid (e.g. a DNA oligonucleotide) may be about 2.9nM or 2.9nM. This may apply for each oligonucleotide used.

[0444] The same targeting nucleic acid may be used in any repeat of this step.

[0445] Targeting nucleic acids (e.g. labelled oligonucleotides) may be designed to bind to any sequence within the genome of the organism being studied. Thus, the targeting nucleic acid (e.g. a labelled oligonucleotide) may be sited (e.g., designed to bind) within a nucleosome- depleted region of a promoter of a gene or non-coding RNA of interest, or a regulatory element (e.g. an enhancer, repressor or CTCF binding site).

[0446] In particular, the targeting nucleic acid (e.g. a labelled oligonucleotide) may be sited (i.e. designed to bind) within the central region of a nucleosome-depleted region of a promoter of a gene or non-coding RNA of interest, or a regulatory element (e.g. an enhancer, repressor or CTCF binding site). As used herein, the term “central” region refers to the middle 50% (preferably middle 30%, 20% or 10%) of the sequence of the nucleosome-depleted region. The term “central” region may also refer to the middle 500, 400, 300, 200, 100 or 50 base pairs of the sequence of the nucleosome-depleted region.

[0447] In this way, a very strong and high-resolution picture of the functional interactions controlling gene (or RNA) expression may be obtained. The nucleosome-depleted regions can readily be defined using assays including DNasel hypersensitivity, ATAC-seq and Chromatin immunoprecipitation. The targeting nucleic acid may be designed to bind to (or overlap with) a DNasel hypersensitive site or an ATAC sequence of a promoter of a gene or non-coding RNA of interest, or a regulatory element in the nucleic acid.

[0448] In contrast, when the targeting nucleic acid (e.g. a labelled oligonucleotide) is moved 1000 bp to the left or right of the central region, the physical interaction profile may be attenuated and it may become more difficult to define the regulatory contacts precisely.

[0449] At loci where gene regulation is already well-defined (such as the alpha and beta globin loci, HBA and HBB), the profiles obtainable from methods of the invention from the central nucleosome-depleted regions at the promoters define all of the known regulatory elements down to almost single base pair resolution.

[0450] The transcription factor binding sites at the distal regulatory elements can also be defined from the signal from the central part of a promoter. This can be achieved by using the junction site between the part of the capture read (at the promoter) and the reporter read (at the enhancer). The transcription factor binding sites can be defined because where they bind to the DNA there is a reduction in cut-site density. Thus the strongest signals may occur at the unprotected sites in between the transcription factor binding sites. This is similar to DNasel hypersensitivity foot printing assays.

[0451] Examples of binding pairs include biotin with streptavidin. Accordingly, the first half of the binding pair may be biotin.

[0452] The methods of identifying chromatin regions within a nucleic acid sample which interact with one another comprise a step of isolating the subgroup of nucleic acid fragments which have been bound by the targeting nucleic acid using the second half of the binding pair. In this step, the second half of the binding pair may be allowed to bind to the first half of the binding pair. In particular, the first half of the binding pair may be biotin and the second half of the binding pair may be a streptavidin-coated bead. In order to aid isolation of the targeted nucleic acid fragments, the second half of the binding pair may be bound to a physical support, for example a column or a bead (e.g. a magnetic bead).

[0453] The targeted nucleic acid fragments may then be isolated from the background by virtue of the fact that they will be bound to the column or magnetic beads, wherein the background nucleic acids may then be removed. The method may be carried out on a microarray. The method may not carried be out on a microarray.

[0454] The methods of identifying chromatin regions within a nucleic acid sample which interact with one another comprise a step of amplifying the isolated subgroup of nucleic acid fragments. In this step, the isolated nucleic acid fragments (e.g. DNA fragments) may be amplified in order to enrich the desired nucleic acid fragments. In particular, the isolated subgroup of nucleic acid fragments may be amplified by PCR. For example, the amplification may comprise IQ- 40 cycles of PCR amplification, such as 12-14 cycles.

[0455] The PCR primers used in PCR amplification may bind to the sequencing adaptors. For example, where the sequencing adaptors comprise P5 or P7 sequences, PCR primers which bind to the latter sequences may be used.

[0456] The steps of fragmenting the 3C library; contacting the nucleic acid fragments with a targeting nucleic acid which binds to a subgroup of the nucleic acid fragments, wherein the targeting nucleic acid is labelled with the first half of a binding pair; isolating the subgroup of nucleic acid fragments which have been bound by the targeting nucleic acid using the second half of the binding pair; and amplifying the isolated subgroup of nucleic acid fragments may result in an enrichment of approximately 5-20,000 fold over the corresponding method without performing these steps.

[0457] The steps of contacting the nucleic acid fragments with a targeting nucleic acid which binds to a subgroup of the nucleic acid fragments, wherein the targeting nucleic acid is labelled with the first half of a binding pair; isolating the subgroup of nucleic acid fragments which have been bound by the targeting nucleic acid using the second half of the binding pair; and amplifying the isolated subgroup of nucleic acid fragments may be repeated (e.g. in this order) one or more times. This may result in greater enrichment of the desired nucleic acid fragments over the corresponding method without performing these steps. For example, a greater enrichment may be achieved such that >90% of the reads contain a sequence targeted by the oligonucleotide capture. These steps may be repeated (in this order), for example, 1-5 times, e.g. 1 , 2, 3, 4 or 5 times.

[0458] The steps of the method may be carried out in the order specified. The methods of identifying chromatin regions within a nucleic acid sample which interact with one another comprise an optional step of sequencing the amplified subgroup of nucleic acid fragments. The skilled person will be well aware of numerous DNA sequencing methods which may be used. In particular, the sequencing is performed using an Illumina platform, e.g. Miseq, HiSeq, NextSeq or NovoSeq, using 150 bp paired end sequences (i.e. 300 bp in total).

[0459] The methods of the invention may be carried out in vitro or ex vivo.

[0460] The invention also relates to a method of identifying allele-specific interaction profiles in SNP- containing regions of nucleic acids.

[0461] Accordingly, the invention provides a method of identifying allele-specific interaction profiles in SNP-containing regions of nucleic acids, the method comprising sequencing the amplified isolated subgroup of nucleic acid fragments which has been obtained by a method of the invention, in order to identify allele-specific interaction profiles in SNP-containing regions.

[0462] The invention also provides a method of identifying allele-specific interaction profiles in SNP- containing regions of nucleic acids, the method comprising a step for sequencing the amplified isolated subgroup of nucleic acid fragments which has been obtained by a method of the invention, in order to identify allele-specific interaction profiles in SNP-containing regions.

[0463] The invention also provides a method of identifying allele-specific interaction profiles in SNP- containing regions of nucleic acids, the method comprising carrying out a method of identifying chromatin regions within a nucleic acid sample which interact with one another of the invention, and sequencing the amplified isolated subgroup of nucleic acid fragments in order to identify allele-specific interaction profiles in SNP-containing regions.

[0464] The invention also relates to a method of identifying a chromatin region that is indicative of a disease. Accordingly, the invention provides a method of identifying a chromatin region that is indicative of a disease, the method comprising:

[0465] (a) quantifying a frequency of interaction between a first chromatin region and a second chromatin region within a nucleic acid sample from a subject with the disease, wherein the first chromatin region and second chromatin region have been identified as interacting with one another by a method of the invention; and (b) comparing the frequency of the interaction in the nucleic acid sample from the subject with the disease with the frequency of interaction in a nucleic acid sample from a subject without the disease, wherein a difference in the frequency of interaction between the nucleic acid samples is indicative of the disease.

[0466] The invention also provides a method of identifying a chromatin region that is indicative of a disease, the method comprising:

[0467] (a) carrying out a method of identifying chromatin regions within a nucleic acid sample which interact with one another of the invention wherein the 3C library is from a subject with a disease;

[0468] (b) quantifying a frequency of interaction between a first chromatin region and a second chromatin region in the nucleic acid sample; and

[0469] (c) comparing the frequency of the interaction between the first chromatin region and the second chromatin region from the subject with the disease with the frequency of interaction between the first chromatin region and the second chromatin region from a subject without the disease, wherein a difference in the frequency of interaction is indicative of the disease.

[0470] The invention also provides a method of identifying a chromatin region that is indicative of a disease, the method comprising:

[0471] (a) a step for quantifying a frequency of interaction between a first chromatin region and a second chromatin region within a nucleic acid sample from a subject with the disease, wherein the first chromatin region and second chromatin region have been identified as interacting with one another by a method of the invention; and

[0472] (b) a step for comparing the frequency of the interaction in the nucleic acid sample from the subject with the disease with the frequency of interaction in a nucleic acid sample from a subject without the disease, wherein a difference in the frequency of interaction between the nucleic acid samples is indicative of the disease.

[0473] The invention also provides a method of identifying a chromatin region that is indicative of a disease, the method comprising:

[0474] (a) a step for carrying out a method of identifying chromatin regions within a nucleic acid sample which interact with one another of the invention, wherein the 3C library is from a subject with a disease; (b) a step for quantifying a frequency of interaction between a first chromatin region and a second chromatin region in the nucleic acid sample; and

[0475] (c) a step for comparing the frequency of the interaction between the first chromatin region and the second chromatin region from the subject with the disease with the frequency of interaction between the first chromatin region and the second chromatin region from a subject without the disease, wherein a difference in the frequency of interaction is indicative of the disease.

[0476] The invention also provides a method comprising producing a means for targeting, or a means which has been identified as targeting, an expression product of a nucleotide sequence, wherein the nucleotide sequence, or a nucleotide sequence having at least 90%, 95%, 97%, or 99% identity thereto, is comprised by a chromatin region which has been identified as being indicative of a disease by a method of the invention.

[0477] The invention also provides a method comprising a step for producing a means for targeting, or a means which has been identified as targeting, an expression product of a nucleotide sequence, wherein the nucleotide sequence, or a nucleotide sequence having at least 90%, 95%, 97%, or 99% identity thereto, is comprised by a chromatin region which has been identified as being indicative of a disease by a method of the invention

[0478] The invention also provides a method comprising producing an agent which targets, or has been identified as targeting, an expression product of nucleotide sequence, wherein the nucleotide sequence, or a nucleotide sequence having at least 90%, 95%, 97%, or 99% identity thereto, is comprised by a chromatin region which has been identified as being indicative of a disease by a method of the invention.

[0479] The invention also provides a method comprising:

[0480] (a) identifying a chromatin region that is indicative of a disease by carrying out a method of identifying a chromatin region that is indicative of a disease of the invention; and

[0481] (b) producing a means for targeting, or a means which has been identified as targeting, an expression product of a nucleotide sequence, wherein the nucleotide sequence, or a nucleotide sequence having at least 90%, 95%, 97%, or 99% identity thereto, is comprised by the chromatin region identified in step (a). The invention also provides a method comprising:

[0482] (a) identifying a chromatin region that is indicative of a disease by carrying out a method of identifying a chromatin region that is indicative of a disease of the invention; and

[0483] (b) a step for producing a means for targeting, or a means which has been identified as targeting, an expression product of a nucleotide sequence, wherein the nucleotide sequence, or a nucleotide sequence having at least 90%, 95%, 97%, or 99% identity thereto, is comprised by the chromatin region identified in step (a).

[0484] The method may comprise identifying a means for targeting the expression product of the nucleotide sequence, wherein the nucleotide sequence, or a nucleotide sequence having at least 90%, 95%, 97%, or 99% identity thereto, is comprised by the chromatin region identified in step (a).

[0485] The method may comprise a step for identifying a means for targeting the expression product of the nucleotide sequence, wherein the nucleotide sequence, or a nucleotide sequence having at least 90%, 95%, 97%, or 99% identity thereto, is comprised by the chromatin region identified in step (a).

[0486] The invention also provides a method comprising:

[0487] (a) identifying a chromatin region that is indicative of a disease by carrying out a method of identifying a chromatin region that is indicative of a disease of the invention; and

[0488] (b) producing an agent which targets, or which has been identified as targeting, an expression product of a nucleotide sequence, wherein the nucleotide sequence, or a nucleotide sequence having at least 90%, 95%, 97%, or 99% identity thereto, is comprised by the chromatin region identified in step (a).

[0489] The method may comprise identifying an agent which targets the expression product of the nucleotide sequence, wherein the nucleotide sequence, or a nucleotide sequence having at least 90%, 95%, 97%, or 99% identity thereto, is comprised by the chromatin region identified in step (a).

[0490] The method may comprise a step for identifying an agent which targets the expression product of the nucleotide sequence, wherein the nucleotide sequence, or a nucleotide sequence having at least 90%, 95%, 97%, or 99% identity thereto, is comprised by the chromatin region identified in step (a). The expression product may be any suitable expression product. In particular, the expression product may be an expression product of a gene. The gene may comprise the nucleotide sequence.

[0491] The invention also provides a method comprising producing a means for targeting, or a means which has been identified as targeting, an expression product regulated by a nucleotide sequence, wherein the nucleotide sequence, or a nucleotide sequence having at least 90%, 95%, 97%, or 99% identity thereto, is comprised by a chromatin region which has been identified as being indicative of a disease by a method of the invention.

[0492] The invention also provides a method comprising a step for producing a means for targeting, or a means which has been identified as targeting, an expression product regulated by a nucleotide sequence, wherein the nucleotide sequence, or a nucleotide sequence having at least 90%, 95%, 97%, or 99% identity thereto, is comprised by a chromatin region which has been identified as being indicative of a disease by a method of the invention.

[0493] The invention also provides a method comprising producing an agent which targets an expression product regulated by a nucleotide sequence, wherein the nucleotide sequence, or a nucleotide sequence having at least 90%, 95%, 97%, or 99% identity thereto, is also comprised by a chromatin region which has been identified as being indicative of a disease by a method of the invention.

[0494] The invention also provides a method comprising:

[0495] (a) identifying a chromatin region that is indicative of a disease by carrying out a method of identifying a chromatin region that is indicative of a disease of the invention; and

[0496] (b) producing a means for targeting, or a means which has been identified as targeting, an expression product regulated by a nucleotide sequence, wherein the nucleotide sequence, or a nucleotide sequence having at least 90%, 95%, 97%, or 99% identity thereto, is comprised by the chromatin region identified in step (a).

[0497] The invention also provides a method comprising:

[0498] (a) identifying a chromatin region that is indicative of a disease by carrying out a method of identifying a chromatin region that is indicative of a disease of the invention; and

[0499] (b) a step for producing a means for targeting, or a means which has been identified as targeting, an expression product regulated by a nucleotide sequence, wherein the nucleotide sequence, or a nucleotide sequence having at least 90%, 95%, 97%, or 99% identity thereto, is comprised by the chromatin region identified in step (a).

[0500] The method may comprise identifying a means for targeting the expression product regulated by the nucleotide sequence, wherein the nucleotide sequence, or a nucleotide sequence having at least 90%, 95%, 97%, or 99% identity thereto, is comprised by the chromatin region identified in step (a).

[0501] The method may comprise a step for identifying a means for targeting the expression product regulated by the nucleotide sequence, wherein the nucleotide sequence, or a nucleotide sequence having at least 90%, 95%, 97%, or 99% identity thereto, is comprised by the chromatin region identified in step (a).

[0502] The invention also provides a method comprising:

[0503] (a) identifying a chromatin region that is indicative of a disease by carrying out a method of identifying a chromatin region that is indicative of a disease of the invention; and

[0504] (b) producing an agent which targets, or which has been identified as targeting, an expression product regulated by a nucleotide sequence, wherein the nucleotide sequence, or a nucleotide sequence having at least 90%, 95%, 97%, or 99% identity thereto, is also comprised by the chromatin region identified in step (a).

[0505] The method may comprise identifying an agent which targets the expression product regulated by the nucleotide sequence, wherein the nucleotide sequence, or a nucleotide sequence having at least 90%, 95%, 97%, or 99% identity thereto, is comprised by the chromatin region identified in step (a).

[0506] The method may comprise a step for identifying an agent which targets the expression product regulated by the nucleotide sequence, wherein the nucleotide sequence, or a nucleotide sequence having at least 90%, 95%, 97%, or 99% identity thereto, is comprised by the chromatin region identified in step (a).

[0507] In particular, the methods of the invention may comprise producing an agent which targets, or which has been identified as targeting, an expression product regulated by a nucleotide sequence, wherein the nucleotide sequence is comprised by a chromatin region which has been identified as being indicative of a disease. The expression product regulated by the nucleotide sequence may be regulated by a regulatory element comprising the nucleotide sequence. The expression product may be regulated by a regulatory element comprising the nucleotide sequence. The regulatory element may be a cis-regulatory element, e.g., a cis-regulatory element described herein.

[0508] The means for targeting, or the means which has been identified as targeting, an expression product described herein may be any suitable means, e.g., an agent. Thus, the means (e.g., agent) may target an expression product of a chromatin region that is involved in the disease. The chromatin region may be causative of the disease. The means (e.g., agent) may agonise the expression product. For example, the means (e.g., agent) may agonise the expression product if the expression product is under-expressed in the disease. The means (e.g., agent) may antagonise the expression product. For example, the means (e.g., agent) may antagonise the expression product if the expression product is over-expressed in the disease.

[0509] The means may be a therapeutic means, e.g., a therapeutic agent. The means (e.g., therapeutic agent) may be a chemically synthesised compound (e.g. a small molecule); a botanically available molecule isolated from a plant, fungi, or mould; a biotherapeutic molecule (e.g. an antibody or antigen-binding fragment thereof); or a nucleic acid molecule (e.g., an antisense oligonucleotide). The means (e.g., therapeutic agent) may be an antibody, or antigen-binding fragment thereof, a small molecule, a polypeptide, a peptide, a peptidomimetic, or a nucleic acid (e.g., an antisense oligonucleotide). In particular, the means (e.g., therapeutic agent) may be an antisense oligonucleotide, a small molecule, or an antibody or antigen-binding fragment thereof. The antibody or antigen-binding fragment thereof described herein may be an antibody, a Fab fragment, a Fab' fragment, a F(ab)2fragment, a single chain Fv (scFv) fragment, a single domain antibody, a multispecific antibody (e.g., a bispecific antibody, a trispecific antibody, a BiTE, a DART, a BiKE, or a TriKE), or a non-lg binding domain such as an adnectin, an affbody, or a fibronectin domain. In particular, the antibody or antigen-binding fragment thereof described herein may be an antibody. The means (e.g., therapeutic agent) may be a drug, a prodrug, or an immunogen.

[0510] The expression product (e.g., of the gene) may be a protein (e.g., encoded by the gene). The protein may be any suitable protein. The protein may be a cell signalling protein (e.g., a ligand such as a growth factor or cytokine, or a receptor), a structural protein (e.g., an extracellular matrix protein), a hormonal protein, an enzyme, or a transport protein. The expression product (e.g., of the gene) may be an RNA sequence (e.g., encoded by the gene), such as an mRNA sequence. The RNA sequence, such as the mRNA sequence, may encode for a protein, e.g., a protein described herein. The RNA sequence may be a non-coding RNA sequence. For example, the RNA sequence may be a microRNA sequence, or an enhancer RNA sequence.

[0511] The expression product may be encoded by a human gene. The expression product may be encoded by an animal gene. For example, the expression product may be encoded by a mammalian gene. The expression product may be any expression product of a cell from a population of cells described herein.

[0512] The invention also provides an agent obtainable, or which has been obtained, by a method of the invention.

[0513] The invention also provides a mean for targeting, or which has been identified as targeting, an expression product described herein.

[0514] The invention also provides a method of producing a pharmaceutical composition, the method comprising combining:

[0515] (a) an agent of the invention; or

[0516] (b) an agent obtainable, or which has been obtained, by a method of the invention, with a pharmaceutically acceptable carrier, excipient or diluent.

[0517] The invention also provides a method of producing a pharmaceutical composition, the method comprising combining:

[0518] (a) a means for targeting an expression product described herein; or

[0519] (b) a means which has been identified as targeting an expression product described herein, with a pharmaceutically acceptable carrier, excipient or diluent.

[0520] The invention also provides a pharmaceutical composition obtainable, or which has been obtained, by a method of producing a pharmaceutical composition of the invention.

[0521] The methods described herein may comprise a step of administering a means for targeting, or a means which has been identified as targeting, an expression product described herein (e.g., administering an agent described herein) to a subject to treat a disease or condition. For example, the methods of identifying a chromatin region that is indicative of a disease of the invention may further comprise a step of administering a means for targeting, or a means which has been identified as targeting, an expression product described herein (e.g., administering an agent described herein) to the subject to treat the disease. Such methods may additionally comprise a step of producing a means for targeting, or a means which has been identified as targeting, an expression product described herein (e.g., producing an agent described herein).

[0522] Thus, the invention provides a method comprising:

[0523] (a) carrying out a method of identifying a chromatin region that is indicative of a disease of the invention;

[0524] (b) producing an agent which targets, or which has been identified as targeting, an expression product comprising a nucleotide sequence, wherein the nucleotide sequence, or a nucleotide sequence having at least 90%, 95%, 97%, or 99% identity thereto, is also comprised by the chromatin region identified in step (a); and

[0525] (c) administering the agent to the subject to treat a disease or condition.

[0526] The invention provides a method comprising:

[0527] (a) carrying out a method of identifying a chromatin region that is indicative of a disease of the invention;

[0528] (b) producing a means for targeting, or a means which has been identified as targeting, an expression product comprising a nucleotide sequence, wherein the nucleotide sequence, or a nucleotide sequence having at least 90%, 95%, 97%, or 99% identity thereto, is also comprised by the chromatin region identified in step (a); and

[0529] (c) administering the means to the subject to treat a disease or condition.

[0530] The invention also provides a method comprising:

[0531] (a) carrying out a method of identifying a chromatin region that is indicative of a disease of the invention; and

[0532] (b) producing an agent which targets an expression product regulated by a nucleotide sequence, wherein the nucleotide sequence, or a nucleotide sequence having at least 90%, 95%, 97%, or 99% identity thereto, is also comprised by the chromatin region identified in step (a); and

[0533] (c) administering the agent to the subject to treat a disease or condition.

[0534] The invention also provides a method comprising:

[0535] (a) carrying out a method of identifying a chromatin region that is indicative of a disease of the invention; and (b) producing a means for targeting, or a means which has been identified as targeting, an expression product regulated by a nucleotide sequence, wherein the nucleotide sequence, or a nucleotide sequence having at least 90%, 95%, 97%, or 99% identity thereto, is also comprised by the chromatin region identified in step (a); and

[0536] (c) administering the means to the subject to treat a disease or condition.

[0537] The invention also provides a method of treating or preventing a disease or condition in a human or animal subject, the method comprising administering an agent of the invention or a pharmaceutical composition of the invention to the subject.

[0538] The invention also provides an agent of the invention, or a pharmaceutical composition of the invention, for use in a method of treating or preventing a disease or condition in a human or animal subject, the method comprising administering the agent or pharmaceutical composition to the subject.

[0539] The invention also provides use of an agent of the invention, or a pharmaceutical composition of the invention in the manufacture of a medicament for treating or preventing a disease or condition in a human or animal subject.

[0540] The agent described herein may have been produced by a method of the invention. The pharmaceutical composition described herein may have been produced by a method of producing a pharmaceutical composition of the invention.

[0541] The invention also provides a method of treating or preventing a disease or condition in a human or animal subject, the method comprising administering a means for targeting, or a means which has been identified as targeting, an expression product described herein to the subject.

[0542] The invention also provides a means for targeting, or a means which has been identified as targeting, an expression product described herein, for use in a method of treating or preventing a disease or condition in a human or animal subject, the method comprising administering the agent or pharmaceutical composition to the subject.

[0543] The invention also provides use of a means for targeting, or a means which has been identified as targeting, an expression product described herein, in the manufacture of a medicament for treating or preventing a disease or condition in a human or animal subject. The means for targeting, or the means which has been identified as targeting, an expression product described herein described herein may have been produced by a method of the invention.

[0544] The disease or condition described herein may be any suitable disease, disorder or condition. Optionally, the disease or condition of a human or animal subject or patient is selected from (a) A neurodegenerative disease or condition; (b) A brain disease or condition; (c) A CNS disease or condition; (d) Memory loss or impairment; (e) A heart or cardiovascular disease or condition, eg, heart attack, stroke or atrial fibrillation; (f) A liver disease or condition; (g) A kidney disease or condition, eg, chronic kidney disease (CKD); (h) A pancreas disease or condition; (i) A lung disease or condition, eg, cystic fibrosis or COPD; (j) A gastrointestinal disease or condition; (k) A throat or oral cavity disease or condition; (I) An ocular disease or condition; (m) A genital disease or condition, eg, a vaginal, labial, penile or scrotal disease or condition; (n) A sexually-transmissible disease or condition, eg, gonorrhea , HIV infection, syphilis or Chlamydia infection; (o) An ear disease or condition; (p) A skin disease or condition; (q) A heart disease or condition; (r) A nasal disease or condition; (s) A haematological disease or condition, eg, anaemia, eg, anaemia of chronic disease or cancer; (t) A viral infection; (u) A pathogenic bacterial infection; (v) A cancer; (w) An autoimmune disease or condition, eg, SLE; (x) An inflammatory disease or condition, eg, rheumatoid arthritis, psoriasis, eczema, asthma, ulcerative colitis, colitis, Crohn's disease or IBD; (y) Autism; (z) ADHD; (aa) Bipolar disorder; (bb) ALS [Amyotrophic Lateral Sclerosis]; (cc) Osteoarthritis; (dd) A congenital or development defect or condition; (ee) Miscarriage; (ff) A blood clotting condition; (gg) Bronchitis; (hh) Dry or wet AMD; (ii) Neovascularisation (eg, of a tumour or in the eye); (jj) Common cold; (kk) Epilepsy; (II) Fibrosis, eg, liver or lung fibrosis; (mm) A fungal disease or condition, eg, thrush; (nn) A metabolic disease or condition, eg, obesity, anorexia, diabetes, Type I or Type II diabetes; (oo) Ulcer(s), eg, gastric ulceration or skin ulceration; (pp) Dry skin; (qq) Sjogren's syndrome; (rr) Cytokine storm; (ss) Deafness, hearing loss or impairment; (tt) Slow or fast metabolism (ie, slower or faster than average for the weight, sex and age of the subject); (uu) Conception disorder, eg, infertility or low fertility; (vv) Jaundice; (ww) Skin rash; (xx) Kawasaki Disease; (yy) Lyme Disease; (zz) An allergy, eg, a nut, grass, pollen, dust mite, cat or dog fur or dander allergy; (aaa) Malaria, typhoid fever, tuberculosis or cholera; (bbb) Depression; (ccc) Mental retardation; (ddd) Microcephaly; (eee) Malnutrition; (fff) Conjunctivitis; (ggg) Pneumonia; (hhh) Pulmonary embolism; (iii) Pulmonary hypertension; (jjj) A bone disorder; (kkk) Sepsis or septic shock; (III) Sinusitus; (mmm) Stress (eg, occupational stress); (nnn) Thalassaemia, anaemia, von Willebrand Disease, or haemophilia; (ooo) Shingles or cold sore; (ppp) Menstruation; (qqq) Low sperm count.

[0545] In an example, the neurodegenerative or CNS disease or condition is selected from the group consisting of Alzheimer disease, geriopsychosis, Down syndrome, Parkinson's disease, Creutzfeldt-jakob disease, diabetic neuropathy, Parkinson syndrome, Huntington's disease, Machado-Joseph disease, amyotrophic lateral sclerosis, diabetic neuropathy, and Creutzfeldt Creutzfeldt-Jakob disease. For example, the disease may be Alzheimer disease. For example, the disease may be Parkinson syndrome.

[0546] In an example, wherein the method of the invention is practised on a human or animal subject for treating a CNS or neurodegenerative disease or condition, the method causes downregulation of Treg cells in the subject, thereby promoting entry of systemic monocyte- derived macrophages and / or Treg cells across the choroid plexus into the brain of the subject, whereby the disease or condition (eg, Alzheimer's disease) is treated, prevented or progression thereof is reduced. In an embodiment the method causes an increase of IFN- gamma in the CNS system (eg, in the brain and / or CSF) of the subject. In an example, the method restores nerve fibre and / or reduces the progression of nerve fibre damage. In an example, the method restores nerve myelin and / or reduces the progression of nerve myelin damage. In an example, the method of the invention treats or prevents a disease or condition disclosed in WO2015136541 and / or the method can be used with any method disclosed in WO2015136541 (the disclosure of this document is incorporated by reference herein in its entirety, eg, for providing disclosure of such methods, diseases, conditions and potential therapeutic agents that can be administered to the subject for effecting treatment and / or prevention of CNS and neurodegenerative diseases and conditions, eg, agents such as immune checkpoint inhibitors, eg, anti-PD-1 , anti-PD-L1 , anti-TIM3 or other antibodies disclosed therein).

[0547] Cancers that may be treated include tumours that are not vascularized, or not substantially vascularized, as well as vascularized tumours. The cancers may comprise non-solid tumours (such as haematological tumours, for example, leukaemias and lymphomas) or may comprise solid tumours. Types of cancers to be treated with the invention include, but are not limited to, carcinoma, blastoma, and sarcoma, and certain leukaemia or lymphoid malignancies, benign and malignant tumours, and malignancies e.g., sarcomas, carcinomas, and melanomas. Adult tumours / cancers and paediatric tumours / cancers are also included. Haematologic cancers are cancers of the blood or bone marrow. Examples of haematological (or haematogenous) cancers include leukaemias, including acute leukaemias (such as acute lymphocytic leukaemia, acute myelocytic leukaemia, acute myelogenous leukaemia and myeloblasts, promyeiocytic, myelomonocytic, monocytic and erythroleukaemia), chronic leukaemias (such as chronic myelocytic (granulocytic) leukaemia, chronic myelogenous leukaemia, and chronic lymphocytic leukaemia), polycythemia vera, lymphoma, Hodgkin's disease, non-Hodgkin's lymphoma (indolent and high grade forms), multiple myeloma, Waldenstrom's macroglobulinemia, heavy chain disease, myeiodysplastic syndrome, hairy cell leukaemia and myelodysplasia.

[0548] Solid tumours are abnormal masses of tissue that usually do not contain cysts or liquid areas. Solid tumours can be benign or malignant. Different types of solid tumours are named for the type of cells that form them (such as sarcomas, carcinomas, and lymphomas). Examples of solid tumours, such as sarcomas and carcinomas, include fibrosarcoma, myxosarcoma, liposarcoma, chondrosarcoma, osteosarcoma, and other sarcomas, synovioma, mesothelioma, Ewing's tumour, leiomyosarcoma, rhabdomyosarcoma, colon carcinoma, lymphoid malignancy, pancreatic cancer, breast cancer, lung cancers, ovarian cancer, prostate cancer, hepatocellular carcinoma, squamous eel! carcinoma, basal cell carcinoma, adenocarcinoma, sweat gland carcinoma, medullary thyroid carcinoma, papillary thyroid carcinoma, pheochromocytomas sebaceous gland carcinoma, papillary carcinoma, papillary adenocarcinomas, medullary carcinoma, bronchogenic carcinoma, renal cell carcinoma, hepatoma, bile duct carcinoma, choriocarcinoma, Wilms' tumour, cervical cancer, testicular tumour, seminoma, bladder carcinoma, melanoma, and CNS tumours (such as a glioma (such as brainstem glioma and mixed gliomas), glioblastoma (also known as glioblastoma multiforme) astrocytoma, CNS lymphoma, germinoma, medulloblastoma, Schwannoma craniopharyogioma, ependymoma, pineaioma, hemangioblastoma, acoustic neuroma, oligodendroglioma, menangioma, neuroblastoma, retinoblastoma and brain metastases).

[0549] Autoimmune Diseases for T reatment or Prevention by the Method o Acute Disseminated Encephalomyelitis (ADEM) o Acute necrotizing hemorrhagic leukoencephalitis o Addison's disease o Agammaglobulinemia o Alopecia areata o Amyloidosis o Ankylosing spondylitis Demyelinating neuropathies Dermatitis herpetiformis Dermatomyositis Devic's disease (neuromyelitis optica) Discoid lupus Dressier's syndrome Endometriosis Eosinophilic esophagitis Eosinophilic fasciitis Erythema nodosum Experimental allergic encephalomyelitis Evans syndrome Fibromyalgia Fibrosing alveolitis Giant cell arteritis (temporal arteritis) Giant cell myocarditis Glomerulonephritis Goodpasture's syndrome Granulomatosis with Polyangiitis (GPA) (formerly called Wegener's Granulomatosis) Graves' disease Guillain-Barre syndrome Hashimoto's encephalitis Hashimoto's thyroiditis Hemolytic anemia Henoch-Schonlein purpura Herpes gestationis Hypogammaglobulinemia Idiopathic thrombocytopenic purpura (ITP) IgA nephropathy lgG4-related sclerosing disease Immunoregulatory lipoproteins Inclusion body myositis Interstitial cystitis Juvenile arthritis Juvenile diabetes (Type 1 diabetes) Juvenile myositis Kawasaki syndrome Lambert-Eaton syndrome Leukocytoclastic vasculitis Lichen planus Lichen sclerosus Ligneous conjunctivitis Linear IgA disease (LAD) Lupus (SLE) Lyme disease, chronic Meniere's disease Microscopic polyangiitis Mixed connective tissue disease (MCTD) Mooren's ulcer Mucha-Habermann disease Multiple sclerosis Myasthenia gravis Myositis Narcolepsy Neuromyelitis optica (Devic's) Neutropenia Ocular cicatricial pemphigoid Optic neuritis Palindromic rheumatism PANDAS (Pediatric Autoimmune Neuropsychiatric Disorders Associated with Streptococcus ) Paraneoplastic cerebellar degeneration Paroxysmal nocturnal hemoglobinuria (PNH) Parry Romberg syndrome Parsonnage-Turner syndrome Pars planitis (peripheral uveitis) Pemphigus Peripheral neuropathy Perivenous encephalomyelitis Pernicious anemia POEMS syndrome o Tolosa-Hunt syndrome o Transverse myelitis o Type 1 diabetes o Ulcerative colitis o Undifferentiated connective tissue disease (UCTD) o Uveitis o Vasculitis o Vesiculobullous dermatosis o Vitiligo o Wegener's granulomatosis (now termed Granulomatosis with Polyangiitis (GPA).

[0550] Inflammatory Diseases for Treatment or Prevention by the Method o Alzheimer's o ankylosing spondylitis o arthritis (osteoarthritis, rheumatoid arthritis (RA), psoriatic arthritis) o asthma o atherosclerosis o Crohn's disease o colitis o dermatitis o diverticulitis o fibromyalgia o hepatitis o irritable bowel syndrome (IBS) o systemic lupus erythematous (SLE) o nephritis o Parkinson's disease o ulcerative colitis.

[0551] The disease may be an infectious disease, a deficiency disease, a hereditary disease, or a physiological disease. The disease may be an immune-related disease. For example, the disease may be an autoimmune disease. In the treatments of the invention, the disease may be a disease or disorder associated with a chromatin region that has been identified as indicative of a disease by a method of the invention. The subject may be a human. The subject may be an animal. In particular, the animal may be a mammal.

[0552] The means or agent may be administered to the subject via any suitable route. The means or agent may be administered enterally (e.g., orally, sublingually, or rectally), parenterally (e.g., intravenously, intramuscularly, subcutaneously, or intraarterially), transnasally, by inhalation, transdermally, or intraosseously.

[0553] The invention also provides a kit comprising:

[0554] (a) a CATH Superfamily 3.40.570.10 Extracellular Endonuclease, subunit A nuclease, or a nuclease variant thereof;

[0555] (b) a nuclease comprising a cd00091 active site;

[0556] (c) a nuclease comprising an amino acid sequence of A / S / D / Q / T / E-K / R / Y / l / H-X- H / Q ' / F-X(a)-N / F / G-X(7)-E / L / M / H / Q-X(3)-R / M / LA / / F, where X is any amino acid and a is 23 to 37;

[0557] (d) a nuclease having a ball and chain morphology;

[0558] (e) a nuclease comprising an amino acid sequence having at least 40% identity to SEQ ID NO: 1 or SEQ ID NO: 2; or

[0559] (f) a means for fragmenting chromatin of the invention; and instructions for use of the nuclease to fragment chromatin within a population of permeabilised cells.

[0560] The invention also provides an apparatus comprising:

[0561] (a) means to fragment chromatin in a population of permeabilised cells with:

[0562] (i) a CATH Superfamily 3.40.570.10 Extracellular Endonuclease, subunit A nuclease, or a nuclease variant thereof

[0563] (ii) a nuclease comprising a cd00091 active site;

[0564] (iii) a nuclease comprising an amino acid sequence of A / S / D / Q / T / E-K / R / Y / l / H- X-H / Q / Y / F-X(a)-N / F / G-X(7)-E / L / M / H / Q-X(3)-R / M / UV / F, where X is any amino acid and a is 23 to 37;

[0565] (iv) a nuclease having a ball and chain morphology;

[0566] (v) a nuclease comprising an amino acid sequence having at least 40% identity to SEQ ID NO: 1 or SEQ ID NO: 2; or

[0567] (vi) a means for fragmenting chromatin of the invention; and

[0568] (b) instructions which, when executed, cause the apparatus to carry out a method of producing nucleic acid fragments of the invention, and optionally to carry out: (i) a method of producing a 3C library of the invention;

[0569] (ii) a method of identifying chromatin regions within a nucleic acid sample which interact with one another of the invention;

[0570] (iii) a method of identifying allele-specific interaction profiles in SNP-containing regions of nucleic acids of the invention; and / or

[0571] (iv) a method of identifying a chromatin region that is indicative of a disease of the invention.

[0572] The apparatus may further comprise a machine. Where the apparatus comprises a machine, the instructions may be executed by the machine.

[0573] The invention also provides a method of identifying a nuclease suitable for fragmenting chromatin within a population of permeabilised cells, the method comprising identifying:

[0574] (a) a CATH Superfamily 3.40.570.10 Extracellular Endonuclease, subunit A nuclease, or a nuclease variant thereof;

[0575] (b) a nuclease comprising a cd00091 active site;

[0576] (c) a nuclease comprising an amino acid sequence of A / S / D / Q / T / E-K / R / Y / l / H-X- H / Q / Y / F-X(a)-N / F / G-X(7)-E / L / M / H / Q-X(3)-R / M / L / V / F, where X is any amino acid and a is 23 to 37;

[0577] (d) a nuclease having a ball and chain morphology; or

[0578] (e) a nuclease comprising an amino acid sequence having at least 40% identity to SEQ ID NO: 1 or SEQ ID NO: 2; and determining whether the nuclease fragments chromatin within a population of permeabilised cells.

[0579] The invention also provides a method of fragmenting chromatin within a population of permeabilised cells with a nuclease which has been identified as suitable for fragmenting chromatin within a population of permeabilised cells by a method of the invention.

[0580] This disclosure is not limited by the exemplary methods and materials disclosed herein, and any methods and materials similar or equivalent to those described herein can be used in the practice or testing of embodiments of this disclosure. Numeric ranges are inclusive of the numbers defining the range. Unless otherwise indicated, any nucleic acid sequences are written left to right in 5' to 3' orientation; amino acid sequences are written left to right in amino to carboxy orientation, respectively. Amino acids are referred to herein using the name of the amino acid, the three letter abbreviation or the single letter abbreviation. The term “protein", as used herein, includes proteins, polypeptides, and peptides. As used herein, the term “amino acid sequence” is synonymous with the term “polypeptide” and / or the term “protein”. In some instances, the term “amino acid sequence” is synonymous with the term “peptide”. In some instances, the term “amino acid sequence” is synonymous with the term “enzyme”. The terms "protein" and "polypeptide" are used interchangeably herein. In the present disclosure and claims, the conventional one-letter and three-letter codes for amino acid residues may be used. The 3- letter code for amino acids as defined in conformity with the IUPACIUB Joint Commission on Biochemical Nomenclature (JCBN). It is also understood that a polypeptide may be coded for by more than one nucleotide sequence due to the degeneracy of the genetic code.

[0581] Other definitions of terms may appear throughout the specification. Before the exemplary embodiments are described in more detail, it is to be understood that this disclosure is not limited to particular embodiments described, and as such may vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting, since the scope of the present disclosure will be defined only by the appended claims.

[0582] It must be noted that as used herein and in the appended claims, the singular forms “a”, “an”, and “the” include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to “an agent” includes a plurality of such candidate agents and reference to “the agent” includes reference to one or more agents and equivalents thereof known to those skilled in the art, and so forth.

[0583] The publications discussed herein are provided solely for their disclosure prior to the filing date of the present application. Nothing herein is to be construed as an admission that such publications constitute prior art to the claims appended hereto. Each publication discussed herein is incorporated by reference in its entirety.

[0584] INFORMAL SEQUENCE LISTING

[0585] SEQ ID NO 1 : C9YSL6

[0586] MIGRTTFIALFVKVLTIWSFTKGEDCVWDNDVDYPEYPPLILDSSFQLVLPVLEGDQRITSVQ SGSELILACPGREISALGSEDAQATCLGGKLVEVDGKEWNIVELGCTKMASETIHRNLGQCG DQDLGIYEVIGFDLPTTGHFYELIRVCFDPANETTIFSENIVHGASIAAKDIDPGRPSFKTSTGF FSVSMISVYSQRNQLELMKNLLGDDELAATIIDPSKQFYFAKGHMAPDADFVTWEQDATYY YINALPQWQAFNNGNWKYLEYDTRDLAEKHGTDLTVYSGGWGVLELEDINGNPVEIYLGLA QDKKWPAPALTWKVIYEKDTNRAAAIVGINNPHITTAPEPLCTDICSSLTWLDFDFGDLVHG

[0587] YTYCCSVADLRAAIPNVPDLGDVDILDE

[0588] SEQ ID NO 2: heat labile variant of C9YSL6 MIGRTTFIALFVKVLTIWSFTKGEDCVWDNDVDYPEYPPLILDSSFQLVLPVLEGDQRITSVQ SGSKLILACPGRGISALGSEDAQATCLGGKLVEVDGKEWNIVELGCTKMASETIHRNLGQC GDQDLGIYEVIGFDLPTTGHFYELIRVCFDPANETTIFSENIVHGASIAAKDIDPGRPSFKTST GFFSVSMISVYSQRSQLELMKNLLGDDELAATIIDPSEQFYFAKGHMAADADFVTWEQDAT YYYINALPQWQAFNNGNWKYLEYDTRDLAEKHGTDLTVYSGGWGVLELEDINGNPVEIYLG LAQDKKWPAPALTWKVIYEKDTNRAAAIVGINNPHITTAPEPLCTDICSSLTWLDFDFGDLV HGYTYCCSVADLRAAIPNVPDLGDVDILDE

[0589] SEQ ID NO 3: B6ZLK3 MANMESKQGIMVLGFLIVLLFVSVNGQDCVWDKDTDFPEDPPLIFDSNLELIRPVLENGKRIV SVPSGSSLTLACSGSELINLGMEAVEAKCAGGVMLAIEGTEWEIWSLGCSNHVKETIRRNLG TCGEADQGDRHSIGFEYYGGSIYYELISVCFEPVSETTLYTEHVLHGANIAAKDIETSRPSFK TSTGFFSVSMSTVYSQASQLQLMTDILGDSDLANNIIDPSQQLYFAKGHMSPDADFVTVAEQ DATYYFINALPQWQAFNNGNWKYLEYATRDLAESHGSDLRVYSGGWSLLQLDDINGNPVDI LLGLSEGKEWPVPSLTWKWYEESSSKAAAIVGINNPHITTAPSPLCSDLCSSLTWIDFNLD DLAHGYTYCCAVDDLRQAIPYIPDLGNVGLLTN

[0590] SEQ ID NO 4: A0A3R7PEK4 MATAVLYLLCGVFVLAAGEDCVFDKNADFPKISPLIMDSGLKFVLPAKENDSVLIRIPTGTALT LACPGSSIKGQDTKESVSATCVGGNVLSVDGKEKKLSDLGCKKKPKSTILSRLDFCGADAA GVLASVGFQFGEDEFHEVISVCYEGNNETTLYSKHTIHGANIDAKDVDPKRPSFKTSKGFFN VSMKTCYEKKNQRQLLQDLLQDNSLAETYFFPKKQYYFAKGHLAPDADFVTEAEQDATYYY INAVPQWQAFNNGNWKYLEMATRTLAEYLGTDLEIYTGAWGVLELDNVKAKPVEVYLGLTV QERWPAPAVTWKVIHDKKNSRAAAWGVNNPHITSPPTPLCSDLCSSLPWIDFDVHDLGH GYTYCCTVEDLRASIPHVPDLGDVSLLDVSLGSSGGGC

[0591] SEQ ID NO 5: Q8I9M9

[0592] MANMESKQGIMVLGFLIVLLFVSVNGQDCVWDKDTDFPEDPPLIFDSNLELIRPVLENGKRIV SVPSGSSLTLACSGSELINLGMEAVEAKCAGGVMLAIEGTEWEIWSLGCSNHVKETIRRNLG TCGEADQGDRHSIGFEYYGGSIYYELISVCFGPVSETTLRTEHVLHGANIAAKDIETSRPSFK TSTGFFSVSMSTVYSQASQLQLMTDILGDSDLANNIIDPSQQLYFAKGHMSPDADFVTVAEQ DATYYFINALPQWQAFNNGNWKYLEYATRDLAESHGSDLRVYSGGWSLLQLDDINGNPVDI LLGLSEGKEWPVPSLTWKWYEESSSKAAAIVGINNPHITTAPSPLCSDLCSSLTWIDFNLD DLAHGYTYCCAVDDLRQAIPYIPDLGNVGLLTN

[0593] SEQ ID NO 6: Q9U5L8 MAGFGLQAFFIVTLLGVGVTGQECMWNKDTDFPEYPPIILDASLEIVRPVAEGEARWRVSA GAKLTLACPGSEIVNLGTTAVDVQCGGGNLLWDGTEWRMDELGCSKKDKESIHRNLGSC GDGGVGVFEGIGFEIFGSDSFYELIRVCFEPKAETTLYSEHVLHGANIAAKDIDSSRPSFKSS TGFFTVSMSTCYTQNSQLASMKILLGDDDLANAIINPHEQYYFAKGHLAPDADFVTEAEQDA TYYYINAVPQWQAFNNGNWKYLEFATRDLAESHSTDLTIYTGGWGVLTLDDINGNPVEIYLG LTEDEMWPAPAITWKWYEESSSRAVGWGVNNPHITSPPTPLCSDLCSSLAWIDFDVNDL GHGYTYCCTVDDLRAAIPHVPDLGSVGLLDK

[0594] SEQ ID NO 7: B8YN15

[0595] MSGFGFRDVLLVALLGLGVTGQECSWNKDTDFPEFPPIMLDASFEIVRPAMEGDLRWRVS AGAQLTLACPGSEIANLGSVAVDVTCGGGDLLWDGTEWRLHELGCTKKDKESIHRNLGSC GDGGAGVFEGIGFEIYGPESFYELIRVCFDPSAETTLYSEHWRGANIAAKDIDSSRPSFKSS SGFFSVSMSTCYTQNSQSALMKTLLGDEDLANTILNPGEQLYFAKGHLAPDADFVTEAEQD ATYYYINAVPQWQVFNNGNWKYLEFATRDLAEAHGTDLTIYTGGWAVLTLDDINANPVEIYL GLTENEMWPAPAVTWKWYEESSSRAAAWGVNNPHITTPPTPLCSDLCSSLSWIDFDVN

[0596] DLAHGYTYCCTVEDLRASVPHVPDLGNVGLLDK

[0597] SEQ ID NO 8: K9LU24

[0598] MELERGIFTHFVLITIFSACVYCSDCVWNKDKDFPQYPPLILDSSFEFVLPVEEDGNRWRVA

[0599] AGETVTLACPGDEIENLHQVEAEARCLENGLLAIENSEWDMASLGCGDDVKEVIFRDQGAC GAGNIGILHEIGFEILSAEKFKKVLRVCFDPSLETTLYTEHLIHGANIAAKDIDSSRPSFKSSTG FFSVSMSTVYSQSSQLELMIKLLGDEDLAHQIINTHKQFYFAKGHMSPDADFVLMANQDATY YYINALPQWQAFNNGNWKYLEFATRDLAEKKRRDLRVYSGGWGVLELIDINGNPVEVFLGL

[0600] AQEKKWPAPAVTWKWHDELTNCAAAIVGVNNPHLTSAPSTLCEDLCSSLPWIDFDVSDL GHGYTYCCSVADLRAAVPHVPELGEVCLLTE

[0601] SEQ ID NO 9: Q0GIJ2

[0602] MDLRRRFSRTLQLWLLFACASNCFGCEWDKDLDFPEHPPLIINNQLDFVLPVLEGVNRWR

[0603] VAEGETVTLACSGSELVNLGEAEVQARCLSSGLLTIGDAEWDLASLGCSSDVKETIFRDLGT CGAGGVGILNGIGFQIFSLNYDKVIINVCFEAASETTLFTDHILHGADIAAKDVEASRPSFKTST GFFSVSMNTVYSQNSQLQLMTSILGDEDPANTIIDPSKQLYFAKGHMSPDAGFVTIASQDAT YYFINALPQWQAFNNGNWKYLETNTRNLAMKKGRDLRVYSGGWDVLELDDINGNPVKVFL

[0604] GLTEGKEWPAPAITWKWHDESTNCAVAWGVNNPHLTAAPATLCEDLCSSLSWITFDVSS LASGYTYRCSVAELRASVPHVPDLGNVCLLTD

[0605] SEQ ID NO 10: A0A8J5NCI5

[0606] VLGDTIVTMAPLTTLWLGLFLLVSCEDCAWNKDKNYPTKNPPLILDASFKLVLPVLEGEDRM

[0607] VRLSAGSKVTLCPGSSLTKFKSEAVKATCIGAHLLAVDGKAFSLKKLRCRSPAKSTIRRELGR CGAGDLGHLQAIGFDIPTLSKFYEVIRVCYDNEADTTLYSEHIIHGANILAKDVDLKRPPFKAS SGFFNVSANTCYTKANQLKLMETLLGDSDLAKTMIDSRKQFYLARGHMAPDADFVTEAEQD ATYYYINAVPQWQAFNNGNWKFLEFATRRLAVESSSSLKVYSGGWGVLELADINNNAVDIY LGLVQKKALLPAPAITWKWYEESRQRAVAIVGVNNPHLKEPPTKLCPDLCNSLAWVDFDVN NLVQGFTYCCTVDDLRQAVPHVPDLGDVDLLTIDTTEGPGGC

[0608] SEQ ID NO 11 : Q0GIJ4

[0609] MAGRRQFFVLSLYFVAFWNLSKGQDCAWDKDADFPLTPPLLLDSSLKMIYPVLEGSLRMVR

[0610] VAAGSTITVACSGTTISCSRSRGCWREPVLAAQLITGDGTDNALNELGCVTPHLRACRRTW

[0611] VPVRCRSWYLHAVGFNIATTGSFHEFESICFDHAAETTLYTKHTLHGANIIAKDVDPSRPPFK PDTGFFTVEVNTVYTQASQLALMEQLLGDSALANQIINPDQELFMSRGHLSPDADHVLIAEQ DATYYFINVMPQWQAFNNGNWKYLEFAGRDLAVAHGTDLTVYDGGWGVLELDDINGNPVQ IYLGLSEGKEWPAPALMYKILHEESTNRAAAVIGINNPHITVAPTPICTDICSSLTWIDFDITDL

[0612] FRGFTYCCTVDDLRAAIPHVPDLGNVGLLDS

[0613] SEQ ID NO 12: A5X381

[0614] MLGFGFLDFLIITLLAFGVTGQECVWNKDADFPEYPPIILDASFEIVRPVMEGELRIVRVSAGA QLTLACPGSEITNLGTVAADLICGGGDLLWDGTEWRLDELGCTVKDKESIHRNLGSCGDG GVGVFEGIGFEIFGPESFYELIRVCFEPKAETTLYSEHWRGANIAAKDIDSSRPSFKSSTGFF TVSMSTIYTQNSQLALMKTLLGDDDLANTILNPQEQLYFARGHLAPDADFVTEAEQDATYYYI

[0615] NAVPQWQAFNNGNWKYLEFATRDLAEAHGTDLTIYTGGFDVLALDDINANPVEIYLGLTENE MWPAPAVTWKVIYEESSSRAVGWGVNNPHITSPPTPLCSDLCSSLSWIDFNVDDLGHGY TYCCTVDDLRAAIPHVPDLGNVGLLDK

[0616] SEQ ID NO 13: D3KYH5

[0617] MEPGRVLSTNFILVGLLSAYAYSSDCVWNRDSDFPVYSPLILDSSFDFVLPVEEDGNKIVRIA

[0618] AGATVTLACPGNEIASLHQVEAEARCLDNGLLAIDNSEWDLASLGCARPVKETIFRDLGTCG AEDIGTLHAIGFEIVSLGKYKEIIRVCFEPSSETTLFTEHVIHGANIAAKDIDTSRPSFRTSSGFF SISMIKAYSQSSQLVLMTNLLGDEDLALSVIDIHKQLYFAKGHMSPDADFVLMANQDASYYYI NALPQWQVFNNGNWRNLEYATRDLAEKKGRDLRVISGGWGILELNDINGNPVEIFLGLIDDK KWPAPAITWKWYDESTNCAAAWGVNNPFLTTAPRKLCEDLCSSLSWIDFDVGDLAHGYT

[0619] YCCSVKDLRASVPHVPDLGDVCLLTD

[0620] SEQ ID NO 14: A0A8J5NAX0 MNKKNKSADFTALFPQLSQSYLPQGVRVSGVRMGSHRCVSVIVINWILAAVTAQDCVWNK DENYPPNPPIILDDSHEIVRPVLEGEDRIVRLAAKSKVTLACPGTKISNLDVAIRVCFDPKAET TLYSENVIHGANIAAKDKDNSRPGFKSSTGFFTVSMSNVYTEKSQLNLMIDLLGDEDLAHKII DVHKQFYLARGHMAPDADFVTEAEQDATYYYINAVPQWQAFNNGNWKFLEFATRDLAVKH GSDLTVYSGGWDLLELDDINNNSVKIYLGLTEGKEWPAPAITWKWYEKSSKRAAAIVGVN NPHLEEPPTKLCPDLCDSLVWVNFDLNDLVHGFTYCCTVDDLRQAVPHVPDLGDVDLLTD

[0621] SEQ ID NO 15: Q7PVB8

[0622] SECTVNIRTQLNAREPLFLRNNQLWAPNGPSLQWNAGETTLIACPGNTIQNTGTVTANIQCV SGTTFNLAGSNVNIADVSCTARSTGSHQTTGQSCGSGGTLLNLGFDVPGVGFVTYIQSCYN MQTASVIYTRHIIPGAAISHSISESYRPSFKTAGTAPHVQPATSYTTAQQAIRFAQLLGSQAQA DRFITSSSYLSRGHLSPDADGIFRPWQWATYFYVNVAPQWQATNGGNWLWENAARNIAG RLNEDVLIFNGAHDILTLPHVNGQQVPITLEAGGIQTPKWYWKIIKSPRTNAAIALVNNNDPFR TSMPAGEMLCQDVCGQYGWGNANYGNFARGFTYCCTVADLRRAIPSIPAEADAANVLRF

[0623] SEQ ID NO 16: A0A182NNC4

[0624] CTINVKYDLNFPEPVFLKNNDLWIPSHGQLVWAPAESNMIACPKGSISTTLLDTALLTCVSGS TFLLNGLPTNISSIVCTKWTGDYQPTNEACENAGTIRNIGFDVSSVGFVKHFSVCYNSQTAS TVYTKHVINGDSIYDSIVESYRPSFKVIGVPTSVLVDTSYTIKSQQTRLTELLGSSTQAAKYVN SMTYLARGHMTPDADGIFRTWQWDTYFYVNVAPQWQKVNGGNWLIVEKAVRNVADRLQR DILVYTGTSGSLTLPHANGTEVPITLNANGITVPKWTWKIVKSASADAAIAIITNNDPYRTSIQY SDLLCNDICRLYGWYVPSFDVFSKGFTYCCTVSSLRSWHSIPAAATAANVLQF

[0625] SEQ ID NO 17: Q16XJ5

[0626] MVILRSDLPARDPVYLQIVNGQAELWAPNGPALIWNSTNEVTALLCSGSNNALMPTGMATN QKTCVDRSTFSVEGNLATSLELQCRNQITGDTINTGQSCGGRGTLLYLGFNSQVYGFIPYIH SCM DMQRASVLWTRH I LPG AAVAGSMVESSRPAWKVAG M PSHVRPSTSYTQASQLERLT TLLGSAEQASRFVFTNSFMARGHMSPDADGIYRSWQFTTYFFTNAVPKWQWNNGNWVR VERVTRDTAIALHEDLVIIQGTDGILTLPHEDGRQIPITLEDGGIEAPRWIWKIIKSPKLDAGIAF VTSNNPFETAIHPMDFICVDVCGETGWAQDEFANFARGYTYCCDPNELIANVPTASPEGHV GRVLLRTML

[0627] SEQ ID NO 18: T1 P9F7

[0628] MQVHAGGTLQMFCPGEFKVKATKLITATCVSGTTFKVDGTSYAFSELVCKSWPGFVAKKKG TTCNGGILVGVGFEVSSTRFVEQMEICYNEQEEVTRYVRHTLGPASNYYQTGVDRITFQTA GFFNGKNVDKLYTQATQLETINAELGGNAGKYFDSSKNIYLARGHMGAKADFMYGTQQRAT FLFINAAPQWQVFNAGNWARVEDGVRAWVSKNSKTVNCYTGVYGVTTLPNKNGVQTPLYL AHDSNNNGLIPVPKLYFRWIEPATKKGIVFVGVNNPHLTLEQIKKDYIICTDVSSKVNYISWK KDDITAGYSYACEVADFLKTVKHLPALTATGGLLV

[0629] SEQ ID NO 19: A0A182V706

[0630] MPQPLFLIPGTDQFKYPSTSSGILTLSAGETLELACQNGFSLFPSETSIWTCVLDDQFNYES KMYAFTEFGCTANWRSVARRTANRCYNDATIVEIGFEMGARFPKIMDVCHDEVTYDNHYLV HEFTPANAGFQTGVPRPGWIQGNFYPGVTVNTLYTVNMQRETIATILDSQPRADELVQTTN NGIYMARGHIAARADFIYATQQNATFWFLNAAPQWQNFNAGNWERIESSVKSFVAARNIRV RVYGGTYGVQTQADGNGDHREIFLDFDPNGRTRLRAPKVYYKILHNEAQNSGIVLIGVNNV HISLEEIQRDYIFCTDVSSRIGWINWDRENISRGYSYACEVNEFNRVTGHLPQLNVASLLI SEQ ID NO 20: D3TPG4

[0631] MWCPKQFKSIASELINVTCIGGSVFMVESKLYNFSDFTCNSWPSFTAQRTGETCNDGVLIRV

[0632] GFEISSKHFVEQMRVCFDEKQEVTRYVHHSLGPASNYFQTGIDRIPFQPGDFFDGKNVDNL

[0633] YTQVKQQETISNALGGDVGSKFFNISKNIYLARGHMAAKADFVFGTQQRATFLFINAAPQWQ VFNAGNWARVEDGLRMWVSKHRINVDCYTGVYGVTSLPDQNGYETSLYLAYDSNNNGLIP VPKIYFRWIEPSTKRGIVFIGVNNPHLTIEQITKDYIFCDDVSDKVTYVNWKKDDITLGYSYAC RVSEFLKNVPILPSLDASGGLLI

[0634] SEQ ID NO 21 : A0A7D3SU36

[0635] MDYGDTLTLSCEGTGNIKHPNEQRPLSVATISCEGGDIFKNDGWLSEPSRFFLFKCSVPPQY

[0636] KSRRTERLCYDGNPIIEVGYWDSHFYPLYESCFNEAALNVLYSKYTQKPYNALYQTKIDRP

[0637] YFVADDQFGYIPVDTLFSPRGQKAAVAKLVGSAIDNYITQTQFLSRGHLAAKTDFVFAFGER ATFHYVNCAPQWVGFNGGNWNTLEVDLRNHIHVTGYNTIIYTGTFGISHLQNAGGDKIDLYL FTDENNNPWPVPLFFYKWYEPQRKCGIAFIGINNPYVTTEEAHQMFFCEDLCRENRDFSW LSWHPDSSSEGYTFCCTIPDFRKTVPHLPTFDVDCVLT

[0638] SEQ ID NO 22: A0A023ERJ8

[0639] DQLWAPDGPELTWEEQESTTVACVKTKLVNVNSNTASLTCVSGQDFIVNGTPVNSTDLQCS

[0640] GRMTGEVEETGETCGSSGGTLLKLGFNVEEVGFLTYIESCYDREEASVIYTRHTIPGKAIEHS

[0641] IKESYRPSFKVAGTSSHVNPATSYTQQAQLNRLTELLGSEEQAKKFLQGGSYYLARGHLAP DADGIYRSWQWATYFYVNVAPQWQWNAGNWLWENLSRSKAAQLQEDLIVYDGVHDVLL LPHVDGQSIPITLEAGGIRAPKWYWKIIVSPATSAGVAFVTNNDPFRTSLPVNEFLCEDVCSQ YGWSDERFQDFTRGYTYCCWADLQTAIEDIPRDLQVNHVLQK

[0642] SEQ ID NO 23: A0A1 L8DXI7

[0643] HLLPTACEIRISGGGLGEPQPLILNSGATEFIEPKDANGIIRLNPGDDISLFCTTGFQSPSTTNN

[0644] IIRATCTTGTIFIIDGNKGNEMTFSNINCKSYPYHTARKSGKTCGNGAGVDIEVGFIVKERFIEL

[0645] FHICHDDIMESTMYVTHSMTPGNEGYQRSFPRPSWLSSGFFGGKNVDNIYTNVRQNARVA EILGSQELADKYIKPTTTSTYLARGHMAAKVDFIFGSQQRATFWLMNVAPQWQSFNAGNWE RVESSTRKMASQRNTHFDVYTGTYGVMTLPDINGQHQEIYLYFDENNNGQIPVPKIYYRVLY EKSTKRGIVLIGVNNIHITAEEAEEQNYIICEDVSDKINWINWERNDQILGYSYACEVEEFTKV

[0646] VEHFPKLEISGLYV

[0647] SEQ ID NO 24: A0A182KAQ0

[0648] CRVYLNGNLTQPHPPLFLKQSGSNQYELLQPTGPFFEWRQAETTWGCSPAKIILSNTESNR

[0649] ANLSCSKGQDFILNGGQKRVSFSELSCSSAVSAAIQPSNMSCGNGEGRIYNIGFEVKDKYFV NYFQVCYNLPKSSAIYSKHQILGKAISHAQINNNRPSFKLGGVKSTVRVASVYTQRHQLEHF STILGSSAQASKFINSSSYLAKGHLTPDGDAVLDSWAGATYFYINAAPEWQWNGGNWLRV ENSARKVATQLNDTVHVYSGVYDVLQLPDKKGVLVSLSLGDDGTVHVPKWLWKWVHQPS

[0650] SKGIVLITLNNPFAVKGETLCEDICSRYGWSQKEFQDIRKGFTYCCSVSEAQKAISLIPKSIKC SGVLASV

[0651] SEQ ID NO 25: A0A6V7LVB2

[0652] GPCILTMTYKNGDLKEPQPLILTRTRTNYEILYPQTKTGELTLEPTEKITLACPGKANFIQGIPY

[0653] LEETSEIEARCLRDKTFFIENDGLKFAFANVSCSRLPYDTAQRTEETCHGDEASHVEIGFTLE

[0654] KGFLKLLEVCRDDDSYSTFWTKFVLSKNIGAFHNPFPRPPWKTGKFYSGYKMDDIYKTKQQ LVSVEKLVGSQELASVYVDNQKGKFMAKGHLTAKADFVYGSTQMSTFWYLNAAPQWQTFN SGNWVALESDVRQFAERKSLDLIVYTGTHGHMSLSDVNGNQVPAYLYADDENQGIPVPRFF WKIIYDPVRQLGTAFIGLNDPYAEVISDDMFICEDIADRIKWLTWRPENILKGISYACSISDLSE

[0655] AIPSVPELEVRGLLY

[0656] SEQ ID NO 26: A0A182R4C6

[0657] MFAYVTILAAWALAHGQCTVNIRTQLNAREPLFLRNNQLWAPNGPALQWNAGETTLIACPG

[0658] NSIENTGTVTANIQCVSGTTFNLGGSNVNLADVSCAARSTGSLQNTGQGCGGSGTLLNLGF DVPSVGFVTYIQSCYNMQTASVIYTRHIIPGTAIDHSISESYRPSFKTVGTASHVQPATSYTTA QQAIRFAQLLGSQAQADRFITTSSYLSRGHLSPDADGIFRPWQWATYFYVNVAPQWQATNA GNWLVIENAARNIAGRLNEDVLIFNGAHDILTLPHVNGQQVPITLEAGGIQTPKWYWKVIKSP RTNAAIALINNNDPFRTSMPAGELLCQDVCAQYGWSNANYGNFARGYTYCCTVADLRRAIP

[0659] SIPAEADAANVLRF

[0660] SEQ ID NO 27: U4TRF2

[0661] MPMVLQPGELNFMYPENGDAVLKISANTTIVLACPGQNITFDSSTIYRNVTSASCVSGSNLN

[0662] VDSKTFAWENITCDGTLSRMAKVTDNKCENNGTEIEIGFQLTQDFFSRVFLICFDENAQRAL YSKVSISGAINKRNIVTNPGWVQTQDIFSIKNPNSYYNRITQRQTINTLLGLPGNSTMYIKDNS NYFLSRGHITAKADNFYPAQQQASFFLLNVAPQWQTCNANNWQTVEISVRDYAEAKRVDLL QWTGVYGLATLPHSKTGQLVQLYLYTQNNTKALPVPELYWKIAYEPIKQKGIVLIWNNPYLE

[0663] TYQRICEDIADKITWINWDRNNQIKGFAYACTVDSFRKWSYFPELTVQGVLL

[0664] SEQ ID NO 28: E9HHG0

[0665] MSKGEHSTCNPTKGCTFNLDTSSPSNPPHLIDPNNNIILPVMVGSSRVITLNTDDKVAVGCL GPGNYIRATGMQLAPATCTATSKLLLVDGQELSYSQLGCLSQNKEILLENGTCANGPGTIIRI GWQAGPEFIPLYDTCHDKAAAANYFSTHIINGKSAAADDTSNDRPSFSQGGYFPGLDVNMA YSQAGQTETITGIVGSASLAAQYWSGSQIFLSRGHMAPDGDFIDAGSQDATYYFLNTAPQ

[0666] WQNFNGANWNVLENAIRNLAISRDVKLTVYTGVFGICTLADVNGVQQPIYLAFDENNNGLIP VPKYYWKLIHDPVSNMATAVLGINNPHNPVGPADVICPDVCDQVPWVTWALTDTYKGYTFC CTAAELFKAIYYAPNLDLPLFV

[0667] SEQ ID NO 29: A0A821 UK65

[0668] MFSFGSREICNLSLKDDFSSPSPVYLHNGEFLAPNSATGDILLRRSEIVQIACPGHKRFWLG DEPTHLDVMPVKCVTGKTFRSDNGWIGELKEVTCNGPPWYSAQETRQYCYGKNKIYSAGY NISGNFHKLYDICFDKSLYTTLYSKHELTPASFYKQITSRPSFIEGDLFGKVKMSKLYKIDHQK KRLREILGDGMDEKYITKIQFLNRGHLSPKADFTLSAEQRASFHYANTAPQWMRGNAGDWA

[0669] SVEDAVRRRVHDLNTTVTLYTGAYGVMTLADSANKHREIYLSTDKNNNGIVPVPLYFFKLVY DSKKQTAIVFVLINSTYYTEARIDELSFCQDVCYENSKYKWLKWRNDGTHSFCCEYHEFVEN INVLPQLKVKGLFY

[0670] SEQ ID NO 30: A0A1 E1WTA0

[0671] ISSNNIQTATAQCVSGALVSGPGWLNGNGAFGGLTCDGHAFHIAEFTDQRCFNNFRVIRVG FWNNVFYPKYFSCFNPNRLEVLYTWATQNPSHAVHQTGVDRPSWLAGSFFPGVNINNLYT QNHQRNSIAQLVGSVELANRYVTSTQFLARGHLVAKSDMIFATGQRSTFYFINAAPQWQPF NAGNWNWLEQNLRARIGAAGYHTTVYTGTWGVTQLRNQAGRLVNIFLNNNQLPVPLYFYK

[0672] WYDASRRQGTAFVSINNPHYTLAEVRALTFCTDRCRNNNAFNWLRWQPDRIDIGYSFCCD VNEFRRTIPHLPAFTVNGLLA

[0673] SEQ ID NO 31 : A0A811 VDE6

[0674] MFEISRTFADPCTISIPTDLPDPQPVFVTQQGLFRPINQVTEVQEGEELTLHCAGKGNVWPL KQQTVTLVCRGGDFYNTETDEQQTLKDLKCTRIPTSELQATETSCADGAGVFYEVGFLVND NFHSVFTICYDSANEHTIYSRSLVNGAAQSFKINDSTRRAFKADGLRFSTTATNNFYVNKNQ KSRFASYFGAKQAFVNRTSFLARGHMAPDADFVFSYEQLATYYYANCAPEWQWNAGNWL

[0675] RVENAVRKLASQLGSDVLTYTSTLGVLELTNPTDNKETQIYLDKTELIPAPEWYYKIVMHPSL AADWFITRNNPFEDVGKEVEFCTNVCDKYDLDLSYYEDSRHGYTFCCELNDFWVAAMNTD SPNFDLPDGWSYKN

[0676] SEQ ID NO 32: A0A6P9AD62

[0677] MLAYWLGLAAISAVSADCSIDMNSDLAKHKAPLFLDGSDMLLPVPQGKSGVLSLSSGQEVT LACPGNKNQLSAAKSAWSATCDSGNKLKVNGKSVAAPNLGCSKTAGSSLRVTEQSCQDG GNLLELGFEVDGQWIKLVDLCHQIEAGNTLWALHTVHGAALAGAEVESKRPSFTRGDKALY KGYNPDHMYKQATHKKMLENSLGKAKADELIGNTFIARGHLAPDADFIFGSGQFLTYFYANV APQWQSINAGNWLATEKNVRKKAIELGRDLTVYTGTQGILTLPNASGEDTPLYLNADDKTLPI

[0678] PDNFWKVLYDPETKQGIALVGSNNPLLASEDNLLCKNVCEANGWSTIRDFRKGLIYCCTVAD

[0679] FQKAVSYAPKLSVSGVLQGPQ

[0680] SEQ ID NO 33: A0A6P8JYJ0

[0681] MKNTGCSIKIRSSELKDPQPLLIKSGTSEIVGFSDTGYVDVDKDKTIEFHCTSSLASPLSGKSV TAKCVGGTTFKIDDKEHDLSAIKCTSWPAFVGKKSGSSCNGGTTLIKYEVCFNEDEEVTRYV YHRLEPGNNYYATGVDRITFGAGGYFAGKNVDKLYTQAVQKETIDKELDMDSSRFFDSAKNI FLARGHMGAKADFVFAPEQRATFLFINAAPQWQTFNAGNWARVEDGVRAWVAKEKKHVE

[0682] CWTGVWGVTTLPNKNGEQRQLYLSHDNNGNGLIPVPKLYFRWIEPSTKKGIVLIGVNNPHL SLEEIKRDYILCTDVSDRINWISWKKTDITAGYSYACEVPDFRKKVTHLPEFSVSGLLV

[0683] SEQ ID NO 34: A0A1 Q3EV84

[0684] IFETVLQVIWGQCMVNIRNDLSAPEPVFFKGHTLMTPNGPSLYWKDGETSLISCDNGKLSIIN RNMVEVKCKSGKTFEILGKTYKSSALTCSRSLTGSVDVASTSCAGGSGKTRNIGYSDGTVF VTYIQSCYNVATASVIYTRHVLPGAAIDYAITDADRPSFKSEGTPTTVSPRTSYEQLQSKARL VALLGPVAGKRVNDADYLARGHMAPNSDGIFRSWRGATFFYVNAVPQWQSINNGNWKRV

[0685] EAAVQNKAASSKEDFVIFTGAHEILTLPNDKGVQVAITLENNLIQVPKWTWKWKSAKQNAGI AFVTSNNPFRTAISAGEILCPDICGTAGWGKLEFADFKKGYTYCCTVDSLMRKIPAIPAEAKV VNVLKF

[0686] SEQ ID NO 35: A0A118M193

[0687] MMKGIAIVALTCGLAIVANACSITLPDDVKGFAPVLLTRTSSNKDYELFKPSGKTTELSDGTKL LLACTGSKNVIVSEKKSTLELTCSGDQFVDANNNAYDLKDLVCKSIPSPTLRVTDEPCSQNR

[0688] GVIHQPGFTVNKKFYGPVFEICYNNDNEHTYYTHNTINGATIDNAISESTRRSFSAKGMKYTT TKTNEYYTQANQVELFKKLFGSKQTYIDGSDYFARGHLTPDADFIFGYEQLATYFYTNVAPE

[0689] YQLINAGNWLRVEELARSVAASYGDDLETYNGYLGQLQLEKSNGQKIDIFLDDGKKIEAPKY YFKVLLHKPSDEGIVFVTVNNPFAADGEAEEICENVCDQADLKHDNFPTLSKGYTFCCRLEE FKLSYDALPEDVEAGKLMTRKN

[0690] SEQ ID NO 36: A0A6G1 SN27

[0691] GNSLWDKTQQEAWECVEDTTFLFKGQQIEWSKLGCKSLAASSARKRGKCGNAGEFSRIE IGFEAADGFAVQIENCFNSVTFDAVYSRMNMTPAITGHESGGKRPQFDPAGFYEGFNPNTL

[0692] YTNKVQKKTVGELLGDPSLGDKYIKRSWYFLARGHLSAKADYILDAQQHATFYYINAAPQW QSFNGKNWENLEDSVRKWVINSGRSVQWTGIWGTATLPNKDGQETELYLGGSKKNLRVP

[0693] KYYWKWYDPATKEGAAFVGMNNPYHVTTEEDVFCKDECARYSWISWSQKDQDKGYSFC CDVNEFKKTVTILPNDIDVQTLL

[0694] SEQ ID NO 37: U5ER46

[0695] SQPGKKQCAVYIKRDLPELQPVFIKDGTPRSLMMPRDGYLTWGHGETSSIACPPSAGSQNY IRATNTIVATITCDNGLMFEMSARRVNISTVTCARKVTGNYRPKPDPQCAGTNKAIGFNVPFP TGDIFFDLFYSCFDETRGSTLFTHHVLFGNEIDHKCIYGSSRPDFKSAGFPDNFFISTAYTQL

[0696] SQKARLTDLFNVRMSNPVAQAEAQRYIFDHSYLQKGHLTPDGDELFTTWQWSTYFFINVAG MWESINIGNWKNLETKVRTLANETKENLEIYTGTYETLALCSTNNYCPDFTLVTGGIPVPKWL WKIVKAPSIDAAIALWSNNPVGAVNPICETGAVNYGWHQNGFDNFSSGKVSYCTVQELQN WGYIPQAALATNVLAFHHHHHH

[0697] SEQ ID NO 38: Q7QK22

[0698] FFLKSSQNIKIAALPAGGCTVPFADLPYPEQPLILIPGTEKYWYPLDETREILVPTGAPIELACQ QGFRLFPTQRSITIQCVTNSTFSFGGSPYPMKSLACTSYWLSSAKTTQARCHNESVIVKVGF ELADARWVNVF DVCYD EQ LYHTH FVRH YM N RANG G YQSG N PRPG WYQG AYYTG VN I NTL YTVNRQRETIATILNSQARADVLVQNTTNGIYMARGHIAARADFVYGTEQNATFWFLNAAPQ WQNFNGVNWERVESSVRDFVGKRDLELTVYSGTYGVQKLADGNGDYREIWLDFDPVEGR RRAPAPMLYYKILHDEASNAGIALVGVNNVHVPVELILREYVLCKDIGDEVEWIDWERKNLTI

[0699] GYCYACEVNAFNDAIGRPHPQLNVAKLLTS

[0700] SEQ ID NO 39: A0A1 L8EBY2

[0701] MTKLLLALTWSCALAIWNACSITLPQDVKGFAPVILTRSSAKGEYQLFKPTGKTTDIPEGTQL

[0702] LLACTGKKNFIESLNKETLELKCSSNKFKDNNNNAYNLKDLVCKSIPSSTLKVTSTLCSKNRG

[0703] KVFEIGFELNKKFYGPVFDICYDNNKETTYYTHNIINGAAINYNIQESTRRPFSAEGMKLKTTK

[0704] TYNLYTKKNQVERFSKLIAPSQKYFDDASFLARGHLTPDADFVFGYEQLSTYYYANVAPEFQ

[0705] NINAGNWLRVEELARKVAGSYGKDIESYNGYIGLVQLPKPNGKLVDIYLDDEKKIEAPKYYFK

[0706] VLLHKPSDAGIVYVTVNDPFIKKGPSHEFCKNVCDSTDLVHPKFPELTKGYTFCCKVDEFKQ KFDGLPSDVKASKLLKSK

[0707] SEQ ID NO 40: A0A7I0ZZH0

[0708] MRCVCVFLFLCFINIVNGGCILSLKKDFPKTSWYLRHGRLLSPDAVTGDVQLTRSETLQVAC

[0709] PGEQKYIVLSNKTTNHSLLDVKCVSDSLFRASRLSWIGNFSEIRCNAPPWTSAEEVGGCGA

[0710] GAKNYRIGYKVSGHFHGLYEACFNKDLLSTLYVKQELGPESVFIQSGGRPNFVEQDFFGKV

[0711] KMSKLYHLTNQKARFKQVLGEGKEEEYLTKKQYLTRGHLSPRADHSLLCSQRASFLYLNTA PQWRRGNAGDWAALEEALRRRVHSYGRPVTVYTGTFGVSTLGDTSSRQQQLYLSVDQNN NGILPVPLYYYKWFDAANNTAAAFVSINSSYYNQTMIEKLTFCEDICGSRNYSWLRWRSSD

[0712] GTHSFCCDYHDFVKTVHDLPGLKVEGLFY

[0713] SEQ ID NO 41 : A0A182GXS5

[0714] MPVNGNLIWAPGQRSWIAFPPSDGTLNSIVGKNLVSALIECVQGQTFRYVHDRTPLSITQVL

[0715] CVRKVTGNVRTSPIPMNQCPGTFKAVGFDVPFPNSQNRFFDLFHVCFDERAAIPIFTRHTVY

[0716] GNEIQHRSTSAAQTSFRVAGFPRQLDVSSAYTKRYQVQRLADLFASDPNPRGSVDMYYAV NSLHRGHLTPRADELLTTWQLSTFFYVNWGMWETINNGNWKYLEESVRTLVDNSKKNLIIY TGIYGTLSLCSDKNHCRDFTLSNGRIPVPKWLWKWKSPDSNAAIALWSNNPFDRESPICG

[0717] LGGESFGWSRKISLNGALGAVSYCSVQDLQRWRYIPQEALAGNILRL

[0718] SEQ ID NO 42: A0A6J0BR40

[0719] MIWLATLLAIAIASTNALPTRSGGCSVSLNGDLNDPQPLILVPKGQEASYVLPVDTSGTLVFA

[0720] ENEELRLACVGENNYLVGVGDEDIQDIEAYCVSDKTFLVNSVEYEFGDLVCSFVPSSTVRRT

[0721] GEACLDKYSQVEIGFDLGDEFLRTIEICRDDDIYFTYWTKFNLTKAIGGYQRSYPRPSWSQG

[0722] DFWGDYNVNTQFTRSVQRATVSLIVNSTELGAAYISSSTNYYFARGHITAKADFVYGSAHW STFWYVNCHPQWQTFNGANWMYLESDVRSYASNSQVDLDVYTGVHGITTLRDVNDVEQPI YFYASGDERALPVPKFFWKIIYDPLTQKGTAFVGVNNPYITEVTSDYYICDDISSKITWLTWTA

[0723] DDIEKGISYACTIDDLREAVPTLPELTWDILT

[0724] SEQ ID NO 43: A0A219YR85

[0725] VGAVCKVSLNSNLPTKEPLYLIHKGSKLDFVYPKSAKGASKRDKGSFELSENEELVFACPGK

[0726] ANKLARTTENAADSHCVSGTRFSLGAKKTSNIEQIECTSSVKATLELDSGKKCAKTGTQMKI

[0727] GFEVERQLLPLMWVCHDIKAADTIFVEHDIPATIGGARIFKARPDFEDGPSHLYEGINVKNVY

[0728] TQKYQRGLFNQLLGKGQGEKFIKGTYYLARGHLAPDGDFLYGSWQWSTYFYVNTAPQWQI INAGHWLALERYLRKFAEQTGEDLHIMTGIVGVLSFESDSGENVEIYLQPDEQKIRVPAAFFK VIRSEVSDRAIVAVCSNNPFEKVPTLCQDIAAEHRWPTSWHDHAKGHIYFCEVNDFLSNASG

[0729] SEQ ID NO 44: A0A195CD05

[0730] MLLVIPFLACYIVWAIANPINIGISCTISVQYKIGDLKEPQPLLLKRVGKHATIWYPDNDNGTLKI

[0731] SAGDSIYLACPGKNNNFKNRSWGNEVEAVCVRDKVFDVKGVHQYISSLVCKSHPEHHAEY

[0732] TGTSPCLGRHSPIEIGFNVSGTFIRTIELCRDEKTYMTYYTKFKMTKMIESYQRSYPRPSKFL

[0733] AGKFYPGINVDYLYKFETQSDTLARILKSTKLTEERLKKSLQFLSRGHLVAKADFVYGIQQRS TFWYLNTAPQWRTFNAGNWNSLEVSVRRFAANRRLDLDVYTGVHGQMTMEDIHNKQQPV YLHAEGAVMSVPKFYWKVIYDPLSKRGTAFVGLNDPFITLVTDDVYLCSDISERIKWLDWRP

[0734] RNITLGISYACSVADLRKAVPWPLLDIIDILM SEQ ID NO 45: A0A164RQ45

[0735] MVSLPWLILIAFAFQTNLLLVKSLPQFKAGCTFNLDSSSPSNPPHLIDANNNIILPEMVGSSRV ITLNTDDKVAVGCLGTGNYIQTTGMQLAPATCLSTSKLLLVDGRELSYSQLGCLKQNKEILLE NGTCANGLGTVIRIGWQAGPEFIPLYDTCHDKVAAANYFSTSFIHGKSAAADDSSNERPSFS QGGYFPGIDVNNAYSQSQQTQTIADIVGSTSLAGQYWAGSQIFLSRGHMAPDGDFIDAAS QDATYYFLNTAPQWQNFNGANWNVLENALRNLAISRDVKLTIYTGTFDILTLADINGVQQKIY LEFDENNNGLIPVPKYYWKLIHDPSTNTATAVLGINNPHIDQVLPADVICPDVCNQVPWVTW DLTDIYKGYTFCCTSAELFKAINFAPDLGDLPLFV

[0736] SEQ ID NO 46: A0A6J0CA36

[0737] MTFNISMQAVACLLIFATSIIAIQAQCSVSINGDLPAPQPLVLNVGEGAGFRHPVDTSGTLSFA AGENVRLVCTGTNNKFVNLDSTESDVLATCVSGSKFTVDGVSIEFSDLACSFLPWHWQRKS GNTCLGSYTDVEIGFDLGTEWITTINACHDETRLETLYTTFKLVKEIGGFQTGFPRPSWVQGI FYPGYTLATQYNRATQIATMGTLLGSADLGTQYIHATNDYFLARGHTTAKADFVYGSQHRAT FYYVNQSPQWQTFNGGNWNSLENDVRNYASNTYQDLTVYTGVYGIAQLPNVDGVLTDLYL YADEAGNRAIPVPKFFWKILYNETTQLGTAFIGVNNPYLTTVTSDYYICSDISSQIQWLTWSA NSVPLG ISYACTVDSLRN WETI PSFTVKG I LS

[0738] SEQ ID NO 47: A0A162PGR9

[0739] MRFFGLVLLILAQLSAGLGYSLQRANKAGCAFNLDISSPTFPPHLIDSNKKIVLPVIEGTARIIN LSEGQQVTVSCLGTGNSLLATGLQLNPATCMASSNLQLSDGTEFSYGQLGCLKQNEETLM EQGTCANGPGTFIRNGWEFGADFIPLFDMCHDEVLALNYYSIDTVYGRSANADDKTNERPS FSQDVYYPGLSVNTLYTQAEQTNTIAMIVGSQELAAQYIRPNTEYYLSRGHMAPDGDFIDAA SQDASYYFMNTAPQFQTFNNGNWKYLEIGVRDVALARQLDLTIYTGTFSVMTLADINGVQQ PIYLAFDANNNGLLPAPKYFWKLIHEPISNTATAVIGINSPYLDPVMPEDIICPDVCDQIPWVN DAIFQLTNIARGYTFCCTAAELHNAVSFAPNMDVPLFI

[0740] SEQ ID NO 48: A0A5N4AZD5

[0741] MFSLVLVLCLLAQSHAAVIPKATGCTIVIGKHLGNPQPVLALQNDATTASEAIILPHDKQGTIAF NKGDSIDIACPGAQVQLYGTSMKEDLVAATCNSGTNFKVNGKLIPFLNITCSRSPKSSARFT GKRCAKNMKEIQIGFQVKDRFIDHILTCYDEKSQNALYAVSYLSHNVAGSQVNFPRPSFQAA SFYKTDQKVEYLYSRKGQASTVNEILGLNANDSSIIHPTSKIFLAVGHFAPKADFVFGSQQRL TFYFVNAAPQWQNINGGNWNTMEKNARKLSIERKLDLTVYTGTHGVMTLPDVDGVEQKLYL WTGKKGQRGIPVPAFFYKVIYEPKSQLGIALVSRNDPHTKISKSDYICRDVCDKVKWLTWKQ DNMDLGLVYCCEVDDFRKAVNVLPAFKVKSLLI

[0742] SEQ ID NO 49: A0A6B9KZ70

[0743] MSGFVQLLGFLLFLRSCVQASPAPGCKLYTNKDLPKKNEPLLLIRDEGKMYRLAMPEAKGK EGFIQLRSGQEIVMACPGRRNFMNETKQELNHARCITGHLLTIKEREFESKYLDCFERAGANI RRTEKKCYKGKGTVLEIGFKSEGWHPLIIVCHDIAGSHTFYSKHTLYGSVLSGKVYRTTGRP GFSRGEKFLFKEFNPERTYSRKHQRYILNNLIGEEKANEYLDDKTKYMARGHLAPDADFLFS SWQLLTYFYINVAPQWQSINGGNWLHVEANSRRIATKLKADLEVATGTNGISKMKDSEGKL KEIYLESNSKVPVPEYYWKLLRNPQDNSCMGFISTNNPYLEEAPKHKCKDVCSQHGWQVM QKDLFKGYVYCCDYDNMKKAIPEMPKFECDKPLNF

[0744] SEQ ID NO 50: A0A1 B6FIU7

[0745] TMLVITLTLWWIFQTAQCKTGSCNLSLQRDLSQNEPLVLTVNKDNLEWVMPEVIGNQGVIS LETGKHLVIACPGSKNNVKANGEETAYIKCDRGSFKIGSKRVAEGGLRCTHSIADSEIWVSQ LRCGSGIHKGTMIQLGYQVKEEWLPLVEVCHNISRGVTFYTYHPLPGHSIEGAVKSNLRGNF KIGPSELFPGISPNTLYTQKRQKEVFKKILGSNSYINGTSFLAKGHLSPDADFIFNSGQLLTYY YVNAAPQWQKFNAGNWISVENAVRRLSAQVDDKLDVYTGTYGVLKLKNNKGLERMIYLEPT KKLVPVPEVFWKLVRHEGSNSCTVWGYNNVFLTRPPSKVCAPVDPTGWPKLTDLRKGYIY YCDYHSFSETVTYVPQLKCKGLLEFPTL SEQ ID NO 51 : A0A7E5VB98

[0746] MTWFGHVLAFSLIFNLICFIKSECVLKINEDLAKIGPVYIRDNDYMDPNNADGTITFDKADTVT

[0747] VACPGTNRWWLGKDTTSSDVLEAACVTGNTFRVDGKVLPFEDISCNSQPYFTAEETREKC

[0748] HGQGIKYRVGYKVRNTFYELYEACFDKYRLHTHYVEHPLTPMSRFMQTGLKRPAFIEGNLF

[0749] GKVKMNQLYKMTHQKTQLDAILGPGMGEEYITHKQFLTRGHLAARADYTTSAETRGTFHYV

[0750] NTAPQWMRGNAGDWAALEEALRRRVQSRGSDVLWTGTHGVMTLPDSKGRMQEVYLSE

[0751] DVNNNLIVPVPMYFYKLVYDTKDKTAAAFISINSSFYNATTINRLAFCEDTCDGNPQYSWLK

[0752] WRSNDGTFSFCCDYHDFIQEIDYLPKREVRGRFF

[0753] SEQ ID NO 52: U5ET74

[0754] MKMLNFLIIIKIIYSIIFIEWVHCQCRINIQTDIAKNDPIFMTNKYQLMIPNGNQLTWKKNEITYLG

[0755] CSEQRTNTFQTPKLIESQTTTSIKCVSDALFEINGQIFNSTDFECKIGITGDLNQTNKACAQGK

[0756] GAIFEIGHDFRKSFNVFIVLIEVCYNMNTASVIYTQNVINGNASSVIESNRPTFKVNGVPSSVS

[0757] PANSYMQTNQKLRFTTLLDNSSLADFYLNRTSYLARGHLSPDADMIYRSSQFSTYFYLNVAP

[0758] QWQNVNAGNWYFIESAVRNKATALNEDLLIFNCVYDILKLPHNSGTLIPITLEDAGIEVPNWFI KIVLSNKLKASIAFVTSNNPFMERFTKSDLLCNDICQPTKWNTVLKQSADITKGYTICCNVNEL RGKINFIGSEWITPNILNY

[0759] SEQ ID NO 53: D6WBD2

[0760] MKLFLVGCVFIAYFVHVLVANQKGCSFSVHGQSEKNSPILLQNDTTLIVPNNGKVSLKRRDT

[0761] VTLLCPDSKNYLVGTQSNITQAQCVQGGILRVANQDLTFRDLQCKKIIRGTVAKTTKKCGEN

[0762] KGRIYKIGYQISSRNFLTLIEVCYDPNSGTTLYTEHALHGQDIKYASKSNYRPAFSPEASAVA

[0763] ASVAYKQTFQKSTFNKLMKSALKAQEYINENSFLSRGHLSPDADFLYAATQYTSYYYINAAP

[0764] QWQTINAGNWKKIELLVRKLADNLQETLTVITGTYGVLTLPDVNDNEVDVYLVSGSKLPVPK FFWKIIYAKHSRQAWLVSLNNPFVKEIGKGDFLCSNVCSKVGWGSGSWSNYERGFVYCCD YKQFVDKVETAPKLSWGVLQGPK

[0765] SEQ ID NO 54: A0A6P7F7E8

[0766] MVRYRCVFSWLILGACSQFGYSASRAGCTFNVYNNANEKYMPVMLTNHSSKYELIVPEQG

[0767] QIHLRSGEGITFVCPEKNYLVLTNSNFTYATCIQGTNLKLFRNTYNFDSFLCSKSIRGQVQKT

[0768] KEKCGNRNGEIINIGYKVTKRYFKPLISVCYDEKNGKALYSQHVIHGEEVAYTSRFKERPNFS

[0769] TDGLAKDVSANLAYKRAYQKSTFSSLLGSSREAEQYINTKSYLSRGHLSPDADFLFASSQLT

[0770] SYFYINTCPQWQSINGGNWVRVESAVRRVADKYHTTFLIITGTHDVLDLPDVHDNPIEMYLV SKRKLPVPKFVWKIMLDENSGKAIAFVSLNNPFVQEITEDEQLCSDICEQYGWGSKYYSDFS KGYIYCCDVNELRETVDTIPVLNVNGILQA

[0771] SEQ ID NO 55: A0A3G5BII3

[0772] DLALILLSVSFEYSNAKCSLDIQKDLPPSDPVFVDKNHNLVKPEGKFIEWNIGDKFGAICSGK

[0773] NNTFISTNQRAAIFECRHDGLHSSGQNIPKLECSKDIVAETLDTNKACIGGGTLKHVGFRTYS

[0774] SNFTRYLEICYNMNEASVIYSHHIIPGKSINYRVQKALRPFFNVAGAPSKAQVSKSYLRKQQT

[0775] KTFTKLLGSEQYMRYFNNSSFLSRGHLSPDADGVFKSWQLATYFYINVAPQWVTTNAGNW

[0776] LAVEKLARRKAARINDDLEIYTGVSGIMSLPDSSGRKVKLYLKNGPTGKYIPTPMWYWKVLK

[0777] YKDEGIAFLSMNNPFANQKLYICPDICQSYNWPNKGFENFRKGFTYCCKVSDLIKRLPEVPP EVLVTRVMKY

[0778] SEQ ID NO 56: B4HZ06

[0779] MKMHELRYLIITLSLFLTVGNAKSFCQITQAETWTDRLFVHIVNNRYELLLTDQLQPNQQISLL

[0780] CDENAQVFTSTCGNNGHFSPPLPRTNCSKAIPPSWPTASNICPHTMYLVGFRYGNTFMEL

[0781] YRSCYDARTMKAYFSINTVYPNNLRSDRPPTVFDKDGIITPADEATFQLNSIYNRFEHLFGSG

[0782] QTYVPSSRSLSFDRGHLTPVADYSFPKILRQTNKYLNWPQYYSINRSNWKIVENWVRGQN DVLNVCTGALGVLQLLNRNQQSVSIYLAPNKNPVPRWTYKIISSLTTGIKYVILTSNNGWETQ QPNPSSVCKATACPRTLNPTGSGYTFCCDPNDFIRRNVPNLTGVC SEQ ID NO 57: A0A6I8V2P4

[0783] MIYVRCLCLSLSLFVLLGSARPDCVLSKHDLSQTNRVFATRDGLSYHLFRKDWPNSAQVHS ICNYNDIVISTCTANGFHPVLKTAGCKKPTTPIVEEVNDASCPYKMYRVGYRFENVAFLEIYR SCYNPDKVQAHFTIYMAQTSTSAVERPMHFNTDRIISGAAAASFQNTQIYSRYNALLGRQTY FAAGTNMFDRGHLTPSLDFAFKASRGQTNKYINLIPQFKTINRGNWKTIENWVRRQLSERHF DALKVCTGVLGVLELYSSTYRADIPMFLINNNKNPIPKWIYKIVSHISGRKFVFLTYNNAHATT RPTGQAVSNVCRELSCRSLGLKLRNEGTAGYTFCCEPHDFIARNLTRLTGIC

[0784] SEQ ID NO 58: A0A482W3B6

[0785] MKLFLAGCVFIGYFVHVLLANQKGCSFSVHGKIEQKSPILLQNDSSLILPSKGKIHLKNREIITL LCPDSNNYLIGIQSNATYAQCIRGKTLRVGNRDVNFHDLQCKHTIRGMVSKTTKKYGNNQG RIYRIGYQLSSRHFITLIAVCYDSNIDSVLYTEHFLHGQDVRHASKSSYRPSFRTEAAAVAVS VAYKQAFQKSTFNKLLRSSRLAQVYINDNSFLVRGQLSPDADFLFAATQYISYYYINTVPQW QAISEGNWRRIEAMVRKVADYLQETLIIITGTHQVLTLPDYNGDEREIYLIADRKLPVPKFIWK AVYSKKSKKAWLVSLNNPFTPGIGKNDFLCESICDQVGWGSLSWADYERGFVHCCDYGEF AAEVGTAPKLAWGVLHGPK

[0786] SEQ ID NO 59: A0A5C1ZYK9

[0787] MNQNICGGIVITVLTSLIMKLWSVIVFALFKLARAQNECTFNVTNPEYEQTLPVLLYNLSSTY ELAIPNQGTIVLTSGQNISLLCSGNRNYVRQTNSNVTTVTCLENQDVKLYRKVYKFHDIKCKH AVRGDVRKTEEKCGNGQGWYEIGYPVSRRDWVTLITVCHVPEIGHTLYTRHLLHGKEIKYA SKSSYRPGFSTAGLDSSVPASVAYKQTFQKATFSSILKSSELSNHFVTNKSYLSRGHLSPDA DFLQAPTQFSTYYYLNTAPQWQNINAASWNSIEFTVRKLAELYGDLEVITGTHDVLTYKDKE GHVQEICLGSEGKLPVPKYFWKLAYGEKSNRAIVFIILNDPYVQSLENEVFCNDICQKEGFTK KSWKDLHSGYVFCCSYPDFLRWKTAPRLTVDDVLSNPQ

[0788] SEQ ID NO 60: A0A0K8TPS5

[0789] FLIWALQSWYVCSTSGQCTLDIRTDLPSNDPVFLNSKGELWIPDGPSLTWKKGESTLILCPG GRNKFNNTADWKTISCDGKNKLQVDGKAYNKNTTTCSNAVSGDVQVTKTACAKRGSMVDI GFDGKKQGFKKSFDVCFDFDKGVPIYTRHTIRGIATPFAASDAGRPPFKPTGIGKHVDPTTS YTQKYQVDRFTKILGDAKLAQKYLNATSYLARGHLTPDHDGIFKTHQSATYFYMNVAPEWQ GINGGNWVRLENAARAKATSLKDDLEVYTGIYGVAQLKGKGKKAVDLYLEDGSKIPVPKFY WKILKYKDSGIGFATINNPFEENIKPICKDICAENGWSAPTFQTATKGYTICCTVADLIKAIPDV PQDIKVKKVLNFK

[0790] SEQ ID NO 61 : A0A0A1WKI0

[0791] MSKLWQLLPFYLVLSSVYANPCTIRVPEEVPEPTPIILKDNELFRPNNNETKIEDGKSITLFCA GRGNKIVILNQAKVTLTCKDGKFRHEKKEYNLKEMKCNDVPDSELRETRKACANGAGEFYE VGFTYENKFLEILTICFNKATQRTIYSRNIIKRNVEKVRMRRACNTPPFKNIGMNFQPNEFYKV ENQTMHFENLFGRNQKFLVSNITRFLARGHLAPNADFTFCYEQFATFYFANCAPEWQQVNA GNWAKVEKATRQLAINTDILTFTGTFDILKLRMPNTPMFTEIYLDHAKQLAAPKWYYKWHNP NWKIKIVFVTLNNPYNTVGEEVEFCKNICREYGLDSRFYDNVNRGYTFCCGIRRFLEK

[0792] SEQ ID NO 62: A0A0K8WFE1

[0793] MPNLWKFLSIFLVFSSVYAKHCTIVMPRDVPQPTPIILTNTGLFRPTSTVTTIVEHDKITLLCTG RDNTVLALNEEIVTLVCENGNFFHNNKTYALQNMKCKSVPTTQLWRKGTTCATGDGVFYEV GIDYKNRHAIFEICFNQRDQRTIYSRHLINGYIQNGRPKYDCRPSFFRSEGMSRNLNDLYTKS TQQKHFETLFGVNQTFNNNASFLARGHIAPFADLIFCYEQFATFYYANVAPEWQIVNAGNW VRVENAVRKIASSQRSDLLVFTGTLGVLELRNPLTSRDTSIYLDEHQTIAVPKWFYKWMHP SFAIDIVFITLNNPFVRNAVGVEFCNNICRQICKKHSLDCTKFTDTKKGYTFCCELKNFWANA MGVGTPYYELPVGWSYKN SEQ ID NO 63: T1 PB16

[0794] MNRNWIFLSNNDLLRTDYIADGGTLTLYCNEHSAAWDLKCTKGTLEQLPYDAVCSERLSLSV VPQYKACSADGANGQLYDMVYSFATGFTVTLYTICYCTSRETVLYSIHKAYGFNLPTGNYQR PQFKLLGNPNRARADSFGSDEIYKTFQTLLGDGQTYVKSNRDYAVQRGHMANSQDFLTYD QMDATFLYMNWPMFRGCNLRNWKRIENWIHRLPDKHTYATWTGTHDVLYLQHSQTHRF VPMYLMPDEKNPMPMWMYKWRYQGECYVFATLNDVSSVPAIQYSNICRITPCPAGLTLDK EPGSCVSYCCSYRHFVQQVGDFANLCSVYSIPPFFIAFL

[0795] SEQ ID NO 64: A0A1 J1 HMG0

[0796] MKTFLLFILVILSISRIIRCQCTIDIKTELNRTATGTVKEPLFLKYTGSEYRLMLPDSTGHLNFRQ GENAMVACTSDQKPNTLTFNNKTSSTIFCLSGDSYNINNMRYKVSDFNECKSLITGNTINLNK KCADQGSMITIGFQLTEGRGFITLIEICYNKKRSSSIYAKHIVQGKIIKNKMITNTRPSLFKKTEV PTKVNPQKSFTKSNQLKRFQQIFGDLETAQEFINKTYLARGHLAPDADFIFASWQSSTYYYIN TVPQWQSINNGNWKHIESAIRSKANQLKADLEIYTGGFDVLKLKNKKISMENDGLDVPKWS WKIVKLPFENLGIAFLTLNNPFALSSPNALCQDVCNESGWNWKERKSISKGWTICCRVSDLM RVILDVPEEAWSENILRYNN

[0797] SEQ ID NO 65: A0A6P8XXQ3

[0798] MYKEGILLLLLVLLESSFHAHGECRLNVNYIQNSLGTLTYYDGHGKMQIQRFASFDEGHKLE LHCLHQRVISVKNLGCHRGIVNPKDPEKCDAPIHANVEVTDADKTCPATMYRIGIVIRNQFFE LYRACYDKVKVQSHFSDAVIYWKPIHPKSPRPFFDADNLITQRQVASFTAQNMYGAFKTIYG NTQTYIKADRSGRAESVIDRGHLTAAGDMTFYDQMYSTFKYLNWPQFKSINNGNWLQIEH WVGNHIPHKNVLHVRTGALGILSLPDFKPRRTLQPAFLIPDKRQNPVPEWMYKIVRSMNNQL LHVFLTYNNIFNPTKPTAHKCCKWSCPLRLEDSATLGFTYCCNPDAFVKCLASH

[0799] SEQ ID NO 66: A0A834KKD8

[0800] MLWIVSLFIHILLVDARSIENDIENDKTTDTEFARNNCVLSMKYKYGDLKEPQPLILTRNGTAS AILYPNINGTLKVQTGQSIYLACPGEQNFLRNTNYAQEVKATCVRNKIFHVNGENRNFSSLTC YHLPEHTARHMKSPCYNRNTHVEIGFNLSTNFLRVIELCHNEKTYNTYYTKFRMTKLIGGYQ RSYPRPSKWMVGDYFGKIDINRQYNFNTQLSTFQKLLNSSELAKSHLAKSRQFLSRGHLTA MVDFVYGALQSLTFWYINAAPQWMSFNAGNWERLENSIRNFSTYRSLDLDVYTGVHGQMT LPNARGKQQDIYLYVNGTNRAVPVPKFYWKIIYDPRSKKGTAFVGLNDPFIKSITEDIYICSDIS SKIKWLLWRPNDIKAGISYACTVDDLRKAVPTIPKFQTVGILT

[0801] SEQ ID NO 67: A0A6P4IS71

[0802] MNESRFLWLSLGLILFLGKARASCLIPKDTMATNHAFLSFENNVWDILHSDTVPVSRSIYLLC GGRAAPQRFDCDVYGSFTPALPTSNCAKPFKPTVREDSNDPTCKTPPLAMYWGYEYGGR FLEIYRSCYHKTQYAAQFSIHKVYRSFDSASRHPLGFTTDRAMRVKEAAAFKNENIYNCFEA DLGQGQSYMAQGSCLFNRGHLTAVADFPFEQLQRSTYKLRNWPQYAAVNNGNWKAVES WVRRLLDRKNYDWKVCTGALEVQQLQNTQSHQMTPIYLLNNKIPVPKWTYKIVSHLSGAK YWITYNDVYATRRPDPRSFCKIIGCDSDLSKDGVGFTFCCKPREFITDNLVHLTGVC

[0803] SEQ ID NO 68: A0A6J0BKG4

[0804] MINAKSTVNGVAVIAALLSVLAEGSPIGKSDEGCSILINGGLAEPQPLILRPGFEGGYVIPEET GSGSVTLAVGESLRLVCLGNSFDLETSEPVEDANVTCVGGTTFTFEGLGITEEFENIGCLSY PTHTARRTGVACPNGEVCEIGFDLGDGDFQRLITLCHDDVDQNTILAHAKVPAVIDAAQTSF PRPGFVKGDFYVGVSMANIYTRVNQRATLAGIIGSEELALEYLPASGTYYWAKGHLVAKTDL IYGAHQRSTFYYLNTVPMWQNINAGNWGIVEANLKNLATNRDVDLDVWTGSTGVLTLADVN GDQQEIHLYVDENNNRAVPVPKLLFKWWEESSGLGVAFVTVNNPYLEELDDEHVICTDVC DQVSYLTWNPTRSDRGLSYCCEVDDFRKAFPDIPEFTTTGILA

[0805] SEQ ID NO 69: A0A1 B6F9M6

[0806] KEPLVLAVKTTQIRGQRKVEQLEIVRPEVENRNGVITIDETQSLIVACPGPRNNNRGTGEQSA

[0807] YVKCKNGKLTVGNRQATKETLSCTNSVAYSSAIVTDEPCGSGDDTGVIVQLGYQLDDKSWF PLIDVCHNIERSVTLYSAHYLYGQSLKGAVKSSDRRKFSRGPKGLYPNVSPDSVYVRKGQE KAFDRILGGDGHKLYINNTSFLSRGHLAPDADFIYNSGQLLTYYYVNVAGQWQNFNAGNWL AIENAVRDLAVKVDETLRVFTGTYGILKLKNANGDQRMIYLEPTKKLIAVPEVFWKLVQHPVT KSCTWVGSNNVFLTEPPTPICSTINPEGWPTLTNLPKGYVYYCDYDSFKKTVPYVPDVDCK DMFEFPEKNK

[0808] SEQ ID NO 70: A0A866U7I6

[0809] MHSVWFLAWFATAVKAADHLDDPSDLALTLEEEEFENFLDEYLTIEELSWISEPEDDEEET VCIFKIRGDLGQPQPVYVHNNKLLEPAGNTGQVRVTAGQEIIIACPGQNNKIRHPKIISDVAFA KATCMNGTIISGVGWLRDEGEFGRLTCARHPESQAVLTNEECFDHNTVIKVGYQVANSFYT SYHSCFNREKLEVLYVKYHLKAENALHQTRVRRPEWMAGDFYPGVNPDQLYSKARQKTQI AGIVGSNMVGRYITRSQYLARGHLTAKSDHPFASAQRATFFFINVAPQWQPFNAGNWNKLE

[0810] QKLRARTAEADYNTWYTGTFGVTELRDSHGVLQKIYLHGPTGNGQIPVPLYFYKWIDEHR GFGTAFVSINNPYYTAAEVRSLQFCTDYCRNSDRFEWIGWDPDRIDLGYSFCCTINDFRRKI SHIPDVRVRGLLN

[0811] SEQ ID NO 71 : A0A834IEZ2

[0812] MSIVFYGFSLSIGCRRLYTNDTHLPVPFYQDGSRYNLAELHNGYLNLFQNQQVLFICPGAKN YLKIANQTCDYFKANATCVASNKLKITGNQFNINDMFCTKSVRADVKETRRRCNSGYIYEIGY NVTARNWIPLIKICHVNTTGNTLYTSHMLGTAMLNVSTKITRISFSLAGFNTNIPVSISYNKAVQ RIRFGSQLRSPSLGQKYVNNTSYLARGHLTPHADFSMASSQVSTYYYLNTAPTWQSLNAGN WKSVEYTVRKLAKQHGNLHVITGTFGTLKYPDIMDNPTEIYLTGNEIPVPKYIWKIAYEPVKLK

[0813] AIVFVMLNDPFIINNNKYSHFCSDVCRDNGFSNTGWRENAKGLVWCCSYDEFEQIVKYRLN LNVRDTLEFKKT

[0814] SEQ ID NO 72: A0A7M7ITG6

[0815] MMTIAGTRRCCLLLLLAFSAADARSIQGARCELSIDYRKGDLKEPQPLILTEAKEPTFLLPEQ HRLLLIPANASILLACPGRNNSLLGLPGSAGSQQRARATCLGGKDFLVESEVRSFSNLTCKR LPKETVRKVAGSTCLGRHDKYEIGFAVDRQLFLPLIEICRDASTLRTYYSKYTIPKEIAGYQSQ YPRPYWKYGSLYPGYNMNYAFKAEVQLRSMTALLGSEEQARKFIQPGKQYLNRGHLAAKA DFVYGAQQTASFWHLNTAAQWASFNSGNWMIAEAAVRNLTSWRQLDLLVYAGVHGVTSL

[0816] PDVEGVEQPLYLLVNATEKAFPVPRFYWKIIHDPVGNRATAFVGLNEPYASEITEDMYLCPD VSQQEGFSWIGWEPRNIEKGVSYVCSVAELRKSVPTIPALGETGLLL

[0817] SEQ ID NO 73: A0A1 B0G3Y0

[0818] MRLLGIAFAFAFINMATTQCILTMPDDINGLNAPVILVMNVPSKGYKLFQPIAKQTHFLDESKIK LICTNEDNFLKRSSGNQLTLTCTDGEFLDASGAKREIRGISCEKLPTYVIKKTTERCAVNFFIY HVGYEIDSRFYGPIYKICYGDNNPAGFYTHHIINGQALKYKLSRGSKFYTSLITTKFMAEELDD LYYLQSQLKVFKYLNHNGRSIVNDFNYFVEGRLTPDTDMITFYEKLCTYDYANVIPQFRTVRD GNVLLVEKRLRDLAMKQNLKLDVYSGYFQTLKMNLEGQDKYIYLDNQAEHPEHPVPQYIYK

[0819] FAFDKFNNSGLVFVILNDPYSARSGKLMKEFCKNVCEEANVKTRRFRLRKNGYTTCCRYSE FKRKIRILPETIEVKHILKWY

[0820] SEQ ID NO 74: A0A182TWX2

[0821] GGCTVPFADLPYPEQPLILIPGTEKYWYPLDETREILVPTGAPIELACQQGFRLFPTQRSITIQ CVTNSTFSFGGSSYPMKSLACTSYWLSSAKTTQARCHNESVIVNVGFELADARWVNVFDV CYDEQLYHTHFVRHHMNRANGGYQSGNPRPGWYQGAYYTGVNINTLYTVNRQRETIATIL NSQARADVLVQNTTNGIYMARGHIAARADFVYGTEQNATFWFLNAAPQWQNFNGVNWER VESSVRDFVGKRDLELTVYSGTYGVQKLADGNGDYREIWLDFDPVEGRRRAPAPMLYYKIL

[0822] HDEASNAGIALVGVNNVHVPVELILREYVLCKDIGDEVEWIDWERKNLTIGYCYACEVNAFN DAIGRPHPQLNVAKLLTSSGGKGAVFALGLMLAGLGLHALLFWRSSALRA SEQ ID NO 75: B3LZE6

[0823] MEHLRCFFVCFCLFLFMGTALAYCEINKNILDSNYLFYYKDVNNKYQLQLSDRVALNHKVYM

[0824] ACGNQPILQLTCILDLSRQKYTFDHAFPRNSCDQNKITIKSDRPWAGACETPNIMHHIGFRL QTEFIELYRTCFNDVLKTVKFTIHLISDSSHTATRSSQFRPDGVLTPAQFDVYRRDGVFACFQ RELGNGQSYVQKSRCSIHRGHLVPNKDFPFHMQQDATFSSRNYVPQSQNRNQGSWKWE NWVRDLAKSNPNRILKVCTGTLNVLELESTSNVMTEVFLWREPTALQIPIPKWMYKIVDYRY VSLTYNDQFSTTAPNPQMLNICNPDSCGGLDLTNRPFTFCCNYQHFITSWPYLYGLC

[0825] SEQ ID NO 76: A0A7R8ZAD0

[0826] MPIDIGLKLSVMYTMLPIVILSLFMGLVGTGDAAAIIPKAGCSVSVNTDLAAPQPLLLIPGGSLD

[0827] VYGFKLPTVPGDVINFLAGEALTLACPGTNNFLNVAGTGNSWTATCVSEKNFRINGVNYAF STLTCRSIPSHTARATGKICYNKNIELEVGFLVGNNFYNLHEICFDSADANPIYTKFELVANIG GLQTGFPRPSWVEGGFYPNLTPNTQYGRNQQIATLSTILGSTDLGAQYVSSTSDYYLSRGH LIAKADYVYGSQQRATFYYMNSAPQWQTFNGGNWNTLEDNVRRYARDKSVDLWYTGTY

[0828] GITTLPNARGVEKELYLYVDENNNNAMPIPKLFWKWYNPLSQAATVFIGVNNPYITSLKNDY QLCNDVSSKVSWLTWDKNSQKKGFSYACEFADFRKSVPAMPALTVKSLLV

[0829] SEQ ID NO 77: A0A2M4BRG7

[0830] VERNCIVPLKPDWSYSSPLIFTQDGALVTPENDSIELEEITLATGDEWLSCSPNYFREFSSE KVLKAKCKKDKTFWNGQDKNFVSALSCKERPIEEIIVTVRGCPSTLRSIEYGFTNPVSNKSYI LGEACYDAKVGRTHFIHTKIKSGTNTIEQLALKVKDNATYFHGYHPTQRYKVDLSKALNINDQ VERFRGVFGAKNAPKIESRRYINEALLTHRQYLSVLKMAWNYQIVKDRDLLLNYDRLLQDIR

[0831] ALEVPEIEIYTGAHGVMTLKDKHNQSAEVFLVRENRFPVPKLHWTWKSADRAIAFAVFSKP QLTEQELEKQGTFCTSVCEQITWLKKLREEEAWRTARAGYVLCCEVEEFRRTIKEMPPIAGV KAMLT

[0832] SEQ ID NO 78: D3TMW5

[0833] MSLLYGLLILAFTRSCLWGQCSINIPDDLKGEEAPVILVKTGNNVKLFRPEEKTTTFPKGTEL LLACTGEGNGLKSNGQETTTLSCNGNQFESAAKEKLKDMSCKSMAKAWEQTTKRCMGDD YTLYEAGYKVNGKFYGSVYDICYDGKAQSNGYTHNFIYGRTWKYKLPEKPYEHYSSRDPQA GKDLDKLYKEQKERFKNTKVNGKPLLDDEHYFTEGQLTPDTSIITGADKLSTYDYANIAPLFK

[0834] DIYDGNIWRYENMTQELADQRQATFEEYTGGFYSYEVEKWKPIGLDDAKYPTHGVPKYIYK LWDTESKDGIVFVTLNDPYHKSPASENLCKDICSEANINEPDFKNVEKGYTICCSYGDFGNR IRTLPKDLQVKGLLKY

[0835] SEQ ID NO 79: B4IYD1

[0836] MKCLSILLLFVAHAAWARVPLREVELPPLVDEHYILPEPLPENRDAACSVTIRGGLPKPEPVY LVTNTEKLYPFNDVGKMDVESGKTLELWCPSGFNTHSENLLTATCVSGTTFKINGDNFEFK DLYCKSWPGFKALKTGAECNGGIVIRVGFEITSSRFAEQMQICFNEEEEVTRYTRHQLFPGS NYYETGVDRITFQTAGYFGGKNVDKLYTQATQLETINKDLGGNADQYFNSGANIYLARGHL

[0837] GAKADFDYAPEQRATFLFINAAPQWQVFNAGNWARVEDGLRSWVSKNRMDVNCYTGVYG VTTLPNTNGVETPLYLARDGNNNGLIPVPKLYFRWIEPSTHRGIVFVGVNNPHLTEAQIKKD YVICNDVSDQVTYINWKTTDIKAGWSYACEVADFLKTVKTLPALTAKGGLLV

[0838] SEQ ID NO 80: N6UD28

[0839] FESGATLDFICPNRNNIIDTTSTGAALLTARCVSGNNFLIGGTSVNWELLSCSGAPTRAIRDT

[0840] GRTCNNGKGQELEAGFAVDDGRFFVSIVICFNRNLQIAYYSYINQTAAINERPTGTARPSWV QGTGIYTIGTVNNLYLRANQRIALNSLLGLPAESTQYVQASGNFYLARGHLTARGDGFYAAQ QNASFYMQNAAPQWQTFNGFNWYQVEIDVRDYAEASGSEVARITGAPTELHLFINATSRAL PVPAISWKVAYNPVTQAGIALIGVNNPYEREFKPICRDVSASIPWFTCRSDLTRGYCY ACAIS

[0841] DLRAWADLSEFPVTSLLT SEQ ID NO 81 : U5EQD9

[0842] MQVKCAIFLLIFINFAIIVIARKISPTKSSKISSNSNCEINLNTEISENEPVYVMDNSIAEPNADKII

[0843] VKSKLKLVCLGKGNSLFQINQQSIELSCDLGKFKSPNSKFSGFRATNCTKEIRADLKPTSKLC

[0844] SKGLATIFEIGFFIDASFSKLFEVCFDQQQASWWTRSYINGKSIKYNIKESKRPTFRQEGIQH

[0845] KKVDAFYKKASQIKTFTDYLHSGDYFSERSFFARGHLTPDADFIFRYEQFATYFYMNVAPQF

[0846] QEVNAGNWLSVEILARKLAEKYQTKLLTFNGVHGLLSLKEKTVYLDWNRKIPAPEWFWKIII DENAKAGIVMISSNNPFGKNLKSSQFCPNVCLKAGYDPLKYKKLENVEKGYTFCCTIGDFTKI VNNLPDEVKSITNLLKINV

[0847] SEQ ID NO 82: A0A226ELJ6

[0848] MTFKLFSTLTAIFLCFQSGDSANSCQLNWGTNLLPAWFPLIWNDTSIVYPKYNGTEWLIEMT

[0849] PKTSLQFGCGEATNLLQNFLLPETLLTCDGYGNLKAGLVSPNFTSMGCLNQVKEKWHLPA

[0850] STCYNGYDNIDITFLSATNSFPLINVCHNTAKDETIFAQHTIHGAALNPYEVSNRRPSFKEGG

[0851] YYPTFTADDAYSQKSQLAWTELTGPSRGPAYIDAARSFYFARGHLAPDGDFVHIYEQNATY

[0852] YFVNWPQWQIINNGNWKAMEMAIRDLATERQIDLKVWTGGLGTLTMQNENGTDVEIYLAK DALKNRRLPVPMISWKVIHNPVSKSAAAIFMVNNPHLVDVPQSLIKCPDICPQLPWVTWPTK TIQKGYTYCCDMKFVQTNGILPMISNLGADVLLLV

[0853] SEQ ID NO 83: A0A1 L8DPK7

[0854] MFIHLGFVILCLKIFIRGTNADCMLDLRRSLPAEKLPIYLTKDNNDNFSWFVPNGQQTTFQTN

[0855] EKLYAYCPDMRANRGITEITCTANGNFPVKSVQCENEIEPELRQTDVDCYEGFGKIWRLGFY

[0856] IPDNSISSTKSFKTTVTICYNVDVLETIYTYHNINGKAIRYQMNRAEQRFSDAGIPTRDLNPRV

[0857] NNADNVYTKNYQSSFFDNLYGSNQNYITNSNFITRGHMAPAGDFIFHPEKRSTFFLVNAAPQ

[0858] FKSVNSGNWETIEKLTRNLASKQNRQVGVITGIFMGISFKNTKTNKDIPATLSSKGRFPVPQY

[0859] YWRIILDTTSKAAIAFITQNNPHSTTPTTLCKNRCNEAGFVDPNFTNYNSGYTICCKYDEFVN GIQLEVPPEFETESFPNLITNV

[0860] SEQ ID NO 84: A0A4Y0BQK8

[0861] MKWLLCLATICAIGSLGEGRDIRHEIPAEVPEIGSFATACSVRTTGDLPRPQPLILIPGTDQFR

[0862] YPSTGNGLLELNAGETLELACQDGFALFPGKTSITITCVINDQFNYDSQMIAFRDFACTENWL

[0863] SSARRTSQQCFNGARIVEIGFNVGQRFPKILDVCHDELSLDNHYWHEFTPANAGFQQGVP

[0864] RPGWYQGDFYPGININGLYTVNTQRATIATILNSQPRADELVQGTDNGIFLARGHIAAPQWQ NFNAGNWERIESSVKTFVASRNIRVRVYGGTYGVQTLANGNGDHREIFLDFDPNGRTRVRA PRVYYKILHNEAQNSGIVLIGVNNVHISLEEIRRDYIFCTDVSSRIGWINWERENLALGYSYAC

[0865] EVNEFNRVTGHLPQLNVASLLV

[0866] SQ ID NO 85: A0A0T6BBY3

[0867] MFIIVLSCVFANVLNLSPSGVNAVTHPYCNISIKNDFNDRPPLLIATKFNRTDFVLPTTSSEIINV

[0868] KEGNFIGVFCPGSNVTLSDVPIRENLTRLECRYDKFYLHNGTSVNFATIACSKSLKSVAQYTG

[0869] KSCLKRYKEFEIGYRYQRDFLTLIRGCFDKVHKITLYTVSAITKAINYAKFAIPRKAYWSKGSF

[0870] FAGVRINRAYIRSNQRNVINRQVGLSNQNSTKYISENDNIYYLSRGHLTPKTDFIYGPHQDVT

[0871] FHFLNAVPQWQLLNGGNWKILEKTLRDLASSRGIDLNIYTGISGILSFRHEKTGRSTELYLHL GDRRKRIPVPKFVWKIAYDSANNKGIAFVGVNNPYLNGNYSKVKICANVCFSASYLHFKKNY GKYGYVYCCKVDEFRRKISTVPDEVARGLDLLT

[0872] SEQ ID NO 86: A0A6P4J7B4

[0873] MAAPAAAAADPQFQPRTEGVKMNHLKYLLLGLSLLLWGNVWTYCQLSKANTWTDRTFAQ

[0874] NVNNRFELLLTDRLQPNQLLYLLCGGGQAIFSTTCLSSGILYPPLPTTNCTVALAPAVEAVRD

[0875] PSCPHTMYRVGFTYQQQFLEIYRSCYQASTMTAYFSITKVYPTYLNSDSPPPTFDRDGLISP

[0876] ADAATFQRSSVFNRFEAILGPHQNYVPTAQTPSFDRGHLSPAGDHTFPRNLRQTNKYLNW AQHQNINRSNWKIVENWVHRLFSEHQFDVLKVCTGTLDVLELNNIRSQPRQVFLAPNKNPV PKWMYKIVSHLSGYKWVLTYNNGWANQSLNPSSVCQMVNCPRSLNPNGNGFTFCCDPA

[0877] HFISHNVPKLTGVC SEQ ID NO 87: A0A6B2EJB1

[0878] MLRIFYCVLGVTVLAVNVTSDCRFDTKRSPPSNKVPIYLRKSNNNQYSLFKPDDMQTNFRTN EKLYAYCPELGPGDGVTEINCEASGNFPIKSHQCKQDLIPNLQQTGLDCLDGWGKIYNLGYY LPGGKFSTNVQVCYNTESSETLYTFHVINGKAIRNRMTGLDRPFFSDAGISRDGLDRNVKVG QIYSYNYQKGIFDGLYGSNQRYVTSDEFLTRGHLAPSADFIFHPEQLSTFFLVNAAPQFKSV NSGNWETIETLTRTLAMKQNRQVGIISGTYGNINFRNLNTNKDVMANLSGRNRFAVPQFYW RAILDPVAKASIVFVTLNNPYSNTVPNLCTNKCAQAKFGFAEFRNFQMGYTICCSYDEFVRL VHINLPGEYDTDDFPNLITNV

[0879] SEQ ID NO 88: T1 P9F9

[0880] MRRDFIFIHNNDILRTDYIPDGDTLTLYCNERTQPVDLTCSRGVLQAMPAGAACTDRLKLTV QATKKPCSASGVSGELYDMVYKWSTGHTITLYSICYSTQRETVLYSTHRTYGFNLATPIYNQ RLRPTFKQLGHMNGARVKSFEADDVFQRFNTLLGPQQTYIKSNRDFALQRGHMANSQDFL TYDQMDATFLYMNWPMARGCNIRNWKHVENWIHKLPSATTYATVLSGTHDVQYLQHSQT HRFVPIYLMAGEKNPMPMWLYKWKYNNRCHVFVTLNDHSKNPAIRATNICQPTQCPQGIS FSSDPEACISFCCDYSEFVKQVGNHAKLC

[0881] SEQ ID NO 89: A0A1Z1 N405

[0882] MPSPQPVLLSSGYAQDFSSYLLPNRKGNLTVRRGQRLTVACPGSSFRILSGSRDHVAADCV GDTVFSVDSVDYTLNGLVCSSAATPSARYTGASCGYRGLFAWEVGFYAGAEFHRLYDVCF SNESSSTAYAHHTIPADWIGAQKNVSLSDVAWTDGSFFGFDVETAYAKQDETIGKLLNVSVE QLFAEGRLERGALAPAADFMLESQQVATFFGVNSAPRWTQLDGGDWGTLEQVLRKSLRRL EVYTGTAGQLALRNSAGVPTAVYLATSKGRKKLPVPKYFWKIVYEPQSKLCAAFVTVNDPT AKLEDLAERYSLCTDICDGINWLPWEKGNQTGGLSFCCEYEEFKRWPGVPELDVSGHRES SSSDVAPVLWLTASLLAAVALTGA

[0883] SEQ ID NO 90: Q9VAU2

[0884] MPDLKYMLTILSLYFFVGSVQANCLIDLAHLNANYVYLSQNNGVYDIQRSDIVEIHQTLYLLCN GGLHRTTFLCRYDSVFSPALSSAACAPPDPWVKVPDTSCSIPSATFAVGFSFNGRFMELYR NCFDGYSLAFQHSIYKAYRYVNTVPRPNPTWQSDQLSGGFDNAYEGRATQACLLTNLGAV QPQCKFDRGHMTPASAFISTELKKSTFRYLNAIPQYRGVNRGKWKAVETWVNNMVRGLYD NPIINNVQIPRTYDVLKVCIGALGVHRLRHNTNNNMIPIYLLDNNKIPVPEWMYKIVSHLSGDK WVMLTYNDVSLPNQQALNQICHVIPCHPGLNLNTKDVGHTVCCDPYRFITINAPHLTGVC

[0885] SEQ ID NO 91 : A0A7M7HFC3

[0886] MTASVTGLLLILLLGLVSAQTESWKSCEFDPAKDLNYTQVLYLHPNETSFLFPVTPDKLLLSG DGPVLRVACPSGTISIRDKVTSLSSVLVECMGGKAVRVSGTFFFGPLSDIGCTGVPADAVAQ LTDRKTATGKKLCEIGYQVGPLDFVPWDSIAVGHAPNGGLFPRWVHYGLVNVLQTRPTDS EAVTLKKGELSQDLAYELVYDIGEQRNSFFEDFEDEEIVEKYLPENGTEYFVAAQLISREDLY YEVQQSATAFYELTTPAWQSVANGNWKLVSQAVRELAQQNLADIQVWSGIYLTLYLPNADD KNVYVKLPERLRPAQFLFKYVLDQSNKRGWFVTVNNPFLTAATPNNVICEPLPSCDLKYPE FADFAKGYTYCCSLKDFKRHAQKLGLPTFQDIETF

[0887] SEQ ID NO 92: Q176L1

[0888] MQYIFAVLLLMAYGIIPNKCRFASEVSTITYETKYNGDFDKNIGCQISLNDDLDQHQPLFIVPG TSRFVTPTANTTDLQFKSGEQVELHCCHGFLISDATSIIATCNGKDKFIHHSKVYDISHLTCRD PVFHIASRVGTLCYNNGTLIKVGFDLGTRFLTLYEICFDEKLLQTHYVKYGLAPWNIKHERAK RDHFVQGDFFPDMEMAEIYSFDVQHATLGLILGSTNRANNLLNRRKDIFIAKGRLAAQADFV YGSQQAATFRYSNVAPQWEKFRTFNWQHIENGVRAFVTRHNLNVTVYTGTYGVIELPDAN GKMQPIYLDYDVRSGGRVPVPKIFYKILHDPQHSAGIALIGVNNPYASLSDIEKDYLFCEDVS RRINWLEWIPDYIPGGYTYACEVNEFNDVIGHHKFDEISNLLV SEQ ID NO 93: A0A6M2DS63

[0889] TLRNTRCSSMTDYLSQLDDHTEISNRILPKGPNSQCNILLQDFTFQQPLILKGINIWRQTDPTL

[0890] NITGDIILACPGGNNVITGTEDRMVQATCANDKLKFSGKSSTPPDVTCKTSITGDVQTTENSC

[0891] AGNKGTWSIVGFDLRNGTFVELYKVCYNQNSGDAIYSHHRLAGRSVLYKQVEPRRPAFKST GTAPGVSPAVAYKGSEQLDQFKKLLGKDQAKDYIDMDKQQVLARGHLSPDADFPLIPLQFS

[0892] TYFYINACPQWQVINAGNWLKVETMVRQLASSRRDDFEIYTGAHDILTLPDENNNKVKIYLA DNKKISVPKYLWKVMTNKRTLEGIAFVTLNNPYAKRRPPHLCTNICDDNNWSYEGFANVDK GFTYCCEVNELRSKIPTIPLLRVEKVLKFEKK

[0893] SEQ ID NO 94: A0A4D6FRS1

[0894] MYLVFKCFLSSAGLIHLVLLVSSFSICTCNPVDSPTNYHLVRAGCMLLTNQDWGEPQSLILNS

[0895] NMTKIVYPETEDSVLLYPCEDIVLACPGSKFKLTEDEVLHAKCERGTQISADHGGSPFDFQT

[0896] ASCEKLPRTTAMATGRCGNDGEMKNIEIGFWKENFISLIDICFDENLLTANFSLYQASYRIAG RQHGFPRMNFIAGKFYGDVEIWKLYSRQKQRETLAKILGSEDLASTYIKDDKSYYLARGHLT

[0897] AKADFVYGAEQMATFYYINVAPQWQIINAGNWAALEDNVRTYIIKNKLEVLIYTIPHGVAVLPD VDGTYQPLYLYFDENNNGLIPVPKLYIKAWDPVSKTGIAFLTVNNPYVTMEEIQEQNYVICE DICDELDWLTWDPTNIKKGYSYCCNIKDLAKSLDFMPEIDVDDILR

[0898] SEQ ID NO 95: A0A6A7G5T0

[0899] NETSPVFVYPEWKLEDMEVRFLLLMNSESIILLCPGRGNALTLTESNSGELKCEDGQLLLDD

[0900] WMVSWSECSCRRHNRSHVRRSEKQCGLPSYSNARLHAIGWMLANSVFVPQLTVCFDHDQ ETTVYVIHRIIGSSLPGRVTDGRRPGFKIGRIFQSRVNLWYKRSSQRKTVGEVFGIKGIGSRR

[0901] GTRFLARGHLSPDADYATLAEQDATYYYSNVIPQWQAFNNGNWKALENSVRRLAEARRGR MLLVYTGTMPQPLQLTDILGDRREVFLEPKEQRIPVPSVLWKLVYDPALHQAWVLQINSLD

[0902] THSPLVSQLSTAATSMCGALPWVHWDTTDPIRGQTFCWRAGDLARILPYAPPLQTFNLLTQ

[0903] SEQ ID NO 96: A0A653DEF2

[0904] MYWLLFSLATFLSVALHGVESECIIPIWSQMSDANRLPHMTYYKNGKYQQVLPSSGTISIDE GGSIILTCSGCSARNYLKFAESESSATVTCRHSSLYYKGIRLDFNRDVKCKKLICNAVTTSGQ

[0905] CGGSGSQYQIGYDVYNEADMLTDWSCYDHSSWIPLYASHVIHGRNLEGSRQIPRPEFKKE KLDNGKLVSDVFQQRNQTEIFKHLTGTDKYITRSNYLEKGHLAPQGDFVYPTSQVSSYFYIN

[0906] WPQWKSINNGNWKSVEYDVRQFANRSLNDVLVYTGSYDVLNLDRKPIHLVQDRYLPVPKY IWKIVYATLAPGKYGAIVLIVLNNPYANQAEPLCRSICGPYGWNKADRSSFSKGHVHCCDYR EFRLKVPFVPQLEVTEVLSGPNPKTLI

[0907] SEQ ID NO 97: A0A1 I8NPD2

[0908] MLTKFVAIVILIIAKDVWGKQCRLNAKNMDRKWIYAVKNPANKYELLHTEMLKDQQKIFLICKS IKMIQLQCKKGKVSSIPPSITCSRRMRATITAIANEAACLQKGGAMYDVRYTLPTGKRVLSLY

[0909] QVCYNKNTEEAIYSRHRAYGFRLSASTYKRPQFATGGWSVARADSFEFGNVYNSFVRLLG SGQDYIKSANLSDRVMERGHLANSQDFLTYDQMDETFKYVNVMPQFGSINRRNWKRIENW IHNLPKNNQYAEWTGGFEVLELPHSKTRKSTPMYLMVNNKNPIPKWTYKWKYNGVCHAF VTYNNPYTNTVASNSPCLAVPCPTGLTFNPDLGSGPSNCCNLRYLVQKVGQQAALC

[0910] SEQ ID NO 98: A0A1 I8PS25

[0911] MIIQFAIIMWLVAKDVWGQQCQLDASNMNRKWIYAIKTPSNQYDLLHTNTLEDGQTLYLICKE

[0912] DDIIQLKCNRGIMDRIPDGSKCMNGIKYQEIQVINEVPCFQRNRGEIYDIRYTLARSGTTLSIY

[0913] QVCYSKLSEEAIYSRHKTYGLSLSSYAYNRPTFAVGSVTGPTRAESFEAANVYASFVSLLGN GQRFITSNSPSNRVIDRGHLVNVQDLLTYDQKDATMNYINVIPQFTSVNIKNWKIIENWVHGL PKNGEYVSVLTGTFEVLELEHSITGQPTKIYLMKNSKNPIPKYIYKVINRNGVCTVIVTHNNPF

[0914] TAQFGNHVACRPIACPDSLVFSRVADSGASNCCDYNQFVHNIGFHAKLCN

[0915] SEQ ID NO 99: A0A1 B6L8N7

[0916] QLSDSVAKEVITTMITFSTTILILGFFGNTIGWSLNPISWFSTAKGCQLRLNSDLNRGHEPIFLT

[0917] KTGRGFELAMPELVDEEGMFKLEKGQTLFVACPGNNNEIRGFTNNGHSLMSRCVIDKTLEM KGKEVISTDLECKNKVFAKLKETEKECGNNIGEEVQLGFEAEGWHTLVTVCYYRSRAETLHA THILFGASLRGAEVKEKKNYFIKGPSSIYPGVNPASAYKQKNQREVLSHLLGEERANYILARS YMSRGHLAPDGDFLLGTWQHLAYFYINTAPQWQSINGGNWLQLENFVRNFASSVKQDFIVT TGTYGILELEDSYGHPQKVYLEPVQESIPVPLLFWKIVADPKKGACIVFVTHNNPFLTKKPETI CNNICHDHGWPADLDDVSKGYTYCCTYPEFKGWDYAPDLDCRSILSNYYV

[0918] SEQ ID NO 100: A0A7M7J3W9

[0919] MFLAVLLWGLLCHATADSAPSHCYINVTKDISDNQPLLLRHNSSEFQYPRVQSPSSFELNAG KAELRVACPNDLLWDGETLGIGSGFLKCHGGSNFGVAGSQVSTPFGRIGCTKPTMLVARV KRNNCPANKCMQVGFHMNVFTFLSVIDTIAYDTENRVPIWAHVKWACIEGRQRNRETREF KMGPLLSNTTVEVADIYNREYQRRRLGAIFDDPERSEWYLSDVIPSGEKYFVEGLLVSPNDL FYKVQQSSTYFYENTVPMWRSVQEGNWRYVSSIVRKLASETMTDLEVWSGAIGVLTLQSS

[0920] TGVQRKVFLARDGSGKSILPVPKLLFKYVYDKKLNRGLVFLWNNPFLDKVYTGDLVICKRQ MVCENLFPQFAHLDKGYTYCCTLDGFLETAKSLGLPTFNGATSIGN

[0921] SEQ ID NO 101 : A0A7D3T376

[0922] MYVYMFLCFFALAIDNTIASCSFNLDSNLDGKSPLVIKNNDVLLPDNELGDVELSNYDQVQLL CSGSKNNLIQYHGYLVTATCRNGKFTLPNGNSYNFKVLKCQNTPMPHAQEIGRCNAGAKLI SLGFTLGETNFLEKITLCFDAQRLTTYWSRNKLRPSSYSCQIHDRGSFKIDFFNAIHSTINPDE VYLHSTFDRGHLTPRCDFFTGPEKRMTFYYINVSPQNRKLNQGNWKTLEGVIRTASKGAKR TLVIYTGTTDTRKYLKKQIPIQRYFWKWIDANSGEGIAFIGVNSVEDVPSPCLAVDCQQIMW

[0923] LGSSRLRFSNISKGAIHCCAIDNLRNIPNIELPTLSEYRGGLLRNM

[0924] SEQ ID NO 102: A0A6P7FKA2

[0925] MHKSIVYFYLVLKLSFFVSLYVKSVHCITSIGCRISIANDTNEKFVPVLLKTGPAGFDLAVSQN STINLKRGEHIAFYCPKKNYLKLTNASLEHGRCLTGNSLNIKGNRFDFPNILCKNQWGELQK TEDTCGNKNGRLFNIGYQITPEDFVTLIKVCYDEYLALPLYAVHTIFGKEIEHATKFANRTQFS TEGIPRDVAAEVTYRGGRRDQKAKISNLLGGTEFGECYINNQSSYLARGHLAPDQDFLFAS GQLSTYFYINTCPQWQSINGGNWVRLEAAIRRNAANYQHNFTIITGTHEIMELPNVNSELTRI

[0926] FLAPKNRYPVPRFVWKIWDDETKNGIAFVNINNPFLTKMSENDMLCKNICKDYDWDYQYFN TSIYKGLIYCCDINELRYWNTIPYIAVRGVLKRKKYKF

[0927] SEQ ID NO 103: A0A7R8UXV7

[0928] MVAHRCVIFKILFIFLLCGSVICANSKNRQQRGCQIKLPTDIPKGEPLYLSGNPGRFRLFKPIN QTNHLAVNSKLTLFCPGRRNTLLNSPNNSTELTCGSGNRFVNENGHEVNLSQFNCTRGIEG DLHMTREKCSSNRGYIYDIGFRIDSRNFLTVFKVCYDNSSETTFFSTHHINGPWKYFVKDSR RRNFKTTGMSATTKADKVYNKKSQIDRFSDIFGTQQIYINNSMHLSRGHLTPDADFVFPNFQ LATYFYVNVAPEFQSVNGGNWARIEFMARELASDYKSDFTIYTGIFETLSLPNRLGQPVEIFL

[0929] DSNHKINAPKWFWKVIKNERLDAGIVMITLNNPFAKRDEIVEFCPNVCKRASLTSKHFEIIKKG YTFCCEVNAFKNWHDLPRNFTAPNLLECAKANYLFDNEVQLQ

[0930] SEQ ID NO 104: X2JC21

[0931] MDRTAASQLLFLACSVILVHQGSSAECSVNVNSPNFPSPQPLLLNRNADKDLSSFLAPDAR GDIWSEGTQLLFACPDSGFKLIDAQTLNATCVSGTVFSVDGTAYTISALACKKLPSASAVNT

[0932] GAQCYVNAVKVQIGFTVEEEFNKLYEVCFDTKELTPLFTKATVIAGIKGFQASFPRPDQWEQ GQLYGSVDMSNQYRTQRDTLVNLLGATPENMSKNYLSRGHLASKADFGLGVQQSATFYYE NSAPQWYAFNSGNWNQLELDVRDFASNNNFDLEVYTGTYGTLQLEDKNGNQTDIYLYVGT SEKKVPVPKFFWKWYEPKSKRAIVFVGVNNPYLSSSKLPEEYRLCKQDACRVNWLHWKP NNQTAGRAYCCEYGEFNDKVPYLHLAVEGTFSLSGVEATSALKLMTIACSVWLLVRHSYN

[0933] SEQ ID NO 105: A0A482VUF8

[0934] MFRSWDPWLSSSTYELIYPPVENAPNDRNITLPVSETIIISCNGGHFDDITTKTLLATCNQDG QFEVQDAIIEFDKLSCVNFQPQVAKVSDTPCGPNAVLIEIGFQIGLNIKLVPQIWCFDVKNLN PLYTNHNLTKSIGFIKSYLTTPYWDPIYNVSIDFDELYSVDYQLSTINRLLGLPANSTRYVDPA SCNPDPIITTPGFCFSRRPLTYRGDFVYISQQDATYRYITVAAQWSFVDYNLEDLQNNVIDYV KKNKLDLEVYTGSYGIVTLPHADTGEDVELYLYVSGDVKAIPVPLLFWKLIYDPLTQRATVFV SVNNPHQTDVSKNIICEDISDGITWLTWEKHNVTKGYSYACAYEDAKATITYLPEIEVKGTLG

[0935] TVRDKYHPPAV

[0936] SEQ ID NO 106: A0A6J2XDR8

[0937] MCSTMSITTPWFVIFIIVKWNGEGCLFNYTNIPGETTLPVPLYNFTSRYELIVPDHGILKLRT

[0938] NEEITLICPGVRNYFTNFYSRNSNITSVTCVDGNTVKIHRNEEHFQDIRCKKSVKGDVLKTSH

[0939] KCSNGYGKIYKIGYQIKRTNFISLITLCHVSPIGTTLYTEHWDGKNLKYSSFSLRPDFSTDGL

[0940] NVSIPASLSYKKVAQKSMFSKLLNSSILAGHYINENSFLARGHLTPQADFLLASSQYSTYFYIN

[0941] VAPQWQTINNGNWKSVESWRTLSRTYGDLTVYTGTHDILTYENNHRIETRIYLALNEILPVP KYFWKIAYDPKSRKGIAFVIKNDPFTEISNNFCTDICTTAGYRNSAWKSRSGGYVWCCDVKE FRNIVHIPPLILSGNGILTNNLFKLAKALFTQNDNQ

[0942] SEQ ID NO 107: A0A1W4WNT4

[0943] MLNPVIVFLISSASLTIPSVSLNKCNVPISSLRTYDVPLFFHAGAKISEKSFVKARGGRMIFRD KENVKWCPIGKTNVNTVTKRIQCLNSKFVDDGQYKDLKDIKCQNSRSPVPITKWGHCKNA FLTVSVGFQIDDVFLNRIESCYDNNTRLPLYTKSILDKSVSTCPTGRTGWYSTGTFRIMLPENI EDIYTIERQRQTVNQLLGLPLNDQSFIGDGHMYFARGHLTPKCDFTYPSEQRLTHNYFNAAP QWQQFNAGNWKVLEARIKDFVTKSTTFDYIWYTGTYGITKFNNIDIYLYNGQNIHAVPVPQY

[0944] FWKIIYEPISKKATVFVINNNPKRIKTQAICTSFCDTLGWLMKAAFQPRNIMKGITFCCAYTELI SIIPNLPDIKVLGSLL

[0945] SEQ ID NO 108: A0A6P4FJE4

[0946] MGNNYLILFTTLLVNQVRGQCKFTPNQLLGSNGAFTYQDANGNLQLQLSSWQVGATLTMY CGGEETHTITCKQDKKIFLFEPKLPITCGTKLSTDTVIRYNDRKEDYSDCKKTMYSIGHKIKGQ FLKLYDACYDTVKLRAVFTQATVFPHVAHLARPKDILFSHEPVMSAFDANSFSQNYIYTRFVK LYGQSQKFVPLITDVKKEVKTEVKTGINRGHLTDSSSFLFEDQISLTYKVINIVPQFASINGGN WKKISEWLNTLSLNSRLTVRTGAMDTLQLLFKQPNRRNVNKPAYLIDDGKNPVPKWIYKIISS

[0947] NAPDGTRTPMIAFLSYNNDKMDPEPSKTLCTEIECPLVFDMSVESGKIKCCNAVTFIRELKQI YIDDK

[0948] SEQ ID NO 109: D6WXP5

[0949] MIRTILLVAVSTSFVLSAPKDNGCYLDWANAPSAPLWNSSLQFPTPIGAKTTQVFFAQGET FYIGCPGGSFNYPILANGEPAVCHQDTTISTLKSQQTVDFNQVYCTLDQTLRPDLFEQVSDD GCENGGKIQSVGYSIDDQFVETYRICYNSSEKSILYVKYLTNQWMNQLRVYDPNEVPKVWS GIFQTDLKQIYNTIQPRIEDLGWGYITNVSYLEGGFLAPYYNFLSPFQRDTTFDMANAPLQW ATVTAGNWQAMEDSLQWLLTYKDFGDVIMLGGTLGVADLPDNDGNPVDLYLVNGTSLPVA

[0950] RWFWKLMHSPSTSSGILFFVYNNPYVAKPQVDKEFFTNLCDEDVIDEAKFLTGVDNLNPYS GYIYACVLDLYEIIDPTMNDIVNNLFNDQGILTGAQSPAPKLPKIED

[0951] SEQ ID NO 110: A0A5N5SW75

[0952] KNPVHFSPEHHFVLDLKPYFIYPDILNSKSQVLKIKTKEHRLVSCPGKKNKFKDGEQKLKSLF CKLGKLVLQDKESIAEDINCKRRPMGSIRHTNIKCHGTKDGVFHVIGWKIDTNWFIPQIIVCW DGTRETTLFTKHIIHGQYLKNRVPSPGRPNFKREEFYSENIGEFYTHKFQENTFKFLLGRVSN IPGSRRYYFARGHLAPDADFVNPSEQDETYFYSNAIPQLQSFNNGNWKKVENSLRDLAMEK NRSLWYSGTFGILTLKDERGKDVPIYLCPEKHLYPIPLIIWKWHDPVAKEAVTILGLNTMTN

[0953] EADKREAFFDTCDNICDQILWLEIDEEKDFRKGSILCCSWNQFKTNVAPLEDLGRPKLLKI

[0954] SEQ ID NO 111 : A0A0L0BML1

[0955] MKALLKKFIKIVFYLNLITLSLSCNLNYRDGHWTNEWRLLLVQEGRNQYKLLRDHNVAANSD VIMLCNGNPRQITVHCDAQNRFNGNVPLRNNCPENLSIRPERIQVPAQNAPHCPYILYRIGF NIEINHHRQFLETYQVCFDHRHLRTVFTINKAYPVAGIRPQVLKFRPDDFFYGDSFHAFDHKV TVKRFNKLLGSTQDKMIEGNLDHIIDRGHLTPSADFTLTNYKRSTFYMINVMPQFKTIDNGNW RVIEEWARDYTRTPTDICTGVLDCDLDSIRPGDLNCWKLQDIRGRWVPMFLYDKRKIPIPLW IYKIVKTRQQSYVFLTLNNIHHQGQVQPPLGKCQVINCPFTLTNTVKLGWTFCCDYNTFINKN VPHLRSVC

[0956] SEQ ID NO 112: A0A182N2I3

[0957] MFAYAAWATLVALAHGQCSVDFHSKLKTPEPVFLRKNNNQIVLWVPNGPLLQWHAGEATL IACPGNKIKLGDVKTETVTAYIQCVSGTLFKIGEQPVDISEVTCIERSTGTHQNTRQACGNGG SGTLLNLGFDIPNVGFITYIQSCYNMQTASVIYTRHTIHGTAIKCESQALKYLFSNNLERGHLT PDADGILRQWQWVTYFYVNAAPQWKTINNGNWKLVEEIARYLADRLREELIIYTGAHGILTLP

[0958] DVNGHQVRIKLEEGGIEVPKWFWKIIWSKNTNRAIAFVTDNNPFTDMPEGEKLCTTDSEHYG WVNFPGWKKYNDGNKLDRGFTYCCTVPVLRLAISDIPDEFRMVDVLRY

[0959] SEQ ID NO 113: N6U1 D0

[0960] MSYESHAVLVALVGVILGFGSADSAHEANSNITTVTCLRGDKVQLYKHSHKFEEVQCRHSV RGNVRATDKECGRESGRIYQIGYPVSRTEWITLISTCYVPAEGRTLYTRHTLYGEEIKYASKA KYRPAWSPSGQYENITASIAYQQVYQKGTLSRILGSSILANNFINNNSFFAKGHLSPHADFIL ASAQFSTYFYINAAPQWQKINGANWKSIESTVRNLGKFYGTLEIITGTHGILALFDKQNNPHDI

[0961] YLGQRNTLPVPKYLWKIVLNKATQEAIVFIILNNPFIEAVNDEVFCENICEQVGFDKTSWKDPS SGLVFCCTIAEFQKWKTAPKLDVSGVLVQQSLLWEPSLLNW

[0962] SEQ ID NO 114: A0A0T6B5D5

[0963] MLILVLGIFAKILSWNSADGRCYMSPSDEVLLIDASKKKPEILMPVPRTTFVLLRDKDSVEIYC PRRNDEAEITGAICRYGTLLNNNLGPLDSVWCAKSPEIDVQYTGENCSENQEKFQIGFPYKE NEFLHVIENCFSNETRISRYTISTISKAINGNLTLDNRSKCETANSFDDIPDLNQLYGKEVQRN TINKQLGFNYNSSKYINDTHYLRQGYLTPRADFIYKSQQDAASQCLNWPQWKILNEGNWKI

[0964] LESNLRDLASSRGIDLGIYTGTSGILSFPHNETGELTELYLNVDGTKKVIPIPKYIWKIAFDDVS QRSIAFVAVNNPYLDNLEDVTICTNICHVLSYITFNSSDYIRNGYVYCCAVQDLPNTILIPEYEL IGGIGLLA

[0965] SEQ ID NO 115: Q3HM52

[0966] DFQNDFDTSNPLPIILRDNKLLNPDPNTCSWLPEGSKILIGCPGDTNSIMYMQPHGRIDTGL KQLEASCNGGSNFISGNSLLDSFALKEISCKEQPKVTVRKTTKCGKDNQHQTYQCGYTIGN VFVWAFESCYDERLYIPSHVKYEIRPHTCTPAPTAPSGNILTGLLEGLLGIVLYILSAVFNVLTS LVAQKARFIELFGTELADVYLRDDNLEPCYLANPFYFATVAEQSGTTHYVNLAPQWKSVTES

[0967] LESLQQGVQNYVNQGQNSLTVFSGTIGVLTLPTSLGSSRALYLVTFLGSNVCPVPQYFYKW IADHGPLVFLISNSIAAVPSGVPFCTDLGSDAHSGISWNRKRSFDLCCDFEEFQKKVSYLQF

[0968] SEQ ID NO 116: A0A5E4N232

[0969] MYWTTLTVLVTFSCLHLTVLCSADEHNIKTYSADYYDAKDCRLSVADNAVALKMPVPFRRD DLHRHSIIYPDSDGRLRVKFLHGFKMSCTTAGKFASADLSNVTEVLWCAGDNSLWYLGRE YAYADFVCNREPRSELTVTDETCQGDGWVAVGFQTESAYLGAYRVCFDKATKNSLYSWY DSRSPYYDQHQSMKKRPSFVASRQLYGNTNVNKLYTVAQQRKTVAKILRSDSLANMYIKND

[0970] NNHSLSRGHYTPKADFYFGFEQSATFYYANVAPQWQSFNGYAWNHLEMWTRYRMDNST GRRVIVTGTYGSCTLPDVDGVEQPLFLDPPTDIPVPLFYWKLYYDADAEDGMVYVGLNNPY MKIDESAYICPNLCPDGYRGNGDTSYVAAAGKNDENDGLIYCCSKESFEEVYGRLDPIVFRS V

[0971] SEQ ID NO 117: Q5WPS9

[0972] MHLQLNLCAILLSVLNGIQGAPKSINSKSCAISFPENVTAKKEPVYLKPSNDGSLSTPLQPSG PFVSLKIGESLAIFCPGDGKDVETITCNTNFDLASYSCNKSTSTDTIETEEVCGGSGKVYKVG FPLPSGNFHSIYQTCFDKKNLTPLYSIHILNGQAVGYHLKHTRGSFRTNGIYGKVNIDKLYKT QIEKFNKLFGPKQTFFRRPLNFLSRGHLSPEVDFTFRREQHATEMYINTAPQYQSINQGNWL

[0973] RVENHVRDLAKVLQKDITWTGILGILRLKSKKIEKEIYLGDDVIAVPAMFWKAVFDPQKQEAI VFVSSNNPHVKTFNPNCKDVCAQAGFGNDNLEYFSNYSIGLTICCKLEEFVKRNKIILPKEVN

[0974] NKNYTKKLLKFPKTRNKEGDKKWRKRAKGA

[0975] SEQ ID NO 118: A0A6L2Q562

[0976] MAELRVTSAWAAVMILGTYTTRGKAECIVNIGSDLGYRSPLILTQNTAEFGGMGFLLPSTST PRASLYNGEEVALLCAGSNNRLNISGTVSQLEEANVRCDTGSSFEYSSQYLPITDFQCTAVP QSTIRGNGTCSSGQKRLQIGYQMSSDRFIRLIDMCFDDVNYSTHYVNYTLANGIQRRQVFEH GSEVDRVDYYNSLPDSPDVYYTCSNQQRTVGRLVGSSRVNRYIKCSNNINYVAKAHLAAVP DFMYFAQQESTKHHVNWPMWYSIKNGNWLNLEEEIRQYASNTSRAPSDLVIYSGTLGITNL DNHRIYLGRDSNNLNWPVPQWVWKLVYEPTTKEGIVFLWNNPYRLSFTCRCVCAQTQWT RAWNRADAHKGYVYCCSVNNFRSVFSGLPNFEVTGLLTKNRPYTPPLFPDIAAQ

[0977] SEQ ID NO 119: A0A194QPJ0

[0978] MSPPLLLLLLLLACAGPALLTDFSRAPPAMSPPLLLLLLLLLLACAGPALLTDSECVIELTCSEC VPEHMPLVTSHRAAAGVLRVPPGDELRLACPRGRFLTYPHHATLAVLCEAGRLRVRHDGVL RHLLHLGCQDDVFEDVPHQVEFCAPPLQGRAYRVQEAHGARHLATLCFDQDRGVASYARA GNAPDGALPLPPHEESSAPRSLLANFNQMFDSATRRAADRLYSDDDRLHRRLREILKHDRF SFAEQTLTSVSLLAPAYFDDQNVRVADFASNRVAAWRSVAAGNLRHLQRDVARLLVAARP HRHLEVYAGTHGVLALRTGGNRTEVFLEAERFPVPRYVWTWQEAGSRRAVAWVLNDPF VSVSEVREAVFCESLCGRVSWLQELRRHRHYESAVYGLTFCCAVHEAARRMPEVPAGALA

[0979] AVPAGDAGLLTDLL

[0980] SEQ ID NO 120: A0A6J2T6C0

[0981] MGHWYREAGLSLLLLALVNGQCHFTLDETKSSKFGFLTYREAAGADHMQRQGWPNGATI YMNCHDPKKPKTGWKDWCNNNQFQKPLPLCEHMKSSEVREVNDNSCPATMFATGVEIR GQFVELFRSCYDKSNLRALFSKNIVYKNTFFGPNPNLSFDCGKWPTRDMEAYKKEKVYDT FFNIYGSNRYIADAKTIAISRGHLSCATEYLYRNLKCASFKYVNWPQFQSINGGNWKNIEKW VNSQVPVNSYLRIETGGIGVLKLPDVEGNLKPAFLLGGNQIPVPEWTYKLIKDARNQPIHAILT SNNKYGEKIPAPAFCTQIRCPGFDYTDASDQGYTYCCDARTFNPPRELIPPTVSKTPSRPNS

[0982] RTRPGGAH

[0983] SEQ ID NO 121 : A0A4C1Z526

[0984] M I DRADASTSNPRYSVDFRWM M I ADCVLSMG LDLTYNQPLI LSSVTGSWYSM DSSPRG M L QLRAEETVNLACPTHPDGVNDFANFPNKNVLQAKCKSGNNFEIEKKVYSFDRLQCRNQVKP TTQNTGLPCSAGLSELIKIGYTISGTFLPVYEACFDKDKKGTIFTEMELSPNYAKNDAPVPEST WVRDNLFGTANLNEEYTCNSQRDTFYKILGKQYFQPGDTCCYAKSKFVDSRDLVYDTQRV ATYHYLNSVPQWSTCNNRANWVEVENRVRRLVSNLQRPLRVWTGTFRIVNLRNKWLLLHS DMYLHYNSKTDRAQPIPLYIWKWHDPQTRRSLAIIYINKPDLVQSNIESYVLCQDVCQNVNW MNFLSRQTVASGYVYCCDLNDFEFGFHITNSPFPKNAGAILRDDSHSSW

[0985] SEQ ID NO 122: A0A2J7RLA6

[0986] MSSSLLLTQLSLFVLLICLLYGHHCAGAGGVLGKADISGRVNRRFGRSGPSETNVNQTQNQ VQNAGCTLRINSDLDKDSPLLLKSSGSLDGNRFFLPTPESDILHFRAKEKFLLACPGTGNSLV

[0987] VNKRQKTFKETSASCDSGKKFKIDGTSTSFPNITCKDTPESSIVKTGKCNMNQYDKHSIGFK VNTGFIGLMEVCYDPQRHTTLYIKFIIVQGVENRQRNVPRQNYFRKSGFFKGLKPPIESVYNC RSGQYAIFANLLGSEELAQRYIDCSSRGGESLEKGHLAAYADFVYYAQQKATLHYINTVPQ WQSFNAGNWKILEEKIRIYAKVHSTNLWYAGFHKVARLRDKSNKEHPIFLTKDLNNNDAIPV PEQMYRIVYDQDMRRSIVFLGINNVEINPNVSKSMICEDVCEHSMSFFTGWNRKNVSKGYIY CCKLHDFLTSSGLKSAFPVNNVPLLN

[0988] SEQ ID NO 123: A0A067QWU8

[0989] MAALKHTILISVTLAVLLSSTHLIRPTVAARTNENDCKILTSDLKENSPLILKTVSRSGESPFHM PKTVGSDLLTFRRGTTVRVACPDVKNNKLLLPKSDSVESEENVICIKDSTFNVNGELQNFSSI RCDHPPRSAIKEDGYCNATNMMRRIIGFNLKDEFIEVIEVCYDLTKQVTVHTRHSLNRNIVYK ESENIRRNDTLLEFFETPDQFYICKNQVATLRNLLQSKNHAQKYIDCDKGGNRYLTKAHLVP

[0990] KGDLLFEFQQKITSYDINTAPQWQTINTGHWRILENRIRRYANRHNADMTIISGTMDVMTLPD

[0991] RFGIDQNVYLTKGEERNITMPVPAIFWKLVHDRARNAGIVFLLVNNPHHQDFETTRGYFICEC VCSETSSWFDGWNRYDIRKGYVYCCTIDDFSQKTGMKSFSFRVRNLLH

[0992] SEQ ID NO 124: A0A6L2Q7U9

[0993] MCADESCGAGQGTVEGCCELGNEIVGFMYYCEKVAVPISRTTTGFSRMTRVRVCCTVHLS

[0994] DISTDSPLILHHSHRKGMSPFRMPSDPSSLSLTFSSGDAVRLACPDVDNNLLAIPNSPLMVY

[0995] EVIAFCIRDSTFSISGAAYNFNSVRCKLPPQSIIKEEGGCGIGNGTKYHIGFKLRRKFLGLIEVC

[0996] YDTAKHSTLNAVYILSKDVGYQDVNSQTNSYISGSDFRSDNTRTPDDFYSCKNQFKTLGHIL

[0997] GSSIQAKKYINCNHKSKTYFDKCHLAPSDDFLFGYQKNATSYYINTAPQWKAISTGNWNILD

[0998] NRIRRYASTHKVDLTIVTGTMNVTTLQDASGTERNLYLSKDLRNKSTVPVPAVFWKLVLDRP

[0999] RSAGIVFVCINNPYHHDIYIRGYVICTNICNSTTSWFDGWNRLDVRLGYVYCCTVDEFRAKS RIKPFPFSARHILR

[1000] SEQ ID NO 125: A0A6G0TEW8

[1001] MISYHLSTASARKSTESLIHLANLVICSHDAIIIIKIICSHTICKQYFNNTFCLLCRNIVENGDCFL

[1002] KASDSDAPKMPIPFTSNGKRHSIIYPDAKGSLNVKSGLSFKLSCGTSKFASNAIRRNGTADAL

[1003] VTCVGNDLLAYRGETYQYTEFRCNDMPKSELRVTGEMCQPANYTVATVGFRTDRAFLQLY

[1004] SMCFDKSTKNSLYTWYDARSPYYDNHQKYSKRPSFINSKELYGKTDVNKKYTIKEQKKSVA

[1005] LILRSEQLADKYIRNDNQHSLSRGHYAAKVDFFFAYEQTATFYYANVAPQWQIFNGNMWAD

[1006] LEQSTRTKLDQDSGTSRHIIITGTYDVCTLADVDDVQQPLYLDLPRSIPVPLFYWKLYYNVDT

[1007] ADGIVYIGLNNPYKTIDDSVFICPNICPNGYHGRGYLDKGSENDDPEPDANNGLIYCCTKESF ENVYGKLDPIVYRPLM

[1008] SEQ ID NO 126: A0A2S2QHN7

[1009] MHLLKLSVLVIISCLCPTEIYSEDNPFIKNENGKLWPELANYCVLSMVNDNEKMPVPFKTYG

[1010] DTQMPMIHYPSLKDMMLLRKDDILRLSCTKSRFSHDRIKHLDELTVRCVGYNNLEFEGHNYH

[1011] FNELECESTPSADFIFVNASCHSGQYDVYAVGFTIRKGLIPVYTICFNPKTKDTLYTWYSARL

[1012] PFMRTSKEYEAKKFFQSEDLYGHMDLDTLYTVENQTYIIGEFLKNKELAEKYVKREEDKYLIK

[1013] GQLASSKDFYYVFEQLAAFYYVNTAPRWKKFDDDRWALLEESLRDMVFQTGNRYMFVSGV

[1014] YGSCKLRDKNNVYRHLSLSENGEITVPMVFWKLLYNVDDKGGIVFLGTNNPYLKDEEFSFLC

[1015] ENLCQNGYRSDNAGAELFKAPNPVYCCDLDSFQRVYGSGGANI

[1016] SEQ ID NO 127: A0A6H5IT40

[1017] MMGFYKFSLLALCVGCCCCLLHLPAAALAKCLMQTGKIGVDNVASNLVLGPLVYQYGPAGT

[1018] EFRYPETGDRCIYRHDGRWLHFACPGSSIDYHIGSKFLRSEASTSYLYYKDNPYGLRIHNHS

[1019] FEDVHFQNDCYPVCRVAPEPLMIDVSRGQRRNFSVGYYVNALDFLEVLRVRRHLGNDLSVA

[1020] ATWPALAEAPRLPHPPKFQRGAIFAETKIKFSEIYSLDYQRKIFQNVWTNETVMDQYLNENN

[1021] HLVMTQMISDDDLYYMTQQKAAYYYENTLPMWLSIARGNWRLISRLIRELADSTDRPLRVKV

[1022] TPTGALRMNGSRRGEPVMMGRSLASSETLWPQTVDKYVIDPGLPQRRSLVYRVLNDPWP SEAEIARAHEGCDLVEECQLRQPQFKDRARGYVICCRVDMADLEILPD

[1023] SEQ ID NO 128: A0A226F5I2

[1024] MGLYKVKAFCVLFVLVILFHFAYQFKNSPSSSNHKPSGDIPLWNKTEHGISTNNILTALPVNRI

[1025] WQGPRGKVEIKYNCEGDSYYIYVYPKKANCAPNDMSFEDRPTYSVCFDDTTKQTLYTTGHI

[1026] LGTQATLVHPARPPWEEAKRLFNTGWDYLYDRKTQKFKFSRPTMLGPASTARLFPEDKYW

[1027] FSRGHLFRRSDANTINGQASTMFYINTAPQWGLLNSGNWNKLEEDIRSYAQGGNTLVKVW

[1028] TGTYGLLTLPNDQGVQIDMYLGNNAVPVPRFFWKAVYDMVSKYGWLVQVNNPHATHEEL FQNGIPCEDICHRVPWLTTLVQVEREKVDFGYTYCCSVEEFRRTVNYPSDLPKARSW

[1029] SEQ ID NO 129: A0A1 B6I107

[1030] MMVLRVWLLFFVHIFQTYSLWPFNSRKSCVLDVNWNLIHQNAPLLLRRSAHGYELASPDFV

[1031] VKVGNYFLGHRLTGVIKIPYEEEIYAACPGYGNTLYGLKEPAQFATFKCLHRRNILRMDKQM FKSHLLKCTQPVATLIRKTDHCCGDGLGMEYEIGYSVTPESNPSARNPQTYGLMWMKMNQ PDAKRVPQEGKSGFYRVMTCCYDSQKVRTLYSVNILQGDKIAGAETRRERPPFIIGQSSLYP AHFDVHNIYKYPNQKKVLQKILGRRKSDSMLTKTFLTRGHLASEHDFLMGNWQASSFSYINL VPQWQSINRGHWYRLEKKLRRLADITKHNLVIASGTYGILTVEDDHNTTREIYLDPDRELLPV PAYLWKLVFSPKNNTCIAFWSNNPFEASRTEPFCQDICHEHNWPEDYIYEWRGRMFCCSA DRLRTVIPYIPKFACYGILDYIPEQTPNTVPFNFWFRTKNRFL

[1032] SEQ ID NO 130: A0A834PC46

[1033] MLWIVSLFIHILLVDARSIENDVENDNTTDTEFARNNCVLSMKYKYGDLKEPQPLILTRNGTA SAILYPNINGTLKVQTDQSIYLACPGEQNSLRNMNYAQEVKATCLSTFQKLLNSSELAKSHLA KSRQFLSRGHLTAMVDFVYGALQSLTFWYINAAPQWMSFNAGNWERLENSIRNFSTYRSL DLDVYTGVHGQMTLPNARGKQQDIYLYVNGTNKAVPVPKFYWKIIYDPRSKKGTAFVGLND PFIKSITEDIYICSDISSEIKWLLWRPNDIKAGISYACTVDDLRKAVPTIPKFQTVGILT

[1034] SEQ ID NO 131 : A0A2S2PWW9

[1035] MNLIKLSVIVIFPCLCLTQPDSDESSIPPDISYEDYNYNYSSLANSSEEIMKRFYNPDEEWNQ VDDCLISTQEENASKMPIPYMKNGTRLVIMYPDIDGLLNIKHDITFKLLCANSRFKHKDLNQTS EVEVKCWDAQLLYNNRLYRYNDFECESMPKSELWTDKKCQSNNYTVAEVGFRSDEGLIV LYKICFDLKTKNALYTWYDARVPYYDISQKYKKRPSFHKSKELYGSIDVNKKYTVKEQRKTLA TILKSDDLANKYIMDDSKHSLSRGHLAAKADFYYAFEQTATFYYANVAPQWQQFNGDKWAD LEKTYREMMDQDEDIHWVTGTYGSCMLPDVNGVLQPLFLDPPKNLPVPLFYWKIMYDLFE

[1036] KKGIVFLGMNNPYKEVDKSMFICKNECEGGYKGLVDIPDNKHEANIFCCSIEGFEQVYGKLD PLIFEKINY

[1037] SEQ ID NO 132: A0A226F143

[1038] MGVTGTFLFATIIMWSTKSAFGASMEKPCTLAYHSRTVRLELLDSSDSLTASIQNDENVAIIP SGKQVRLECINEKQIFRDLEAGTVELFCAGGTLSHTGAGAQDIDQLLECRSPRAAPFGELIR NGTCVNNASNIDVYVTSDSGESSHLYSVCFDEVTSTTLYSKSYIDGKLLAGGKPDPSKRPGE FQDPGFFPFNWLAYKQKQEKITLTKLLGEEKFEEYFPKQTYYLSRGHLFPNGDPYYKYQKD GTFFYINWPQWQILNNGLLCKRNWKAIEDSFRQMSAKLASHVNVYTGTHGWELKNQDEE DVELWMGLPIKSGDKVKRLPVPKFLWKIAYSEVEDAAVALVQINNPWATIAAEDYLCKDICD

[1039] QLDWLSLTKKQRGSIQKGFTYCCDVNELRKKIETIPVFKYSKVLESAQPSKWSNLEEEDED WDDDEESDEDEEDEE

[1040] SEQ ID NO 133: A0A232EDJ9

[1041] NGESKCELHALSSACYDILNIESATGPLIFTYAKQKPEIRYPIIIKNDFCLFFSSNDKEQPNIRLY LSCPGHNLSMNFELSTYVLSIHTTFLTAIDGDFNVHSFGKFSASYLKSIRCTEEPVHIMNKVS EVNLDKMYYSVGFMINAIHFLTVYRLFRSIKLKSHVYVRWSVIQEKPKFDKLMKEYKIGQIYE NKKINYTIMYQTEFQVELFKNILGNENKFVEIYFSDGNYVDKGHLINKDDLFYNAQQISTYYYE NTIPICDSINKGNWDLVSKIVRNLADETVTDMEVSTLSLNDLSIDDIRFDNRRMYNQFKIYHIIR IPRLLIKIVYSSLNATQEWFHTANDPYMTEIDMYNLMYNRVQPENGLCRNLGLCRDEYPQFL DAKKGLTYC

[1042] SEQ ID NO 134: T1GNB7

[1043] MNFHIVLLLFTGSLFFGSGKCHWPPKALSIGRSAEECVITVRTEKVEPQPLYVEPRRSSFW PITDNGHLWKHNEEIEFLCPGHFKYPFEDDKTVIAKCVDGRTFSVDGKHHDFTAFACKTWP SYTARRTGRSCPGGELLESGFSYEDRFFRLTEICFNPTEEATRYVYHTLEPGSDYYQQGMA RVEFITDGFFGGKDVNYLYTQKKQKETISHILGFDAADFIMGAEQRGTFLFVNTAPQWQRFN AGNWQRVEDSTRTMATKRKLNLECYTGTYGVTTLANKHGEQTEIYLYDEGGHKQIPVPKLY FRWIEPRERKGIVLIGVNNPHLTLEEIERDYIPVLCTGPQIGINSLHDGINSLHDT

[1044] SEQ ID NO 135: A0A226D326

[1045] MELKKFQIFFALWLATCEKFPSQHDDRENVILRRIQFGNKTDNEIVRYKPPSHALSAIPQNRI

[1046] FQGASGKLEITDNCEGQTFYLVSAYPRKDDCTETSPSKMQYESLPTYISCFDPILKQTLYSV GHVYGNGGLPRTGTGNFEQAMSLFQFQWKNMYNTPKQRTNLKAMLGPTNYARFFPDQQP YAYTCG H LFARG DVNTRVG QASTM FYLNVAPQQ G I LN AG N WKKI EDDVRVLVDPG N ALLKV WTGTFGLLELPNDNGNMVKMYLGQGAIPVPKFFWKLVYDMVSTKGWIIQFNNPHATPDEIT RDRIQCQDMCDRVPWLTMTQAQRDNILHGYTYCCNVDEFKQIVGYPSNLPTIINY

[1047] SEQ ID NO 136: A0A5N5TD00

[1048] MKYLLAIFLVFLPYVKGEDCLWNRDTDFPEYPSLLMNEATWQSYLPVLEGTERVIRVCQGA

[1049] SIVIACPGSTISATGTEVAEGVCAGGTYVTVDGKNYDMKDLGCESRVSETLASLSGSCGND NDGVYICFDPDLEITIYTKHLLRGESVKATDSNPSRPSFKSSSGIYSVSPNTCYEIDSQKTLM

[1050] KNILGDSSVIDDSNSYYFARGHLSPDADFNTEMESDATYYYVNAVPQVANYQQWKFQETAI RDLAASHQTDFQIWTGPYDILELDDTDGDPVICELSTNKCLGVITSNNPYIDSFPDPICDDICN EISWIDFDVNDLTRGAVYCCKVQDLHAAIPDAPNISGAGLLDN

[1051] SEQ ID NO 137: A0A7R8V306

[1052] MQYTSKRIHLPLFLWLLQATSHVQSECEFSQYSYYHSGIFLQQFGEGYTMVNFQGNPIRIP DQGSMEAVCSSGFRFPGHPQIFGKIEGEPGRDYEEYTNVTITCQNGDMIYHLPGGLKQDVE

[1053] ENSLICKDNKARFYKANVANCGETGFLYGFMVKEVPVILAEVCYNIAEDKTLFVHFIGGKRSV VLENQTQHNPSLMSEHLHNFGKYVFPNHAEMQKSLDKIFVLPGYTYIAEYKLEYLAPMANYS SLIGNLADNFEYPNVIPMWTTLRNYNWRIFQDLLRKESMKESFEIYAGTSGKVQYPYDLRCN

[1054] STTDFSVWEDTLEVNIPLHIWYYLTARGKPDVSTWIAVNSPIVELGPSTVICTDVCDEISWLK PMSRSRKILSLGYIYCCKPQWANLLKGFPLNEGTDSAAKSKNGAAPKSKM

[1055] SEQ ID NO 138: A0A7S2PGC2

[1056] KWWLSNAYSSFQCHKNQLSVGRALKGEDHEHLITGTFPNESTQLLNPLRKTSFVRKSQGVT ELDYGHNVKIKYCTTSKMPLMAQYSLPPSKCGSPKDTTFKEKTNPYIKVDCGLDRRDQASD

[1057] KLYKGSGFHRGKMIPIADAMRVSPELKERCYSYTNVCPKYSDFNQGIWKVLENHTRNYCDIL SKQGCWGIISGPMWDYRVNIFNILTLSLSIGWFGEYPPTHYWRKFCVPTHYWRIICVRDEKT GKLIDETSIVLPHLKKNKKIKNRSCLKEYQSSPEEIQSKITVDLNVEAFRDGDESRNELFFIEI

[1058] SEQ ID NO 139: P13717

[1059] MRFNNKMLALAALLFAAQASADTLESIDNCAVGCPTGGSSNVSIVRHAYTLNNNSTTKFAN

[1060] WVAYHITKDTPASKTRNWKTDPALNPADTLAPADYTGANAALKVDRGHQAPLASLAGVSD WESLNYLSNITPQKSDLNQGAWARLEDQERKLIDRADISSVYTVTGPLYERDMGKLPGTQK

[1061] AHTIPSAYWKVIFINNSPAVNHYAAFLFDQNTPKGADFCQFRVTVDEIEKRTGLIIWAGLPDD VQASLKSKPGVLPELMGCKN

[1062] SEQ ID NO: 140: Benzonase®

[1063] MRFNNKMLALVALLFAAQASADTLESIDNCAVGCPTGGSSNVSIVRHAYTLNNNSTTKFAN WVAYHITKDTPASGKTRNWKTDPALNPADTLAPADYTGANAALKVDRGHQAPLASLAGVSD

[1064] WESLNYLSNITPQKSDLNQGAWARLEDQERKLIDRADISSVYTVTGPLYERDMGKLPGTQK AHTIPSAYWKVIFINNSPAVNHYAAFLFDQNTPKGADFCQFRVTVDEIEKRTGLIIWAGLPDD VQASLKSKPGVLPELMGCKN

[1065] SEQ ID NO 141 : A0A2P8ZMT4

[1066] MFNGCSIGVNSDLPEPQPLLLIPGGSKDGYGFYLPDDNQGLVNINAGEDILLACAGNSNYIAT LGSGTKTALATCRSGTSFTINSKTYDFSNLKCNSYPYHVARKSGEKCYDGTKTHIEIGYEVG

[1067] SDFYRVIDICFDESNLNALYSRFDLVSGIGGYQSGFPRPSFIQDSFYPGLSVNDLYTKNTQRE TISKLLGSTTLGNQYIASSGDYFMARGHLTAKADFVYGSVHRATFHFVNVAPQWQTFNGQN WNSLEMSVRTYADKNKLDLEVYTGTHGIATLPNVNGVETELYLYVDDNNNKAIPVPNLYWK

[1068] AVYNPKTQAGIVFVGINNPYVSNPKGDYLICQDVCSKVSWLQWDQTNIAKGYSYCCEVNDF RSKVDTLPQFTVTSLLT SEQ ID NO 142: A0A7J6Z4A4

[1069] MSYTPFLFLVFIPSQIVGGLLPKAGCAIDTNKDLGDPQPLLILQNDAETDIEAFVLPDDISGVIN FNKGETFDIVCPEGKWIDGTTTNTDVLEAACKSSLNFIIGGKSVLFPKITCTKQPGHTARYTG KKCASKYKEIEIGFHVQDRFIRHITTCFDEDLQHVLYSENYLVFNVAGSQVNFPRPDFQVSDF YNVKPDSVSDLYSKKGQRKTINGILGLDESDPSIIQPSGNTYLATGHYTPKADFIYGSQQRLT FYYVNAAPQWQSFNGGNWNTMEANYRKLAIDRELDFTIYSGSYGIATLPDVNGDEQELYLW IGADGQKGIPVPALFFKWYEPISEAGWLLGFNNPHKEFDESYMICKNVCSKITWLTWKPEN SNLGFGYCCEVNDFRKTVKYLPKFTVRSLLI

[1070] SEQ ID NO 143: A0A6P4G4S6

[1071] MHELRLLIITLSLFLSLGQVKPFCQLGENYIWTNHIFAHWNNRYELLPADRFQYNQQIKLLCG GNAPVFSTTCQSNGQFNPPLPTLNCSKEIKSSIETVARDPTCAFTLYKVGFPFGNTFLEVYR SCYDAKTMTALFSIHKVYPTHLTSYRRKSWDRDNLISPADEALFTKEKIYERFKTLLGKGQTY IPSKSTDSFDRGHLTPSADYTFYKVLGLTNKYLNVIAQSSSINRGRWAQIEKWVRDQVSNGQ YDVLKVCTGGLEVLELNDTHQNPIQIYLGPRKLPVPKWTYKIVSHSSGHNWILMTNNGWEK NALNPSSVCRRVSCPQGINTAGNTFCCDHFDFITRNVPWLTGVC

[1072] SEQ ID NO 144: A0A6P4GEK2

[1073] MMSMVHHTAIKDSSGQYVYVPKGHKPFCQLTQADILTDRIFAFSQYFNGFELLLRDRLYTFE NMFMLCDANGGDTVFMTTCQPDGTFDPPLLRSNCGVPIEPSFQKIADFSCEHTMFLVGFSF GPTFMELYRNCYDIKTATAHFSINKVYPIYLSSDRPLTGFHRDNIISAADADSFQKKNIYQRFK ALLGPRQPYIASTESSSFDRGHLTPVGDYFFPRIMKQTNKYINVAPQYYSINRGNWKTVETW IRKTVKEIQDDVLNVCTGALGVLKLNNNHQQPTQIYLAPSKIRVPRWIYKIIRSDTTYRKYVILT SNNGWETYRPNSSSVCQEVACPLSLYTTATGYTFCCDPTDFIRRNPKLAGVC

[1074] SEQ ID NO 145: A0A6A4K1 N4

[1075] MLRLAILLTAACAAHGATYERERAQVSKGSGGCDIDVRLDLSGKSPPVILVEEKGEWKLAEP TSRTKYSSYMNLDNKKILLVCPGERNMLKLGSQTARDATATCKSANGVKYFWNGKEHKSD SIKCSERVQPYEYNTKESCNGDGTILKLGYLVNNKIADLITTCHRLADGTTYWSRHILNGGAL PERGNGGKRPDEEGRPQFSFRNKDLFRGFPPYKYYKQANELNALAQAIGSSEAREIFDEKT DKFFSKGHLSPNADFLLETWRDVTFLFVNVQPQWQQINGGNWNEVEVANRYNAAATGKKY EVITGAPGNLAPKSKNVYLDYSSQKMPVPSTYTKLIREVESNKCVAMVSTNFPESKVQPKCK DICSAYKWPELKDAQKAGYVYCCSYDEFAKAFPKSAPKVDCSGGILTNMNRG

[1076] SEQ ID NO 146: A0A1W4U4I6

[1077] MKLRKYQCVCLVLFICAENAWTSCLIDQAHLLENFIYLNSNNGVFDIQRSDIVANDKTVYLLC RSGLHPTSFHCGRDNDFHPPLSTARCSSPRASWPMPDTSCLRLYSSYAVGFYYNGHFME LFRDCIDRDRMSVQHSIYKVHRYINSAPRPSSTYFNRDGFMSPQLAAAYKRKVSEDCLTNIL GGPQQNCVFDRGHVTPNAGFIFSELRRSTDRYINVFPQNSAVNIGNWKNVEAWVGKLVTG HYDSPHRTYDLLKVCTGVLGQQQLEHTTSNSLVDIYLAGNQIPVAKWSYKIVGHLSGDKWV MLTHNQVARPTLPEINQVCINVKCPDGLKEDGVGHTVCCEPYDFIQRNVAHLTGVC

[1078] SEQ ID NO 147: A0A6P4GTH4

[1079] MNELKFLCIAQSLFFFVGTASATCIIHEDLLKIQNIYLSNTNWVYDLQRTDIVPHGQMIFLLCH GGMHPLGFQCQHNVFNPPLETVNCSSVLKASWNVPDTSCPTPSLTYAVGYYLKDHFMEL YRNCYDKEKLALQHSIYKTYRYTKSATRPTSVHWQRDLLMSAQEWSFKRQVSQACVSTTL GANQPNCVFDRGHLTPSSAFIFTEFKKATFRYLNAVPQYNLVNVQNWEYIEAWVTRLVKGN YDNRYRTYDVLKVCTGALGVHDLKHSNKRITRRVPIYLHNDIKIPVPKWMYKIVSHLSGDKW VMLTYNDVNQPFPRDLNQICITTACPEGLNRNGVGFTVCCKPLEFIERNIVHLTGIC

[1080] SEQ ID NO 148: A0A7J7A0U5

[1081] MSIQIVFVTVIATSVCVDAITINKGCSILFPKKDDRLQPLLLRNISKWTYDFFIPEGKTLNMEVK EPITFLCPSKGNVIEGVNTNMTSYECNNKELLSKNGNTVPFDSIQCKSIVKGSVITTSRRCGN NMGKIMNIGYQVTRTQFLTLFESCYDSKRVNPIYTNHTIYGQAVPALSRQSFRPPFSEEGTP YNLQFLSRGHLSPDADFIFAPMQFTTYFFVNVNPHWQSINAGNWLKIEAMVRKLAGKLKIPL

[1082] QIFTGSHDVLSLPDINNNPVEIYLDSDNRIPVPKYLWKWYNEATQEGIALIAINNPFLKSVSA

[1083] HDIICPNICDEYYWDSPAFKTIHKGYLFCCNVMELKRWPTIPNIKIRNVLSA

[1084] SEQ ID NO 149: A0A2P8ZAP5

[1085] MRLLLAIIVLHSLLFQEGKTGLYSDHNLFIQLRSGDCTFRITADMGRDSPLILSADGQRFISPK RQTDILLFRHDEKVLLACPGSKNKFWQGATQHYREVIATCAQNNTFSLGGDSANFRYIKCA KTPDSSILDAREDCGGAKRHKKLFVGFRLPVTQRFLKTVELCYDPIRHSTLYTKFEVSKGVK SGQNIPRASYFRTEVFRDNLRPNFASIYKGHDDRLRALLGSEREGEDLQKGHLAAFGHFPY YSQQDATMYAINTVPQWDKVNMGNWKRLEDRIRSLASNSDQDLHVIAGAFGVLKIKGVEIFL APPSAVPVPQLIYAWYDQAMSHGWLIWNDTRNPSPRKICRDVCNRNKSLMKNWPERSD ASKGFLYCCTLREFNQKTKLAFPDKPLLE

[1086] SEQ ID NO 150: A0A7R8UVR0

[1087] MEYFLAGVIPSNTNVIVSCGPNYLKKQFKTEFVNVQCRGDGQFELKNGFEVKTAKDLGCDL RSVQEILTEVKGCPRDKTAVKFGLVNPVDHKAYIIGEACYCETAGKLLFAHIRRGDLLSYSLEI KDKDIFKAPAPSSTYKVDFMKAFRLDEFNNRLKTALKVEQVPLFDMRNFVDDLFLPNRQLYS IKKLAWNYFVSHEPLLNYQLLKQDIADLEGNIDIYTGAHGVTTLKNKSGGKVPIYLDLEEKRFP VPELIWIWRHEKGEAAFMVFNDPSLNSELIEDKLNIRSKCNYMSWLKRLKEHNKHSIGRNG FWCCDVKELAEEIPEFPLHIYMSSKNID

[1088] SEQ ID NO 151 : A0A1 I8NK31

[1089] MDFKSLLLIAILGCYFCAVTHVSGDCYFSFPKNDSNRPVLYKKIGSRKTLIHTDGLSRYDINDG EVITIDCETRILSPIEGEGLRSFDLNCTDSRMPVFGYPVSNIKVWCDSVKWNLYESSEHFDW CPPPMSSYLLARPLDNVYEYLAGVCYNFEQQQIQSLYNAAGYQFSKYKHPSRLENYSPLVEI KDIPKKFIPRRIDASQFSSEETREFMQFAKYENHAVIQDPQLYKDYFDRFGGLLEMDWWPSL RLGNWYFYEKALRQHIEDDKAIYDILAGVSGAVTVPSQNKIHPKNYTMIDMFYGYNQRVPLY VWHYLKSPRENGKDLWIGVNSAFIDFYNDENDLIFCTDICHQIDWLKMVRSTFRYKTMGLIF CCDVNEVRRSSHLEGFPLASE

[1090] SEQ ID NO 152: A0A6P4G5L6

[1091] MDPVTIVLIFGALAALYKTSLPYPKCQINVSSSAPLLVTNIGSQIILTDYYGLIERNKSEEIQFYC GTGFTFKNDGISQLISDNRMETLICQPDGSFLLQNHGIKIKGDARSVECQNGVAAMFESRIGL PNCKDHTTLLLGNDFQDMGSTKSAALCYDIAGTNLKYLTYTTHPTRSRWKKTHLGDLNKLG FDLSVDGSDRFFKKASQADVDAFWNKDKVLSQMFGTGPFDYASLVQDEALGAQLAGYEG MMSWWLHSLRTGNWRHWLAALRSASVSGKQFEVRLGVSGVLEMPESRDGCDLAIDLAD GNSLPVPVHIWAHVRDLQPTGAAQDEFVLVGHNSPFLRGDPSAEFCPSVCDEVSWLKGTL FASLHRYPINGLMLCCRVEDVAQKLDSFYGSTAHAAATTENLKVAEDLVLYELQRK

[1092] SEQ ID NO 153: A0A6B2EAJ4

[1093] MLKAIFSAIFLLGIIFQSQGLSLGSDSNCHLPLKPVAYGFFAPLVFNADDGKFFGQFSFTTNDT ILLSAGDSVIVSCMPGYFKSFPKLKFLKAKCVEGEVFELEDGKRESWSKSFACELRWEEIIA PRLEGCPEAAQSIEFGFINPHTRTSNIVGEACYSVQEGRTIFAHMRDPSWRIEDTRYLSSGR HPEGRDKINLFRALRQDRVNDLLQERMGRHGLPMIGSRPLLTGQMLDYPQLHLITRLTWNY AITHNDDSMQGWNNLQSDVEKQSRKEDVEVWVGSSGVQTLKDLSGQSFDFYLDEKKFPV PKYLWLWKIADKTSGFLFSNTPHNLKTEICSDSCSAIDWISHPEANDLQCCSLEKLRQIVPEI PQNQISPPPNAERSV

[1094] SEQ ID NO 154: B4N2D1

[1095] MRALIILAWFLLMDSSKAECQIEMLNPTPYILNEFGSYHFMAESFGVINRGEGQNLKLYCPN GFKAVEDYNRREWIPDSMLSLRCDDYFRNKNDEPYRTISCVDGKKSEMFESRKKLANCEQ SMTYVMGQNFNNLGSSKSLALCYDIVELRLKYIAYTAYLGNQKIIQNHQIGQLNTLGLDINVA YTNEIFQPVSPIAIDNFNVMYREVFGYNAYEYANLIQDKPLTNQFTEYEDMLSIVWLRNLRTG NWYKWLNALNEATKTGHKFDVRLGVSGELQLPKFANQCLNRTLSIVGDNTVKISVPKLIWAH VTEILPTNDTTNDIVIIGHNTPFITDEMTKFCTDKCNKVSWLKDTMFLNLRQYAAFGLVQCCQ

[1096] VDEIVNQLDNFPMAILEDSGESVETN

[1097] SEQ ID NO 155: A0A1 B0C8D8 MLSVYAILATLAVTFPGHFSAEIGSNCNLPFKPVPYGFFAPLIYGADEKFSEASTNESTLLEP QESVTLSCEPGYFKSYPKTRTLKAKCIGGEELQLEDGTKEQWSKNFACDLRPVEEVLTPNIP GCPPEAEGTEFGYINPFTKKSIIIGEACYNVTKGQIIFVHAKTHPGLKIEDQNYLQRGSHPESR NKLEFFKALRYDAIYDRFKNVFGPPEKVLRIIYRSLLTAQMLSLPQLHLVNRLTWNYAISHDD ESLAGWTALRDGVEKFTAENAAADVWVGTSGVAKWDVAGTAWDFYLQHPKFPVPKYHWI IVKVNNKATGFLFYNIPPFLNTRSYGNEEPKILEKICQRKCKEFSWLEGVEGIRCCSVEEMRQ VIPEIPPISGVLGELPPKQEAPQGESMIS

[1098] SEQ ID NO 156: A0A2P8ZIZ4

[1099] MAAGALRMNKFSTLSLLIILYVSAILPEVTVKATEVKTPVRGDKAQGCAIHVQNDLNWNTPILL NKDMQILLPDVSSGYVTLPVKEEVTFVCPGDGNHFLKPFHPKNVTTWGRCVNGKIFDIGVK MNIKEITCKMLPESILRNISSSCNRKNYTLAEIGFLVNKTTFINQIKVCFDEQHCKPVYTKHELH PTMKFHQTGKSFSFRKVETYLFGELGEGPGDLYPCTRQKEVLAELLGSKKQVHEYDRCPE DRSTQLVKGHLSPNADFTFSFQQNATYYHANAAPQWMFFNNRNWKSLEERIRSQAYIRNR

[1100] SLVIYTGTSDEPTFLPDVNGNNVEIFLSQINNKLPSPLYYWKLVYDPKTEKGIGFIGYNVHFD QDPPFFTQPKYSSICEKSGWFRKSEITLDKGLIYCMSVKDLITLTKLEDFKDKISGDMEYELTS DKYTRKLQEELPKSSCGELWSIGFKIVFVTILAMY

[1101] SEQ ID NO 157: A0A1W4VJ01

[1102] MWKLYSVSLCLSVLCGIVQDVLGSCQLRIDSIGSPPLIVNRFGTKTMLSQSWGVITREEGESI ELLCGGGVTYNKNHGRTTKTGNGEKLTLECSRDGYFRDPVENYNLRELSVGCHEGIYQLFE SNSSLPNCEGDMTLVLGHDLKELGSKNIAALCYDIVASRLKYIAYTTFSASNQVLGAEVGQL NDVELNTKVNYRKSTFKPVRQTDIDAYVANVDQLAGLFESASLVQDNGMEAMVTGYEDMM TTTWLRSLRSGNWRHWIAAMRLAARRGLHFDVRLGASGELQLPPSIGRGPCNATRPMLIPL

[1103] AGSVGGDTVRVPAHIWAHVHALEPTGGVQDEIVIIGHNSPFVRSGSQSDLCSSMCDQVSWL QQDSLFASLHEFAIYGLVHCCRVEDVATKLDHFPGPYAKEKNKLGGARGAEHVASTTDSNS L

[1104] SEQ ID NO 158: A0A7R8V2P4

[1105] MFRVIVLLLCGTGAIASKCTLRYQLSSPPILTQKFGSHNIILNARDGQLSWDDRATVNAYCPG GLRNLRQHPNYGYYNPSSSQQVTFSCSSSQVMFDRPDTGVITSFQNDGQLSCPTTSKFYE TLVPACEYQGLAYGFLLGKQAWLAEVCYNLDLLEPMFLHYVAGARSTILESQTAHNVANITT SSVEVFIDNSLRESENRTLLEAVRKSLPQYLGWQYQIEEFIPLREASGVFGLYVDDFRSVN MIPWWRSLKFGNWKLLEEAIDMVSRNETYDVFAGTSDMYPSNDKCMSTDLLTFKVDSYIR

[1106] NVPKYIWNYVKKREEADDGIVIIGINSPFSDSSKGGDVLCDDICDNIPWLSSLKRSRKMAALG YIFCCKPADVKNKLDHFLI

[1107] SEQ ID NO 159: A0A7R8UH49

[1108] MWRQIVFVCLIIQAVAGTCRLDFKSKAPGVFTRKFGSNNIILSAADGEISWDDNMIVNAHCRN GFTDYSYYTSTPEKDLILSCSGSTVYYKTSEYDATNQVRSGYMSCQTESTYYEATISACKEY GLAYGLFVGKQPVILADVCYNLNTMRTQFVHYVAGPRSPWEIQSTFNPSNATVGSNVSVN AFFDSEAIENENNSLVKAVQDKMPFYSEIVQYQLQELVPLDKSAGILSPYTKDFINTNTIAWW KPLKFGNWKSVEEAIAQVSASSSFDVFAGTSGWMYPLDKRCSFTEKLTYRADQYQRNVPL

[1109] YVWNYVKSRDDKTDKGWIIGVNSPFFESGNKPDIPCEDICDSISWLSSVNNVRKMSSLGYV FCCKPGTIKLENFPDL

[1110] SEQ ID NO 160: A0A7R8UG29

[1111] MNLSLILAVIFVSFVTSDVFAATCTLKFNKAYPGVFVSPYGRYDVILSAGNGTVQFDSNYGLD AYCSGGFKNYASASTPKPESQLKFTCTGSTLRYQIPGSSFTYLAKNDGSLLCDNAAQYYSQ TVQYCKNTGLAYGFNIANTTSVIFGEVCTDFKKYRTDFVHYVAGVRNDLVNEQYLLNPGNVS KPLPANNVYAVNDTTFIKENQTLPTAIRSKSGNYAQLMTYRMDNLIPIKYQAGVMDSYVRNF QNAAAIPWWSNLLDFNWQQLQYLINDLSYSGWYDVYAGTSDDVQMPYETKNVTFTYDVNQ LNRTIPLYVWNLVKKNDKKDTGWVIGVNSPFFPGDKSKTVLCKDTCDKVTWLNPLKNTRKI ASMGYIFCCDPKDVKKILNGFPQM

[1112] SEQ ID NO 161 : A0A7R8V4D0

[1113] MFNLKVLWLLVSAFLAEGCELSLDKNRPGVFLQQFSSKKLVLDVSGGSVSFAEGSTIQGYC SSGFRNLLRNQQSYNQSFTNVTLLCQNGDIFYLSGDSLEPVESRYSLACYESQASFYIGPVD YCKKLGLVYGMNIDGTPIIFAEVCYDIENMEVDFIHFVMGERPIILGKQVDVSPNNITYQLQKD QLFTFAEQNLKTLDSALRAGVKKELPNYEKIAQYSFDDLTPVRSYAGRFRPFTDSFEAVNLIP WWSTLKSDNWKTFYGILEKLSVDGPIDIYAGTSKWRYPADDGCYEMKNFTYKVDATTSET VPLQIWNYIKPRNQNISEWWGINSPFLGFISKSPDIFGTGNCKDTVWLDPFMEVRRMPALG YTFCVTVQEAAENLKGFPAV

[1114] SEQ ID NO 162: A0A7R8YPE0

[1115] MVRLLLFTVLAVNFYIIGVLAETCTLKFNKAKPGLFTAKVDKNDIIQDPQKGFIKFDSSSGLDT YCSTGFRNYAAITTSKLVYQLKFTCSAGSLMYQTPEGTSMELANNDGNLLCEDEAQYYIQTV QYCKNTGLVYGFNVGNTTSISFAEICYDLKNYRAEFVHYVAGVRTKLVDNQFDFNPNNSTKS...

Claims

CLAIMS1 . A method of producing nucleic acid fragments, the method comprising fragmenting chromatin within a population of permeabilised cells with a CATH Superfamily 3.40.570.10 Extracellular Endonuclease, subunit A nuclease, or a nuclease variant thereof.

2. The method according to claim 1 , wherein the nuclease comprises a cd00091 active site.

3. A method of producing nucleic acid fragments, the method comprising fragmenting chromatin within a population of permeabilised cells with a nuclease, wherein the nuclease comprises a cd00091 active site.

4. The method of claim 3, wherein the nuclease is a CATH Superfamily 3.40.570.10 Extracellular Endonuclease, subunit A nuclease, or a nuclease variant thereof.

5. The method according to any one of the previous claims, wherein the nuclease comprises an active site having the amino acid sequence of A / S / D / Q / T / E-K / R / Y / l / H-X- H / Q ' / F-X(a)-N / F / G-X(7)-E / L / M / H / Q-X(3)-R / M / LA / / F, where X is any amino acid and a is 23 to 37.

6. A method of producing nucleic acid fragments, the method comprising fragmenting chromatin within a population of permeabilised cells with a nuclease, wherein the nuclease comprises an active site having the amino acid sequence of A / S / D / Q / T / E-K / R / Y / l / H-X- H / Q ' / F-X(a)-N / F / G-X(7)-E / L / M / H / Q-X(3)-R / M / LA / / F, where X is any amino acid and a is 23 to 37.

7. The method according to any one of the previous claims, wherein the nuclease has a ball portion and optionally a chain portion.

8. The method according to any one of the previous claims, wherein the nuclease has a ball and chain morphology.

9. A method of producing nucleic acid fragments, the method comprising fragmenting chromatin within a population of permeabilised cells with a nuclease, wherein the nuclease has a ball and chain morphology.

10. The method according to any one of the previous claims, wherein the nuclease comprises an active site having the amino acid sequence A / S / D-K / R-X-H-X(29)-N-X(7)-E-X(3)- R, where X is any amino acid.11 . The method of any one of claims 1-9, wherein the nuclease comprises an active site having the amino acid sequence:(a) A-K-X-H-X(29)-N-X(7)-E-X(3)-R, where X is any amino acid;(b) A-R-X-H-X(29)-N-X(7)-E-X(3)-R, where X is any amino acid;(c) S-R-X-H-X(29)-N-X(7)-E-X(3)-R, where X is any amino acid;(d) D-R-X-H-X(29)-N-X(7)-E-X(3)-R, where X is any amino acid;(e) D-R-X-H-X(3i)-N-X(7)-E-X(3)-R, where X is any amino acid;(f) D-R-X-H-X(30)-N-X(7)-E-X(3)-R, where X is any amino acid;(g) S-Y-X-F-X(28)-F-X(7)-Q-X(3)-V, where X is any amino acid;(h) D-R-X-H-X(37)-G-X(7)-E-X(3)-L, where X is any amino acid;(i) Q-R-X-Q-X(26)-F-X(7)-L-X(3)-M, where X is any amino acid;(j) T-H-X-F-X(28)-F-X(7)-H-X(3)-L, where X is any amino acid;(k) T-Y-X-Y-X(23)-F-X(7)-M-X(3)-L, where X is any amino acid;(l) D-K-X-H-X(36)-N-X(7)-E-X(3)-L, where X is any amino acid;(m) D-R-X-H-X(35)-G-X(7)-E-X(3)-L, where X is any amino acid;(n) E-l-X-Q-X(26)-F-X(7)-L-X(3)-M, where X is any amino acid;(o) S-R-X-H-X(30)-N-X(7)-E-X(3)-R, where X is any amino acid;(p) S-Y-X-F-X(29)-F-X(7)-H-X(3)-L, where X is any amino acid;(q) D-R-X-H-X(33)-N-X(7)-E-X(3)-F, where X is any amino acid; or(r) D-R-X-H-X(33)-N-X(7)-E-X(3)-L, where X is any amino acid.

12. The method of claim 11 , wherein the nuclease comprises an active site having the amino acid sequence:(a) A-K-X-H-X(29)-N-X(7)-E-X(3)-R, where X is any amino acid;(b) A-R-X-H-X(29)-N-X(7)-E-X(3)-R, where X is any amino acid;(c) S-R-X-H-X(29)-N-X(7)-E-X(3)-R, where X is any amino acid; or(d) D-R-X-H-X(29)-N-X(7)-E-X(3)-R, where X is any amino acid.

13. The method according to any one of the previous claims, wherein the nuclease comprises an active site having the amino acid sequence A-K-X-H-X(29)-N-X(7)-E-X(3)-R, where X is any amino acid.

14. The method according to any one of the previous claims, wherein the nuclease comprises an extracellular endonuclease subunit A domain.

15. The method according to any one of the previous claims, wherein the nuclease comprises an amino acid sequence having at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 99%, or 100% identity to the amino acid sequence of an extracellular endonuclease subunit A domain defined in Table 1 or Table 2, optionally wherein the nuclease comprises an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 99%, or 100% identity to the amino acid sequence of an extracellular endonuclease subunit A domain defined in Table 1 or Table 2, such as at least 90%, at least 95%, at least 97%, at least 99%, or 100% identity to the amino acid sequence of an extracellular endonuclease subunit A domain defined in Table 1 or Table 2.

16. The method according to any one of the previous claims, wherein the nuclease comprises an endonuclease_NS domain.

17. The method according to claim 16, wherein the nuclease comprises an amino acid sequence having at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 99%, or 100% identity to amino acids 136 to 385 of SEQ ID NO: 1 or SEQ ID NO: 2, optionally wherein the nuclease comprises an amino acid sequence having at least at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 99%, or 100% identity to amino acids 136 to 385 of SEQ ID NO: 1 or SEQ ID NO: 2, such as at least 90%, at least 95%, at least 97%, at least 99%, or 100% identity to amino acids 136 to 385 of SEQ ID NO: 1 .

18. The method according to any one of the previous claims, wherein the nuclease comprises a non-specific endonuclease domain.

19. The method according to claim 18, wherein the nuclease comprises an amino acid sequence having at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 99%, or 100% identity to amino acids 50 to 389 of SEQ ID NO: 1 orSEQ ID NO: 2, optionally wherein the nuclease comprises an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 99%, or 100% identity to amino acids 50 to 389 of SEQ ID NO: 1 or SEQ ID NO: 2, such as at least 90%, at least 95%, at least 97%, at least 99%, or 100% identity to amino acids 50 to 389 of SEQ ID NO: 1 or SEQ ID NO: 2.

20. The method according to any one of the previous claims, wherein the nuclease comprises a DNA / RNA non-specific endonuclease domain.21 . The method according to claim 20, wherein the nuclease comprises:(a) an amino acid sequence having at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 99%, or 100% identity to amino acids 136 to 386 of SEQ ID NO: 1 or SEQ ID NO: 2, optionally wherein the nuclease comprises an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 99%, or 100% identity to amino acids 136 to 386 of SEQ ID NO: 1 or SEQ ID NO: 2, such as at least 90%, at least 95%, at least 97%, at least 99%, or 100% identity to amino acids 136 to 386 of SEQ ID NO: 1 or SEQ ID NO: 2;(b) an amino acid sequence having at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 99%, or 100% identity to amino acids 145 to 386 of SEQ ID NO: 1 or SEQ ID NO: 2, optionally wherein the nuclease comprises an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 99%, or 100% identity to amino acids 145 to 386 of SEQ ID NO: 1 or SEQ ID NO: 2, such as at least 90%, at least 95%, at least 97%, at least 99%, or 100% identity to amino acids 145 to 386 of SEQ ID NO: 1 or SEQ ID NO: 2;(c) an amino acid sequence having at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 99%, or 100% identity to amino acids 136 to 385 of SEQ ID NO: 1 or SEQ ID NO: 2, optionally wherein the nuclease comprises an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 99%, or 100% identity to amino acids 136 to 385 of SEQ ID NO: 1 or SEQ ID NO: 2, such as at least 90%, at least 95%, at least 97%, at least 99%, or100% identity to amino acids 136 to 385 of SEQ ID NO: 1 or SEQ ID NO: 2; and / or(d) an amino acid sequence having at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, atleast 90%, at least 95%, at least 97%, at least 99%, or 100% identity to amino acids 192 to 376 of SEQ ID NO: 1 or SEQ ID NO: 2, optionally wherein the nuclease comprises an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 99%, or 100% identity to amino acids 192 to 376 of SEQ ID NO: 1 or SEQ ID NO: 2, such as at least 90%, at least 95%, at least 97%, at least 99%, or 100% identity to amino acids 192 to 376 of SEQ ID NO: 1 or SEQ ID NO: 2.

22. The method according to any one of the previous claims, wherein the nuclease comprises a His-Me finger superfamily domain.

23. The method according to any one of the previous claims, wherein the nuclease comprises a DNA / RNA non-specific endonuclease superfamily domain.

24. The method according to any one of the previous claims, wherein the nuclease comprises a NUC domain.

25. The method according to claim 24, wherein the nuclease comprises an amino acid sequence having at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 99%, or 100% identity to amino acids 146 to 394 of SEQ ID NO: 1 or SEQ ID NO: 2, optionally wherein the nuclease comprises an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 99%, or 100% identity to amino acids 146 to 394 of SEQ ID NO: 1 or SEQ ID NO: 2, such as at least 90%, at least 95%, at least 97%, at least 99%, or 100% identity to amino acids 146 to 394 of SEQ ID NO: 1 or SEQ ID NO: 2.

26. The method according to any one of claims 7-25, wherein the ball portion of the nuclease has:(a) a length of 50 A to 100 A along a first axis;(b) a length of 25 A to 75 A along a second axis; and(c) a length of 25 A to 75 A along a third axis, wherein the first, second and third axis define the longest height, width, and depth of the ball portion.

27. The method according to claim 26, wherein the ball portion of the nuclease has: (a) a length of 65 A to 90 A along a first axis;(b) a length of 30 A to 60 A along a second axis; and(c) a length of 30 A to 60 A along a third axis, wherein the first, second and third axis define the longest height, width, and depth of the ball portion.

28. The method according to any one of claims 7-27, wherein the ball portion of the nuclease has a first axis, second axis and third axis, wherein the ratio of the lengths of the first axis : second axis : third axis is 50 A to 100 A : 25 A to 75 A : 25 A to 75 A, wherein the first, second and third axis define the longest height, width, and depth of the ball portion.

29. The method according to any one of claims 7-28, wherein the ball portion of the nuclease has a first axis, second axis and third axis, wherein the ratio of the lengths of the first axis : second axis : third axis is 65 A to 90 A : 30 A to 60 A : 30 A to 60 A, wherein the first, second and third axis define the longest height, width, and depth of the ball portion.

30. The method according to any one of claims 7-29, wherein the ball portion of the nuclease comprises an amino acid sequence having at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 99%, or 100% sequence identity to the amino acid sequence of a ball portion as defined in Table 1 or Table 2, optionally wherein the ball portion of the nuclease comprises an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 99%, or 100% sequence identity to the amino acid sequence of a ball portion as defined in Table 1 or Table 2, such as at least 90%, at least 95%, at least 97%, at least 99%, or 100% sequence identity to the amino acid sequence of a ball portion as defined in Table 1 or Table 2.31 . The method according to any one of claims 7-30, wherein the chain portion of the nuclease has:(a) a length of 50 A to 100 A along a first axis;(b) a length of 1 A to 25 A along a second axis; and(c) a length of less than 10 A along a third axis, wherein the first, second and third axis define the longest height, width, and depth of the chain portion.

32. The method according to claim 31 , wherein the chain portion of the nuclease has:(a) a length of 60 A to 80 A along a first axis;(b) a length of 1 A to 15 A along a second axis; and(c) a length of less than 5 A along a third axis, wherein the first, second and third axis define the longest height, width, and depth of the chain portion.

33. The method according to any one of the claims 7-32, wherein the chain portion of the nuclease has a first axis, second axis and third axis, wherein the ratio of the lengths of the first axis : second axis : third axis is 50 A to 100 A : 1 A to 25 A : less than 10 A, wherein the first, second and third axis define the longest height, width, and depth of the chain portion.

34. The method according to any one of claims 7-33, wherein the chain portion of the nuclease has a first axis, second axis and third axis, wherein the ratio of the lengths of the first axis : second axis : third axis is 60 A to 80 A : 1 A to 15 A : less than 5 A, wherein the first, second and third axis define the longest height, width, and depth of the chain portion.

35. The method according to any one of claims 7-34, wherein the chain portion of the nuclease comprises an amino acid sequence having at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 99%, or 100% sequence identity to the amino acid sequence of a chain portion defined in Table 1 or Table 2, optionally wherein the chain portion of the nuclease comprises an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 99%, or 100% sequence identity to the amino acid sequence of a chain portion defined in Table 1 or Table 2, such as at least 90%, at least 95%, at least 97%, at least 99%, or 100% sequence identity to the amino acid sequence of a chain portion defined in Table 1 or Table 2.

36. The method according to any one of the previous claims, wherein the nuclease has a molecular weight of:(a) greater than 17 kDa;(b) at least 20 kDa;(c) at least 30 kDa;(d) at least 40 kDa; and / or(e) up to 50 kDa.

37. The method according to any one of the previous claims, wherein the nuclease is:(a) an endonuclease;(b) a non-specific endonuclease;(c) a DNA / RNA nonspecific endonuclease; and / or(d) an endonuclease which does not have exonuclease activity.

38. The method according to any one of the previous claims, wherein the nuclease is a nuclease selected from Table 1 , Table 2, or Table 3, or a nuclease variant thereof.

39. The method according to any one of the previous claims, wherein the nuclease has an amino acid sequence having at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 99%, or 100% sequence identity to the amino acid sequence of a nuclease shown in Table 1 , Table 2, or Table 3, optionally wherein the nuclease has an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 99%, or 100% sequence identity to the amino acid sequence of a nuclease shown in Table 1 , Table 2, or Table 3, such as at least 90%, at least 95%, at least 97%, at least 99%, or 100% sequence identity to the amino acid sequence of a nuclease shown in Table 1 , Table 2, or Table 3.

40. The method according to any one of the previous claims, wherein the nuclease has an amino acid sequence having at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 99%, or 100% sequence identity to SEQ ID NO: 1 or SEQ ID NO: 2, optionally wherein the nuclease has an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 99%, or 100% sequence identity to SEQ ID NO: 1 or SEQ ID NO: 2, such as at least 90%, at least 95%, at least 97%, at least 99%, or 100% sequence identity to SEQ ID NO: 1 or SEQ ID NO: 2.41 . A method of producing nucleic acid fragments, the method comprising fragmenting chromatin within a population of permeabilised cells with a nuclease comprising an amino acid sequence having at least 40% identity to SEQ ID NO: 1 or SEQ ID NO: 2, optionally wherein the nuclease has an amino acid sequence having at least 70% identity to SEQ ID NO: 1 , such as at least 90% identity to SEQ ID NO: 1 or SEQ ID NO: 2.

42. The method according to any one of the previous claims, wherein the nuclease has the amino acid sequence of SEQ ID NO: 1.

43. The method according to any one of the previous claims, wherein the nuclease has the amino acid sequence of SEQ ID NO: 2.

44. The method according to any one of the previous claims, wherein the method comprises permeabilising a population of cells prior to fragmenting to produce the population of permeabilised cells.

45. The method according to any one of the previous claims, wherein cells in the population of permeabilised cells comprise pores in the cell membrane having a width of up to 10 nm, such as 8-10 nm.

46. The method according to any one of the previous claims, wherein the number of cells in the population of permeabilised cells is 1-10,000; 10,000-100,000; 100,000-250,000;250,000-500,000; 500,000-1 million; or 1 million to 100 million.

47. The method according to any one of the previous claims, wherein the chromatin is contained by permeabilised nuclei of the cells.

48. The method according to any one of the previous claims, wherein the method comprises fragmenting the chromatin with a fixed concentration of the nuclease.

49. The method according to any one of the previous claims, wherein the method comprises fragmenting the chromatin with 0.5-40 Units of the nuclease, optionally 20 Units of the nuclease.

50. The method according to any one of the previous claims, wherein the nucleic acid fragments comprise mono-nucleosomes with inter-nucleosome linker nucleic acids attached.51 . The method according to any one of the previous claims, wherein at least 70% of the nucleic acid fragments are at least 145 bp and up to 290bp fragments, optionally at least 145 bp and up to 190 bp fragments, or at least 180 bp and up to 200 bp fragments.

52. The method according to any one of the previous claims, wherein up to 10% of the nucleic acid fragments are smaller than 150bp fragments.

53. The method according to any one of the previous claims, wherein the nuclease is used to fragment the chromatin with a cut site preference profile such that:(a) 15-35% of nucleic acid residues immediately before a cut site are C;(b) 15-35% of nucleic acid residues immediately before a cut site are G;(c) 15-35% of nucleic acid residues immediately before a cut site are T; and(d) 15-35% of nucleic acid residues immediately before a cut site are A, wherein the sum of (a), (b), (c) and (d) is 100%.

54. The method according to any one of the previous claims, wherein the nuclease is used to fragment the chromatin with a cut site preference profile such that:(a) 10-40% of nucleic acid residues immediately after a cut site are C;(b) 10-40% of nucleic acid residues immediately after a cut site are G;(c) 10-40% of nucleic acid residues immediately after a cut site are T; and(d) 10-40% of nucleic acid residues immediately after a cut site are A, wherein the sum of (a), (b), (c) and (d) is 100%.

55. The method according to any one of the previous claims, wherein the chromatin is cross-linked chromatin, optionally wherein the chromatin is formaldehyde cross-linked chromatin.

56. The method according to any one of the previous claims, wherein the method comprises cross-linking the chromatin prior to fragmenting, optionally wherein the method comprises cross-linking the chromatin prior to permeabilising the cells.

57. A pool of nucleic acid fragments obtainable, or obtained by, the method of any one of the previous claims.

58. A pool of nucleic acid fragments, wherein the pool of nucleic acid fragments is present in a cell, and wherein:(a) 15-35% of the nucleic acid fragments have a 3’ terminal C residue;(b) 15-35% of the nucleic acid fragments have a 3’ terminal G residue;(c) 15-35% of the nucleic acid fragments have a 3’ terminal T residue; and(d) 15-35% of the nucleic acid fragments have a 3’ terminal A residue,wherein the sum of (a), (b), (c) and (d) is 100%; optionally wherein:(e) 10-40% of the nucleic acid fragments have a 5’ terminal C residue;(f) 10-40% of the nucleic acid fragments have a 5’ terminal G residue;(g) 10-40% of the nucleic acid fragments have a 5’ terminal T residue; and(h) 10-40% of the nucleic acid fragments have a 5’ terminal A residue, wherein the sum of (e), (f), (g) and (h) is 100%.

59. A pool of nucleic acid fragments, wherein the pool of nucleic acid fragments is present in a cell, and wherein:(a) 10-40% of the nucleic acid fragments have a 5’ terminal C residue;(b) 10-40% of the nucleic acid fragments have a 5’ terminal G residue;(c) 10-40% of the nucleic acid fragments have a 5’ terminal T residue; and(d) 10-40% of the nucleic acid fragments have a 5’ terminal A residue, wherein the sum of (a), (b), (c) and (d) is 100%; optionally wherein:(e) 15-35% of the nucleic acid fragments have a 3’ terminal C residue;(f) 15-35% of the nucleic acid fragments have a 3’ terminal G residue;(g) 15-35% of the nucleic acid fragments have a 3’ terminal T residue; and(h) 15-35% of the nucleic acid fragments have a 3’ terminal A residue, wherein the sum of (e), (f), (g) and (h) is 100%.

60. The pool of nucleic acid fragments according to any one of claims 57-59, wherein the cell is a mammalian or human cell, optionally wherein the cell is a THP-1 cell or a monocyte.61 . The pool of nucleic acid fragments according to any one of claims 57-60, wherein the nucleic acid fragments are cross-linked nucleic acid fragments.

62. A method of producing a 3C library, the method comprising:(a) (i) ligating cross-linked nucleic acid fragments obtainable, or which have been obtained, by the method of any of claims 1-56; or(a) (ii) ligating the pool of cross-linked nucleic acid fragments according to any one of claims 57-61 ; or(a) (iii) producing cross-linked nucleic acid fragments by the method according to any one of claims 1-56 and ligating the cross-linked nucleic acid fragments; and(b) de-crosslinking the ligated nucleic acid fragments.

63. A 3C library of nucleic acid fragments obtainable, or obtained by, the method of claim 62.

64. A 3C library of nucleic acid fragments, wherein:(a) 15-35% of the nucleic acid fragments have a 3’ terminal C residue;(b) 15-35% of the nucleic acid fragments have a 3’ terminal G residue;(c) 15-35% of the nucleic acid fragments have a 3’ terminal T residue; and(d) 15-35% of the nucleic acid fragments have a 3’ terminal A residue, wherein the sum of (a), (b), (c) and (d) is 100%; optionally wherein:(e) 10-40% of the nucleic acid fragments have a 5’ terminal C residue;(f) 10-40% of the nucleic acid fragments have a 5’ terminal G residue;(g) 10-40% of the nucleic acid fragments have a 5’ terminal T residue; and(h) 10-40% of the nucleic acid fragments have a 5’ terminal A residue, wherein the sum of (e), (f), (g) and (h) is 100%.

65. A 3C library of nucleic acid fragments, wherein:(a) 10-40% of the nucleic acid fragments have a 5’ terminal C residue;(b) 10-40% of the nucleic acid fragments have a 5’ terminal G residue;(c) 10-40% of the nucleic acid fragments have a 5’ terminal T residue; and(d) 10-40% of the nucleic acid fragments have a 5’ terminal A residue, wherein the sum of (a), (b), (c) and (d) is 100%; optionally wherein:(e) 15-35% of the nucleic acid fragments have a 3’ terminal C residue;(f) 15-35% of the nucleic acid fragments have a 3’ terminal G residue;(g) 15-35% of the nucleic acid fragments have a 3’ terminal T residue; and(h) 15-35% of the nucleic acid fragments have a 3’ terminal A residue, wherein the sum of (e), (f), (g) and (h) is 100%.

66. A method of identifying chromatin regions within a nucleic acid sample which interact with one another, the method comprising:(a) (i) fragmenting a 3C library obtainable, or which has been obtained, by the method of claim 62; or(a) (ii) fragmenting a 3C library according to any one of claims 63-65; or(a) (iii) producing a 3C library by a method according to claim 62 and fragmenting the3C library; and(b) optionally, adding sequencing adaptors to the ends of the nucleic acid fragments and / or amplifying the nucleic acid fragments;(c) contacting the nucleic acid fragments with a targeting nucleic acid which binds to a subgroup of the nucleic acid fragments, wherein the targeting nucleic acid is labelled with the first half of a binding pair;(d) isolating the nucleic acid fragments which have been bound by the targeting nucleic acid using the second half of the binding pair;(e) amplifying the isolated subgroup of nucleic acid fragments;(f) optionally repeating steps (c), (d), and (e) one or more times; and(g) optionally sequencing the amplified isolated subgroup of nucleic acid fragments, in order to identify chromatin regions within the nucleic acid sample which interact with one another.

67. A method of identifying allele-specific interaction profiles in SNP-containing regions of nucleic acids, the method comprising:(a) sequencing the amplified isolated subgroup of nucleic acid fragments which has been obtained by the method of claim 66 in order to identify allele-specific interaction profiles in SNP-containing regions; or(b) identifying chromatin regions within a nucleic acid sample which interact with one another according to claim 66 including sequencing the amplified isolated subgroup of nucleic acid fragments in order to identify allele-specific interaction profiles in SNP-containing regions.

68. A method of identifying a chromatin region that is indicative of a disease, the method comprising:(a) (i) quantifying a frequency of interaction between a first chromatin region and a second chromatin region within a nucleic acid sample from a subject with the disease, wherein the first chromatin region and second chromatin region have been identified as interacting with one another by the method of claim 66; or(a) (ii) identifying chromatin regions within a nucleic acid sample which interact with one another according to claim 66 wherein the nucleic acid sample is obtained from a subject with a disease, and quantifying a frequency of interaction between a first chromatin region within the nucleic acid sample and a second chromatin region within the nucleic acid sample; and(b) comparing the frequency of the interaction in the nucleic acid sample from the subject with the disease with the frequency of interaction in a nucleic acid sample from a subject without the disease,wherein a difference in the frequency of interaction between the nucleic acid samples is indicative of the disease.

69. The method according to claim 68, wherein the disease is an autoimmune disease.

70. A method comprising producing an agent which targets an expression product of a nucleotide sequence, wherein the nucleotide sequence, or a nucleotide sequence having at least 90%, 95%, 97%, or 99% identity thereto, is also comprised by a chromatin region which has been identified by the method according to claim 68 or claim 69 as being indicative of a disease.71 . A method comprising producing an agent which targets an expression product regulated by a nucleotide sequence, wherein the nucleotide sequence, or a nucleotide sequence having at least 90%, 95%, 97%, or 99% identity thereto, is also comprised by a chromatin region which has been identified by the method according to claim 68 or claim 69 as being indicative of a disease.

72. A method comprising:(a) identifying a chromatin region that is indicative of disease by the method according to claim 68 or claim 69; and(b) producing an agent which targets an expression product of a nucleotide sequence, wherein the nucleotide sequence, or a nucleotide sequence having at least 90%, 95%, 97%, or 99% identity thereto, is also comprised by the chromatin region identified in step (a).

73. A method comprising:(a) identifying a chromatin region that is indicative of disease by the method according to claim 68 or claim 69; and(b) producing an agent which targets an expression product regulated by a nucleotide sequence, wherein the nucleotide sequence, or a nucleotide sequence having at least 90%, 95%, 97%, or 99% identity thereto, is also comprised by the chromatin region identified in step (a).

74. The method according to any one of claims 70-73, wherein the agent agonises the expression product.

75. The method according to any one of claims 70-73, wherein the agent antagonises the expression product.

76. The method according to any one of claims 70-75, wherein the agent is a therapeutic agent, optionally wherein the therapeutic agent is a chemically synthesised compound (e.g. a small molecule); a botanically available molecule isolated from a plant, fungi, or mould; a biotherapeutic molecule (e.g. an antibody or antigen-binding fragment thereof); or a nucleic acid molecule (e.g., an antisense oligonucleotide).

77. The method according to any one of claims 70-76, wherein the expression product is an expression product of a gene, optionally wherein the expression product of the gene is a protein or RNA encoded by the gene.

78. The method according to claim 77, wherein the protein is a cell signalling protein (e.g., a ligand such as a growth factor or cytokine, or a receptor), a structural protein (e.g., an extracellular matrix protein), a hormonal protein, an enzyme, or a transport protein.

79. An agent which has been obtained by the method of any one of claims 70-78.

80. A method of producing a pharmaceutical composition, the method comprising combining:(a) the agent according to claim 79; or(b) an agent which has been obtained by the method of any one of claims 70-78, with a pharmaceutically acceptable carrier, excipient or diluent.81 . A pharmaceutical composition which has been obtained by the method of claim 80.

82. A method of treating or preventing a disease or condition in a human or animal subject, the method comprising administering an agent according to claim 79 or a pharmaceutical composition according to claim 81 to the subject.

83. An agent according to claim 79, or a pharmaceutical composition according to claim 81 , for use in a method of treating or preventing a disease or condition in a human or animal subject, the method comprising administering the agent or pharmaceutical composition to the subject.

84. Use of an agent according to claim 79, or a pharmaceutical composition according to claim 81 in the manufacture of a medicament for treating or preventing a disease or condition in a human or animal subject.

85. The method according to claim 82, the agent for use or pharmaceutical composition for use according to claim 83, or the use according to claim 84, wherein:(a) the agent has been produced by the method of any one of claims 70-78; or(b) the pharmaceutical composition has been produced by the method of claim 80, and the disease is an autoimmune disease.

86. A kit comprising:(a) a CATH Superfamily 3.40.570.10 Extracellular Endonuclease, subunit A nuclease, or a nuclease variant thereof;(b) a nuclease comprising a cd00091 active site;(c) a nuclease comprising an active site having the amino acid sequence of A / S / D / Q / T / E-K / R / Y / I / H-X-H / Q / Y / F-X(a)-N / F / G-X(7)-E / L / M / H / Q-X(3)-R / M / L / V / F, where X is any amino acid and a is 23 to 37;(d) a nuclease having a ball and chain morphology;(e) a nuclease comprising an amino acid sequence having at least 40% identity to SEQ ID NO: 1 or SEQ ID NO: 2; or(f) a means for fragmenting chromatin; and instructions for use of the nuclease to fragment chromatin within a population of permeabilised cells.

87. An apparatus comprising:(a) means to fragment chromatin in a population of permeabilised cells with:(i) a CATH Superfamily 3.40.570.10 Extracellular Endonuclease, subunit A nuclease, or a nuclease variant thereof(ii) a nuclease comprising a cd00091 active site;(iii) a nuclease comprising an active site having the amino acid sequence of A / S / D / Q / T / E-K / R / Y / l / H-X-H / Q / Y / F-X(a)-N / F / G-X(7)-E / L / M / H / Q-X(3)-R / M / UV / F, where X is any amino acid and a is 23 to 37;(iv) a nuclease having a ball and chain morphology;(v) a nuclease comprising an amino acid sequence having at least 40% identity to SEQ ID NO: 1 or SEQ ID NO: 2; or(vi) a means for fragmenting chromatin; and(b) instructions which, when executed, causes the apparatus to carry out the method of any one of claims 1-56, 62 or 66-69.

88. A method of identifying a nuclease suitable for fragmenting chromatin within a population of permeabilised cells, the method comprising identifying:(a) a CATH Superfamily 3.40.570.10 Extracellular Endonuclease, subunit A nuclease, or a nuclease variant thereof;(b) a nuclease comprising a cd00091 active site;(c) a nuclease comprising an active site having the amino acid sequence of A / S / D / Q / T / E-K / R / Y / I / H-X-H / Q / Y / F-X(a)-N / F / G-X(7)-E / L / M / H / Q-X(3)-R / M / L / V / F, where X is any amino acid and a is 23 to 37;(d) a nuclease having a ball and chain morphology; or(e) a nuclease comprising an amino acid sequence having at least 40% identity to SEQ ID NO: 1 or SEQ ID NO: 2; and determining whether the nuclease fragments chromatin within a population of permeabilised cells.

89. A method of fragmenting chromatin within a population of permeabilised cells with a nuclease which has been identified according to the method of claim 88.