A molecular detection method for chromosomal structural variation
By employing chromosome conformation capture and low-depth sequencing, this method addresses the issues of low resolution and high cost in existing chromosome structural variation detection technologies, enabling rapid and accurate detection of chromosome structural variations, which is suitable for prenatal screening and cancer-assisted diagnosis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- YIKON GENOMICS SHANGHAI CO LTD
- Filing Date
- 2023-12-12
- Publication Date
- 2026-07-21
AI Technical Summary
Existing chromosome karyotype analysis and chromosome conformation capture technologies suffer from problems such as low resolution, high cost, long cycle time, susceptibility to human influence, and high data noise when detecting chromosome structural variations, making it difficult to meet the needs of clinical diagnosis.
Using a method based on chromosome conformation capture and low-depth sequencing, a sequencing library was constructed by cross-linking, enzymatic digestion and ligation of chromosomal DNA in sample cells, followed by low-depth paired-end sequencing and spatial contact matrix analysis to detect chromosomal structural variations.
It enables rapid, accurate, and cost-effective detection of chromosomal structural variations, can identify structural rearrangements and copy number abnormalities, shortens the detection cycle to 2.5 days, and reduces detection and labor costs.
Smart Images

Figure SMS_1 
Figure SMS_2 
Figure SMS_3
Abstract
Description
Technical Field
[0001] This invention relates to the field of chromosomal structural variation detection technology. More specifically, this invention relates to methods for detecting chromosomal structural variations using chromosome conformation capture technology and high-throughput sequencing technology, as well as detection products for using such methods, and diagnostic applications of said methods and products, particularly in prenatal screening. Background Technology
[0002] Chromosomal disorders are diseases caused by chromosomal abnormalities. Chromosomal abnormalities include numerical and structural abnormalities. Widespread prenatal screening and diagnosis of chromosomal disorders can effectively prevent the occurrence of various abnormal phenotypes such as malformations, intellectual disabilities, and growth retardation, and is of great significance in reducing birth defects.
[0003] Chromosomal number abnormalities include increases or decreases in the number of individual chromosomes, such as chromosomes 13, 21, 18, and X and Y chromosomes; they also include copy number variations (CNVs) in chromosomal segments. Chromosomal structural abnormalities include, but are not limited to, rearrangements, translocations, inversions, deletions, and duplications.
[0004] For many years, in clinical applications, chromosome abnormalities have primarily been detected through karyotype analysis of cells. Chromosomal karyotype analysis focuses on chromosomes in metaphase of cell division. Based on characteristics such as chromosome length, centromere location, long / short arm ratio, and presence or absence of satellites, and using banding techniques, chromosomes are analyzed, compared, sorted, and numbered. Diagnosis is then made based on variations in chromosome structure and number. See also... Figure 1B Chromosomal karyotype analysis has long been considered the "gold standard" for diagnosing chromosomal structural variations and a first-line method for prenatal diagnosis of chromosomal diseases. However, this technique has drawbacks such as low resolution, long analysis time, and the susceptibility of results interpretation to the experience and subjective judgment of the analyst. For example, the commonly used high-resolution 550-band karyotype analysis is limited by the resolution of the naked eye, making it difficult to distinguish chromosomal structural variations between 5M and 10M, and completely unable to distinguish chromosomal variations below 5M. Furthermore, karyotype analysis requires fresh "live cell" culture; if cell culture fails, the test fails.
[0005] In recent years, the application of chromosome conformation capture technology in chromosome structure analysis has been proposed. See, for example, Jon-Matthew Belton et al.; Hi-C: A comprehensive technique to capture the conformation of genomes, Methods. 2012 November; 58(3):.doi:10.1016 / j.ymeth.2012.05.001.
[0006] Classic chromosome conformation capture techniques, such as Hi-C techniques, primarily focus on the whole cell, utilizing high-throughput sequencing technology combined with bioinformatics analysis methods to study the spatial relationships of entire chromosomes across the entire genome. This technique typically involves covalently cross-linking chromosomal DNA / proteins, nuclease digestion, and religation to capture spatially interacting chromosomal DNA fragments. See also Figure 1C (Cited from Jon-Matthew Belton et al., 2012, Figure 1 of the above citation). Combined with high-throughput sequencing technology, chromosome conformation capture technology allows for fine interpretation of chromosome interactions across the entire genome and provides a 3D structural map of the whole genome.
[0007] However, chromosome conformation capture technology, due to its inherent complexity, also suffers from drawbacks such as high experimental costs, significant data noise, and long experimental cycles. These limitations restrict its clinical application. For example, for the approximately 3Gb genome of mice and humans, it is generally considered that at least 600M sequencing reads are needed for high-resolution compartment, topologically related domain (TAD), and chromosome loop identification; for low-resolution compartment and TAD analysis, at least 100M to 300M reads are required. (See Arima-HiC Kit, User Guide for Mammalian Cell Lines, Document Part Number: A160134v01, October 2019). Acquiring and analyzing large sequencing data volumes requires significant sequencing costs and places high demands on computational power and algorithmic tools. Furthermore, the use of restriction endonucleases in chromosome conformation capture technology means that only genomic regions located near the selected restriction endonuclease can be captured, introducing human bias into the resulting sequencing data and affecting the interpretation of subsequent results.
[0008] In view of the foregoing, when applying chromosome conformation capture technology to clinical diagnosis, there is still an urgent need in the field to optimize the methodological steps to promote accurate, rapid and cost-effective clinical detection of chromosomal structural variations (SVs). Summary of the Invention
[0009] To meet the above needs, the inventors explored the application of chromosome conformation capture technology in the detection of chromosomal structural variants (SVs); and ultimately established a new method for detecting chromosomal structural variants based on chromosome conformation capture and low-depth sequencing. This method overcomes the shortcomings of karyotype analysis, such as low resolution, low accuracy, and dependence on live cells; it can obtain more objective and accurate automated and visualized SV interpretation results without relying on manual interpretation. Moreover, compared with existing chromosome conformation capture methods, the method of this invention for detecting chromosomal structural variants not only shortens the detection cycle (to approximately 2.5 days) and simplifies the operation steps, but also requires only 10-20M low-depth sequencing data to ensure accurate detection of chromosomal variants, thereby significantly reducing detection costs and labor costs. As demonstrated in the embodiments of this application, the method of this invention is not only applicable to the detection of chromosomal structural variants, but also to the detection of chromosomal copy number abnormalities; and can accurately determine the type, length, and breakpoint of chromosomal structural variants.
[0010] Therefore, in a first aspect, the present invention provides a method for detecting chromosomal structural variations in a sample based on chromosome conformation capture, comprising:
[0011] (a) Cross-linking, digestion, and ligation of chromosomal DNA in sample cells, wherein the digestion is performed in the presence of DNase I;
[0012] (b) Using the enzyme digestion and ligation products from step (a), construct a sequencing library;
[0013] (c) Perform low-depth paired-end sequencing on the sequencing library, wherein the sequencing depth does not exceed 3x (e.g., 2x-3x), wherein preferably, pair-ended sequencing is used, and the sequencing data volume is 10M to 30M reads, preferably not more than 25M reads, and more preferably not more than 20M reads;
[0014] (d) Based on sequencing data, detect chromosomal structural variations in the sample. Preferably, the structural variations (SVs) are selected from: structural rearrangements and copy number abnormalities (CNVs). Preferably, structural rearrangements are selected from translocations (balanced translocations, unbalanced translocations, reciprocal translocations, non-reciprocal translocations, Roche translocations, and complex translocations), inversions (intra-arm inversions and inter-arm inversions), and insertions; copy number abnormalities are selected from deletions and duplications.
[0015] In some embodiments, the sample used in the method of the present invention is a mammalian cell, body fluid, or tissue sample, preferably an isolated peripheral blood mononuclear cell (PBMC) sample. In some embodiments, the sample according to the present invention contains 1 million to 5 million nucleated cells, preferably not exceeding 2 million nucleated cells.
[0016] In some embodiments, the method according to the invention includes using the enzyme digestion ligation product to construct a sequencing library, wherein when using the enzyme digestion ligation product to construct the sequencing library, no separation of unligated DNA fragments from ligated DNA fragments is performed.
[0017] In some embodiments, the method according to the invention includes: performing end repair on the DNA fragment in a reaction system containing dNTPs and Klenow polymer before ligating the enzyme-digested chromosomal DNA fragment. In some preferred embodiments, the end repair reaction lasts for 1-3 hours (e.g., about 1 hour).
[0018] In some embodiments, the method according to the invention includes: in step (a), the enzyme-digested chromosomal DNA fragments are ligated in the presence of DNA ligase for no more than 4 hours. In some embodiments, the ligation is carried out in a reaction system containing PEG and T4 DNA ligase. In some embodiments, the ligation reaction time is 1-3 hours, and the ligation temperature is 20-30°C, preferably about 22°C. In some embodiments, the ligation reaction system is 0.5-1.0 ml, containing 1%-10% (preferably 5%) of PEG8000 and 5U-50U (preferably 0.02U / ul-0.2U / ul) of T4 DNA ligase.
[0019] In some embodiments, step (a) of the method according to the invention further includes decrosslinking the enzyme-digested ligation product to separate the DNA from the protein in the crosslinked complex; and purifying the separated DNA, preferably using magnetic beads.
[0020] In some embodiments, in step (a) of the method according to the invention, the enzymatic digestion of chromosomal DNA is carried out in a reaction system containing 1 U-5 U DNase I and 200-500 μL, and the digestion reaction time does not exceed 5 minutes. Preferably, the DNA is purified using magnetic beads after enzymatic digestion.
[0021] In some embodiments, step (a) of the method according to the invention further includes lysing the cells prior to enzymatic digestion. In some embodiments, the cell lysate contains IGEPA and a protease inhibitor, preferably containing 8-15 mM pH 8.0 Tris-HCl, 7-15 mM NaCl, 0.1-0.5% IGEPA and 5-20 μl of protease inhibitor, more preferably containing 10 mM pH 8.0 Tris-HCl, 10 mM NaCl, 0.2% IGEPA and 10 μl of protease inhibitor.
[0022] In some embodiments, the method according to the invention includes:
[0023] 1) Optionally, extract the target nucleated cells from the sample;
[0024] 2) Use cross-linking agents to fix cells so that chromosomal DNA can be cross-linked;
[0025] 3) Lyse cells and digest chromosomal DNA in the presence of DNase I;
[0026] 4) Use biotin-labeled nucleotides to repair the ends of enzyme-digested chromosomal DNA fragments and ligate the ends of adjacent DNA fragments in the cross-linked complex;
[0027] 5) By decrosslinking, the DNA and protein in the crosslinked complex are separated.
[0028] 6) Preferably, biotin removal from the ends of unconnected chromosomal DNA fragments is not performed; biotin-tagged chromosomal DNA fragments are captured for use in constructing sequencing libraries.
[0029] In some embodiments, the method according to the invention also has one or any combination of the following features:
[0030] (i) The capture was performed using streptavidin magnetic beads;
[0031] (ii) The covalent cross-linking is carried out using a formaldehyde cross-linking agent, preferably with a formaldehyde concentration of 1%-3%;
[0032] (iii) Treat the chromosomes with a combination of 0.2% SDS and 2% Triton prior to enzymatic digestion; and / or
[0033] (iv) Decrosslinking was performed in a buffer containing SDS and proteinase K.
[0034] In some embodiments, the sequencing library construction of step (b) according to the method of the present invention includes: fragmenting the DNA in the enzyme digestion product and screening for fragments of 150bp-600bp (preferably 300-500bp) for library construction. In some preferred embodiments, the DNA fragment screening is performed by adding magnetic beads to the fragmented DNA product in two steps at a magnetic bead ratio of 0.6x and 0.4x.
[0035] In some embodiments, step (d) of the method according to the invention includes: constructing a spatial contact matrix based on sequencing data; and examining changes in chromosome interaction patterns relative to a control reference system to detect chromosome structural variations. In some preferred embodiments, step (d) includes:
[0036] (d1) Transform the sequencing data into a spatial contact matrix;
[0037] (d2) The spatial contact matrix is corrected and normalized, wherein the correction is performed using a matrix reference frame constructed using a self-comparison reference frame;
[0038] (d3) Examine chromosome structural variations on the corrected and normalized spatial contact matrix.
[0039] (d4)Optionally, visualize the detected chromosomal structural variations and output variation information.
[0040] In some embodiments, the sequencing data obtained according to the method of the present invention contains at least 50%, preferably at least 60% or 65% of valid contacts.
[0041] In a second aspect, the present invention provides a detection product for implementing the method according to the present invention, comprising the following modules:
[0042] (1) Cell fixation module: used to fix cells so that the protein / DNA complex in the cells can be covalently cross-linked;
[0043] (2) Chromosome digestion module: used to digest chromosomal DNA in the presence of DNase I;
[0044] (3) DNA linking module: used to link adjacent DNAs in covalently cross-linked complexes to generate a DNA linker fragment library;
[0045] (4) Sequencing library construction module: used to generate sequencing libraries from DNA ligation fragment libraries, preferably the DNA ligation fragment libraries are not separated from unligated fragments when used for sequencing library construction;
[0046] (5) Sequencing module: used for low-depth paired-end sequencing (e.g., pair-ended sequencing) of sequencing libraries, wherein the sequencing depth does not exceed 3x and the sequencing data volume is 10M to 30M reads, preferably not more than 25M reads, and more preferably not more than 20M reads.
[0047] In some embodiments, the detection product of the present invention further includes one or more modules selected from the following:
[0048] (i) End repair module: used to repair the ends of DNA in the enzyme digestion products from the chromosome digestion module (and preferably end biotin labeling);
[0049] (ii) Decrosslinking module: used to decrosslink DNA from the DNA linking module to separate DNA from proteins and release linear DNA fragments from the crosslinked complex.
[0050] In some embodiments, the detection product of the present invention further includes one or more modules selected from the following:
[0051] (i) Cell extraction module: used to extract target cells from samples;
[0052] (ii)SV assessment module: used to assess chromosomal structural variations in samples from sequencing data.
[0053] In some preferred embodiments, the detection product of the present invention includes the following modules:
[0054] 1) Optional, cell extraction module: used to extract target cells from the sample;
[0055] 2) Cell fixation module: Used to fix cells so that the protein / DNA complex in the cells undergoes covalent cross-linking;
[0056] 3) Chromosome digestion module: used to digest chromosomal DNA in the presence of DNase I;
[0057] 4) End repair and DNA ligation module: used to repair the ends of chromosomal DNA produced by enzymatic digestion and label the ends with biotin, and to ligate neighboring DNA in covalently cross-linked complexes;
[0058] 5) Decrosslinking module: Used to decrosslink DNA to separate it from the protein, thereby releasing the linear DNA fragments in the crosslinked complex;
[0059] 6) Sequencing library construction module: used to capture biotin from the obtained linear DNA fragments without removing biotin from the ends of unligated fragments, and to generate sequencing libraries from the captured linear DNA fragments;
[0060] 7) Sequencing module: used for low-depth paired-end sequencing (e.g., pair-ended sequencing) of sequencing libraries, wherein the sequencing depth does not exceed 3x, and the sequencing data volume is 10M to 30M reads, preferably not exceeding 25M reads, and more preferably not exceeding 20M reads;
[0061] 8) SV assessment module: used to assess chromosomal structural variations in samples from sequencing data.
[0062] In a third aspect, the present invention provides the use of a detection method according to the invention, or a detection product according to the invention, for prenatal screening and / or cancer-aided diagnosis, or for use in the preparation of products for prenatal screening and / or cancer-aided diagnosis. In some embodiments, the product is used to detect chromosomal structural variations in a sample of an identified individual. In some embodiments, the structural variations (SVs) are selected from: structural rearrangements and copy number abnormalities (CNVs), preferably, wherein structural rearrangements are selected from translocations (balanced translocations, unbalanced translocations, reciprocal translocations, non-reciprocal translocations, Roche translocations, and complex translocations), inversions (intra-arm inversions and inter-arm inversions), and insertions; and copy number abnormalities are selected from deletions and duplications. Attached Figure Description
[0063] Figure 1A An exemplary process of the method of the present invention is illustrated schematically.
[0064] Figure 1B The illustration shows the traditional core analysis process and the time required.
[0065] Figure 1C This illustration shows the standard HiC sequencing library construction process.
[0066] Figure 2 The illustration shows an exemplary chromosome structural abnormality detection process according to the present invention.
[0067] Figure 3 This shows the fragment size analysis results of the pre-sequencing library amplification products obtained according to the method in Example 1.
[0068] Figure 4 An example is shown: spatial contact matrix diagrams before and after correction obtained from a test sample.
[0069] Figure 5 This shows the corrected spatial contact matrix diagram of a positive sample 1 of a balanced translocation of chromosomes obtained according to the method of Example 1.
[0070] Figure 6 This displays the corrected spatial contact matrix diagram of positive sample 2 for balanced translocation of chromosomes obtained according to the method of Example 1. The upper figure shows the interactions between all chromosomes; the lower figure shows an enlarged view of the interaction region of chromosomes 8 and 14, which is presented as a "butterfly-shaped" color block in the upper figure.
[0071] Figure 7 The output of the visual variants of balanced translocation samples one and two from Example 1 is displayed.
[0072] Figure 8 The output shows the visualized karyotype of normal chromosomes.
[0073] Figure 9This shows the corrected spatial contact matrix diagram of the Roche translocation sample obtained according to Example 4.
[0074] Figure 10 This shows the corrected spatial contact matrix diagram of the chromosome inversion sample obtained according to Example 5.
[0075] Figure 11 This displays the corrected spatial contact matrix diagram of the chromosome insertion sample obtained according to Example 6, as well as a magnified view of the breakpoint location region.
[0076] Figure 12 This displays a CNV-plot of the copy number variation sample according to Example 7.
[0077] Figure 13 Displays the visual variation output results of the complex translocation sample according to Example 8. Invention Details
[0078] definition
[0079] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. See, for example, Lackie, DICTIONARY OF CELL AND MOLECULARBIOLOGY, Elsevier (4th edition, 2007); Sambrook et al., MOLECULARCLONING, ALABORATORY MANUAL, Cold Spring Harbor Laboratory Press (Cold Spring Harbor, New York, 1989).
[0080] In practice, any methods, apparatus, and materials similar to or equivalent to those described herein may be used. The definitions provided herein are intended to aid in understanding certain terms frequently used herein and do not constitute a limitation on the scope of the invention.
[0081] In this article, the term " Include When the terms “comprising” and its variations, such as “including” and “containing”, are placed before a described step or element, they are used to indicate that adding other steps or elements is optional and non-exclusive. It should be understood herein that the term “comprising” or its variations also cover schemes consisting of the stated steps or elements.
[0082] In this article, when referring to specific numerical values, the term " About "Approximately X" refers to the typical range of error for that value as known to those skilled in the art. It should be understood herein that this expression encompasses references to the specific value itself. Therefore, for example, when referring to "approximately X," it also encompasses a specific reference to the specific value X itself.
[0083] In this article, the term " chromosome The term "chromatin" is used interchangeably with "chromatin" and refers to the complex of chromosomes that contain all or part of a cell's genome. A cell's genome is typically characterized by its karyotype, which is the collection of all chromosomes that contain the cell's genome. A cell's genome may contain one or more chromosomes. In humans, each chromosome has a short arm (called "p") and a long arm (called "q"). Each chromosome arm is divided into regions or bands, which can be seen using a microscope in routine karyotype analysis. Chromosomal bands are labeled p1, p2, p3, etc., and are counted from the centromere towards the telomere. Higher-resolution subbands within a band are sometimes also used to identify regions within a chromosome. Subbands are also numbered from the centromere towards the telomere. Information on chromosome bands and chromosome nomenclature can be found in Strachan, T. and Read, AP 1999. Human Molecular Genetics, 2nd ed., New York: John Wiley & Sons, pp. 37-39.
[0084] In this article, the term " Chromosomal structural variations Chromosomal structural variation (SV), also known as "chromatin structural variation," refers to the structural differences in chromosomes of an individual organism relative to the chromosomes of other individuals within the same species. These structural differences can be chromosomal rearrangements and copy number variations; and encompass chromosomal structural changes of various sizes, such as 1-5M, 5-10M, over 10M, or even large portions of individual chromosomes, such as half, one-third, or three-quarters of the structure. Non-restrictive examples of types of chromosomal structural variation include translocations, reciprocal translocations, non-reciprocal translocations, balanced translocations, unbalanced translocations, complex translocations, inversions, insertions, deletions, duplications, amplifications, and combinations of two or more variations of the same or different types.
[0085] In this article, the term " Chromosomal structural rearrangement "Structural rearrangement" refers to changes in the order and position of DNA sequences on chromosomes. Structural rearrangements can be changes in the position of chromosomal DNA sequences, such as translocations; changes in orientation, such as inversions; or changes in type, such as the insertion of non-homologous chromosomal segments. Structural rearrangements can occur between chromosomes or within chromosomes.
[0086] In this article, the term " transpositionTranslocation refers to the breakage of a chromosome and its complete or partial reattachment to different chromosomal segments, such as the exchange of chromosomal segments between non-homologous chromosomes and / or between two or more locations on the same chromosome. Translocations typically result in abnormal proximity of two non-adjacent chromosomal regions. Depending on the location of the break and reattachment, a translocation may have no effect on gene expression, affect the expression of a single gene, or affect the expression of multiple genes. Types of translocations include, but are not limited to, reciprocal translocations, non-reciprocal translocations, Robertsonian translocations, balanced translocations, and unbalanced translocations.
[0087] In this article, the term " Balance Transposition "Translocation" refers to a form of translocation in which no increase or decrease of genetic material occurs during the exchange of genetic material.
[0088] In this article, the term " Non-equilibrium translocation "Translocation" refers to a form of translocation in which genetic material is lost during the exchange of genetic material.
[0089] In this article, the term " mutual transposition "Reciprocal translocation" refers to the exchange of segments between two chromosomes. Carriers of balanced reciprocal translocations are often phenotypically normal, but may cause unbalanced chromosomal translocations in gametes, leading to infertility, miscarriage, or phenotypic abnormalities in offspring.
[0090] In this article, the term " Non-mutual transposition "Transfer of genetic material" refers to the unidirectional transfer of genetic material from one chromosome to another non-homologous chromosome.
[0091] In this article, the term " Luo changed his position Robertsonian translocation, also known as acrocentromere translocation, refers to a translocation caused by the breakage of two chromosomes with acrocentromeres at or near the centromere. The long arms of the two chromosomes rejoin to form the translocated chromosome, while the two short arms form a very small chromosome, which is often lost. Robertsonian translocations typically occur between chromosomes 13, 14, 15, 21, and 22 in humans, with an incidence of approximately 1.23 per 1000 in the general population, accounting for about 2-3% of infertile individuals. Carriers of Robertsonian translocations have only 45 chromosomes but often exhibit a normal phenotype.
[0092] In this article, the term " Complex translocation "Complex translocations" refers to translocations involving two or more chromosomes and two or more breakpoints. Complex translocations can be balanced or unbalanced.
[0093] In this article, the term " ReversedAn inversion is a structural rearrangement within a chromosome that alters the orientation of the DNA sequence. In this process, fragments resulting from two breaks on the same chromosome are rejoined after being inverted 180 degrees. Inversions can be intraarm or interarm. An intraarm inversion occurs on the same arm of a chromosome (e.g., the long or short arm). An interarm inversion occurs on a segment of the chromosome containing the centromere.
[0094] In this article, the term " insert Insertion mutations refer to the process by which a segment of DNA on a chromosome moves from its original location to another location on the same or different chromosomes. Insertion mutations usually do not cause an increase or decrease in genetic material; therefore, the genetic material resulting from insertion mutations is usually in equilibrium.
[0095] In this article, the term " breakpoint A "breakpoint" is a site or region where a chromosome breaks during a structural change, such as a translocation or inversion.
[0096] In this article, the term " Copy number anomaly "(CNV)" refers to a copy number variation in a chromosomal DNA sequence. A CNV can be a copy number abnormality of an entire chromosome or a segment of a chromosome, such as an increase or decrease in the copy number. In this disclosure, a CNV specifically refers to a copy number abnormality of a chromosomal segment (e.g., less than 1M, 1-5M, 5M-10M or more), such as a deletion or duplication.
[0097] In this article, the term " Read the passage "Read" refers to a short nucleotide sequence produced by any sequencing process described herein or known in the art. In paired-end sequencing, a pair of reads (also called read pairs) with a physical relationship are generated from both ends of the sequenced nucleic acid. Pair-end reads obtained by pair-ended sequencing are also simply referred to as PEreads.
[0098] In this disclosure, " Sequencing data volume "The data is given in units of reads. For paired-end sequencing, a pair of PE reads is counted as only one read. For example, when 20 million sequencing reads are obtained through paired-ended sequencing, the sequencing data volume can be expressed as 20M."
[0099] In this article, the term " Read longer "" refers to the nucleotide length of the read. The length of a sequence read is usually associated with a specific sequencing technology. For example, next-generation sequencing (NGS) can provide sequence reads ranging from tens to hundreds of base pairs (bp).
[0100] In this article, the term " Sequencing depth"Sequencing depth" refers to the ratio of the total number of bases obtained from sequencing to the size of the sequenced genome. Sequencing depth can be calculated using the following formula:
[0101] Sequencing depth = read length × total number of reads / length of the genome sequence being sequenced.
[0102] In this article, the term "chromosome conformation capture" is used. Connecting DNA fragments "or" Connecting fragments "" refers to a linker DNA fragment containing a linker point, formed during chromosome conformation capture, after enzymatic digestion, by the end-joining of adjacent DNA fragments in a covalently cross-linked DNA / protein complex. Accordingly, in this document, the term "" not yet Connecting DNA fragments "or" Unconnected fragments "Linked fragments" refers to enzymatically digested DNA fragments that do not have internal linker sites, generated when the DNA fragments have not undergone end ligation. In conventional chromosome conformation capture techniques, it is generally considered necessary to distinguish and separate linked DNA fragments from unlinked fragments, typically by biotin-labeling the ends of digested DNA and removing the biotin-labeled ends of unlinked fragments after ligation. However, the inventors have surprisingly discovered that, as shown in the embodiments of this application, in some embodiments of the method according to the present invention, when constructing sequencing libraries using the enzymatic digestion products, it is not necessary to distinguish and remove such unlinked fragments from the linked fragments, yet good sequencing library quality can still be obtained. It should be understood herein that "linked fragments" also includes derivative fragments generated from said fragments during library construction. For example, a small fragment generated by mechanical or enzymatic action, provided that the derived fragment retains the internal linker site characteristic of the linker fragment from which it originated. Similarly, "unlinked fragment" also covers derived fragments generated from said fragments during library construction (e.g., small fragments generated by mechanical or enzymatic action during library construction), provided that the derived fragment retains at least one enzyme-digested end characteristic of the unlinked fragment from which it originated, or an enzyme-digested end that has been modified, labeled, and / or repaired accordingly in the case of end modification, labeling, and / or end repair. Therefore, as those skilled in the art will appreciate, the term "unlinked fragment" covers the terminal small DNA fragment that retains the enzyme-digested end and its derivatives generated by breaking the fragment, but does not cover the internal small DNA fragments or their derivatives generated by breaking the fragment.
[0103] In this article, Chromosome conformation capture technologyFor example, Hi-C technology refers to a technique that captures information about interactions between chromosomal DNA fragments in cells based on chromosome conformation capture and sequencing (e.g., NGS). This technique typically involves two stages: the first stage of sequencing library construction and the second stage of sequencing and bioinformatics analysis. In the first stage, sequencing library construction usually involves capturing spatially interacting chromosomal DNA fragments through covalent cross-linking, nuclease digestion, and religation of chromosomal DNA / proteins using chromosome conformation capture, generating a sequencing library. In the second stage, the library is sequenced to obtain chromosome conformation capture sequencing data, which will also be referred to as sequencing data in this paper.
[0104] In this article, Spatial contact matrix In this paper, it is sometimes referred to as the genome interaction matrix or Hi-C contact matrix. It is a two-dimensional matrix generated from chromosome conformation capture sequencing data to represent the interactions between DNA sequences in the genome. It is commonly used to study the three-dimensional structure of chromosomes and the spatial organization of the genome. The rows and columns of the matrix represent chromosome position intervals on chromosome coordinates; the matrix elements are the number of paired end sequence fragments falling into the corresponding row and column interaction intervals (also referred to as "contact positions" in this paper), also known as contact strength / density or interaction frequency (IF). The genome interaction matrix can be a numerical matrix or presented as a visualized spatial contact map. In this spatial contact map, color gradients can reflect changes in contact strength, thus presenting color blocks of different intensities at matrix positions with different contact strengths.
[0105] In some implementations, generating a spatial contact matrix (or simply contact map) from chromosome conformation capture sequencing data may include the following steps:
[0106] (1) Quality control and preprocessing of sequencing data: The raw sequencing data is quality controlled, pruned, and filtered to remove low-quality reads and sequencing adapter sequences. Available tools include, but are not limited to, FastQC, Trimmomatic, or Cutadapt.
[0107] (2) Sequence mapping: The preprocessed reads are mapped onto the reference genome through sequence alignment. Available alignment tools include, for example, Bowtie2, BWA, or HiC-Pro. This step aligns the paired-end reads obtained from sequencing onto the reference genome and determines their corresponding genomic locations.
[0108] (3) Filtering and processing of mapped reads: Tools known in the art, such as, but not limited to, HiC-Pro and Juicer, can be used to process the mapped data to, for example, remove invalid data such as duplicate reads and PCR artifacts; and filter reads mapped to repetitive regions of chromosomes to overcome the bias they introduce in the spatial contact map.
[0109] (4) Generating the contact matrix: Tools known in the art, such as Juicer or HiC-Pro, can be used to transform the quality-controlled and filtered mapping data into a contact matrix. This typically involves: binning the genome into equal- or unequal-sized intervals (e.g., 10-100 kb bins) and counting the number of interactions between each pair of bins. The size of the bins determines the resolution of the contact matrix map; the smaller the value, the higher the resolution. The value of each matrix element represents the “interaction” frequency between the two bins corresponding to that value. Typically, cis-type Hi-C contact matrices have high-density contacts near the diagonal bins, while the contact density decreases exponentially with increasing distance between row and column bins. This is represented on the matrix map as an “interaction” frequency that decreases exponentially with increasing distance between two bins.
[0110] (5) Contact matrix normalization: This normalization process aims to eliminate systematic bias in order to preserve interaction frequencies that reflect the underlying genomic structure as much as possible. Commonly available normalization methods include, for example, HiCNorm, sequential component normalization (SCN), iterative correction and eigenvector decomposition (ICE), Knight-Ruiz (KR), chromosomeR, and multiHiCcompare. See, for example, Hongqiang Lyu, BioTechniques Vol.68, No.2, Comparison of normalization methods for Hi-C data, 2019, https: / / doi.org / 10.2144 / btn-2019-0105.
[0111] (6) Visualization of the contact matrix: The normalized contact matrix can be visualized as a contact diagram, such as a Hi-C contact thermal map, using tools known in the art, such as HiCExplorer, HiGlass, or Juicebox.
[0112] As used herein, the term "module" refers to a reagent, reagent kit, component, assembly, and / or device or system that can be used to achieve the function of the module. It should be understood that the composition of a module is not limited to a particular physical form, as long as it can achieve the desired function. Depending on the intended function, a module can be a combination of reagents and / or devices to achieve the function; or it can be a software object or routine (e.g., as a separate thread) that executes centrally on a single computing system (e.g., a computer program, a tablet computer (PAD), one or more processors). For example, a chromosome digestion module according to the invention can consist of reagents or reagent kits for performing the enzymatic digestion and optionally devices for ensuring that the enzymatic digestion is performed under specific conditions. As another example, an SV evaluation module implementing the invention can be a program stored on a computer-readable medium containing computer program logic or code portions for performing the SV evaluation.
[0113] The following describes the steps and modules / components involved in the method and product of the present invention. It should be understood that the features described in these steps and modules / components may exist in the method and product of the present invention in any combination unless explicitly inappropriate, as if such combinations of features were individually specified.
[0114] Method of the present invention
[0115] High-throughput chromosome conformation capture technologies (such as Hi-C technology) can determine spatial interactions between different sites on chromosomes by digesting and reconnecting spatially close chromosome segments and performing high-throughput sequencing on them.
[0116] Conventional Hi-C technology typically includes the following main steps: A) Covalent cross-linking: Cells are momentarily fixed with formaldehyde, causing covalent cross-linking between chromosomal DNA and proteins; B) Enzymatic digestion: DNA is cleaved using restriction endonucleases (such as HindIII); C) Labeling and re-ligation: Biotin is labeled at the ends of the digested DNA, and neighboring DNA molecules in the covalently cross-linked complex are ligated to form chimeric DNA molecules; D) Uncross-linking: The cross-links between DNA and proteins are broken to obtain linear DNA fragments carrying biotin; E) Biotin removal and DNA fragment purification: Biotin is removed from the ends of the unligated linear fragments, and DNA is purified to enrich the biotin-labeled DNA chimeric molecules inserted at the internal ligation sites; F) Fragmentation and bio-linking. Biotin-tagged molecule enrichment: Purified DNA chimeric molecules are fragmented (e.g., by mechanical or enzymatic shearing) to reduce their overall size to a level suitable for sequencing library construction; streptavidin-coated magnetic beads are used to capture / pull-down biotin-tagged DNA fragments; G) Sequencing library construction: DNA fragment end modification and sequencing library adapter addition are performed, followed by PCR amplification to generate Hi-C libraries; H) Sequencing: High-throughput paired-end sequencing of the Hi-C libraries; and I) Spatial association map creation: Sequencing data are mapped onto a reference genome, and based on the genomic location mapped to each cross-linked DNA segment, the spatial interactions between the two cross-linked DNA sequences are determined, generating a spatial association map. See also Figure 1C .
[0117] While conventional Hi-C sequencing (Hi-C) offers advantages in acquiring whole-genome spatial structure information, it suffers from drawbacks such as high experimental costs, significant data noise, and long experimental cycles. To improve the applicability of the Hi-C method in clinical SV detection, the inventors conducted in-depth research, optimizing the SV detection system and steps while maintaining the accuracy of SV variant detection. By using DNase I, which has random nucleic acid cleavage properties, combined with low-depth high-throughput sequencing, the time, reagent costs, and labor costs for Hi-C library construction for SV detection were significantly reduced, while sequencing costs were also substantially lowered.
[0118] During the optimization of the detection steps, the inventors also surprisingly discovered that, compared to the conventional Hi-C technology process, in the Hi-C library construction process,
[0119] - Shorten the duration of biotin labeling and binding (i.e., step C of the conventional Hi-C method above);
[0120] - Omit the step of removing biotin from the ends of unconnected fragments (i.e., step E as described in the conventional Hi-C method above); and / or
[0121] -Simplify the DNA extraction process after protein decrosslinking (i.e., step E of the conventional Hi-C method above).
[0122] It can simplify the process while maintaining the quality of sequencing libraries and the accurate detection of chromosomal structural variations.
[0123] Based on the above optimizations and findings, the inventors have proposed the SV detection method of this invention. Referring to Figure 1 and the embodiments of this application, this invention demonstrates at least the following advantages over traditional karyotype techniques and existing Hi-C techniques in SV detection.
[0124] 1. Compared with traditional karyotype techniques, the detection time of the method of this invention is reduced from 4-8 days to 2.5 days; and it does not rely on cell culture.
[0125] 2. Compared with conventional Hi-C technology (more than 4 days), the method of this invention significantly shortens the time by optimizing the detection system and detection steps (including chromosome enzyme digestion step, biotin labeling and ligation step, and protein decrosslinking step) while ensuring the detection results. It also reduces the reagent cost and labor cost of constructing Hi-C library for SV detection.
[0126] 3. Compared with conventional Hi-C technology, the Hi-C library obtained by the method according to the present invention has comparable sequencing library quality and only requires 10M-20M of sequencing data, instead of the traditional 100M or more (see, for example, Mallard C, Johnston MJ, Bobyn A, et al. Hi-C detects genomic structural variants in peripheral blood of pediatric leukemia patients[J]. Molecular Case Studies, 2022, 8(1):a006157.), which can accurately detect chromosomal structural variations and copy number variations, thereby greatly reducing sequencing costs.
[0127] 4. Compared with foreign commercial Hi-C sequencing library kits (approximately 4000-5000 RMB), the method of this invention has a much lower cost, about 1 / 4 of the price of foreign commercial kits, which is very beneficial for controlling detection costs.
[0128] Therefore, in one aspect, the present invention provides a method for detecting chromosomal structural variations in a sample based on chromosome conformation capture, comprising:
[0129] (a) Cross-linking, digestion, and ligation of chromosomal DNA in sample cells, wherein the digestion is performed in the presence of DNase I;
[0130] (b) Using the enzyme digestion and ligation products from step (a), construct a sequencing library;
[0131] (c) Perform low-depth paired-end sequencing on the sequencing library, wherein the sequencing depth does not exceed 3x (e.g., 2x-3x), wherein preferably, pair-ended sequencing is used, and the sequencing data volume is 10M to 30M reads, preferably not more than 25M reads, and more preferably not more than 20M reads;
[0132] (d) Based on sequencing data, detect chromosomal structural variations in the samples.
[0133] Sequencing library construction
[0134] sample
[0135] The method according to the invention is applicable to any sample containing nucleated cells. Such samples can be obtained from suitable subjects, such as eukaryotic individuals and mammalian individuals. In some embodiments, the subject is a mammal, particularly a human. In some embodiments, the subject is a diseased or non-diseased individual. In some embodiments, the subject is a human individual to be subjected to prenatal screening.
[0136] A sample can be a sample directly isolated or obtained from a single or multiple objects or their parts. Examples of samples include, but are not limited to, bodily fluids or tissues from an individual mammal, including but not limited to blood or blood products (e.g., serum, plasma, platelets, erythrocyte sedimentation rate, serotonin, etc.), cord blood, chorionic villi, amniotic fluid, cerebrospinal fluid, lavage fluid (e.g., lung, stomach, peritoneum, catheter, ear, arthroscopy), biopsy samples, isolated cells (lymphocytes, placental cells, stem cells, bone marrow-derived cells, embryonic or fetal cells), or combinations thereof.
[0137] In some embodiments, the sample is a mammalian cell, body fluid, or tissue sample. In some preferred embodiments, the sample is an isolated peripheral blood mononuclear cell (PBMC) sample. In some embodiments, the sample contains 0.1 x 102 6 Up to 10x10 6 Nucleated cells, for example, 0.5 x 102 6 1 x 10 6 1, 2x10 6 1, 3x10 6 1, 4x10 6 5 x 10 6 6x106 1, 7x10 6 8x10 6 1, 9x10 6 10x10 6 One. In some implementations, the sample contains 1x10. 6 Up to 5x10 6 Nucleated cells, for example, 1x10 6 Up to 3x10 6 Nucleated cells, for example, no more than 2 x 10n 6 A nucleated cell.
[0138] Crosslinking
[0139] In this step, the DNA within the chromosome is cross-linked (e.g., immobilized), such that regions within the nucleic acids that interact with each other are maintained or immobilized in spatial proximity. Chromosomal DNA can be cross-linked in a cell or an intact nucleus. In some cases, a reversible cross-linking agent, or a cross-linking agent or method that leads to reversible cross-linking, can be used for the cross-linking. Exemplary cross-linking agents include, but are not limited to, formaldehyde, paraformaldehyde, formalin, and other cross-linking agents, provided they can achieve covalent cross-linking of protein-protein, DNA-DNA, RNA-RNA, RNA-DNA, protein-RNA, protein-DNA, or protein-RNA-DNA molecules.
[0140] The cross-linking agent may contact the cell or intact cell nucleus for at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 minutes or longer. The concentration of the cross-linking agent may be about 0.5, 0.6, 0.7, 0.9, 1.0, 1.2, 1.5, 2, 3, 4, or about 5%. In some cases, the cross-linking may be quenched after a suitable period of time to inhibit or stop the cross-linking reaction. Suitable quenching agents include those containing one or more amines, such as glycine, aspartic acid, arginine, and glutamic acid. The cross-linking and quenching reactions may be carried out at room temperature.
[0141] In some embodiments, formaldehyde is used to fix the cells to achieve cross-linking. In some embodiments, glycine is used to quench the formaldehyde. In some embodiments, the amount of formaldehyde used for fixation is 1% to 3%, for example, about 2%, and the amount of glycine used to quench the formaldehyde is about 0.125M. In some embodiments, the formaldehyde fixation time is about 5-15 minutes, for example, about 10 minutes.
[0142] Chromosome digestion
[0143] This step aims to enzymatically digest chromosomal DNA, causing spatially interacting neighboring DNA fragments immobilized in the cross-linking complex to generate free ends, which can then be ligated together. In conventional chromosome capture techniques, 4 bp or 6 bp restriction endonucleases are typically used for this digestion. However, the application of restriction endonucleases can confine the captured fragments to chromosomal regions containing only the corresponding restriction sites, thus introducing bias into the sequencing data. Furthermore, restriction enzyme digestion is usually time-consuming. In the method according to the present invention, the inventors proposed using DNase I, which has no sequence-specific cleavage bias, to digest chromosomal DNA, and surprisingly found that using DNase I to construct sequencing libraries not only significantly shortens library construction time but also provides high-quality sequencing data, which is beneficial for detecting chromosomal structural variations with low-depth sequencing data volumes not exceeding 20M.
[0144] In some embodiments of the method of the present invention, the enzymatic digestion involves digesting the chromosome with DNase I for 1-6 minutes, preferably not exceeding 5 minutes. In some embodiments, the enzymatic digestion is carried out using 0.5-5 U of DNase I in a reaction system of 100-500 μL. In some embodiments, the enzymatic digestion is carried out using approximately 1 U of DNase I in a reaction system of 200 μL for approximately 4 minutes.
[0145] Cell lysis is preferred before chromosome digestion. Hypotonic buffer with a low concentration of nonionic surfactant can be used for lysis, or isotonic buffer with a higher concentration of nonionic surfactant can be used. During lysis, protease inhibitors can be added to inhibit intracellular DNA digestion activity. The cell lysis buffer can be in contact with cells for at least approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 minutes or longer, for example, 5–20 minutes or 10–15 minutes.
[0146] In some embodiments, the cell lysis buffer is a buffer containing the nonionic surfactant NP-40 or IGEPA. In some embodiments, the cell lysis buffer also contains a protease inhibitor. In some embodiments, the buffer is a Tris buffer. In some embodiments, the cell lysis buffer is a hypotonic Tris buffer.
[0147] In some embodiments, the cell lysis buffer contains 8-15 mM pH 8.0 Tris-HCl, 7-15 mM NaCl, 0.1-0.5% IGEPA, and preferably also contains 5-20 μl of protease inhibitor. In some embodiments, the cell lysis buffer contains 10 mM pH 8.0 Tris-HCl, 10 mM NaCl, 0.2% IGEPA, and 10 μl of protease inhibitor.
[0148] In some embodiments, permeabilization of the cell nucleus is preferably performed prior to enzyme digestion. In some embodiments, the permeabilization is performed using 0.1-0.5% SDS, followed by quenching of the SDS with 1-3% Triton X-100. In some embodiments, the permeabilization comprises treating cells with 0.1-0.5% SDS (e.g., 0.2%) for 5-15 minutes (e.g., 10 minutes), followed by treating cells with 1-3% (e.g., 2%) Triton X-100 for 5-15 minutes (e.g., 10 minutes).
[0149] In some implementations, DNA purification is preferably performed after enzyme digestion, for example, using magnetic beads.
[0150] In a preferred embodiment, the enzymatic digestion step includes:
[0151] (a) Treat 1–5 x 10 cells in a cell lysis buffer containing 10 mM pH 8.0 Tris-HCl, 10 mM NaCl, 0.2% IGEPA, and 10 μL of protease inhibitor for 1–5 x 10⁻⁵ hours. 6 It takes approximately 10-20 minutes per cell to lyse the cells;
[0152] (b) Treat cells in the presence of 0.2% SDS for approximately 5–15 minutes to permeate the nucleus;
[0153] (c) Quench SDS with 2% Triton X-100 for approximately 5-15 minutes;
[0154] (d) Digest the chromosomes in a 200 μL reaction system containing 1 U DNase I for approximately 3-5 minutes, preferably approximately 4 minutes;
[0155] (e) Purification of digestion products using magnetic beads.
[0156] connect
[0157] This step aims to ligate the free ends of adjacent DNA fragments generated by enzyme digestion together to produce a ligation fragment library. DNase I digestion produces a DNA product with sticky ends. End repair can be performed prior to ligation. Reaction systems suitable for end repair are known in the art. For example, Klenow enzyme can be used for end repair. In one embodiment, the end repair reaction is carried out in a reaction system containing dCTP, dTTP, dGTP, and dATP, as well as Klenow polymerase. In some embodiments, the amount of Klenow polymer in the reaction system can be 0.1-0.5 U / ul, for example, 0.2-0.3 U / ul, preferably about 0.25 U / ul. Preferably, the reaction time is 1-3 hours, more preferably the reaction temperature is 20-30°C, preferably about 23°C. In some embodiments, to label the ligation fragments for library construction, end repair can be performed using labeled nucleotides. In some embodiments, the label is biotin-labeled. In some embodiments, a reaction system comprising dCTP, dTTP, dGTP, and dATP, wherein one of the dNTPs (preferably dATP) is biotin-labeled, is used for end repair to simultaneously repair the sticky ends generated by DNase I and to perform biotinylated end labeling. In some embodiments, the biotin-labeled reaction system comprises 25 U of Klenow enzyme fragment, 10 mM each of dCTP, dTTP, and dGTP, and 1 mM of biotin-labeled dATP per 100 μL of reaction system.
[0158] The end-repaired products can be ligated in the presence of DNA ligase. The inventors have found that although conventional Hi-C technology typically uses long ligation times, the shortened ligation time in the method of this invention does not significantly affect the quality of the sequencing library, but rather reduces detection cycles and labor costs. Therefore, in some embodiments, preferably, the enzyme-digested chromosomal DNA fragments are ligated in the presence of DNA ligase for no more than 5 hours and no more than 4 hours. In some embodiments, the ligation reaction is carried out for approximately 0.5 to 3.5 hours, or approximately 1 to 3 hours, or approximately 1 to 2 hours, for example, approximately 1 hour. In some embodiments, the ligation reaction system is carried out in a reaction system containing PEG and DNA ligase, preferably PEG 6000-8000, more preferably PEG 8000. In some embodiments, the ligation reaction system contains 1-10% PEG 8000, for example, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, or 10%. In some embodiments, the ligation reaction system comprises T4 DNA ligase, for example, 0.02 U / µl to 0.2 U / µl of T4 DNA ligase, for example, 5 U-50 U per 1 ml of reaction system, such as 8 U, 10 U, 12 U, 15 U, 18 U, 20 U, 23 U, 25 U, 30 U, 35 U, 40 U, 45 U, or 50 U. In some embodiments, the ligation reaction system is 0.5-1.0 ml, containing 1%-10% PEG8000 and 5 U-50 U of T4 DNA ligase. In some embodiments, preferably, the ligation reaction system is 1.0 ml, containing approximately 5% PEG8000 and 20 U of T4 DNA ligase.
[0159] Decrosslinking
[0160] After generating the ligation fragment library, the linear DNA in the cross-linked complex can be released by decrosslinking to facilitate library construction. In some embodiments, the crosslinks are removed by heating the ligation product to a high temperature, such as 50°C, 60°C, 70°C, 80°C, or higher. In some embodiments, proteinase K is used for decrosslinking. In some embodiments, SDS and proteinase K are used for decrosslinking. In some embodiments, the decrosslinking reaction includes treating the digested product with 1-5% (e.g., about 3%) SDS and proteinase K at about 50-70°C (e.g., about 60°C) for about 0.5-2 hours, such as about 1 hour. After the decrosslinking reaction, DNA purification is preferably performed. DNA purification can be performed using methods known in the art, such as sodium acetate / isopropanol precipitation and / or magnetic bead purification. In some preferred embodiments, DNA purification is performed using magnetic beads.
[0161] Library Construction
[0162] After harvesting the enzyme digestion products containing the ligation fragment library, sequencing libraries can be constructed using methods known in the art.
[0163] To produce fragments of suitable size for sequencing library construction, various mechanical, chemical, and / or enzymatic methods can be used to fragment or cut the DNA in the digestion products to the desired length. Fragmentation can be performed by sonication, or by brief exposure to a DNase or a mixture of one or more restriction enzymes, or by transposases or nickases. Many enzyme digestion kits are commercially available, such as those from Takara. After fragmentation, fragment size screening and the addition of sequencing adapters can be performed using methods known in the art.
[0164] In some embodiments, the method of the present invention therefore includes: fragmenting the DNA in the enzymatic digestion product according to the present invention and screening for fragments of 150 bp-600 bp (preferably 300-500 bp) for library construction. Preferably, fragment screening uses a two-step magnetic bead addition to recover fragments of suitable size. Fragment screening can be performed by adjusting the magnetic bead ratio. In one embodiment, the DNA fragment screening is carried out by adding magnetic beads to the fragmented DNA product in two steps at a magnetic bead ratio of 0.6x and 0.4x. In the case that the ligation fragment is biotin-tagged, the ligation fragment can be further screened and enriched by capturing the biotin-tagged DNA fragment. In some embodiments, preferably, streptavidin magnetic beads are used for the capture. During the screening of library fragments, however, as shown in the embodiments of this application, since it is not necessary to distinguish and remove unligated fragments from the ligation fragments to obtain good sequencing library quality, in some embodiments, preferably, a biotin removal step at the ends of unligated fragments is not performed before capturing the biotin-tagged ligation fragments. In some embodiments, according to the method of the present invention, approximately 0.4-5 ng / µl of DNA fragment product is obtained after capturing the biotin-labeled fragment. In some embodiments, the sequencing library constructed using the method of the present invention contains DNA fragments between 300-500 bp.
[0165] sequencing
[0166] Sequencing can be performed using paired-end sequencing methods known in the art. In paired-end sequencing, typically, the ends of a nucleic acid fragment are sequenced with appropriate read lengths sufficient to map each read (e.g., reads at both ends of a fragment template) to a reference genome. For example, the read length of a paired-end read can be approximately 10 to approximately 500 consecutive nucleotides, approximately 10 to approximately 400 consecutive nucleotides, approximately 10 to approximately 300 consecutive nucleotides, approximately 50 to approximately 200 consecutive nucleotides, approximately 100 to approximately 200 consecutive nucleotides, or approximately 100 to approximately 150 consecutive nucleotides.
[0167] In some implementations, a paired-end sequencing method is used. In some implementations, the sequencing depth is less than 6x, for example, sequencing depths of approximately 0.1x, 0.3x, 0.5x, 0.8x, 1x, 1.5x, 2x, 2.5x, 3x, 3.5x, 4x, 4.5x, and 5x. In some implementations, the sequencing depth is 2x to 3x. In some implementations, the read length is approximately 100-300 bp, for example, approximately 150 bp. In some implementations, the sequencing data volume is 10M to 30M reads, for example, 10M, 15M, 20M, 25M, or 30M, preferably not exceeding 25M reads, more preferably not exceeding 20M reads.
[0168] Detecting chromosome structural variations
[0169] In this paper, chromosomal structural variations (SVs) refer to structural changes on the genome. These changes can include large-scale rearrangements, inversions, insertions, translocations, duplications, or deletions, as well as small-scale structural variation events, such as copy number abnormalities (CNVs, e.g., duplications or deletions) ranging from 1M to 50M. Since primary structure is the basis of spatial structure, and abnormalities in the primary structure of DNA often lead to changes in the 3D structure of chromosomes, chromosome capture techniques can be used to identify these chromosomal structural variations.
[0170] In chromosome conformation capture sequencing data (e.g., Hi-C sequencing data), the interaction frequency between chromosomal regions exhibits a power-law distribution with genomic distance. When sigma mutates (SVs) alter genomic structure, two previously distant chromosomal regions move closer together, causing their interaction frequency to increase to the level of interaction between closely spaced chromosomal regions, significantly deviating from the situation in normal chromosomes. For example, genomic fragment deletions lead to spatial proximity between the ends of the deleted fragment, resulting in an increased interaction frequency in chromosomal regions flanking the deletion breakpoint. When the corresponding Hi-C reads are mapped to the reference genome, this increased interaction frequency is represented on the spatial contact matrix heatmap as noticeably darker patches in the interaction intervals of the chromosomal regions on either side of the breakpoint. Similarly, any other form of chromosomal structural change (including but not limited to duplication, insertion, translocation, and inversion) will lead to an increase or decrease in the interaction frequency of the corresponding chromosomal regions, and quantitative assessment of this change is a key indicator for detecting the presence of SVs in chromatin.
[0171] After establishing a sequencing library and obtaining sequencing data according to the method of this invention, a genome interaction matrix (e.g., a Hi-C contact map) recording spatial contact information can be generated from the sequencing data using methods known in the art, to determine the type and location of SV variants and breakpoints on the test sample. In some cases, a reference frame can be introduced, and chromosomal structural variations related to biological conditions such as disease or developmental progression can be detected by comparing genome interaction matrices from samples from the reference frame. The construction of the reference frame can be determined by those skilled in the art for the specific application of the method according to this invention. For example, in some cases, the reference frame can be composed of samples from healthy control individuals, or samples from individuals with a specific disease type or stage.
[0172] In some cases, the processing of chromosome conformation capture sequencing data includes sequence mapping, data filtering, and data correction. The reference genome used for sequence mapping can be any one of the human reference genomes GRCh37, GRCh38, T2T-CHM13, etc. For example, the PE reads obtained from pair-ended sequencing can be aligned with the reference genome to identify only paired reads whose ends uniquely match the reference sequence. After data filtering, the final valid data (i.e., valid contact data) for contact analysis is obtained. The percentage of valid contact data to all sequencing data is an important indicator for evaluating the quality of chromosome conformation capture sequencing libraries. In some embodiments, the chromosome conformation capture library sequencing data obtained according to the method of the present invention contains at least 50%, preferably at least 60% or 65% valid contacts.
[0173] In some implementations, SV detection includes two main steps: (a) constructing a control reference system using healthy samples, and (b) detecting variants in the sample to be tested. In some implementations, according to... Figure 2 As shown, SV detection is performed.
[0174] In some implementation schemes, constructing a control reference system includes the following steps:
[0175] (1) Transform the sequencing data into a numerical matrix that records spatial contact information, for example, use Juicer (https: / / doi.org / 10.1016 / j.cels.2016.07.002) to process the sequencing data and standardize the data to obtain a numerical matrix representing spatial contact information;
[0176] (2) For each sample used to construct the reference frame, its spatial contact matrix is divided into equal-sized intervals (bins) of 1Mb by rows and columns. The interaction strength value between each pair of bins is recorded. The interaction strength of each bin is divided by the long range ratio of the sample to obtain the spatial contact strength value of that bin.
[0177] (3) After performing the above processing on all control samples, for each bin, the mean contact intensity of all samples with non-zero values in that bin is taken as the baseline contact intensity value under the matrix reference system for that bin. In this paper, "long rangeratio" refers to the proportion of reads >20K.
[0178] In some implementation schemes, variant detection of the sample to be tested includes the following steps:
[0179] (1) Transform the sequencing data into a numerical matrix that records spatial contact information. For example, use Juicer (https: / / doi.org / 10.1016 / j.cels.2016.07.002) to preprocess and normalize the sequencing data to obtain the spatial contact numerical matrix, which contains the contact locations and contact intensities after normalization.
[0180] (2). Correct the spatial numerical matrix obtained in step (1). For example, in some embodiments, this step includes:
[0181] (i) The spatial contact numerical matrix is processed by quality control information. Preferably, in some embodiments, for each sample to be tested, its spatial contact matrix is divided into equal-sized intervals (bins) of 1Mb by rows and columns, the interaction strength value between each pair of bins is recorded, and the interaction strength of each bin is divided by the long range ratio of the sample to obtain the spatial contact strength value of that bin.
[0182] (ii) Correcting the spatial contact numerical matrix using a matrix reference system constructed from a self-referenced reference system, preferably, in some embodiments, by subtracting the contact intensity baseline value at the corresponding position in the matrix reference system from each position of the numerical matrix; and
[0183] (iii) Convert the data format of the corrected spatial contact matrix, for example, by using HiSV ( https: / / doi.org / 10.1371 / journal.pcbi.1010760 The spatial contact information is converted into a numerical matrix for subsequent data analysis.
[0184] (3). On the standardized and corrected spatial contact matrix, examine the chromosome structural variations. In some embodiments, the examination includes determining the location of the variations and identifying the direction of the variations. In other embodiments, the chromosome structural variations are chromosome structural rearrangements and / or copy number variations, such as those selected from: translocations, balanced and unbalanced translocations, Roche translocations, intra- and inter-arm inversions, insertions, and CNVs.
[0185] (4). Optionally, copy number variations (CNVs) are examined based on sequencing data and spatial contact numerical matrices;
[0186] (5) Optionally, visualize the detected SV results and output variation information.
[0187] In some implementations, the test for balanced chromosomal translocations includes the following steps:
[0188] (i) On the spatial contact matrix (e.g., the spatial contact numerical matrix before correction using a matrix reference frame), for example via hic_breakfinder ( https: / / github.com / dixonlab / hic_breakfinder (and doi:10.1038 / s41588-018-0195-8), preliminary calculation of the location of the equilibrium translocation breakpoint;
[0189] (ii) On the corrected spatial contact matrix, based on the initially obtained equilibrium dislocation breakpoint location, set the hypothetical breakpoint location, and calculate the difference index score based on the difference in contact density on both sides of the hypothetical breakpoint.
[0190] (iii) Move the hypothetical breakpoint upstream and downstream (e.g., move 5Mb with a step size of 100kb) and calculate the difference index score again for the new hypothetical breakpoint position obtained after each step. The position of the highest difference index score at the end of the step that meets the following condition is determined as the actual breakpoint position: If the highest score is higher than the preset score threshold (e.g., the median of the highest difference index score obtained by the sample that actually carries translocation and the highest difference index score obtained by the sample that does not carry translocation), then the breakpoint is judged as positive and reported as a test result; otherwise, the breakpoint is judged as a false positive and not reported as a test result.
[0191] In some implementations, the examination for balanced chromosomal translocations also includes the following steps:
[0192] (iv) On the standardized and corrected spatial contact matrix, for each combination of two chromosomes, infer whether the centromere translocation hotspot region (e.g., starting from the short arm side of the centromere and extending 5Mb towards the short arm telomere, and extending 5Mb from the long arm side of the centromere towards the long arm telomere) conforms to the balanced translocation characteristics (e.g., the contact intensity between the short arm of chromosome A and the short arm of chromosome B, and the contact intensity between the long arm of chromosome A and the long arm of chromosome B are significantly greater than the contact intensity between the long arm of chromosome A and the short arm of chromosome B, and the contact intensity between the short arm of chromosome A and the long arm of chromosome B). For chromosome regions that conform to this characteristic, infer the difference index score as described in steps (i)-(iii) above, and calculate the actual breakpoint position.
[0193] In some implementation schemes, the Roche translocation test includes the following steps: on a standardized and corrected spatial contact matrix, for all combinations of chromosomes 13, 14, 15, 21, and 22, the spatial contact intensity in the matrix interaction interval corresponding to each chromosome combination and the difference between the spatial contact intensity of each combination and other combinations are determined, and a difference index score is calculated; if the score is higher than a preset score threshold (for example, the median of the highest difference index score obtained from a sample that actually carries a Roche translocation and the highest difference index score obtained from a sample that does not carry a Roche translocation), then the corresponding chromosome combination is determined to have a Roche translocation; otherwise, it is determined to be a false positive and is not reported as a test result.
[0194] In some implementations, the examination of chromosomal structural abnormalities includes the examination of inversions. In some implementations, the examination includes: searching for two matrix positions with high contact intensity within the same chromosome on a standardized and corrected spatial contact matrix; determining whether these positions meet the characteristics of inversion breakpoints (e.g., if the contact intensity upstream of breakpoint 1 and upstream of breakpoint 2, and downstream of breakpoint 1 and downstream of breakpoint 2 are all significantly greater than the contact intensity upstream of breakpoint 1 and downstream of breakpoint 2, and downstream of breakpoint 1 and upstream of breakpoint 2, then an inversion is identified); inferring breakpoints from positions meeting the inversion breakpoint characteristics and calculating a difference index score. If the score is greater than a threshold, the breakpoint is identified as an inversion breakpoint.
[0195] In some implementations, chromosome structural abnormality screening includes insertion screening. This screening involves searching for matrix positions on a standardized and corrected spatial contact matrix where contact strength between different chromosomes is high, determining whether these positions match insertion characteristics (e.g., a significantly enhanced contact strength between a segment within chromosome A and chromosome B), and then inferring a difference index score as a criterion for identifying insertion variations, following the same steps described above.
[0196] In some implementations, CNV inspection includes the following steps:
[0197] (i) Determine the sequencing depth at different genomic locations from the sequencing data of the sample to be tested (e.g., raw fastQ data) to make a preliminary judgment on CNV interval information;
[0198] (ii) Apply the spatial contact numerical matrix to determine whether the contact density of the CNV region matches the CNV copy number determined based on sequencing depth, and filter out inconsistent CNV results; in some preferred embodiments, the determination can be made by calculating the ratio of the spatial contact intensity of the CNV region to the spatial contact intensity on both sides of the CNV region.
[0199] In some implementation schemes, the SV detection method of the present invention further includes: merging and statistically analyzing all detected SV results, inferring the fusion chromosome state and the origin of the fusion fragment after the occurrence of SV based on the consistency of SV breakpoints, determining whether there is a combination of multiple variant results, and outputting in a standardized format; and visualizing and outputting variant information for the standardized output results.
[0200] The following embodiments are described to aid in understanding the invention. It is not intended, and should not be construed in any way, as limiting the scope of the invention. Example
[0201] Example 1: Detection of chromosomal structural variations in two carriers of balanced chromosomal translocations and one normal individual.
[0202] 1) According to the karyotype report, the karyotypes of the two balanced translocation carriers were t(2;7)(q34;q31) and t(8;14)(p21.1;q24.3), respectively, while the other carrier was karyotype negative. Approximately 2 ml of non-frozen anticoagulated peripheral blood was collected from each of the three subjects and mixed with 2 ml of PBS. 3 ml of Ficoll-Paque solution was added to a new 15 ml Falcon tube. The mixed peripheral blood was slowly added along the tube wall to the top layer of the Ficoll-Paque solution, carefully avoiding blood from entering the lower layer. After turning off the centrifuge brake, the tube was centrifuged at 400 g for 40 min to separate the solutions. The top layer of plasma was removed. The PBMC layer was placed in a new Falcon tube, mixed gently with 3 times its volume of PBS, and centrifuged at 400 g for 10 min. The supernatant was discarded, and the cell pellet was resuspended in 3 ml of PBS. After centrifugation at 100 g for 10 min, the cells were resuspended in 500 μl-1 ml of PBS, and the number of PBMC cells was counted using a hemocytometer.
[0203] 2) Fixed PBMC: Take 2x10 6 One PBMC cell was prepared and brought to a final volume of 946 μl with PBS. 54 μl of 37% formaldehyde was added to achieve a final formaldehyde concentration of approximately 2%. The mixture was fixed at room temperature for 10 min, with the sample inverted and mixed every 2 min during this period. Then, 54 μl of 2.5 M glycine was added to quench the reaction. The mixture was incubated at room temperature for 5 min, then on ice for 15 min, and centrifuged at 2000g for 10 min. The supernatant was discarded, and the cells were resuspended in 1 ml of PBS and centrifuged at 2000g for 10 min. The cell pellet was retained after discarding the supernatant.
[0204] 3) PBMC lysis and chromosome digestion: Add 1 ml of ice-cold lysis buffer (10 mM pH 8.0 Tris-HCl, 10 mM NaCl, 0.2% IGEPAL (Sigma, I8896) and 10 μl of protease inhibitor (Clotech, RM02916) to the cell pellet obtained in step 2, incubate on ice for 15 min, then centrifuge at 5000g for 5 min; after discarding the supernatant, add 100 μl of DNase I buffer (Thermo, EN0521) containing 0.2% SDS and lyse the cells at 37°C for 10 min, then add 100 μl of DNase I buffer (2% Triton X-100) and incubate at 37°C for 10 min to stop the PBMC lysis reaction; then add 1 U of DNase I (Thermo, EN0521) to digest the chromosomes for 4 min, and use stop buffer (125 mM EDTA, 2.5%). The chromatin digestion reaction was terminated with SDS. The mixture was centrifuged at 5000g for 5 min to remove the supernatant. The precipitate was resuspended in 50 μl of nuclease-free water, and then 100 μl of AMPure magnetic beads (Beckman, A63880) was added to recover the DNA product. After purification, the magnetic beads were not removed. The product was dissolved in 55 μl of buffer 1 (10 mM pH 8.0 Tris-HCl, 50 mM NaCl, 10 mM MgCl2, 1 mM DTT) to obtain product A. Product A was quantified using Qubit.
[0205] 4) Biotin labeling and ligation: The sticky ends of the chromosomes obtained in step 3 were labeled with biotin-labeled dATP, and the reaction system used for biotin labeling is shown in Table 1 below.
[0206] Product A 55 dCTP / dTTP / dGTP 10mM each mix(Thermo,R0181) 1.13 biotin dATP 1mM (Thermo, EP0052) 5 Klenow polymerase 5 U / μl (Thermo, 18012096) 5 <![CDATA[H2O]]> 33.9 Total 100
[0207] After bathing in a metal oscillating mixer (Eppendorf ThermoMixer C) at 23°C for 1 hour at a medium-low speed, the connection system shown in Table 2 below was added.
[0208] Table 2: Connection System
[0209] 10x T4 ligase buffer(NEB,M0202V) 100 50% PEG 8000 (Sigma, 89510) 100 T4 DNA ligase 5U / ul(NEB,M0202V) 4 <![CDATA[H2O]]> 696 Biotin-labeled products 100 Total 1000
[0210] The linking system of Table 2 was used in a metal oscillating mixer at 22°C for 1 hour to link chromosomes that are spatially close together.
[0211] 5) Protein digestion, cross-linking, and DNA extraction: Centrifuge (1000g for 10min) to process the product from step 4, collect the precipitate and resuspend it in 150μl of buffer 1, add 5μl of PBS buffer with 10% SDS and 8μl of proteinase K (800U / ml) (NEB, P8107S), incubate at 60℃ for 1h, add 82μl of AMPure magnetic beads (0.5X) to purify and extract DNA to obtain product B. Use Qubit to quantify the DNA concentration, which should be above 1ng / ul.
[0212] 6) NGS Library Construction and Biotin Capture: A commercially available enzyme digestion kit was used to break down the DNA of product B to 400-600 bp. After the reaction, 30 μl of AMPure magnetic beads (0.6x) (Beckman, A63880) was added to 50 μl of the reaction product to remove large fragments, followed by 20 μl of magnetic beads (0.4x) to recover smaller fragments for fragment selection. The selected fragments were then subjected to end repair, A-tailing, and sequencing adapter ligation using a commercially available kit to obtain product C.
[0213] Biotin-containing DNA in product C was captured using streptavidin magnetic beads (Dynabeads MyOne Streptavidin C1). 30 μl of C1 beads were washed twice with 100 μl of 1X BW buffer (5 mM pH 8.0 Tris-HCl, 1 M NaCl, 0.5 mM EDTA), and then resuspended in 100 μl of 2X BW buffer. Product C was added to the bead suspension and incubated at room temperature for 15 min. The product was then washed three times with 1X BW buffer containing 0.1% Tween 20, followed by two washes with 200 μl of 10 mM pH 8.0 Tris-HCl. The product was dissolved in 25 μl of Tris-HCl and quantified using Qubit to obtain product D at a concentration of 0.4–5 ng / μl.
[0214] 7) Library amplification and sequencing: Product D was amplified using a commercial library amplification kit. After PCR amplification, the above system was purified using 50 μl of AMPure magnetic beads (1X) to obtain product E before sequencing, which was quantified using Qubit. Fragment size was analyzed using a Qsep Automated Nucleic Acid Fragment Analyzer (Qsep100) to assess whether it met the sequencing criteria. Figure 3 As shown in the figure, Qsep fragment quality inspection results indicate that the above method can produce fragments with a size between 300-500bp, which meets the requirements for instrumentation.
[0215] 8) Sequencing and data quality control
[0216] Table 3 below shows the quality control parameters for the sample sequencing data. As can be seen from Table 3, when applying the improved hi-C protocol to 1.7-2 million cells, with a "pair-ended" sequencing length of 2x150 bp and 20 million sequencing data entries, the effective contacts usable for analysis reached 67.22%, and duplicates were less than 1%, which is superior to conventional quality control data. This indicates that although the experimental procedure was shortened from approximately 4 days using the classic hi-C technology to 2.5 days, good hi-C library quality and sequencing data quality were still obtained.
[0217] Table 3: Summary of quality control parameters for Hi-C library sequencing data
[0218]
[0219] 9) Offline data processing and standardization
[0220] The sequencing data is processed using Juicer (v1.6) to output spatial contact matrix information in HIC format and sequence information in BAM format. The reference genome used for sequence mapping can be any of the human reference genomes GRCh37, GRCh38, T2T-CHM13, etc. This example uses GRCh38 as an example for detailed explanation. The HIC format spatial contact matrix is converted to Cool format using HiCLift (v1.0), and further converted to HICPro format using HICExplorer (v3.7.2). This format contains a matrix file recording contact strength and a BED file recording contact locations. The spatial contact matrix is divided into 1Mb equal-sized bins by rows and columns, and the interaction strength value between each pair of bins is recorded. Based on the long range ratio in the current sample's Juicer processing results, the interaction strength of each bin is divided by the long range ratio of that sample to obtain the spatial contact strength value for that bin.
[0221] 10) Construction of reference system and contact strength correction
[0222] Twenty control samples without large-fragment spatial contact (SV) fragments were collected and processed using the same method as the test samples to obtain a standardized spatial contact matrix. For each bin in the spatial contact matrix, the mean contact intensity of all samples with non-zero values in that bin was taken as the baseline contact intensity value for that bin in the matrix reference system. This baseline value was subtracted from all subsequent test samples after standardization to generate a calibration matrix, as shown below. Figure 4 As shown.
[0223] 11) Balance displacement detection
[0224] Using `hic_breakfinder`, the uncorrected spatial contact matrix of the test sample is input to initially calculate the location of the equilibrium translocation breakpoint. Further, the corrected spatial contact matrix is converted to a numerical matrix using HiSV format conversion. The initial equilibrium translocation breakpoint location obtained from `hic_breakfinder` is set as the hypothetical breakpoint, and a difference index score is calculated based on the difference in contact density on both sides of this hypothetical breakpoint. Then, the hypothetical breakpoint is moved upstream and downstream in 100kb steps by 5Mb. The difference index score is recalculated for each new hypothetical breakpoint obtained after each movement. The location of the highest final score that meets the following conditions is determined as the actual breakpoint location: if the highest score is higher than a preset score threshold, the breakpoint is considered positive and reported as a detection result; otherwise, the breakpoint is considered a false positive and not reported as a detection result. Furthermore, for each combination of two chromosomes, the average contact density within 5Mb of the centromeres on both sides is statistically analyzed. For results that meet the characteristics of balanced translocation, the difference index score is calculated as described above to determine whether a balanced translocation exists and to determine the breakpoint of the balanced translocation.
[0225] The spatial contact matrix diagrams obtained after standardization and correction for samples one and two are shown below. Figure 5 and Figure 6 In. Figure 5 In the matrix diagram (the horizontal axis represents the coordinates of chromosome 2, and the vertical axis represents the coordinates of chromosome 7), a significantly darker color patch can be observed in the interaction region between chromosomes 2 and 7, corresponding to the translocation position in sample 1. Figure 6 In the study, it can be observed that in the interaction region between chromosomes 8 and 14, which correspond to the translocation position of sample 2, there are patches of color that are significantly darker. Figure 6 The matrix diagram above shows the interactions between all chromosomes, with "butterfly-shaped" color blocks that are clearly visible in the interaction regions corresponding to chromosomes 8 and 14. Figure 6 The matrix below (horizontal axis represents the coordinates of chromosome 8, and vertical axis represents the coordinates of chromosome 14) is a magnified display of the "butterfly-shaped" color block.
[0226] Using the determined actual breakpoint locations as reference points, the relative contact strengths upstream and downstream are calculated. In this example, the breakpoint locations of the balanced translocation positive sample 1 are 203100000 of chr2 and 111600000 of chr7. The contact strengths upstream of the chr2 breakpoint and downstream of the chr7 breakpoint, as well as upstream of the chr7 breakpoint and downstream of the chr2 breakpoint, are higher than the contact strengths in the other two directions. Therefore, it is determined that after the balanced translocation, the new chromosome composition should be one chr7(0-111600000)+chr2(203100000-242193529) and another chr2(0-203100000)+chr7(111600000-159345973). Therefore, the corresponding karyotype can be determined to be t(2;7)(q34;q31), which is consistent with the karyotype report. In this example, the breakpoints of the balanced translocation positive sample 2 are at 24400000 of chr8 and 80400000 of chr14. The contact strength downstream of the chr8 breakpoint and downstream of the chr14 breakpoint, as well as the contact strength upstream of the chr8 breakpoint and upstream of the chr14 breakpoint, are higher than the contact strengths in the other two directions. Therefore, it is determined that after the balanced translocation, the new chromosome composition should be one chr14(107043718-80400000)+chr8(24400000-145138636) and another chr14(0-80400000)+chr8(24400000-0). The karyotype can be determined as t(8;14)(p21.1;q24.3), which is consistent with the karyotype report.
[0227] 12) Standardized output of mutation results
[0228] Based on the inferred balanced translocation results, using the structure and chromosomal banding diagram of normal chromosomes as a template, the variant chromosomes are disassembled and reassembled according to the inferred structure, and the resulting visualization is output as follows: Figure 7 As shown. When no balanced translocation is detected, a normal chromosome diagram is output as the result, such as... Figure 8 As shown.
[0229] Example 2: System Step Test
[0230] This example is a comparative example, designed to examine the impact of adjusting the method steps based on conventional Hi-C technology on Hi-C sequencing libraries and sequencing data.
[0231] Specifically, compared to Example 1, Example 2 adjusted several system steps and extended the overall experimental time by referring to the traditional hi-C technology scheme to examine the impact on Hi-C library quality and sequencing data. Specifically, step three in Example 1 only required 2 hours, while in this example, biotin labeling and ligation in step three were extended to 4 hours and overnight treatment, respectively. In addition, in this example, in step five, after decrosslinking and before purifying DNA with magnetic beads, a step of removing biotin from the ends of unligated fragments was added to ensure that only ligated DNA fragments were captured and sequenced (this step increased the time by 2 hours and increased the experimental cost); and in step five, before purifying DNA with magnetic beads, an additional DNA purification and extraction operation of precipitating DNA with sodium acetate / isopropanol was added (this step further increased the operation time and experimental cost). In short, basically following the description in Belaghzal H, Dekker J, Gibcus J H. Hi-C 2.0: An optimized Hi-C procedure for high-resolution genome-wide mapping of chromosome conformation[J]. Methods, 2017, 123: 56-65, in step five, after treating the DNA with proteinase K, the unligated biotin at the ends is removed using T4 DNA polymerase and a low concentration of dNTPs; the DNA is precipitated using sodium acetate / isopropanol at -80°C for 20 min, centrifuged for 40 min to collect the precipitated DNA, and then purified with magnetic beads to extract the DNA.
[0232] The balanced translocation carrier samples used in this embodiment have a karyotype t(6;13)(q23;q14). Hi-C library construction and sequencing data acquisition were performed as follows.
[0233] Table 4: Summary of quality control parameters for Hi-C library sequencing data
[0234]
[0235]
[0236] As shown in Table 4 above, with the same total reads as in Example 1, the effective Contacts data available for analysis is 52.14%, lower than the corresponding data in Example 1. Furthermore, the proportion of multimap (unnecessary multimapped reads) is also significantly higher than in Example 1 (34.59% vs. 14.87%). This indicates that adding the above steps or extending the operation time does not improve the performance of SV detection.
[0237] Example 3: Comparison of sequencing library quality
[0238] In this embodiment, the experimental steps of Example 1 were followed, and the same samples (Trans6-13) as in Example 3 were used to compare and test the quality of sequencing libraries constructed by the method of the present invention and those constructed using a commercial Hi-C sequencing library construction kit. The Arima Hi-C kit was commercially available. Hi-C library construction was performed according to Example 1 and the Arima kit's instruction manual. Table 5 below summarizes the sequencing data quality control parameters for the hi-C libraries constructed using the two methods.
[0239] Table 5: Summary of quality control parameters for Hi-C library sequencing data
[0240]
[0241] As can be seen from the table above, the quality control indicators of the sequencing data obtained by the method of this invention are comparable to those of Arima. However, compared with the method of this invention, the Arima kit requires the simultaneous use of a mixture of multiple restriction endonucleases in the chromosome digestion system, which makes the entire process much more expensive and difficult to meet the cost control requirements of detection.
[0242] Example 4: Detection of Roche translocation samples
[0243] Following the experimental steps of Example 1, the detection performance of the method of the present invention on Roche translocation samples was tested in this example. The samples used had a karyotype 45,XY,rob(13;14)(q10;q10). The Hi-C contact matrix (partial) obtained by the method described in Example 1 is shown in [the diagram / image]. Figure 9 In the figure, in the region corresponding to chr13-chr14, a significantly deepened Roche translocation positive signal can be observed in sample 45,XY,rob(13;14)(q10;q10).
[0244] From a quantitative perspective, the average contact strength between the long arms of chromosomes 13 and 14 should be higher than that of chromosomes without Robertsonian translocations. In this embodiment, the average contact strength within a 30Mb range near the centromere of the long arm and the average contact strength within a 30Mb range near the telomere of the long arm were first calculated for the chromosome 13 / 14 combination. Further, the average contact strength within a 30Mb range near the centromere of the long arm was calculated for chromosome combinations 13 / 15, 13 / 21, 13 / 22, 14 / 15, 14 / 21, and 14 / 22. A difference index score was calculated using the ratio of the centromere contact strength of chromosome 13-14 to the telomere contact strength, and the product of the squares of the centromere contact strength of chromosome 13-14 and the average centromere contact strength of other chromosome combinations. This calculation was then performed on other chromosome combinations that may have Robertsonian translocations (chromosomes 13, 14, 15, 21, and 22). Ultimately, only the combination of chromosomes 13 and 14 scored above the threshold and was reported as the result of the Robertson translocation.
[0245] Example 5: Detection of chromosome inversion samples
[0246] Following the experimental steps of Example 1, the detection performance of the method of the present invention on inverted samples was tested in this example. The samples used had a karyotype of 46,XX,inv(6)(p22q22). The Hi-C contact matrix obtained by the method described in Example 1 is shown in... Figure 10 In the middle. For example Figure 10 As shown, the inverted positive signal of sample 46,XX,inv(6)(p22q22) was successfully detected.
[0247] To determine the breakpoint location, after correcting for the contact density matrix within the chromosome, breakpoints with contact strength greater than a threshold (e.g., the median of the highest difference index score obtained from samples actually carrying inversions and samples not carrying inversions) are searched. The actual breakpoint location is inferred at the searched breakpoints using the method described in Example 1 (11).
[0248] Example 6: Detection of chromosome insertion samples
[0249] Following the experimental steps of Example 1, the detection performance of the method of the present invention on inserted samples was tested in this example. The samples used had a karyotype of 46,XX,ins(14;8)(q24;q13q11.2). The Hi-C contact matrix obtained by the method described in Example 1 is shown in... Figure 11 In the middle. For example Figure 11 As shown, insertion-positive signals were successfully detected at both the fragment origin (chromosome 8) and the insertion location (chromosome 14) of the sample insertion variant.
[0250] To determine the breakpoint location of the inserted fragment's origin, the two boundary locations with the greatest difference in contact intensity on chromosome 8 were searched as the two breakpoints on either side of the inserted fragment. To determine the fragment insertion location, the location with the strongest signal (e.g., in the region where there is a significant enhancement of the contact signal between chromosome 14 and chromosome 8) was searched. Figure 7 (As shown in the blue box), the average of the two positions on chr14 is the fragment insertion position.
[0251] To determine the direction of the inserted segment, based on the already determined insertion position, the relative direction between the region of strongest signal and the insertion position is evaluated. For example... Figure 7 As shown, the two blue boxes correspond to the combination of the chr8L direction and the chr14R direction, and the combination of the chr8R direction and the chr14L direction, respectively. Therefore, it can be determined that the L side of chr8 is closer to the R side of chr14, and the R side of chr8 is closer to the L side of chr14. Thus, the insertion sequence should be a reverse insertion. The actual breakpoint location is inferred at the found breakpoint using the method described in Example 1 (11).
[0252] Example 7: Detection of Copy Number Variation Samples
[0253] Following the experimental steps of Example 1, the detection performance of the method of the present invention on samples with copy number variations was tested in this example. The samples used had a CNV-seq result of 16P13.11 (15126890-16308351) x 3 repeats. After mapping the sequencing data obtained using the method described in Example 1 to a reference genome, the sequencing depth was determined by the number of reads mapped to each segment of the reference genome, and a CNV-plot was plotted. The results are shown in... Figure 12 In the CNV results, regions matching CNV characteristics (copy numbers of 0, 1, 3, 4, etc.) are verified using a spatial contact matrix to determine if the ratio of the spatial contact intensity of this region to the spatial contact intensity on both sides of the CNV region matches the CNV copy number determined based on sequencing depth. CNVs with inconsistent results are filtered out, and those meeting the criteria are reported as CNV results.
[0254] like Figure 12 As shown in the arrow, a positive signal of sample 16P13.11(15126890-16308351)x3 was successfully detected on the CNV-plot, which is a 1.18Mb duplication of chromosome 16.
[0255] Example 8: Determining Complex Mutant Combinations
[0256] In this embodiment, the same experimental steps as in Example 1 were used, and complex variant carrier samples (with insufficient karyotype detection capability, reported as 46,XX,der(8)t(8;18)(q11.2;q21)?,der(14)ins(14;8)(q24;q13q11.2)?,der(18)t(8;18)(q13;q21)?) were used to detect chromosomal structural variations. After sequencing and analysis, using the methods described in Examples 1, 5, and 6, a total of 2 translocation breakpoints (chr8:70100000, chr18:42800000), 2 inversion breakpoints (chr18:40100000, chr18:42600000), 4 insertion fragments (chr8:70100000-48000000, chr18:40500000-41900000, chr8:31900000-32600000, chr18:44300000-47800000), and 2 insertion sites (chr14:80450000, chr15:61600000) were detected. All detected variants were combined and statistically analyzed based on breakpoint consistency. The combined chromosome structure is as follows:
[0257] chr8(0-31900000)+chr8(32600000-48000000)+chr18(40500000-0100000)+chr18(42800000-44300000)+chr18(47800000-80373285)
[0258] chr18(0-40100000)+chr18(41900000-42600000)+chr8(70100000-145138636)
[0259] chr14(0-80450000)+chr8(70100000-48000000)+chr14(80450000-107043718)
[0260] chr15(0-61600000)+chr18(40500000-41900000)+chr8(31900000-32600000)+chr18(44300000-47800000)+chr15(61600000-101991189).
[0261] The karyotype results are reported as: 46,XX,der(8)(8pter->8q12::18q13->18ter),der(14)(14pter->14q31.1::8q13->8q11::14q31.1->14ter),der(15)(15pter->15q22::18q12::8p12::18q12q21::15q22->15qter),der(18)(18pter->18q12.3->18q12.3::8113.3->8qter). The visualized variant structure is as follows: Figure 13 As shown.
Claims
1. A method for non-diagnostic purposes based on chromosome conformation capture for detecting chromosomal structural variations in a sample, comprising: (a) Cross-linking, digestion, and ligation of chromosomal DNA in sample cells, wherein the digestion is performed in the presence of DNase I; (b) Using the enzyme digestion and ligation products from step (a), construct a sequencing library; (c) Perform low-depth paired-end sequencing on the sequencing library, with a sequencing depth not exceeding 3x, using pair-ended sequencing, and the sequencing data volume is 20M reads; (d) Based on sequencing data, assess chromosomal structural variations in the sample, wherein the structural variations are selected from: structural rearrangements and copy number abnormalities (CNVs). Steps (a) and (b) are performed as follows: (i) Immobilize 1 to 5 million nucleated cells using a cross-linking agent to cross-link chromosomal DNA; (ii) Lyse cells using cell lysis buffer; (iii) Perform enzymatic digestion in a 200-500 μL reaction system for 1-6 minutes in the presence of 1-5 U DNase I; (iv) In a reaction system containing dNTPs and Klenow polymer, end repair of the digested DNA fragments is performed at a temperature of 20-30°C, wherein one of the dNTPs is biotin-labeled, and wherein the end repair reaction lasts for 1 hour. (v) Ligate the digested fragments at 20-30°C for 1 hour in a reaction system containing 1-10% PEG8000 and 5 U / ml to 50 U / ml T4 DNA ligase. (vi) By decrosslinking, the DNA and protein in the crosslinked complex are separated. (vii) Without removing biotin from the ends of unconnected chromosomal DNA fragments, biotin-tagged chromosomal DNA fragments are captured for use in constructing sequencing libraries.
2. The method of claim 1, wherein the structural rearrangement is selected from balanced translocation, unbalanced translocation, mutual translocation, non-mutual translocation, Rogowski translocation, complex translocation, intraarm inversion, interarm inversion, and insertion.
3. The method of claim 1, wherein the sequencing depth is 2x-3x.
4. The method of claim 1, wherein the sample is a mammalian cell, body fluid, or tissue sample.
5. The method of claim 1, wherein the sample is an isolated peripheral blood mononuclear cell (PBMC) sample.
6. The method of claim 4, wherein the sample comprises no more than 2 million nucleated cells.
7. The method of claim 1, wherein the ligation reaction system is 1.0 ml, containing 5% PEG8000 and 20 U of T4 DNA ligase.
8. The method of claim 1, wherein, Step (vi) includes decrosslinking the enzyme-digested ligation product to separate the DNA from the protein in the crosslinked complex; and purifying the separated DNA.
9. The method of claim 8, wherein the DNA purification is performed using magnetic beads.
10. The method of claim 1, wherein, In step (iii), the digestion reaction time shall not exceed 5 minutes.
11. The method of claim 10, wherein the DNA is purified using magnetic beads after enzymatic digestion.
12. The method of claim 1, wherein, In step (ii), the cell lysate contains IGEPA and a protease inhibitor.
13. The method of claim 12, wherein the cell lysis buffer comprises 8-15 mM pH 8.0 Tris-HCl, 7-15 mM NaCl, 0.1-0.5% IGEPAL and 5-20 μl of protease inhibitor.
14. The method of claim 13, wherein the cell lysis buffer comprises 10 mM pH 8.0 Tris-HCl, 10 mM NaCl, 0.2% IGEPAL and 10 μl of protease inhibitor.
15. The method of claim 1, wherein: (i) The capture was performed using streptavidin magnetic beads; (ii) The crosslinking is performed using a formaldehyde crosslinking agent; (iii) Treat the chromosomes with a combination of 0.2% SDS and 2% Triton prior to enzyme digestion; and / or (iv) Decrosslinking in a buffer containing SDS and proteinase K.
16. The method of claim 15, wherein the crosslinking is performed using formaldehyde at a concentration of 1%-3%.
17. The method of claim 1, wherein library construction comprises: The DNA in the enzyme digestion and ligation product was fragmented and selected for fragments of 150bp-600bp for library construction.
18. The method of claim 17, wherein the DNA fragment screening is performed by adding magnetic beads to the fragmented DNA product in two steps at a ratio of 0.6x and 0.4x magnetic beads.
19. The method of claim 1, wherein, Step (d) includes creating a spatial contact matrix based on sequencing data and examining changes in chromosome interaction patterns relative to a control reference system to detect chromosome structural variations.
20. The method of claim 19, wherein step (d) comprises: (d1) Transform the sequencing data into a spatial contact matrix; (d2) The spatial contact matrix is corrected and normalized, wherein the correction is performed using a matrix reference frame constructed using a self-comparison reference frame; (d3) Examine chromosome structural variations on the corrected and normalized spatial contact matrix. (d4) Visualize the detected chromosomal structural variations and output variation information.
21. The method of any one of claims 1-20, wherein the sequencing data contains at least 50% valid contacts.
22. The method of claim 21, wherein the sequencing data contains at least 60% or 65% valid contacts.