A method for detecting chromosomal abnormalities using inter-nucleic acid fragment distance information.
By calculating the distance between nucleic acid fragments in a biological sample aligned with a reference genome, the method addresses the limitations of existing chromosomal abnormality detection techniques, offering high sensitivity and accuracy in detecting fetal and tumor-related abnormalities.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- GREEN CROSS GENOME CORP
- Filing Date
- 2020-08-19
- Publication Date
- 2026-05-15
AI Technical Summary
Current methods for detecting chromosomal abnormalities, such as karyotyping, FISH, DNA microarrays, and next-generation sequencing, are time-consuming, costly, and lack sufficient sensitivity and accuracy, while invasive prenatal tests carry risks and have low sensitivity.
A method involving the extraction of nucleic acids from a biological sample, alignment with a reference genome, and calculation of the distance between nucleic acid fragments to determine chromosomal abnormalities using Fragments Distance Index (FDI) or Read Distance Index (RDI), enabling high sensitivity and accuracy.
The method achieves high sensitivity and accuracy in detecting chromosomal abnormalities, including fetal and tumor-related abnormalities, with reduced invasiveness and cost, and provides a reliable diagnostic tool.
Smart Images

Figure 0007859967000010 
Figure 0007859967000011 
Figure 0007859967000012
Abstract
Description
[Technical Field]
[0001] The present invention relates to a method for detecting chromosomal abnormalities using information on the distance between nucleic acid fragments, and more specifically, to a method for detecting chromosomal abnormalities that uses a method of extracting nucleic acids from a biological sample to obtain sequence information, and then calculating the distance between reference values of nucleic acid fragments. [Background technology]
[0002] Chromosomal abnormalities are associated with genetic defects and tumor diseases. Chromosomal abnormalities can also refer to chromosomal deletions or duplications, partial deletions or duplications of chromosomes, or chromosomal damage (break), translocation, or inversion. Chromosomal abnormalities are a type of genetic imbalance that can induce fetal death or serious physical and mental defects and tumor diseases. For example, Down's syndrome is a common form of chromosomal number abnormality caused by the presence of three copies of chromosome 21 (trisomy 21). Edwards syndrome (trisomy 18), Patau syndrome (trisomy 13), Turner syndrome (XO), and Klinefelter syndrome (XXY) are also chromosomal number abnormalities. Chromosomal abnormalities are also found in tumor patients. For example, duplication of the 4q, 11q, and 22q regions and deletion of the 13q region were observed in liver cancer patients (liver adenomas and adenocarcinomas), while duplication of the 2p, 2q, 6p, and 11q regions and deletion of the 6q, 8p, 9p, and 21 chromosome regions were observed in pancreatic cancer patients. These regions are associated with tumor-related oncogenes and tumor suppressor gene regions.
[0003] Chromosomal abnormalities can be detected using karyotyping and FISH (Fluorescent In Situ Hybridization). However, these detection methods are disadvantageous in terms of time, effort, and accuracy. DNA microarrays can also be used for detecting chromosomal abnormalities. In particular, genomic DNA microarray systems are easy to prepare probes for and can detect chromosomal abnormalities not only in extended regions of chromosomes but also in intron regions of chromosomes. However, it is difficult to prepare a large number of DNA fragments whose location and function within the chromosome have been confirmed.
[0004] In recent years, next-generation sequencing technology has been used for chromosome number aberration analysis (Park, H., Kim et al., Nat Genet 2010, 42, 400-405; Kidd, J. Met al., Nature 2008, 453, 56-64). However, this technology requires high coverage reading for chromosome number aberration analysis, and CNV measurement also requires independent validation. As a result, it is very costly and the results are difficult to understand, making it unsuitable as a general gene search analysis at the time.
[0005] Currently, real-time qPCR is used as a cutting-edge technique for quantitative gene analysis because it exhibits a wide dynamic range (Weaver, S. et al., Methods 2010, 50, 271-276) and a reproducibly observed linear correlation between the threshold cycle and the initial target amount (Deepak, S. et al., Curr Genomics 2007, 8, 234-251). However, the sensitivity of qPCR analysis is not sufficiently high to distinguish differences in replication numbers.
[0006] On the other hand, existing prenatal tests for fetal chromosomal abnormalities include ultrasound, blood labeling, amniocentesis, chorionic villus sampling, and percutaneous umbilical cord blood testing (Mujezinovic F, et al. Obstet Gynecol. 2007, 110(3): 687-94). Of these, ultrasound and blood labeling are classified as screening tests, while amniocentesis is classified as a definitive diagnostic test. Ultrasound and blood labeling are non-invasive methods that do not involve direct sample collection from the fetus and are therefore safe, but their sensitivity is low, at less than 80% (ACOG Committee on Practice Bulletins. 2007). Amniocentesis, chorionic villus sampling, and percutaneous umbilical cord blood testing are invasive methods that can definitively diagnose fetal chromosomal abnormalities, but they have the disadvantage of carrying a risk of fetal loss due to the invasive medical procedure.
[0007] In 1997, Lo et al. successfully performed Y-chromosome sequencing analysis of fetal genetic material from maternal plasma and serum, making fetal genetic material in the mother available for prenatal testing (Lo YM, et al. Lancet. 1997, 350(9076):485-7). Fetal genetic material in maternal blood is actually derived from the placenta, as it is a portion of trophoblast cells that have undergone cell death during placental reformation and entered the bloodstream through a substance exchange mechanism. This is defined as cff DNA (cell-free fetal DNA).
[0008] Cff DNA can be detected in most maternal blood as early as day 18 after embryo transfer, and no later than day 37. Because cff DNA is a short strand of less than 300 bp and is present in small amounts in maternal blood, large-scale parallel base sequencing (NGS) techniques are used to detect fetal chromosomal abnormalities. Non-invasive fetal chromosomal abnormality detection using large-scale parallel base sequencing techniques shows a detection sensitivity of 90-99% or more depending on the chromosome, but false positive and false negative results amount to 1-10%, and corrective techniques are currently needed to address this (Gil MM, et al. Ultrasound Obstet Gynecol. 2015, 45(3):249-66).
[0009] Therefore, the inventors of the present invention made intensive efforts to solve the above problems and develop a method for detecting chromosomal abnormalities with high sensitivity and accuracy. As a result, after grouping nucleic acid fragments aligned in chromosomal regions and calculating the distances between the reference values of the nucleic acid fragments and comparing them with the normal group, it was confirmed that chromosomal abnormalities can be detected with high sensitivity and accuracy, and the present invention was thus completed. Summary of the Invention
[0010] An object of the present invention is to provide a method for determining chromosomal abnormalities using distance information between nucleic acid fragments.
[0011] Another object of the present invention is to provide an apparatus for determining chromosomal abnormalities using distance information between nucleic acid fragments.
[0012] Still another object of the present invention is to provide a computer-readable storage medium including instructions configured to be executed by a processor that determines chromosomal abnormalities by the above method.
[0013] To achieve the above object, the present invention provides a method for detecting chromosomal abnormalities by calculating the distances between reference values of nucleic acid fragments extracted from a biological sample.
[0014] The present invention also provides a chromosomal abnormality detection apparatus including: a decoding unit that extracts nucleic acids from a biological sample and decodes sequence information; an alignment unit that aligns the decoded sequences with a standard chromosomal sequence database; and a chromosomal abnormality determination unit that measures the distances between reference values of nucleic acid fragments aligned with selected nucleic acid fragments to calculate an FD value (Fragments Distance), calculates an FDI value (Fragments Distance Index) for each chromosomal whole region or specific gene region based on the calculated FD value, and determines that there is a chromosomal abnormality when the FDI value does not belong to the reference value range.
[0015] The present invention also relates to a computer-readable storage medium including instructions configured to be executed by a processor for detecting chromosomal abnormalities, (A) extracting nucleic acids from a biological sample to obtain nucleic acid fragments and acquiring sequence information; (B) aligning the nucleic acid fragments to a reference genome database based on the acquired sequence information (reads); (C) measuring the distance between reference values of the selected nucleic acid fragments (fragments) and calculating a FD value (Fragments Distance); and (D) calculating a FDI value (Fragments Distance Index) for each whole chromosomal region or specific genetic region based on the FD value calculated in step (C), and determining that there is a chromosomal abnormality when the FDI value does not belong to a reference value or range, and providing a computer-readable storage medium including instructions configured to be executed by a processor for detecting chromosomal abnormalities by the above steps.
[0016] The present invention also provides a method for detecting chromosomal abnormalities, including: (A) extracting nucleic acids from a biological sample and acquiring sequence information; (B) aligning the acquired sequence information (reads) to a reference genome database; (C) measuring the distance between the aligned reads of the aligned sequence information and calculating a RD value (Read Distance); and (D) calculating a RDI value (Read Distance Index) for each whole chromosomal region or specific genetic region based on the RD value calculated in step (C), and determining that there is a chromosomal abnormality when the RDI value does not belong to a reference value range.
[0017] The present invention also provides a chromosomal abnormality detection device that includes a decoding unit for extracting nucleic acids from a biological sample and decoding sequence information; an alignment unit for aligning the decoded sequences in a standard chromosome sequence database; and a chromosomal abnormality determination unit that measures the distance between aligned reads for the selected sequence information (reads) to calculate the RD value (Read Distance), calculates the RDI value (Read Distance Index) for the entire chromosome region or for specific genetic regions based on the calculated RD value, and determines that a chromosomal abnormality exists if the RDI value does not fall within the reference range.
[0018] The present invention also provides a computer-readable storage medium comprising instructions configured to be executed by a processor that detects chromosomal abnormalities, the instructions comprising: (A) extracting nucleic acids from a biological sample and obtaining sequence information; (B) aligning the obtained sequence information (reads) in a reference genome database; (C) measuring the distance between aligned reads for the selected sequence information (reads) and calculating the RD value (Read Distance); and (D) calculating the RDI value (Read Distance Index) for the entire chromosome region or for specific genetic regions based on the RD value calculated in step (C), and determining that there is a chromosomal abnormality if the RDI value does not fall within a reference range. [Brief explanation of the drawing]
[0019] [Figure 1] This is an overall flowchart for determining chromosomal abnormalities based on FD values according to one embodiment of the present invention.
[0020] [Figure 2] This is a conceptual diagram illustrating a method for calculating the FD value of the present invention from reads produced by a single-end sequencing method.
[0021] [Figure 3] This is a conceptual diagram illustrating a method for calculating the FD value of the present invention from reads produced by a paired-end sequencing method.
[0022] [Figure 4] This is a conceptual diagram of a method for correcting FD values using lead-external position information in the present invention.
[0023] [Figure 5] This graph shows the difference in FD values calculated using lead data produced by a paired-end sequencing method in one embodiment of the present invention, with and without using lead-external position information.
[0024] [Figure 6] This is an overall flowchart for determining chromosomal abnormalities based on RD values according to one embodiment of the present invention.
[0025] [Figure 7] This diagram illustrates the concept of Read Distance calculated using an RD value-based method according to one embodiment of the present invention. In the case of Reads used in calculating Reads Distance, the Reads may be used regardless of their alignment direction (Figure 7(A)), or they may be used considering their alignment direction. (Positive direction: Figure 7(B), Negative direction: Figure 7(C))
[0026] [Figure 8] This diagram illustrates the distribution of X chromosome read count and RepRD using an RD value-based method according to one embodiment of the present invention, confirming that the relationship between the two values is nonlinear rather than linear.
[0027] [Figure 9]This diagram illustrates the distribution of chromosome-specific read counts and RepRD in an RD value-based method according to one embodiment of the present invention, where (A) represents a normal chromosome, (B) represents trisomy chromosome 21, (C) represents trisomy chromosome 18, and (D) represents trisomy chromosome 13.
[0028] [Figure 10] This shows the relationship between the RDI value calculated by the RD value-based method according to one embodiment of the present invention and the fetal fraction (Figure 10A), gestational age (Figure 10B), and G-score value (Korean Patent No. 10-1686146, Figure 10C).
[0029] [Figure 11] These are the results of ROC analysis on the normal group and samples confirmed to have aneuploidy in each chromosome, using an RD value-based method according to one embodiment of the present invention.
[0030] [Figure 12] This is the result of confirming the accuracy based on the number of reads in an RD value-based method according to one embodiment of the present invention, where the X axis represents the number of reads and the Y axis represents AUC.
[0031] [Figure 13] This is the result of confirming the correlation between the RD value-based method according to one embodiment of the present invention and the number of reads and chromosomal abnormalities.
[0032] [Figure 14] This is a comparison of the RD value-based method according to one embodiment of the present invention with the results of microarray analysis.
[0033] [Figure 15] This is the result of confirming the RDI value distribution in an RD value-based method according to one embodiment of the present invention, where the RepRD of normal individuals and aneuploidy samples on chromosome 21 was set to the reciprocal of the median.
[0034] [Figure 16] This is the result of confirming the RDI value distribution in an RD value-based method according to one embodiment of the present invention, where the RepRD of normal individuals and aneuploidy samples of chromosome 21 was set as the mean value.
[0035] [Figure 17] This is the result of confirming the RDI value distribution in an RD value-based method according to one embodiment of the present invention, where the RepRD of normal individuals and aneuploidy sample of chromosome 21 was set to the reciprocal of the mean value. [Modes for carrying out the invention]
[0036] Unless otherwise specified, all technical and scientific terms used herein have the same meaning as those commonly understood by skilled experts in the art to which this invention pertains. In general, the nomenclature and experimental methods described below herein are well known and commonly used in the art.
[0037] In this invention, we aimed to confirm that chromosomal abnormalities can be detected with high sensitivity and accuracy when comparing representative values of the chromosome to be analyzed between a normal human population and an experimental subject by aligning sequence information (read) data obtained from a sample onto a reference gene, grouping the aligned nucleic acid fragments, and then calculating the distance between reference values of the nucleic acid fragments.
[0038] The chromosomal abnormality detection method according to the present invention can be used not only for detecting fetal chromosomal abnormalities such as aneuploidy, but also for detecting tumors, i.e., for diagnosing tumors or predicting their prognosis.
[0039] In other words, in one embodiment of the present invention, after sequencing DNA extracted from blood and aligning it with a reference chromosome, the nucleic acid fragments are grouped into a whole group, a forward group, and a reverse group. The distance between reference values of nucleic acid fragments (fragment distance, FD) is calculated for each group, a representative value (RepFD) of the distance between reference values of nucleic acid fragments per genetic region is derived, the RepFD ratio is calculated using a normalization factor, and the group-specific FDI (Fragment Distance Index) value is derived by comparing it with the RepFD ratio in a normal reference population. If all group-specific FDI values are below or above the reference value, a method has been developed to determine that the experimental subject has a chromosomal abnormality (Figure 1).
[0040] Therefore, in one aspect, the present invention relates to a method for detecting chromosomal abnormalities by calculating the distance between reference values of nucleic acid fragments extracted from a biological sample.
[0041] In the present invention, any nucleic acid fragment extracted from a biological sample can be used, but preferably it may be a cell-free nucleic acid or an intracellular nucleic acid fragment, although it is not limited thereto.
[0042] In the present invention, the nucleic acid fragment may be characterized by being obtained by direct sequence analysis, sequence analysis by next-generation nucleotide sequence analysis, or sequence analysis by non-specific whole genome amplification.
[0043] In the present invention, any existing known technique can be used for the method of directly analyzing the sequence of the nucleic acid fragment.
[0044] In the present invention, the method of sequence analysis by nonspecific full-length gene amplification refers to any method of performing sequence analysis after amplifying nucleic acids using random primers.
[0045] In the present invention, a method for calculating the distance between nucleic acid fragment reference values using sequence analysis by next-generation base sequence analysis and determining the presence or absence of chromosomal abnormalities based on this is: (A) The step of extracting nucleic acids from a biological sample to obtain nucleic acid fragments and acquire sequence information; (B) The step of confirming the position of nucleic acid fragments from a standard chromosome sequence database (reference genome database) based on the acquired sequence information (reads); (C) A step of grouping the sequence information (reads) into the overall sequence, the forward sequence, and the reverse sequence; (D) Using the grouped sequence information, define a reference value for each nucleic acid fragment, measure the distance between the reference values, and calculate the FD value (Fragments Distance) for each group; and (E) Based on the FD values for each group calculated in step (D) above, the FDI value (Fragments Distance Index) is calculated for the entire chromosome region or for specific regions, and if none of the FDI values fall within the reference range, it is determined that there is a chromosomal abnormality; It may be characterized by being carried out in a manner that includes, but is not limited to, the above.
[0046] In this invention, the term "chromosomal abnormality" refers to various mutations that occur in chromosomes, but can be broadly classified into numerical abnormalities, structural abnormalities, microdeletions, and chromosomal instability.
[0047] Chromosomal number abnormalities refer to cases where there is an abnormality in the number of chromosomes, and can include any case where the abnormality occurs from the total number of chromosomes, which is 23 pairs and 46, such as Down syndrome (where there is an extra chromosome 21, resulting in a total of 47 chromosomes), Turner syndrome (where there is a single X chromosome, resulting in a total of 45 chromosomes), and Klinefelter syndrome (where there are chromosome numbers such as XXYY, XXXY, XXXXY, etc.).
[0048] Chromosomal structural abnormalities refer to any case in which the number of chromosomes does not change, but the structure of the chromosome changes, such as deletions, duplications, inversions, translocations, fusions, and microsatellite instability (MSI-H). Examples include partial deletion of chromosome 5 (Cat's Cry syndrome), partial deletion of chromosome 7 (Williams syndrome), and partial duplication of chromosome 12 (Wolf-Hirschhorn syndrome). Chromosomal structural abnormalities found in tumor patients include translocation between chromosomes 9 and 22 (chronic myeloid leukemia), duplication of the 4q, 11q, and 22q regions and deletion of the 13q region (liver cancer), duplication of the 2p, 2q, 6p, and 11q regions and deletion of the 6q, 8p, 9p, and 21 chromosome regions (pancreatic cancer), TMPRSS2-TRG gene fusion (prostate cancer), and microsatellite instability throughout the chromosome (colorectal cancer). These regions are associated with tumor-related oncogenes and tumor suppressor gene regions, but are not limited to those mentioned above.
[0049] In the present invention, The aforementioned (A) stage is, (Ai) The step of obtaining nucleic acids from blood, semen, vaginal cells, hair, saliva, urine, oral cells, placental cells or fetal cells, amniotic fluid, tissue cells and mixtures thereof; (A-ii) The step of removing proteins, fats, and other residues from the collected nucleic acids using a salting-out method, a column chromatography method, or a beads method to obtain purified nucleic acids; (A-iii) The step of preparing a single-end sequencing or pair-end sequencing library from purified nucleic acids or nucleic acids randomly fragmented by enzymatic cleavage, pulverization, or hydroshear method; (A-iv) The step of reacting the prepared library with a next-generation sequencer; and (Av) This may be characterized by including a step of obtaining nucleic acid sequence information (reads) from a next-generation sequencer.
[0050] In the present invention, the next-generation sequencer may be used in any sequencing method known in the art. Sequencing of nucleic acids separated by the selection method is typically performed using next-generation sequencing (NGS). Next-generation sequencing includes any sequencing method that determines the sequence of one nucleotide from individual nucleic acid molecules or from a cloned, extended proxy for individual nucleic acid molecules in a highly similar manner (e.g., 10 5 (More than 100 molecules are sequenced simultaneously.) In one embodiment, the relative abundance of nucleic acid species in a library can be estimated by measuring the relative number of its congeneral sequences from the data produced by the sequencing experiment. Next-generation sequencing methods are publicly known in the art and are described, for example, in the literature incorporated herein by reference (Metzker, M. (2010) Nature Biotechnology Reviews 11:31-46).
[0051] In one embodiment, next-generation sequencing is performed to determine the nucleotide sequence of individual nucleic acid molecules (e.g., HeliScope Gene Sequencing system from Helicos BioSciences and PacBio RS system from Pacific BioSciences). In other embodiments, sequencing methods, such as high-volume parallel short-read sequencing that generates more sequence bases per sequencing unit compared to other sequencing methods that generate fewer but longer reads (e.g., Illumina Inc. Solexa sequencer in San Diego, California), determine the nucleotide sequences of cloned and extended proxies for individual nucleic acid molecules (e.g., Illumina Inc. Solexa sequencer in San Diego, California; Life Sciences (located in Branford, Connecticut) and Ion Torrent). Other methods or machines for next-generation sequencing are provided by, but are not limited to, 454 Life Sciences (located in Branford, Connecticut), Applied BioSeaSystems (located in Foster City, California; SOLiD sequencers), Helicos Biosciences Corporation (located in Campbridge, Massachusetts), and emulsion and microfluidic sequencing techniques such as nanoinfusions (e.g., GnuBio infusions).
[0052] Platforms for next-generation sequencing include, but are not limited to, Roche / 454's Genome Sequencer (GS) FLX system, Illumina / Solexa Genome Analyzer (GA), Life / APG's Support Oligonucleotide Ligation Detection (SOLiD) system, Polonator's G.007 system, Helicos BioSciences' HeliScope Gene Sequencing system, Pacific Biosciences' PacBio RS system, and MGI's DNBseq.
[0053] NGS Technologies may include, for example, one or more of the following stages: mold making, sequencing, and imaging and data analysis.
[0054] The template preparation step and the method for preparing the template may include steps such as randomly disrupting nucleic acids (e.g., genomic DNA or cDNA) into small pieces, and creating a sequencing template (e.g., a fragment template or a mate-to-mate template). Spatially separated templates may be attached to or fixed to a solid surface or support, which allows for a large number of sequencing reactions to occur simultaneously. Types of templates usable for NGS reactions include, for example, templates from which clones derived from a single DNA molecule have been amplified, and single DNA molecule templates.
[0055] Methods for producing a cloned template include, for example, emulsion PCR (emPCR) and solid-phase amplification.
[0056] EmPCR can be used to produce templates for NGS. Typically, a library of nucleic acid fragments is prepared, and adapters containing typical priming sites are ligated to the ends of the fragments. The fragments are then denatured into single strands and captured by beads. Each bead captures a single nucleic acid molecule. After amplification and enrichment of emPCR beads, a large amount of template can be attached and immobilized on a polyacrylamide gel on a standard microscope slide (e.g., Polonator), chemically crosslinked to an amino-coated glass surface (e.g., Life / APG; Polonator), or deposited onto individual picoTiterPlate (PTP) wells (e.g., Roche / 454), at which point the NGS reaction can be performed.
[0057] Solid-phase amplification can also be used to generate templates for NGS. Typically, forward and backward primers adhere covalently to a solid support. The surface density of the amplified fragments is defined as the primer-to-template ratio on the support. Solid-phase amplification can generate millions of spatially separated template clusters (e.g., Illumina / Solexa). The ends of the template clusters may be hybridized with conventional primers for the NGS reaction.
[0058] Other methods for producing cloned amplified templates include, for example, multiple displacement amplification (MDA) (Lasken RSCurr Opin Microbiol. 2007;10(5):510-6). MDA is a non-PCR-based DNA amplification technique. The reaction involves the steps of randomly annealing hexamer primers to the template and synthesizing DNA at a constant temperature using a high-fidelity enzyme, typically a Φ29 polymerase. MDA can produce large-sized products with a lower error rate.
[0059] Template amplification methods such as PCR can either bind an NGS platform to a target or enrich specific regions of the genome (e.g., exons). Representative template enrichment methods include, for example, microinfusion PCR (Tewhey R. et al., Nature Biotech. 2009, 27:1025-1031), custom-designed oligonucleotide microarrays (e.g., Roche / NimbleGen oligonucleotide microarrays), and solution-based hybridization methods (e.g., molecular inversion probe (MIP)) (Porreca GJ et al., Nature Methods, 2007, 4:931-936; Krishnakumar S. et al., Proc. Natl. Acad. Sci. USA, 2008, 105:9296-9310; Turner EHet al., Nature Methods, 2009, 6:315-316), and biotinylated RNA capture sequences (Gnirke A. et al.). This includes al., Nat. Biotechnol. 2009; 27(2): 182-9).
[0060] Single-molecule templates are another type of template available for NGS reactions. Spatially isolated single-molecule templates may be immobilized on a solid support by various methods. In one approach, individual primer molecules covalently attach to the solid support. Adapters are added to the template, and the template is then hybridized with the immobilized primers. In another approach, the single-molecule template covalently attaches to the solid support by priming and extending a single-chain single-molecule template from the immobilized primers. Subsequently, the usual primers are hybridized with the template. In yet another approach, a single polymerase molecule attaches to the solid support to which the primed template is bound.
[0061] Sequencing and imaging. Typical sequencing and imaging methods for NGS include, but are not limited to, cyclic reversible termination (CRT), sequencing by ligation (SBL), single-molecule addition (pyrosequencing), and real-time sequencing.
[0062] CRT uses a reversible terminator in a cyclic method that minimizes nucleotide incorporation, fluorescence imaging, and cleavage steps. Typically, DNA polymerase includes a primer containing a single fluorescently modified nucleotide complementary to the complementary nucleotide of the template base. DNA synthesis is terminated after the addition of the single nucleotide, and any missing nucleotides are washed away. Imaging is performed to determine the identity of the included labeled nucleotide. Subsequently, in the cleavage step, the terminator / inhibitor and fluorescent dye are removed. Representative NGS platforms using the CRT method include, but are not limited to, the Illumina / Solexa genome analyzer (GA), which uses a clone-amplified template method coupled with a four-color CRT method detected by total internal reflection fluorescence (TIRF); and the Helicos BioSciences / HeliScope, which uses a single-molecule template method coupled with a one-color CRT method detected by TIRF.
[0063] SBL uses DNA ligase and either a single-nucleotide-encrypted probe or a double-nucleotide-encrypted probe for sequencing.
[0064] Typically, a fluorescently labeled probe is hybridized to a complementary sequence adjacent to a primed template. A DNA ligase is used to ligate the dye-labeled probe to the primer. After the unligated probe is washed, fluorescence imaging is performed to determine the identity of the ligated probe. The fluorescent dye may be removed using a cleavable probe that regenerates the 5'-PO4 group for subsequent ligation cycles. Alternatively, a new primer may be hybridized to the template after the old primer has been removed. Typical SBL platforms include, but are not limited to, Life / APG / SOLiD (Support Oligonucleotide Ligation Detection), which uses a 2-base encoded probe.
[0065] Pyrosequencing methods are based on detecting the activity of DNA polymerase with other chemiluminescent enzymes. Typically, the method sequences a single strand of DNA by synthesizing a complementary strand along one base pair at a time and detecting the base actually added at each step. The template DNA is fixed, and solutions of A, C, G, and T nucleotides are added sequentially and removed from the reaction. Light is generated only when the nucleotide solution replenishes the bases that do not form a template pair. The sequence of the solution that produces the chemiluminescent signal determines the sequence of the template. Representative pyrosequencing platforms include, but are not limited to, Roche / 454, which uses a DNA template prepared by emPCR with 1 to 2 million beads deposited in PTP wells.
[0066] Real-time sequencing involves imaging the sequential uptake of dye-labeled nucleotides during DNA synthesis. Representative real-time sequencing platforms include, but are not limited to, the Pacific Biosciences platform, which uses DNA polymerase molecules attached to the surface of individual zero-mode waveguide (ZMW) detectors to obtain sequence information as phosphate-linked nucleotides are incorporated into the growing primer strand; the Life / VisiGen platform, which uses genetically engineered DNA polymerase with attached fluorescent dyes to produce an enhanced signal after nucleotide uptake by fluorescence resonance energy transfer (FRET); and the LI-COR Biosciences platform, which uses dye-quencher nucleotides in the sequencing reaction.
[0067] Other sequencing methods of NGS include, but are not limited to, nanopore sequencing, hybridization sequencing, nanotransistor array-based sequencing, Polony sequencing, scanning electron tunneling microscopy (STM)-based sequencing, and nanowire molecular sensor-based sequencing.
[0068] Nanopore sequencing involves the electrophoresis of nucleic acid molecules in solution through nanoscale pores that provide highly sealed spaces from which single nucleic acid polymers can be analyzed. A typical method of nanopore sequencing is described, for example, in the literature [Branton D. et al., Nat Biotechnol. 2008;26(10):1146-53].
[0069] Hybridization sequencing is a non-enzymatic method that uses DNA microarrays. Typically, a single pool of DNA is fluorescently labeled and hybridized into an array containing known sequences. The hybridization signal from a given spot on the array can be used to identify the DNA sequence. In DNA double helix, the binding of one DNA strand to its complementary strand is sensitive even to single base mismatches, provided the hybridization region is short or an embodiment of the mismatch detection protein is present. Representative methods of hybridization sequencing are described, for example, in the literature (Hanna GJet al., J.Clin.Microbiol.2000;38(7):2715-21; and Edwards JRet al., Mut.Res.2005;573(1-2):3-12).
[0070] Polony sequencing is based on sequencing through Polony amplification and multiple single-base extension (FISSEQ). Polony amplification is a method for amplifying DNA in situ on a polyacrylamide film. A typical Polony sequencing method is described, for example, in U.S. Patent Application Publication 2007 / 0087362.
[0071] Nanotransistor array-based devices, such as carbon nanotube field-effect transistors (CNTFETs), may also be used for NGS. For example, DNA molecules are stretched and driven across nanotubes by microfabricated electrodes. The DNA molecules sequentially come into contact with the carbon nanotube surface, and differences in current flow from each base are generated due to charge transfer between the DNA molecules and the nanotubes. The DNA is sequenced by recording these differences. A typical nanotransistor array-based sequencing method is described, for example, in U.S. Patent Publication 2006 / 0246497.
[0072] Scanning electron tunneling ring microscopes (STMs) may also be used for next-generation sequencing (NGS). The STM uses a piezo-electron-controlled probe to perform a raster scan of the specimen and form an image of its surface. The STM may be used, for example, to image the physical properties of a single DNA molecule, creating consistent electron tunneling ring imaging and spectroscopy by integrating a actuator-driven flexible gap with a scanning electron tunneling ring microscope. A typical sequencing method using an STM is described, for example, in U.S. Patent Application Publication No. 2007 / 0194225.
[0073] Molecular analyzers composed of nanowire molecular sensors may also be used for NGS. Such devices can detect interactions between nanowires, such as DNA, and nitrogenous substances positioned on nucleic acid molecules. Molecular guides are positioned to guide molecules near the molecular sensor to allow for interaction and subsequent detection. A typical sequencing method using nanowire molecular sensors is described, for example, in U.S. Patent Application Publication 2006 / 0275779.
[0074] Double-end sequencing methods may be used for NGS. Double-end sequencing uses blocking and unblocking primers to sequence both the sense and antisense strands of DNA. Typically, these methods include the steps of: annealing an unblocking primer to the first strand of nucleic acid; annealing a second blocking primer to the second strand of nucleic acid; extending the nucleic acid along the first strand with a polymerase; terminating the first sequencing primer; deblocking the second primer; and extending the nucleic acid along the second strand. A typical double-strand sequencing method is described, for example, in U.S. Patent No. 7,244,567.
[0075] Data analysis phase. After NGS reads are generated, they are aligned against a known reference sequence or de novo assembled.
[0076] For example, identifying genetic modifications such as single nucleotide polymorphisms and structural variants from a sample (e.g., a tumor sample) may be done by aligning NGS reads relative to a reference sequence (e.g., a wild-type sequence). Methods for sequence alignment relative to NGS are described, for example, in the literature (Trapnell C. and Salzberg SLNature Biotech., 2009, 27:455-457).
[0077] Examples of de novo assemblies are described, for example, in the literature (Warren R. et al., Bioinformatics, 2007, 23:500-501; Butler J. et al., Genome Res., 2008, 18:810-820; and Zerbino DRand Birney E., Genome Res., 2008, 18:821-829).
[0078] Sequence sorting or assembly may be performed using read data from one or more NGS platforms, for example, by mixing Roche / 454 and Illumina / Solexa read data. In the present invention, the sorting step may be performed using the BWA algorithm and the hg19 sequence, although this is not limited thereto.
[0079] In the present invention, the step of confirming the position of the nucleic acid fragment in step (B) is preferably characterized by being performed through sequence alignment, the sequence alignment being a computer algorithm that includes a computer method or approach used for identity from the case where read sequences (e.g., short read sequences from next-generation sequencing) in the genome may largely derive from evaluating the similarity between the read sequences and the reference sequence. Various algorithms may be applied to the sequence alignment problem. Some algorithms are relatively slow but allow for relatively high specificity. These include, for example, dynamic programming-based algorithms. Dynamic programming is a method of solving complex problems by breaking them down into simpler steps. Other approaches are relatively more efficient but typically less thorough. These include, for example, heuristic algorithms and probabilistic methods designed for large database searches.
[0080] Typically, the sorting process can have two stages: candidate screening and sequence sorting. Candidate screening reduces the search space for sequence sorting from the entire genome to a shorter enumeration of possible sorting locations. As the term suggests, sequence sorting involves the stage of sorting sequences having the sequences provided in the candidate screening stage. This may be done using global sorting (e.g., Needleman-Wunsch sorting) or local sorting (e.g., Smith-Waterman sorting).
[0081] Most attribute sorting algorithms can feature one of three types based on their indexing method: hash table algorithms (e.g., BLAST, ELAND, SOAP), suffix tree algorithms (e.g., Bowtie, BWA), and merge sorting algorithms (e.g., Slider). Short read sequences are typically used for sorting. Examples of sequence sorting algorithms / programs for short read sequences are, but are not limited to, BFAST (Homer N. et al., PLoS One. 2009; 4(11): e7767), BLASTN (from blast.ncbi.nlm.nih.gov on the World Wide Web), BLAT (Kent WJ Genome Res. 2002; 12(4): 656-64), Bowtie (Langmead B. et al., Genome Biol. 2009; 10(3): R25), BWA (Li H. and Durbin R. Bioinformatics, 2009, 25: 1754-60), BWA-SW (Li H. and Durbin R. Bioinformatics, 2010; 26(5): 589-95), CloudBurst (Schatz MCBioinformatics.2009;25(11):1363-9), Corona Lite (Applied Biosystems, Carlsbad, California, USA), CASHX (Fahlgren N. et al., RNA, 2009;15,992-1002), CUDA-EC (Shi H. et al., J Comput Biol.2010;17(4):603-15), ELAND (bioit.dbi.udel.edu / howto / eland on the World Wide Web), GNUMAP (Clement N. et al., Bioinformatics.2010;26(1):38-45), GMAP (Wu TD and Watanabe CKBioinformatics.2005;21(9):1859-75), GSNAP (Wu TD and Nacu S., Bioinformatics.2010;26(7):873-81), Geneious Assembler (Biomatters Ltd., Auckland, New Zealand), LAST, MAQ (Li H. et al., Genome Res. 2008;18(11):1851-8), Mega-BLAST (on the World Wide Web at ncbi.nlm.nih.gov / blast / megablast.shtml), MOM (Eaves HLand Gao Y. Bioinformatics. 2009;25(7):969-70), MOSAIK (on bioinformatics.bc.edu / marthlab / Mosaik on the World Wide Web), Novoalign (on novocraft.com / main / index.php on the World Wide Web), PALMapper (on fml.tuebingen.mpg.de / raetsch / suppl / palmapper on the World Wide Web), PASS (Campagna D. et al., Bioinformatics. 2009;25(7):967-8), PatMaN (Prufer K. et al., Bioinformatics. 2008;24(13):1530-1), PerM (Chen Y. et al., Bioinformatics, 2009, 25(19):2514-2521), ProbeMatch (Kim YJet al.,Bioinformatics.2009;25(11):1424-5), QPalma(de Bona F.et al.,Bioinformatics,2008,24(16):i174), RazerS(Weese D.et al.,Genome Research,2009,19:1646-1654), RMAP(Smith ADet al.,Bioinformatics.2009;25(21):2841-2), SeqMap(Jiang H.et al.Bioinformatics.2008;24:2395-2396.), Shrec(Salmela L.,Bioinformatics.2010;26(10):1284-90), SHRiMP(Rumble SMet al.,PLoS Comput.Biol.,2009,5(5):e1000386), SLIDER(Malhis N.et al.,Bioinformatics,2009,25(1):6-13), SLIM Search(Muller T.et al.,Bioinformatics.2001;17 Suppl 1:S182-9), SOAP(Li R.et al.,Bioinformatics.2008;24(5):713-4), SOAP2(Li R.et al.,Bioinformatics.2009;25(15):1966-7), SOCS(Ondov BDet al.,Bioinformatics,2008;24(23):2776-7), SSAHA(Ning Z.et al.,Genome This includes Res.2001;11(10):1725-9, SSAHA2 (Ning Z. et al., Genome Res.2001;11(10):1725-9), Stampy (Lunter G. and Goodson M. Genome Res.2010, epub ahead of print), Taipan (on the World Wide Web at taipan.sourceforge.net), UGENE (on the World Wide Web at ugene.unipro.ru), XpressAlign (on the World Wide Web at bcgsc.ca / platform / bioinfo / software / XpressAlign), and ZOOM (Bioinformatics Solutions Inc., located in Waterloo, Ontario, Canada).
[0082] Sequence alignment algorithms may be selected based on a number of factors, including, for example, sequencing method, read length, number of reads, available computing resources, and sensitivity / scoring requirements. Different sequence alignment algorithms can achieve different speed levels, alignment sensitivity, and alignment specificity. Alignment specificity refers to the percentage of target sequence residues aligned as accurately as typically found in a submission compared to a predicted alignment. Alignment sensitivity also refers to the percentage of target sequence residues aligned as accurately as typically found in a generally predicted alignment in a submission.
[0083] Alignment algorithms, such as ELAND or SOAP, may be used to align short reads (e.g., from Illumina / Solexa sequencers) to a reference genome when speed is the primary factor. Alignment algorithms like BLAST or Mega-BLAST may be used for similarity investigations with short readouts (e.g., from Roche FLX) when specificity is the most important factor, although these methods are relatively slow. Alignment algorithms like MAQ or Novoalign may be used for single or paired-end data when quality scores are considered and therefore accuracy is essential (e.g., in high-speed, high-volume SNP searches). Alignment algorithms like Bowtie or BWA utilize the Burrows-Wheeler Transform (BWT) and therefore require a relatively small memory footprint. Alignment algorithms such as BFAST, PerM, SHRiMP, SOCS, or ZOOM map color space reads and may therefore be used in conjunction with the ABI SOLiD platform. In some applications, results from two or more alignment algorithms may be combined.
[0084] In the present invention, the length of the sequence information (reads) in step B) is 5 to 5000 bp, and the number of sequence information entries used may be 5,000 to 5 million, but is not limited to these values.
[0085] In the present invention, the step of grouping sequence information in step (C) above can be performed based on the adapter sequence of the sequence information (reads). The FD value can be calculated for the sequence information that has been separated into nucleic acid fragments aligned in the forward direction and nucleic acid fragments aligned in the reverse direction, or the FD value can be calculated for the entire group.
[0086] The present invention may further include a step of separately classifying nucleic acid fragments that satisfy the mapping quality score of the aligned nucleic acid fragments prior to performing step (C) above.
[0087] In the present invention, the mapping quality score may vary depending on the desired criteria, but is preferably 15 to 70 points, more preferably 50 to 70 points, and most preferably 60 points.
[0088] In the present invention, the FD value in step (D) may be defined as the distance between the reference value of the i-th nucleic acid fragment and the reference value of any one or more nucleic acid fragments selected from the i+1 to n-th nucleic acid fragments, for the n nucleic acid fragments obtained.
[0089] In the present invention, the FD value can be calculated by determining the distance between the reference value of the first nucleic acid fragment and the reference value of any one or more nucleic acid fragments selected from the group consisting of the second to n nucleic acid fragments, and the FD value can be a statistical value that includes, but is not limited to, one or more inverse values and weighted values of the sum, difference, product, mean, log of product, log of sum, median, quantile, minimum, maximum, variance, standard deviation, median absolute deviation, and coefficient of variation, as well as one or more of these values.
[0090] In this invention, the phrase "one or more values and / or one or more inverse values thereof" is interpreted to mean that one or more of the above-mentioned numerical values can be used in combination.
[0091] In the present invention, the "reference value of nucleic acid fragments" may be characterized by being a value obtained by adding or subtracting an arbitrary value from the median value of nucleic acid fragments.
[0092] The aforementioned FD value can be defined as follows for the n nucleic acid fragments obtained. FD = Dist(Ri~Rj) (1 <i<j<n)、 Here, the Dist function calculates a statistical result that includes, but is not limited to, one or more values selected from the group consisting of the sum, difference, product, mean, logarithm of the product, logarithm of the sum, median, quantile, minimum, maximum, variance, standard deviation, median absolute deviation, and coefficient of variation, and / or one or more inverses thereof, weighted values.
[0093] In other words, in this invention, the FD value (Fragment Distance Value) means the distance between aligned nucleic acid fragments. Here, the number of nucleic acid fragments to be selected for distance calculation can be defined as follows: When there are a total of N nucleic acid fragments, JPEG0007859967000001.jpg 10128 combinations of distances between nucleic acid fragments are possible. That is, when i is 1, i+1 is 2, and the distance to any one or more nucleic acid fragments selected from the 2nd to nth nucleic acid fragments can be defined.
[0094] In the present invention, the FD value may be characterized by calculating the distance between a specific position within the i-th nucleic acid fragment and a specific position within any one or more of the i+1 to n-th nucleic acid fragments.
[0095] For example, if a nucleic acid fragment is 50 bp long and aligned at position 4,183 on chromosome 1, then the genetic position values that can be used to calculate the distance of this nucleic acid fragment are from 4,183 to 4,232 on chromosome 1.
[0096] When a 50 bp long nucleic acid fragment adjacent to the aforementioned nucleic acid fragment is aligned at position 4,232 of chromosome 1, the genetic position value usable for calculating the distance between these nucleic acid fragments is between 4,232 and 4,281 on chromosome 1, and the FD value between the two nucleic acid fragments can be between 1 and 99.
[0097] Furthermore, when other adjacent 50bp long nucleic acid fragments are aligned at position 4123 of chromosome 1, the genetic position values available for calculating the distance of these nucleic acid fragments are 4,123 to 4,172 on chromosome 1, the FD value between the two nucleic acid fragments is 61 to 159, and the FD value with respect to the first example nucleic acid fragment is 12 to 110. The FD value can be used as a calculation result and / or a statistical value including one or more values selected from the group consisting of the sum, difference, product, mean, logarithm of product, logarithm of sum, median, quantile, minimum, maximum, variance, standard deviation, median absolute deviation, and coefficient of variation of one of the values in both FD value ranges, and / or one or more inverse values and weighted values thereof. Preferably, it may be characterized by being the inverse value of one of the values in both FD value ranges, but is not limited thereto.
[0098] Preferably, in the present invention, the FD value may be characterized by being a value obtained by adding or subtracting an arbitrary value from the median value of the nucleic acid fragment.
[0099] In this invention, the median of FD is the value that is furthest in the middle when the calculated FD values are arranged in order of magnitude. For example, if there are three values such as 1, 2, and 100, 2 is the value that is furthest in the middle, so 2 is the median. If there is an even number of FD values, the median is determined to be the average of the two values that are furthest in the middle. For example, if there are FD values of 1, 10, 90, and 200, the median is the average of 10 and 90, which is 50.
[0100] In the present invention, any of the above values can be used as long as they can indicate the position of the nucleic acid fragment, but are preferably 0 to 5 kbp, or 0 to 300% of the nucleic acid fragment length, 0 to 3 kbp, or 0 to 200% of the nucleic acid fragment length, 0 to 1 kbp, or 0 to 100% of the nucleic acid fragment length, more preferably 0 to 500 bp or 0 to 50% of the nucleic acid fragment length, but are not limited thereto.
[0101] In the present invention, the FD value may be characterized in that, in paired-end sequencing, it is derived based on the position values of the forward and reverse sequence information (reads).
[0102] For example, in a paired-end read pair of 50 bp length, if the forward read is aligned to position 4183 on chromosome 1 and the reverse read is aligned to position 4349, then the ends of this nucleic acid fragment will be 4183 and 4349, and the usable reference range for nucleic acid fragment distance will be 4183 to 4349. In this case, in another paired-end read pair adjacent to the aforementioned nucleic acid fragment, if the forward read is aligned to position 4349 on chromosome 1 and the reverse read is aligned to position 4515, then the positional value of this nucleic acid fragment will be 4349 to 4515. The distance between these two nucleic acid fragments can be 0 to 333, and most preferably it can be 166, which is the median distance between each nucleic acid fragment.
[0103] In the present invention, when acquiring sequence information by paired-end sequencing, the invention may further include a step of excluding nucleic acid fragments whose number of alignment points in the sequence information (reads) is less than a reference value from the calculation process.
[0104] In the present invention, the FD value may be characterized in that, in single-end sequencing, it is derived based on one type of position value of the forward or reverse sequence information (read).
[0105] In the present invention, the single-ended sequencing is characterized by adding an arbitrary value when deriving a position value based on sequence information aligned in the forward direction, and subtracting an arbitrary value when deriving a position value based on sequence information aligned in the reverse direction. The arbitrary value can be any value that allows the FD value to clearly indicate the position of the nucleic acid fragment, preferably 0 to 5 kbp or 0 to 300% of the nucleic acid fragment length, 0 to 3 kbp or 0 to 200% of the nucleic acid fragment length, 0 to 1 kbp or 0 to 100% of the nucleic acid fragment length, and more preferably 0 to 500 bp or 0 to 50% of the nucleic acid fragment length, but is not limited thereto.
[0106] In this invention, the nucleic acid to be analyzed may be sequenced and expressed in units called reads. These reads can be classified into single-end sequencing reads (SE) and paired-end sequencing reads (PE) depending on the sequencing method. An SE read means that one of the 5' and 3' ends of the nucleic acid molecule is sequenced in a random direction for a fixed length, while a PE read is sequenced in both the 5' and 3' ends for a fixed length. Due to these differences, it is a well-known fact to the average technician that when sequencing in SE mode, one read is generated from one nucleic acid fragment, while in PE mode, two reads are generated as a pair from one nucleic acid fragment.
[0107] The most ideal method for calculating the precise distance between nucleic acid fragments is to sequence the nucleic acid molecule from beginning to end, align the reads, and use the median (center) of the aligned values. However, technically, this method is currently limited in terms of the limitations of sequencing technology and cost. Therefore, sequencing is performed using methods such as SE and PE. In the PE method, the start and end positions of the nucleic acid molecule are known, and the precise position (median) of the nucleic acid fragment can be determined by combining these values. However, in the SE method, only the end information of one side of the nucleic acid fragment is available, which limits the accuracy of the position (median) calculation.
[0108] Furthermore, when calculating the distance of nucleic acid molecules using the end information of all reads sequenced (aligned) in both the forward and reverse directions, the sequencing direction can sometimes result in inaccurate values.
[0109] Therefore, due to technical reasons in the sequencing method, the 5' end of the forward read has a positional value smaller than the center position of the nucleic acid molecule, while the 3' end of the reverse read has a larger value. Using this characteristic, by adding an arbitrary value (Extended bp) to the forward read and subtracting it from the reverse read, a value close to the center position of the nucleic acid molecule can be estimated.
[0110] In other words, the arbitrary value (Extended bp) can vary depending on the sample used. Since the average length of cell-free nucleic acids is known to be around 166 bp, it can be set to approximately 80 bp. If the experiment is performed using fragmentation equipment (e.g., sonication), the extended bp can be set to about half of the target length set during the fragmentation process.
[0111] In the present invention, the step of determining the chromosomal abnormality in step (E) is (Ei) The stage of determining representative FD values (RepFD) for the entire chromosome or specific regions; (E-ii) The step of calculating one or more values and / or one or more inverse values thereof selected from the group consisting of the sum, difference, product, mean, log of product, log of sum, median, quantile, minimum, maximum, variance, standard deviation, absolute median deviation, and coefficient of variation of RepFD values for the entire chromosome region to be analyzed or a specific region within the sample other than the specific region, and deriving the Normalized Factor; (E-iii) The step of calculating the representative ratio (RepFD ratio) based on the following formula 1; Formula 1: RepFD ratio = RepFD Target genomic region / Normalized Factor (E-iv) Steps to compare the RepFD ratio values of the sample with those of a normal reference population and calculate the FDI (Fragments Distance Index): It may be characterized by being carried out including [this].
[0112] In the present invention, the representative value (RepFD) in step (Ei) is characterized by being one or more values selected from the group consisting of the sum, difference, product, mean, median, quantile, minimum, maximum, variance, standard deviation, median absolute deviation, and coefficient of variation of the FD values, and / or one or more inverse values thereof, and preferably being the median, mean, or inverse value thereof of the FD values, but is not limited thereto.
[0113] In the present invention, the entire chromosome region or the specific genetic region can be any set of human nucleic acid sequences, but preferably it may be a specific region of a chromosome unit or a part of a chromosome. For example, the specific region for detecting the presence or absence of numerical abnormalities may be an autosome that is considered to be euploid, and the specific region for detecting the presence or absence of structural abnormalities may be any genetic region other than regions with low uniqueness (centromere, telomere), but is not limited thereto.
[0114] The specific region within the sample other than the entire chromosome region or specific genetic region to be analyzed in step (E-ii) above is: a) The step of randomly selecting regions other than the entire chromosome or specific genetic region to be randomly analyzed; b) A step in which representative values of the RepFD values of the genetic regions selected in step a) are determined as pre-normalized factors (PNF); c) The step of calculating the representative ratio (RepFD ratio) based on the following formula 2: Formula 2: RepFD ratio = RepFD Target genomic region / PNF d) The step of calculating the coefficient of variation (SD / Mean) of the RepFD ratio values for the normal reference population; and e) The method may be characterized by selecting a genetic region that has the smallest coefficient of variation among the coefficients of variation obtained by repeatedly performing steps a) to d) above, and determining it as a specific region within the sample other than the entire chromosome or a specific genetic region.
[0115] In the present invention, the repeated execution of step e) may be characterized by being 100 times or more, preferably in the range of 10,000 to 1,000,000 times, and most preferably 100,000 times, but is not limited thereto.
[0116] In the present invention, step (E-iv) may be characterized by comparing the RepFD ratio of a normal reference population with the RepFD ratio of a sample.
[0117] In the present invention, any method that can confirm that there is a statistically significant difference between the RepFD ratio values of the normal reference population and the RepFD ratio of the sample is acceptable for comparing the RepFD ratio of the normal reference population with the RepFD ratio of the sample. Preferably, the method selected may be a mean and standard deviation-based Z-score or a median-based Log ratio, or a likelihood ratio calculated through a classification algorithm. Most preferably, a mean and standard deviation-based Z-score calculation method is used, but the invention is not limited thereto.
[0118] In the present invention, the FDI (Fragments Distance Index) is calculated by comparing the Rep FD ratio values of the sample to be analyzed with those of the normal reference population. A standard score system such as the Z score can be used for the comparison method, and the critical value can be an integer or range such as an infinite positive or negative number, preferably 3, but is not limited to this.
[0119] In another aspect, the present invention is a decoding unit for extracting nucleic acids from a biological sample and decoding the sequence information; An alignment unit that aligns the decoded sequences in a standard chromosome sequence database; and This invention relates to a chromosomal abnormality detection device that includes a chromosomal abnormality determination unit that measures the distance between aligned nucleic acid fragments for selected nucleic acid fragments to calculate the FD value (Fragments Distance), calculates the FDI value (Fragments Distance Index) for the entire chromosome region or for specific genetic regions based on the calculated FD value, and determines that a chromosomal abnormality exists if the FDI value is below or above a reference value or interval.
[0120] In yet another aspect, the present invention includes a computer-readable storage medium comprising instructions configured to be executed by a processor that detects chromosomal abnormalities, (A) The step of extracting nucleic acids from a biological sample to obtain nucleic acid fragments and acquire sequence information; (B) The step of aligning nucleic acid fragments in a standard chromosome sequence database (reference genome database) based on the acquired sequence information (reads); (C) A step of measuring the distance between selected nucleic acid fragments and calculating the FD value (Fragments Distance); and (D) A step in which the Fragments Distance Index (FDI) is calculated for the entire chromosome region or for specific genetic regions based on the FD value calculated in step (C) above, and if the FDI value does not fall within the reference range, it is determined that there is a chromosomal abnormality; This relates to a computer-readable storage medium containing instructions configured to be executed by a processor that detects chromosomal abnormalities.
[0121] Specifically, the computer-readable storage medium according to the present invention includes instructions configured to be executed by a processor that detects chromosomal abnormalities, (A) The step of extracting nucleic acids from a biological sample to obtain nucleic acid fragments and acquire sequence information; (B) The step of aligning nucleic acid fragments in a standard chromosome sequence database (reference genome database) based on the acquired sequence information (reads); (C) A step of grouping the nucleic acid fragments aligned based on the sequence information (reads) into a whole sequence, a forward sequence, and a reverse sequence; (D) A step of measuring the distance between aligned nucleic acid fragment reference values for each of the grouped nucleic acid fragments and calculating the FD value (Fragments Distance) for each group; and (E) Based on the FD values for each group calculated in step (D) above, the FDI value (Fragments Distance Index) is calculated for the entire chromosome region or for specific regions, and if none of the FDI values fall within the reference range, it is determined that there is a chromosomal abnormality; It may be characterized by including, but is not limited to, instructions configured to be executed by a processor that detects chromosomal abnormalities.
[0122] In another embodiment of the present invention, it was confirmed that the positional values of both ends of the reads can be calculated by adding or subtracting 50% of the average length of the nucleic acid fragment to be analyzed from the median of the aligned nucleic acid fragments, and the distance between the reads can be calculated (Figure 6).
[0123] Therefore, in yet another view, the present invention (A) The step of extracting nucleic acids from a biological sample and obtaining sequence information; (B) The step of aligning the acquired sequence information (reads) with the standard chromosome sequence database (reference genome database); (C) A step of measuring the distance between aligned reads and calculating the RD value (Read Distance) for the aligned sequence information (reads); and (D) A chromosomal abnormality detection method that includes a step of calculating the RDI (Read Distance Index) for the entire chromosome region or for specific regions based on the RD value calculated in step (C) above, and determining that there is a chromosomal abnormality if the RDI value does not fall within the reference range.
[0124] In the present invention, The aforementioned stage A) is, (Ai) The step of obtaining nucleic acids from blood, semen, vaginal cells, hair, saliva, urine, oral cells, placental cells or fetal cells, amniotic fluid, tissue cells, and mixtures thereof; (A-ii) The step of removing proteins, fats, and other residues from the collected nucleic acids using a salting-out method, a column chromatography method, or a beads method to obtain purified nucleic acids; (A-iii) The step of preparing a single-end sequencing or pair-end sequencing library from purified nucleic acids or nucleic acids randomly fragmented by enzymatic cleavage, pulverization, or hydroshear method; (A-iv) The step of reacting the prepared library with a next-generation sequencer; and (Av) This method may be characterized by including a step of obtaining nucleic acid sequence information (reads) using a next-generation sequencer.
[0125] In the present invention, the length of the sequence information (reads) in step (B) is 5 to 5000 bp, and the number of sequence information items used may be 5,000 to 5 million, but is not limited to these.
[0126] In the present invention, step (C) may be further characterized by the availability of a step of grouping the aligned leads by their aligned direction.
[0127] In this invention, a step of grouping the reads may be further utilized, in which case the grouping criterion may be based on the adapter arrangement of the aligned reads. The RD value can be calculated for the arrangement information that has been sorted separately into reads aligned in the forward direction and reads aligned in the reverse direction.
[0128] The present invention may further include a step of separately classifying reads that satisfy the mapping quality score of the aligned reads prior to performing step (C) above.
[0129] In the present invention, the mapping quality score may vary depending on the desired criteria, but is preferably 15 to 70 points, more preferably 50 to 70 points, and most preferably 60 points.
[0130] In the present invention, the RD value in step (C) above may be characterized by calculating the distance between the values obtained by adding or subtracting 50% of the average nucleic acid length to one of the end values of the i-th read and one or more reads selected from i+1 to n-th reads for the n reads obtained.
[0131] In the present invention, the RD value can be calculated by determining the distance between the first read and one or more reads selected from the group consisting of the second to the nth reads, and the RD value can be a statistical value that includes, but is not limited to, the sum, difference, product, mean, log of product, log of sum, median, quantile, minimum, maximum, variance, standard deviation, median absolute deviation, coefficient of variation, their inverse values and combinations, and weighted values, calculated from the distance between the n acquired reads and the first read.
[0132] In this invention, the median of RD refers to the value that is furthest in the middle when the calculated RD values are arranged in order of magnitude. For example, when there is an odd number of values, such as 1, 2, and 100, 2 is the value that is furthest in the middle, so 2 is the median. If there is an even number of RD values, the median is determined by the average of the two middle values. For example, if there are RD values of 1, 10, 90, and 200, the median is the average of 10 and 90, which is 50.
[0133] In the present invention, the RD value may be characterized by calculating the distance between the 5' or 3' end inside the i-th read and the 5' or 3' end of one or more reads from i+1 to n.
[0134] For example, in a paired-end read pair of 50 bp in length, if the forward read is aligned to position 4183 on chromosome 1 and the reverse read is aligned to position 4349, then the ends of this nucleic acid fragment will be 4183 and 4349, and the reference range usable for nucleic acid fragment distance is 4183 to 4349. In this case, in another paired-end read pair adjacent to the aforementioned nucleic acid fragment, if the forward read is aligned to position 4349 on chromosome 1 and the reverse read is aligned to position 4515, then the positional value of this nucleic acid fragment is 4349 to 4515. The distance between these two nucleic acid fragments can be 0 to 333, and most preferably it can be 166, which is the median distance between each nucleic acid fragment. In the above example, if the average length of the nucleic acid fragments is 166, and the 50th percentile of the average length of the nucleic acid fragments is subtracted from the median (4266), the position value of the first nucleic acid fragment becomes 4183 and the position value of the second nucleic acid fragment becomes 4349. In this case, the distance between the reads is 166 (4349~4183). On the other hand, if the 50th percentile is added to the median, the position value of the first nucleic acid fragment becomes 4349 and the position value of the second nucleic acid fragment becomes 4515. In this case, the distance between the reads is 166 (4515~4349).
[0135] In the present invention, the step of determining the chromosomal abnormality in step (D) above is (Di) The stage of determining a representative value (RepRD) for each chromosome region or specific genetic region; (D-ii) The step of calculating one or more values and / or one or more inverse values thereof selected from the group consisting of the sum, difference, product, mean, log of product, log of sum, median, quantile, minimum, maximum, variance, standard deviation, median absolute deviation, and coefficient of variation of RepRD values for the entire chromosome region to be analyzed or for a sample region other than a specific genetic region, and deriving the Normalized Factor; (D-iii) The step of calculating the representative ratio (RepRD ratio) based on the following formula 10; Equation 10: RepRD ratio = RepRD Target genomic region / Normalized Factor (D-iv) Step 1: Compare the RepRD ratio values of the normal reference population and the sample, and calculate the RDI (Read Distance Index): It may be characterized by being carried out including [this].
[0136] In the present invention, the representative value (RepRD) in the (Di) stage is characterized by being one or more values selected from the group consisting of the sum, difference, product, mean, log of product, log of sum, median, quantile, minimum, maximum, variance, standard deviation, median absolute deviation, and coefficient of variation of RD values, and / or one or more reciprocals thereof, and statistical values not limited thereto. Preferably, it may be characterized by being the median, mean, or reciprocal value of the RD values, but is not limited thereto.
[0137] In this invention, the median of the RepRD values is the value that is closest to the center when the calculated RepRD values are arranged in order of magnitude. For example, when there are an odd number of values, such as 1, 2, and 100, 2 is the closest to the center, so 2 is the median. If there are an even number of RepRD values, the median is determined to be the average of the two central values. For example, if there are RepRD values of 1, 10, 90, and 200, the median is the average of 10 and 90, which is 50.
[0138] In the present invention, the entire chromosome region or the specific genomic region can be any set of human nucleic acid sequences, but preferably it may be a specific region of a chromosome unit or a part of a chromosome. For example, the specific region for detecting the presence or absence of numerical abnormalities may be an autosome that is considered to be euploid, and the specific region for detecting the presence or absence of structural abnormalities may be any genetic region other than regions with low uniqueness (centromere, telomere), but is not limited thereto.
[0139] The specific region within the sample other than the entire chromosome region or specific genetic region to be analyzed in step (D-ii) above is: a) The step of selecting regions other than the entire chromosome or specific genetic region to be randomly analyzed; b) A step in which representative values of the RepRD values of the genetic regions selected in step a) are determined as pre-normalized factors (PNF); c) The step of calculating the representative ratio (RepRD ratio) based on the following formula 11: Equation 11: RepRD ratio = RepRD Target genomic region / PNF d) The step of calculating the coefficient of variation (SD / Mean) of the RepRD ratio values for the normal reference population; and e) The method may be characterized by selecting the genetic region having the smallest coefficient of variation among the coefficients of variation obtained by repeatedly performing steps a) to d) above, and determining it as the entire chromosome region or a specific genetic region or other region.
[0140] In the present invention, the repeated execution of step e) may be 100 times or more, preferably in the range of 10,000 to 1,000,000 times, and most preferably in the range of 100,000 times, but is not limited thereto.
[0141] In the present invention, step (iv) may be characterized by comparing the RepRD ratio of a normal reference population with the RepRD ratio of a sample.
[0142] In the present invention, any method that can confirm that there is a statistically significant difference between the RepRD ratio values of the normal reference population and the RepRD ratio of the sample can be used. Preferably, a method such as a mean and standard deviation-based Z-score or a median-based Log ratio, or a likelihood ratio calculated by a classification algorithm, is selected, and most preferably a mean and standard deviation-based Z-score calculation method is used, but the invention is not limited thereto.
[0143] In the present invention, the RDI value (Reads Distance Index) is calculated by comparing the Rep RD ratio value of the sample to be analyzed with that of the normal reference population. A standard score system such as the Z score can be used as the comparison method, and the critical value can be an integer or range such as an infinite positive or negative number, preferably -3 or 3, but is not limited thereto.
[0144] In yet another aspect, the present invention is a decoding unit for extracting nucleic acids from a biological sample and decoding their sequence information; An alignment unit that aligns the decoded sequences in a standard chromosome sequence database; and This invention relates to a chromosomal abnormality detection device that includes a chromosomal abnormality determination unit that measures the distance between aligned reads in selected sequence information (reads) to calculate the RD value (Read Distance), calculates the RDI value (Read Distance Index) for each genetic region based on the calculated RD value, and determines that a chromosomal abnormality exists if the RDI value is below or above a reference value or interval.
[0145] In yet another aspect, the present invention includes a computer-readable storage medium comprising instructions configured to be executed by a processor that detects chromosomal abnormalities, (A) The step of extracting nucleic acids from a biological sample and obtaining sequence information; (B) The step of aligning the acquired sequence information (reads) with the standard chromosome sequence database (reference genome database); (C) A step of measuring the distance between aligned reads and calculating the RD value (Read Distance) for the selected sequence information (reads); and (D) The present invention relates to a computer-readable storage medium that includes instructions configured to be executed by a processor that detects chromosomal abnormalities, wherein the processor calculates an RDI value (Read Distance Index) for each genetic region based on the RD value calculated in step (C) above, and determines that there is a chromosomal abnormality if the RDI value does not fall within the reference range. [Examples]
[0146] The present invention will be described in more detail below with reference to examples. It will be obvious to those of ordinary skill in the art that these examples are merely illustrative and should not be construed as limiting the scope of the present invention.
[0147] Example 1. DNA is extracted from blood and next-generation sequencing analysis is performed. Blood samples were collected in 10 mL increments and stored in EDTA tubes. Within 2 hours of collection, the plasma portion was primarily centrifuged at 1200 g, 4°C, and 15 minutes. The purified plasma was then subjected to secondary centrifugation at 16000 g, 4°C, and 10 minutes to separate the plasma supernatant from the precipitate. Cell-free DNA was extracted from the separated plasma using the Tiangen micro DNA kit (Tiangen). Paired-end (PE) data were produced using the MGI E-Asy Cell-free DNA Library Prep set kit (MGI) for library preparation, followed by 50 cycles x 2 using a DNBseq G400 (MGI). Single-end (SE) data were produced using the Illumina Nextseq 500 after library preparation using the Illumina Truseq Nano DNA HT library prep kit (Illumina).
[0148] The PE data provided sequence information for approximately 10 million nucleic acid fragments, while the SE data provided sequence information for approximately 1.3 million nucleic acid fragments.
[0149] Example 2. Quality control of sequence information data and FD value calculation Before preprocessing the nucleotide sequence information and calculating the FD value, the following steps were performed: The fastq files generated by the next-generation sequencing (NGS) system were aligned using the BWA-mem algorithm based on the reference chromosome Hg19 sequence. There is a probability of errors occurring during the alignment of the library sequences, so two steps were taken to correct these errors. First, duplicate library sequences were removed, and then, among the library sequences aligned by the BWA-mem algorithm, sequences whose Mapping Quality Score did not reach 60 were removed.
[0150] After the selected reads were grouped into forward and reverse reads according to their alignment direction, the distance to the nearest adjacent read was calculated as the FD value using Equation 3 below, the concept of which is shown in Figures 2 and 3. The D function in Equation 3 below is a function that calculates the difference in genotype positions. In Equation 3 below, a and b are position values of nucleic acid fragments, which in PE sequencing can be any one of the minimum to maximum values of the aligned position values of two sequence information, and in SE sequencing they can be the aligned position values of the sequence information or a value obtained by extending a specific value to the position value.
[0151] Equation 3: Fragment Distance (FD) = D(a,b) | a ∈ Fi , b ∈ Fi)
[0152] Example 3. Confirmation of the difference in FD value due to extension. Data produced by PE (Peripheral Emission) provides information on the start and end positions of nucleic acid fragments, allowing for the calculation of inter-fragment distances based on intermediate positions. The PE-produced data was randomly grouped into For and Rev reads. For reads classified as For, the FD (Functional Distance) was calculated based on the 5' position of the For read, and for reads classified as Rev, it was calculated based on the 3' position of the Rev read. Then, extension was performed by adding 80 bp to the For reads and subtracting 80 bp from the Rev reads.
[0153] As a result of comparing the difference between the FD value of the process described above and the FD value of the extended process, as shown in Figure 5, it was confirmed that the FD value calculated after extension was similar to the centered FD value of PE, and that the FD value without extension had a difference of +166 and -166.
[0154] Example 4. FDI Value Calculation 4-1. Calculation of FDI value for detecting chromosomal numerical abnormalities The FDI value was calculated using SE sequencing data, and the extension value was set to 80 bp. For each chromosome whose presence or absence of aneuploidy was to be checked, the median ratio of the RepFD values of the selected set of chromosomes was defined as the Normalized Factor and calculated using Equation 4 below.
[0155] [Table 1]
[0156] Equation 4: Normalized Factor = Median of RepFD selected chromosome set
[0157] (Selected chromosome set: This is the portion corresponding to a chromosome set in the table above.)
[0158] The RepFD ratio was calculated using Equation 5, with the FD and Normalized Factor calculated using Equations 3 and 4.
[0159] Equation 5: RepFD ratio = RepFD Target chromosome / Normalized Factor
[0160] The mean and standard deviation of the RepFD ratio were calculated in a reference group of 2000 normal individuals, and the FDI (Fragments Distance Index) value of the sample to be analyzed was calculated using Equation 6.
[0161] Formula 6: FDI = (MEAN(RepFD Ratio reference - RepFD Ratio sample -) / SD(RepFD Ratio reference )
[0162] The process of the above Formula 6 was performed respectively when using all nucleic acid fragments (Formula 7), when using nucleic acid fragments aligned in the forward direction (Formula 8), and when using nucleic acid fragments aligned in the reverse direction (Formula 9).
[0163] Formula 7: FDI all = mean(RepFD all Ratio reference ) - RepFD all Ratio sample / SD(RepFD all Ratio reference ) Formula 8: FDI For = mean(RepFD For Ratio reference ) - RepFD For Ratio sample / SD(RepFD For Ratio reference ) Formula 9: FDI Rev = mean(RepFD Rev Ratio reference ) - RepFD Rev Ratio sample / SD(RepFD Rev Ratio reference )
[0164] 4-2. Performance confirmation of FDI value for detecting chromosomal abnormalities As a result of analyzing 2000 samples of the normal standard population and 88 clinical samples including trisomy 6 samples, 100% sensitivity and 100% specificity were confirmed.
[0165] The cut-off value for positive discrimination is FDI21 all , FDI21 ForFDI21 Rev In both cases, the value 3 was used.
[0166] For each sample, three FDI values were calculated to determine a positive result, and a final positive result was determined if all three values were 3 or higher. Of the 88 samples analyzed, three samples (G19NIPT261-3, G19NIPT261-10, G19NIPT261-13) were determined to be positive based on one FDI value, but the remaining two FDI values were determined to be negative, and the final result was determined to be negative.
[0167] [Table 2]
[0168] Example 5. DNA is extracted from blood for RD value-based analysis, and the next-generation base sequence is analyzed. Blood samples were collected in 10 mL increments from 400 healthy individuals, 175 individuals with trisomy 21, 67 individuals with trisomy 18, and 26 individuals with trisomy 13. These samples were stored in EDTA tubes. Within 2 hours of collection, the plasma portion was primarily centrifuged at 1200 g, 4°C, and 15 minutes. The purified plasma was then subjected to secondary centrifugation at 16000 g, 4°C, and 10 minutes to separate the plasma supernatant from the precipitate. Cell-free DNA was extracted from the separated plasma using the Tiangen micro DNA kit (Tiangen), and the library was prepared using the Trueq nanoDNA HT library preparation kit (Illumina). Sequencing was then performed using a Nextseq500 equipped with Illumina in 75SE (Single-end) mode.
[0169] As a result, we confirmed that approximately 13 million leads are produced per sample.
[0170] Example 6. Quality control of sequence information data and calculation of RD values Before preprocessing the nucleotide sequence information and calculating the RD value, the following steps were performed: After converting the Bcl file (containing nucleotide sequence information) generated by the next-generation sequencing (NGS) system to fastq format, the library sequences in the fastq file were aligned based on the reference chromosome Hg19 sequence using the BWA-mem algorithm. Since there is a probability of errors occurring during library sequence alignment, two error correction processes were performed. First, duplicate library sequences were removed, and then, among the library sequences aligned by the BWA-mem algorithm, sequences whose mapping quality score did not reach 60 were removed.
[0171] The selected reads were grouped into forward-direction reads and reverse-direction reads according to their alignment direction. The distance to the nearest adjacent read was then calculated as the RD value using Equation 12, and this concept is shown in Figure 7. The D function in Equation 12 below is a function that calculates the difference in gene location.
[0172] Equation 12: Read Distance (RD) = D(a,b) | a ∈ Ri , b ∈ Ri)
[0173] Example 7. RDI Value Calculation 7-1. Calculation of RDI value for detecting chromosomal numerical abnormalities For each chromosome, the median RD value was defined as RepRD. For each chromosome whose presence or absence of aneuploidy was to be checked, the ratio of the median RepRD values of the selected set of chromosomes was defined as the Normalized Factor and calculated using Equation 13 below.
[0174] [Table 3]
[0175] Equation 13: Normalized Factor = Median of RepRD selected chromosome set
[0176] (Selected chromosome set: This is the portion corresponding to a chromosome set in the table above.)
[0177] The RepRD ratio was calculated using Equation 14, employing the RD and Normalized Factor calculated in Equations 12 and 13.
[0178] Equation 14: RepRD ratio = RepRD Target chromosome / Normalized Factor
[0179] The mean and standard deviation of the RepRD ratio were calculated in a reference group of 400 healthy individuals, and the Reads Distance Index (RDI) value for the sample to be analyzed was calculated using Equation 15.
[0180] Equation 15: RDI = RepRD Ratio sample - MEAN(RepRD Ratio reference ) / SD(RepRD Ratio reference )
[0181] 7-2. Calculation of RDI values for chromosomal structural abnormalities After dividing the chromosomes into 50kbase units, the median RD value for each region was defined as RepRD. The normalized factor was the median autosomal RD value. The reference population consisted of data from 437 normal women, and the mean and standard deviation of the RepRD ratio were calculated for each chromosomal region. The RDI value was calculated using Equation 15.
[0182] 7-3. Performance verification using the RD representative value (RepRD) calculation method (using the reciprocal of the median) After calculating the RD values of the sequence information arranged for each genetic region (chromosome by chromosome), the reciprocal of the median of these values was defined as the RD representative value (RepRD). Here, the median is the value that is in the middle when the calculated RD values are arranged in order of magnitude. For example, if there are three values such as 1, 2, and 100, 2 is in the middle, so 2 is the median.
[0183] If there are an even number of RD values, the median is determined by taking the average of the two central values. For example, if there are RD values of 1, 10, 90, and 200, the median is 50, which is the average of the central values of 10 and 90. The analysis samples used were 49 samples confirmed as trisomy 21 and 3,448 samples confirmed as normal. The RepRD value was calculated using the reciprocal of the median RD value. The analysis method calculated the RDI value using the Z-score method, which uses the mean and standard deviation of the RepRD values of the 3,448 normal samples. As a result of the analysis, the presence or absence of chromosomal number abnormalities in the samples could be detected with an accuracy of approximately 0.999 (Table 4, Figure 15).
[0184] [Table 4]
[0185] 7-4. Performance verification using the RD representative value (RepRD) calculation method (using the average) After calculating the RD values of sequence information aligned for each genetic region (chromosome-specific), the average of these values was defined as the RD representative value (RepRD). Here, the mean is the arithmetic mean of the calculated RD values; for example, if there are RD values of 10, 50, and 90, the RD representative value is (10 + 50 + 90) / 3, which is 50. Using 1,999 normal human samples and 163 T21 samples, the RDI value was calculated using the Z-score method with the RepRD mean and standard deviation of the normal human population. The chromosomes used as normalized factors were 2, 7, 9, 12, and 14. The results of the analysis showed that the presence or absence of chromosomal number abnormalities in the samples could be detected with an accuracy of approximately 0.9995, and when the critical value was set to 4.0, the sensitivity was confirmed to be 0.999 and the specificity 1.000 (Table 5, Figure 16).
[0186] [Table 5]
[0187] 7-5. Performance verification using the RD representative value (RepRD) calculation method (using the inverse value of the mean) After calculating the RD values of sequence information aligned for each genetic region (chromosome-specific), the reciprocal of the mean of these values was defined as the RD representative value (RepRD). Here, the mean is the arithmetic mean of the calculated RD values. For example, if there are RD values of 10, 50, and 90, the mean is (10 + 50 + 90) / 3, which is 50. The reciprocal of this value, 1 / 50 = 0.02, was used as the RD representative value. Using 1,999 normal human samples and 163 T21 samples, the RDI value was calculated using the Z-score method with the RepRD mean and standard deviation of the normal population. The chromosomes used as normalized factors were 2, 7, 8, 9, 12, and 14. The results of the analysis showed that the presence or absence of chromosomal number abnormalities in the samples could be detected with an accuracy of approximately 0.9995, and when the critical value was set to 4.3, the sensitivity was confirmed to be 0.993 and the specificity 1.000 (Table 6, Figure 17).
[0188] [Table 6]
[0189] Example 8. Performance verification of RDI value for detecting chromosomal numerical abnormalities 8-1. Distribution of Read Count and Read Distance In an analysis using the distance concept of aligned reads, the greater the number of reads produced, the shorter the distance between each read will likely remain. To confirm this, we analyzed the distribution of read number and RepRD value for each chromosome.
[0190] As a result, we confirmed that RepRD decreases as the number of reads increases overall. In particular, we confirmed that the relationship between the number of reads and the RepRD value is nonlinear rather than linear, which means that the read distance concept reflects not only the number of simple reads but also their aligned positions (Figure 8).
[0191] Compared to normal samples, the RepRD values of the chromosomes identified as aneuploid in trisomy 13, 18, and 21 were found to be lower (Figure 9).
[0192] 8-2. Performance of RDI (Reads Distance Index) and its relationship with fetal fractions, clinical information, and existing G-score. In fetal aneuploidy screening using maternal blood, fetal fraction and gestational age significantly affect the accuracy of the test. Fetal fractions tend to be higher with increasing gestational age, and higher fetal fractions indicate higher test accuracy. (RDI of 21 Trisomy Samples) chr21 Analysis of the distribution of values, gestational age of the mother, and fetal fraction revealed that as the fetal fraction increases, RDI chr21 We confirmed that the value decreased. We also confirmed the relationship between gestational age and RDI. chr21 Regarding the relationship with the values, a tendency for the values to decrease was observed in samples of 15 weeks or more. When the relationship with the G-score (Korean Patent No. 10-1686146), which is a value based on the number of existing reads, was examined, a similar trend was confirmed (Figure 10).
[0193] 8-3. Results of analysis of positive clinical specimens The analytical performance was validated using RDI values for the normal group and for samples identified as having aneuploidy on each chromosome. After setting the RDI value to a constant cutoff (-3), the analytical performance between normal and aneuploid samples was compared, and accuracy was confirmed to be 0.991 for trisomy 13, 0.989 for trisomy 18, and 0.998 for trisomy 21 (Table 7). In addition, the AUC values for trisomy 13, 18, and 21 were confirmed to be 0.999, 0.984, and 1.000, respectively (Figure 11).
[0194] [Table 7]
[0195] 8-4. Performance verification using the RDI calculation method We examined the results of a Log ratio analysis using the median, which differs from the Z-score method that utilizes the mean and standard deviation of the normal reference population RDI ratio. Equation 16 was used for the Log ratio analysis.
[0196] Equation 16: RDI = log 10 (RepRD Ratio sample / Median(RepRD Ratio reference ))
[0197] Using the same samples as in Example 8-3, the analytical performance was compared after setting the RDI value to a fixed cutoff (-0.0045). There were slight differences in performance depending on the type of positive result, and the accuracy was confirmed to be 0.976 for trisomy 21, 0.994 for trisomy 18, and 0.991 for trisomy 13 (Table 8).
[0198] [Table 8]
[0199] 8-5. Downsampling Performance Verification In non-invasive fetal aneuploidy testing using next-generation sequencing technology, the amount of data produced (number of reads) is known to be an important factor in accuracy. In this example, the analytical performance of the RDI method was calculated based on the number of reads. The analytical performance criterion was the AUC value of ROC analysis, and the number of reads was determined using an in-silico random read selection method. 1 million to 10 million reads were randomly selected. Analysis using an aneuploidy sample with lead number 21 confirmed that analytical performance decreased as the number of reads decreased (Figure 12).
[0200] Example 9. Performance verification of RDI values for detecting chromosomal structural abnormalities. 9-1. Distribution of lead count and lead distance To investigate the presence or absence of chromosomal structural abnormalities using RDI, it is necessary to divide the chromosome into appropriately sized segments. In this embodiment, the chromosome was divided into segments of 50 k bases. Read distances are distributed such that they become shorter as the number of reads increases, and longer as the number of reads increases. When the relationship between the number of reads and distance in the divided segments was examined, it was confirmed that the read distances in regions where deletions, which are structural abnormalities of the chromosome, were confirmed were longer than those in regions without structural abnormalities (Figure 13).
[0201] 9-2. Comparison with Microarray Results The results of microarray testing to detect chromosomal structural abnormalities were compared with RDI analysis results. The analyzed sample was one in which a 3,897,640 bp deletion was confirmed at the end of chromosome 1. Analysis using RDI confirmed the detection of structural abnormalities (deletions) in a similar region of 3,700,000 bp (Figure 14).
[0202] Having described in detail certain aspects of the present invention, it will be clear to those with ordinary skill in the art that such specific descriptions merely represent preferred embodiments and do not limit the scope of the invention. Therefore, the substantial scope of the present invention is defined by the appended claims and their equivalents. [Industrial applicability]
[0203] Unlike existing methods that determine chromosome quantity based on the number of reads, the chromosomal aberration detection method according to the present invention groups aligned nucleic acid fragments and then uses a distance concept between reference values of nucleic acid fragments. While existing methods lose accuracy as the number of reads decreases, the method of the present invention can improve detection accuracy even when the number of reads decreases, and is also useful because it maintains high detection accuracy even when analyzing the distance between nucleic acid fragments in a certain interval rather than all chromosomal intervals.
Claims
1. (A) The step of extracting nucleic acids from a biological sample to obtain nucleic acid fragments and acquire sequence information; (B) The step of confirming the position of nucleic acid fragments in the reference genome database based on the acquired sequence information (reads); (C) A step of grouping the sequence information (reads) into a whole sequence group, a forward sequence group, and a reverse sequence group based on the aligned direction; (D) A step of defining a reference value for each nucleic acid fragment using the grouped sequence information, measuring the distance between the reference values, and calculating the FD value (Fragments Distance) for each group; and (E) Based on the FD values for each group calculated in step (D) above, the FDI value (Fragments Distance Index) is calculated for the entire chromosome region or for specific chromosomes in each group, and if all group-specific FDI values are below or above the reference value, it is determined that there is a chromosomal abnormality. An in vitro method for detecting numerical abnormalities of chromosomes, including, The overall array group consists of a forward-direction array group and a reverse-direction array group. The FD value in step (D) above is calculated from the distance between the reference value of the i-th nucleic acid fragment and the reference value of the adjacent nucleic acid fragment. The reference value for the nucleic acid fragment is obtained by adding 30-70% of the average length of the nucleic acid fragment to the median value of the nucleic acid fragment, or by subtracting 30-70% of the average length of the nucleic acid fragment, and The aforementioned (E) stage includes the following steps: (E-i) The step of determining a representative value (RepFD) for the entire chromosome region or a specific chromosome; (E-ii) A step in which one or more values selected from the group consisting of the sum, difference, product, mean, log of product, log of sum, median, quantile, minimum, maximum, variance, standard deviation, median absolute deviation, coefficient of variation, reciprocals thereof, and combinations thereof are calculated to derive a normalized factor; (E-iii) The step of calculating the representative ratio (RepFD ratio) based on the following formula 1; Formula 1: RepFD ratio = RepFD Target genomic region / Normalized Factor (E-iv) The step of calculating the Fragments Distance Index (FDI) by comparing the RepFD ratio of the sample with that of a normal reference population; The method is characterized in that the representative value (RepFD) at the (E-i) stage is one or more values selected from the group consisting of the sum, difference, product, mean, log of product, log of sum, median, quantile, minimum, maximum, variance, standard deviation, median absolute deviation, and coefficient of variation, and / or one or more inverse values thereof.
2. The in vitro method for detecting numerical abnormalities of chromosomes according to claim 1 is characterized in that step (A) is performed by a method comprising the following steps: (A-i) The step of obtaining nucleic acids from blood, semen, vaginal cells, hair, saliva, urine, oral cells, placental cells or fetal cells, amniotic fluid, tissue cells and mixtures thereof; (A-ii) The step of removing proteins, fats, and other residues from the collected nucleic acids using a salting-out method, a column chromatography method, or a beads method to obtain purified nucleic acids; (A-iii) The step of preparing a single-end sequencing or pair-end sequencing library from purified nucleic acids or nucleic acids randomly fragmented by enzymatic cleavage, grinding, or hydroshear method; (A-iv) The step of reacting the prepared library with a next-generation sequencer; and (A-v) The stage of obtaining nucleic acid sequence information (reads) using a next-generation sequencer.
3. The in vitro method for detecting numerical abnormalities of chromosomes according to claim 1, characterized in that the reference value of the nucleic acid fragment is the minimum or maximum position value of the forward and reverse sequence information (reads) in paired-end sequencing.
4. An in vitro method for detecting numerical abnormalities in chromosomes according to claim 3, further comprising a step of excluding nucleic acid fragments from the calculation process if the alignment match score of the sequence information (reads) is less than a reference value, characterized in that the alignment match score is obtained in step (b).
5. The in vitro method for detecting numerical abnormalities of chromosomes according to claim 1, characterized in that the reference value of the nucleic acid fragment is one of either the positional value of the forward or reverse sequence information (read) in single-end sequencing.
6. An in vitro method for detecting numerical abnormalities in chromosomes according to claim 5, characterized in that when deriving a position value based on sequence information aligned in the forward direction, 30 to 70% of the average length of the nucleic acid fragment is added, and when deriving a position value based on sequence information aligned in the reverse direction, 30 to 70% of the average length of the nucleic acid fragment is subtracted.
7. The in vitro method for detecting numerical abnormalities of chromosomes according to claim 1, characterized in that the representative value (RepFD) of the (E-i) stage is the median, mean, or reciprocal value of the FD values.
8. The in vitro method for detecting numerical abnormalities in chromosomes according to claim 1 is characterized by selecting the entire chromosome region to be analyzed in step (E-ii) above, or a specific region within a sample other than the specific chromosome, by a method including the following steps: a) The step of randomly selecting the entire chromosome region to be analyzed or a region other than a specific chromosome; b) A step in which the representative value (RepFD) of the genetic region selected in step a) above is determined as a Pre-Normalized Factor (PNF); c) The step of calculating the representative ratio (RepFD ratio) based on the following formula 2: Equation 2: RepFD ratio = RepFD Target genomic region / PNF d) The step of calculating the coefficient of variation (SD / Mean) of the RepFD ratio values for the normal reference population; and e) A step in which the genetic region with the smallest coefficient of variation among the coefficients of variation obtained by repeatedly performing steps a) to d) above is determined as the entire chromosome region or a specific region within the sample other than the specific genetic region.
9. The (E-iv) stage above uses FDI as shown in equation 6: FDI = (MEAN(RepFD Ratio reference - RepFD Ratio sample -) / SD(RepFD Ratio reference ) An in vitro method for detecting numerical abnormalities of chromosomes according to claim 1, characterized by using a calculation.
10. A computer-readable storage medium comprising instructions configured to be executed by a processor that detects numerical abnormalities in chromosomes, (A) The step of extracting nucleic acids from a biological sample to obtain nucleic acid fragments and acquire sequence information; (B) The step of confirming the position of nucleic acid fragments in the reference genome database based on the acquired sequence information (reads); (C) A step of grouping the sequence information (reads) into a whole sequence group, a forward sequence group, and a reverse sequence group based on the aligned direction; (D) A step of defining a reference value for each nucleic acid fragment using the grouped sequence information, measuring the distance between the reference values, and calculating the FD value (Fragments Distance) for each group; and (E) Based on the FD values for each group calculated in step (D) above, the entire chromosome region Each region or group of specific chromosomes has its own FDI value (Fragments D The `stance Index` is calculated, and all group-specific FDI values are compared to the baseline. The stage at which a chromosomal abnormality is determined if the value falls below or exceeds a certain level. A computer-readable storage medium containing instructions configured to be executed by a processor that detects chromosomal abnormalities, The overall array group consists of a forward-direction array group and a reverse-direction array group. The FD value in step (D) above is calculated from the distance between the reference value of the i-th nucleic acid fragment and the reference value of the adjacent nucleic acid fragment, and The reference value for the nucleic acid fragment is obtained by adding 30-70% of the average length of the nucleic acid fragment to the median value of the nucleic acid fragment, or by subtracting 30-70% of the average length of the nucleic acid fragment. The aforementioned (E) stage includes the following steps: (E-i) The step of determining a representative value (RepFD) for the entire chromosome region or a specific chromosome; (E-ii) A step in which one or more values selected from the group consisting of the sum, difference, product, mean, log of product, log of sum, median, quantile, minimum, maximum, variance, standard deviation, median absolute deviation, coefficient of variation, reciprocals thereof, and combinations thereof are calculated to derive a normalized factor; (E-iii) The step of calculating the representative ratio (RepFD ratio) based on the following formula 1; Formula 1: RepFD ratio = RepFD Target genomic region / Normalized Factor (E-iv) The step of calculating the Fragments Distance Index (FDI) by comparing the RepFD ratio of the sample with that of a normal reference population. The computer-readable storage medium is characterized in that the representative value (RepFD) of the (E-i) stage is one or more values selected from the group consisting of the sum, difference, product, mean, log of product, log of sum, median, quantile, minimum, maximum, variance, standard deviation, median absolute deviation, and coefficient of variation, and / or one or more inverse values thereof.
11. An in vitro method for detecting numerical abnormalities of chromosomes according to claim 1, characterized in that when the reference value of the nucleic acid fragment is obtained by adding or subtracting 50% of the average length of the nucleic acid fragment, the calculated FD value is the RD value (Read Distance).
12. (A) The step of extracting nucleic acids from a biological sample and obtaining sequence information; (B) The step of aligning the acquired sequence information (reads) with the reference genome database; (C) A step of measuring the distance between aligned reads and calculating the RD value (Read Distance) with respect to the aligned sequence information (reads); and (D) Based on the RD value calculated in step (C) above, the RDI value (Read Distance Index) is calculated for the entire chromosome region or for specific regions, and if the RDI value is below or above the reference value, it is determined that there is a chromosomal abnormality. The RD value in step (C) above is calculated for the n acquired reads from the distance between the i-th read and one of the end values of the adjacent read, plus or minus 50% of the average nucleic acid length. The aforementioned (D) stage includes the following steps: (D-i) The step of determining a representative value (RepRD) for the entire chromosome region or a specific region; (D-ii) A step of calculating one or more values and / or one or more inverse values thereof selected from the group consisting of the sum, difference, product, mean, log of product, log of sum, median, quantile, minimum, maximum, variance, standard deviation, median absolute deviation, and coefficient of variation of RepRD values for the entire chromosome or a specific region within the sample other than the specific region, and deriving a Normalized Factor; (D-iii) The step of calculating the representative ratio (RepRD ratio) based on the following formula 10; Formula 10: RepRD ratio = RepRD Target genomic region / Normalized Factor (D-iv) The step of comparing the RepRD ratio of the sample with that of a normal reference population and calculating the RDI (Read Distance Index). An in vitro method for detecting numerical abnormalities in chromosomes, characterized in that the representative value (RepRD) of the (D-i) stage is one or more values selected from the group consisting of the sum, difference, product, mean, log of product, log of sum, median, quantile, minimum, maximum, variance, standard deviation, median absolute deviation, and coefficient of variation, and / or one or more inverse values thereof.
13. The in vitro method for detecting numerical abnormalities of chromosomes according to claim 12, characterized in that step (A) is carried out by a method comprising the following steps: (A-i) The step of obtaining nucleic acids from blood, semen, vaginal cells, hair, saliva, urine, oral cells, placental cells or fetal cells, amniotic fluid, tissue cells and mixtures thereof; (A-ii) The step of removing proteins, fats, and other residues from the collected nucleic acids using a salting-out method, a column chromatography method, or a beads method to obtain purified nucleic acids; (A-iii) The step of preparing a single-end sequencing or pair-end sequencing library from purified nucleic acids or nucleic acids randomly fragmented by enzymatic cleavage, grinding, or hydroshear method; (A-iv) The step of reacting the prepared library with a next-generation sequencer; and (A-v) The stage of obtaining nucleic acid sequence information (reads) using a next-generation sequencer.
14. An in vitro method for detecting numerical abnormalities of chromosomes according to claim 12, characterized in that a step of grouping reads aligned prior to step (C) by their aligned direction is further available.
15. The in vitro method for detecting numerical abnormalities of a chromosome according to claim 12, characterized in that the RD value is obtained by calculating the distance between the 5' or 3' end inside the i-th read and the 5' or 3' end of an adjacent read.
16. The in vitro method for detecting numerical abnormalities of chromosomes according to claim 12, characterized in that the representative value (RepRD) of the (D-i) stage is the median, mean, or reciprocal value of the RD values.
17. The in vitro method for detecting numerical abnormalities in chromosomes according to claim 12 is characterized by selecting specific regions within a sample other than the entire chromosome region or specific genetic region to be analyzed in step (D-ii) above by a method including the following steps: a) The step of randomly selecting regions other than the entire chromosome region or a specific genetic region to be analyzed; b) A step in which a representative value of the RepRD value of the genetic region selected in step a) above is determined as a Pre-Normalized Factor (PNF); c) The step of calculating the representative ratio (RepRD ratio) based on the following formula 11: Formula 11: RepRD ratio = RepRD Target genomic region / PNF d) The step of calculating the coefficient of variation (SD / Mean) of the representative ratio (RepRD ratio) of the normal reference population; and e) A step in which the genetic region with the smallest coefficient of variation among the coefficients of variation obtained by repeatedly performing steps a) to d) above is determined as the entire chromosome region or a specific region within the sample other than the specific genetic region.
18. The (D-iv) step uses FDI as shown in equation 15: RDI = RepRD Ratio sample - MEAN(RepRD Ratio reference ) / SD(RepRD Ratio reference ) An in vitro method for detecting numerical abnormalities of chromosomes according to claim 12, characterized by using a calculation.