Method for detecting chromosomal abnormality by using nucleic acid fragment end sequence pattern
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- GREEN CROSS GENOME CORP
- Filing Date
- 2026-01-23
- Publication Date
- 2026-07-30
Smart Images

Figure KR2026001411_30072026_PF_FP_ABST
Abstract
Description
Detection of chromosomal abnormalities using nucleic acid fragment terminal sequence patterns
[0001] The present invention relates to a method for detecting chromosomal abnormalities using nucleic acid fragment terminal sequence patterns, and more specifically, to a method for detecting chromosomal abnormalities using a method of extracting nucleic acids from a biological sample, obtaining sequence information, calculating motif frequencies of nucleic acid fragment terminal sequences for each chromosome based on aligned reads, inputting this into a classification model, and comparing the calculated values with reference values.
[0002]
[0003] Chromosomal abnormalities are associated with genetic defects and tumor diseases. Chromosomal abnormalities may refer to deletions or duplications of chromosomes, deletions or duplications of parts of chromosomes, or damage (breaks), translocations, or inversions within chromosomes. Chromosomal abnormalities are a disorder of genetic balance that causes fetal death or severe defects in physical and mental conditions, as well as tumor diseases. For example, Down's syndrome is a common form of chromosomal number abnormality caused by the presence of three copies of chromosome 21 (trisomy 21). Edwards syndrome (trisomy 18), Patau syndrome (trisomy 13), Turner syndrome (XO), and Klinefelter syndrome (XXY) are also examples of chromosomal number abnormalities.
[0004] Chromosomal abnormalities can be detected using karyotype analysis and Fluorescent In Situ Hybridization (FISH). These detection methods are disadvantageous in terms of time, effort, and accuracy. Additionally, DNA microarrays can be used to detect chromosomal abnormalities. In particular, while genomic DNA microarray systems allow for easy probe fabrication and can detect chromosomal abnormalities in intron regions as well as extended regions of chromosomes, it is difficult to produce a large number of DNA fragments with confirmed localization and function within the chromosome.
[0005] Recently, next-generation sequencing technology has been used for the analysis of chromosomal number anomalies (Park, H., Kim et al., Nat Genet 2010, 42, 400-405.; Kidd, JM et al., Nature 2008, 453, 56-64). However, this technology requires high coverage readings for the analysis of chromosomal number anomalies, and CNV measurements also require independent validation. Consequently, it was not suitable as a general gene search analysis at the time because the cost was very high and the results were difficult to understand.
[0006] Meanwhile, existing prenatal screening tests for fetal chromosomal abnormalities include ultrasound, blood marker tests, amniocentesis, chorionic villus sampling, and percutaneous cord blood testing (Mujezinovic F, et al. Obstet Gynecol. 2007, 110(3):687-94.). Among these, ultrasound and blood marker tests are classified as screening tests, while amniocentesis is classified as a confirmatory test. Ultrasound and blood marker tests are non-invasive methods that are safe because they do not involve direct sample collection from the fetus, but their sensitivity drops below 80% (ACOG Committee on Practice Bulletins. 2007). Invasive methods such as amniocentesis, chorionic villus sampling, and percutaneous cord blood testing can confirm fetal chromosomal abnormalities, but they have the disadvantage of a risk of fetal loss due to the invasive medical procedure.
[0007] In 1997, Lo et al. succeeded in sequencing the Y chromosome of fetal-derived genetic material from maternal plasma and serum, enabling the use of fetal genetic material within the mother for prenatal testing (Lo YM, et al. Lancet. 1997, 350(9076):485-7). Fetal genetic material in maternal blood is defined as cell-free fetal DNA (cff DNA), which is a portion of trophoblast cells that have undergone apoptosis during placental remodeling and entered the maternal blood through a material exchange mechanism, and is actually derived from the placenta.
[0008] CFF DNA is detected in most maternal blood as early as day 18 after embryo transfer, and by day 37. Since CFF DNA is characterized by being a short strand of less than 300 bp and existing in small quantities in maternal blood, large-scale parallel sequencing technology using next-generation sequencing (NGS) is being used to detect fetal chromosomal abnormalities. Although the non-invasive fetal chromosomal abnormality detection performance using large-scale parallel sequencing technology shows a detection sensitivity of over 90-99% depending on the chromosome, false positive and false negative results amount to 1-10%, making it necessary to develop correction technology for this.
[0009] In this regard, Korean Patent Publication No. 10-2017-0036648 provides a non-invasive fetal chromosome analysis method capable of reducing the probability of false positives by improving the reliability and specificity of results, by determining fetal chromosomal aneuploidy from DNA sequence information obtained from a maternal blood sample, and by eliminating inter-experimental variation by comparing the average read number of a specific chromosome to be determined with the average read number of other chromosomes excluding said chromosome, and by using the ratio of read numbers between chromosomes weighted by the coefficient of variation (CV).
[0010] In addition, in 2020, Peiyong Jiang et al. reported that nucleic acid terminal sequence motif information of plasma cfDNA can be used for cancer diagnosis (Peiyong Jiang et al., cancer discovery, Vol. 10, 2020, pp. 664-673).
[0011] Patents (US 12374462B, US 2024-0209455A) related to cancer diagnosis / treatment methods using such nucleic acid terminal sequence motif information have been disclosed, but there is a lack of research on methods to detect chromosomal abnormalities using this.
[0012] Accordingly, the inventors have made diligent efforts to solve the aforementioned problems and develop a method for detecting chromosomal abnormalities with high sensitivity and accuracy. As a result, they confirmed that chromosomal abnormalities can be detected with high sensitivity and accuracy by calculating the frequency of nucleic acid terminal sequence motifs per chromosome based on reads aligned to chromosomal regions and analyzing this using a classification model, thereby completing the present invention.
[0013]
[0014] Summary of the Invention
[0015] The objective of the present invention is to provide a method for detecting chromosomal abnormalities using nucleic acid fragment terminal sequence patterns.
[0016] Another objective of the present invention is to provide a device for determining chromosomal abnormalities using nucleic acid fragment terminal sequence patterns.
[0017] Another objective of the present invention is to provide a computer-readable storage medium comprising instructions configured to be executed by a processor that determines chromosomal abnormalities by the above method.
[0018] To achieve the above objective, the present invention provides a method for detecting chromosomal abnormalities comprising: (a) a step of extracting nucleic acids from a biological sample to obtain sequence information; (b) a step of aligning the obtained sequence information (reads) to a standard chromosome sequence database (reference genome database); (c) a step of calculating the frequency of terminal sequence motifs of nucleic acid fragments by chromosome using the aligned sequence information (reads); and (d) a step of determining chromosomal abnormalities by inputting the calculated frequency of terminal sequence motifs of nucleic acid fragments into a classification model and comparing the output result analyzed with a cut-off value.
[0019] The present invention also comprises a decoding unit that extracts nucleic acids from a biological sample and decodes sequence information;
[0020] An alignment unit that aligns the decoded sequence to a standard chromosome sequence database; and
[0021] A chromosomal abnormality detection device is provided, comprising: a nucleic acid fragment analysis unit that calculates the frequency of terminal sequence motifs of nucleic acid fragments based on aligned sequences for each chromosome; and a chromosomal abnormality determination unit that inputs the calculated frequency of terminal sequence motifs of nucleic acid fragments for each chromosome into a classification model for analysis and determines whether there is a chromosomal abnormality by comparing it with a reference value.
[0022] The present invention also provides a computer-readable storage medium comprising instructions configured to be executed by a processor for detecting chromosomal abnormalities, wherein the instructions include: (a) a step of extracting nucleic acids from a biological sample to obtain sequence information; (b) a step of aligning the obtained sequence information (reads) to a standard chromosomal sequence database (reference genome database); (c) a step of calculating the frequency of terminal sequence motifs of nucleic acid fragments by chromosome using the aligned sequence information (reads); and (d) a step of determining chromosomal abnormalities by inputting the calculated frequency of terminal sequence motifs of nucleic acid fragments into a classification model and comparing the output result value analyzed with a cut-off value.
[0023]
[0024] Effects of the invention
[0025] The nucleic acid fragment terminal sequence motif frequency-based chromosomal abnormality detection method according to the present invention is useful because it can improve the efficiency and reliability of prenatal diagnosis by having higher accuracy and sensitivity compared to the existing method using a step of determining chromosomal amount based on z-scores.
[0026]
[0027] Figure 1 is an overall flowchart for performing the method for detecting chromosomal abnormalities using nucleic acid fragment terminal sequence motif frequencies of the present invention.
[0028] FIG. 2a is an overall flowchart for performing a method for detecting chromosomal abnormalities using nucleic acid fragment terminal sequence motif frequencies using a machine learning model performed according to an embodiment of the present invention, and FIG. 2b is an overall flowchart for performing a method for detecting chromosomal abnormalities using nucleic acid fragment terminal sequence motif frequencies using a statistical analysis model performed according to an embodiment of the present invention.
[0029] FIG. 3 is a conceptual diagram showing a nucleic acid fragment terminal sequence motif obtained according to one embodiment of the present invention.
[0030] Figure 4 shows the results of verifying the performance of the T21 model constructed according to one embodiment of the present invention, where 4a is the confusion matrix, (A) represents the train group, (B) represents the validation group, and (C) represents the test group, and 4b is the distribution of output values for each data set.
[0031] Figure 5 shows the results of verifying the performance of the T18 model constructed according to one embodiment of the present invention, where 5a is the confusion matrix, (A) represents the train group, (B) represents the validation group, and (C) represents the test group, and 5b is the distribution of output values for each data set.
[0032] Figure 6 shows the results of verifying the performance of the T13 model constructed according to one embodiment of the present invention, where 6a is the confusion matrix, (A) represents the train group, (B) represents the validation group, and (C) represents the test group, and 6b is the distribution of output values for each data set.
[0033] Figure 7 is the verification result of an artificially positive sample based on genomic DNA of a T21 model constructed according to one embodiment of the present invention.
[0034] FIG. 8 is a result of comparing a T21 model constructed according to an embodiment of the present invention with a conventional Z-score based method, where (A) is the result of the conventional method and (B) is the result of the T21 model, the left panel represents the ROC-curve, and the right panel represents the remaining output values of each model.
[0035] FIG. 9 is a result of comparing a T21 model constructed according to an embodiment of the present invention with a conventional Z-score-based method by 50% downsampling, where (A) is the result of the conventional method and (B) is the result of the T21 model, the left panel represents the ROC-curve, and the right panel represents the remaining output values of each model.
[0036] FIG. 10 is a result of comparing a T21 model constructed according to an embodiment of the present invention with a conventional Z-score-based method by downsampling by 25%, where (A) is the result of the conventional method and (B) is the result of the T21 model, the left panel represents the ROC-curve, and the right panel represents the remaining output values of each model.
[0037] Figure 11 shows the results of analyzing the frequency of 256 nucleic acid fragment terminal sequence motifs on chromosome 21 of normal samples and Down syndrome samples using a heatmap.
[0038] FIG. 12 is a conceptual diagram illustrating a method for converting the Z-score of a statistical analysis model constructed according to one embodiment of the present invention.
[0039] FIG. 13 is a graph showing the integrated Z-score distribution for each chromosome of a T21 sample analyzed according to a statistical analysis model constructed according to one embodiment of the present invention.
[0040] FIG. 14 is a graph showing the integrated Z-score distribution of each sample analyzed according to a statistical analysis model constructed according to one embodiment of the present invention, where (A) is the result of analysis based on the whole nucleic acid fragment and (B) is the result based on a nucleic acid fragment of 150 bp or less.
[0041]
[0042] Detailed Description of the Invention and Preferred Embodiments
[0043] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by a skilled expert in the art to which this invention pertains. In general, the nomenclature used herein and the experimental methods described below are well known and commonly used in the art.
[0044] Terms such as first, second, A, B, etc., may be used to describe various components, but such components are not limited by the said terms and are used solely for the purpose of distinguishing one component from another. For example, without departing from the scope of rights of the technology described below, the first component may be named the second component, and similarly, the second component may be named the first component. The term "and / or" includes a combination of multiple related described items or any of the multiple related described items.
[0045] In terms used in this specification, singular expressions should be understood to include plural expressions unless the context clearly indicates otherwise, and terms such as “includes” should be understood to mean that the described features, number, step, action, component, part, or combination thereof exist, and not to exclude the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.
[0046] Before providing a detailed description of the drawings, it is to clarify that the classification of components in this specification is merely based on the primary function each component is responsible for. That is, two or more components described below may be combined into a single component, or a single component may be divided into two or more components based on more subdivided functions. Furthermore, each component described below may additionally perform some or all of the functions of other components in addition to its own primary function, and it goes without saying that some of the primary functions of each component may be exclusively performed by other components.
[0047] Furthermore, in performing the method or operation method, each process constituting the method may occur differently from the specified order unless a specific order is clearly indicated in the context. That is, each process may occur in the same order as specified, may be performed substantially simultaneously, or may be performed in the reverse order.
[0048]
[0049] In this invention, we aimed to confirm that chromosomal abnormalities can be detected with high sensitivity and accuracy by aligning sequence analysis data obtained from a sample with a reference genome, deriving the frequency of terminal sequence motifs of nucleic acid fragments per chromosome based on the aligned sequence information, and inputting this into a classification model for analysis.
[0050] In the present invention, a method was developed in which DNA extracted from blood is sequenced and aligned to a reference chromosome, the frequency of terminal sequence motifs of nucleic acid fragments per chromosome is derived using this, and the output value is calculated by inputting it into a classification model and comparing it with a reference value; if the output value is greater than or equal to the reference value, it is determined that there is a chromosomal abnormality (Fig. 1).
[0051]
[0052] Therefore, the present invention, in one aspect,
[0053] This relates to a method for detecting chromosomal abnormalities comprising the following steps:
[0054] (a) A step of extracting nucleic acids from a biological sample to obtain sequence information;
[0055] (b) A step of aligning the acquired sequence information (reads) to a standard chromosome sequence database (reference genome database);
[0056] (c) a step of calculating the frequency of terminal sequence motifs of nucleic acid fragments by chromosome using the above-mentioned aligned sequence information (reads); and
[0057] (d) A step of determining chromosomal abnormalities by comparing the output value analyzed by inputting the frequency of terminal sequence motifs of the calculated nucleic acid fragment into a classification model with a cut-off value.
[0058]
[0059] In the present invention, the nucleic acid fragment may be any fragment of nucleic acid extracted from a biological sample without limitation, and preferably may be a fragment of cell-free nucleic acid or intracellular nucleic acid, but is not limited thereto.
[0060] In the present invention, the nucleic acid fragment can be obtained by any method known to a person skilled in the art, preferably by direct sequencing, by next-generation sequencing, by non-specific whole genome amplification, or by probe-based sequencing, but is not limited thereto.
[0061] In the present invention, the nucleic acid fragment may refer to a read when using next-generation sequencing analysis.
[0062] In the present invention, the term "reads" refers to a single nucleic acid fragment obtained by analyzing sequence information using various methods known in the art. Accordingly, in this specification, the terms "sequence information" and "reads" have the same meaning in that they refer to the result obtained by acquiring sequence information through a sequencing process.
[0063] In the present invention, the term “chromosomal abnormality” refers to various variations occurring in chromosomes, which can be broadly classified into numerical abnormalities, structural abnormalities, microdeletions, chromosomal instability, etc.
[0064] An abnormality in the number of chromosomes is a case where an abnormality occurs in the number of chromosomes, and can include all cases where an abnormality occurs in the total number of chromosomes, such as Down Syndrome (47 total chromosomes with one extra 21st chromosome), Turner Syndrome (45 total chromosomes with a single X), and Klinefelter Syndrome (XXYY, XXXY, XXXXY, etc.).
[0065] Structural abnormalities of chromosomes refer to all cases where there is no change in the number of chromosomes but changes occur in their structure, such as deletions, duplications, inversions, and translocations. For example, there may be a deletion of a portion of chromosome 5 (Cat-Cry syndrome), a deletion of a portion of chromosome 7 (Phillias syndrome), a duplication of a portion of chromosome 12 (Wolf-Hirschhorn syndrome), and a translocation between chromosomes 9 and 22 (chronic myeloid leukemia). It may also include microduplications and microdeletions in certain chromosomal regions found in patients with tumors. However, it is not limited to the contents described above.
[0066]
[0067] In the present invention, the step of obtaining sequence information of step (a) may be characterized by obtaining the obtained DNA through whole-genome sequencing at a depth of 100,000 to 100 million reads.
[0068]
[0069] The term “NGS” in the present invention refers to Next Generation Sequencing, and also refers to next-generation sequencing and next-generation DNA sequencing. It refers to a technology that fragments a whole genome and performs high-throughput sequencing based on chemical reactions (hybridization) of the fragments, or performs high-throughput sequencing by amplifying target gene regions using multiplex PCR. It includes technologies from Agilent, Illumina, Roche, and Life Technologies, and in a broader sense, is defined to include third-generation technologies such as those from Pacificbio and Nanopore Technology, as well as fourth-generation technologies.
[0070] In the present invention, the next-generation sequencer may be any sequencing method known in the art. Sequencing of nucleic acids separated by a selection method is typically performed using next-generation sequencing (NGS). Next-generation sequencing includes any sequencing method that determines the nucleotide sequence of an individual nucleic acid molecule or one of a proxy cloned for an individual nucleic acid molecule in a highly similar manner (e.g., 10⁵ or more molecules are sequenced simultaneously). In one embodiment, the relative abundance of nucleic acid species within a library may be estimated by counting the relative number of occurrences of their cognate sequences in the data generated by the sequencing experiment. Next-generation sequencing methods are known in the art and are described, for example, in the literature incorporated herein by reference (Metzker, M. (2010) Nature Biotechnology Reviews 11:31-46).
[0071] Platforms for next-generation sequencing include, but are not limited to, Roche / 454’s Genome Sequencer (GS) FLX system, Illumina / Solexa’s Genome Analyzer (GA), Life / APG’s Support Oligonucleotide Ligation Detection (SOLiD) system, Polonator’s G. 007 system, Helicos BioSciences’ HeliScope Gene Sequencing system, and Pacific Biosciences’ PacBio RS system.
[0072]
[0073] In the present invention, the sequence information obtained after performing the alignment step of step (b) may be in the bam file format, but is not limited thereto, and the standard reference chromosome sequence database may be a reference chromosome registered with a public health institution such as NCBI.
[0074] In the present invention, the length of the sequence information (reads) in step (b) is 5 to 5000 bp, and the number of sequence information used can be 5,000 to 5,000,000, but is not limited thereto.
[0075]
[0076] In the present invention, the nucleic acid fragment terminal sequence motif of step (c) may be characterized as being a pattern of 2 to 30 base sequences at both ends of the nucleic acid fragment.
[0077] In the present invention, “nucleic acid fragment terminal sequence motif frequencies may be similar to those described in US Patent No. 12,374,462, but are not limited thereto.”
[0078]
[0079] In the present invention, the frequency of the nucleic acid fragment terminal sequence motif in step (c) may be characterized as the number of each motif detected in the entire nucleic acid fragment.
[0080] That is, when analyzing the nucleic acid fragment end sequence motif based on the four bases at both ends (4-mer motif), since there are four types of base combinations (A, T, G, and C) possible at the 1st, 2nd, 3rd, and 4th positions respectively, a total of 256 (4*4*4*4) combinations of motif values are subject to analysis.
[0081] Motif frequency is the number of times each motif is observed in all nucleic acid fragments produced by sequencing, and the relative frequency of each motif is calculated by dividing this value by the total number of produced nucleic acid fragments.
[0082] For example, if the total number of nucleic acid fragments is 126,430,124 and the number of nucleic acid fragments analyzed as AAAA as a nucleic acid fragment terminal sequence motif is 125,071, the frequency of the AAAA nucleic acid fragment terminal sequence motif is 125,071, and the relative frequency of the nucleic acid fragment terminal sequence motif calculated by dividing this by the total number of nucleic acid fragments is 0.00099.
[0083]
[0084] In the present invention, prior to performing step (c), the invention may further include a step of classifying nucleic acid fragments that satisfy the mapping quality score of the aligned nucleic acid fragments.
[0085] In the present invention, the mapping quality score may vary depending on the desired criteria, but preferably it may be 15-70 points, more preferably 50-70 points, and most preferably 60 points.
[0086]
[0087] In the present invention, the classification model may be characterized as being a learned artificial intelligence model or a statistical analysis model.
[0088] In the present invention, the artificial intelligence model may be used without limitation as long as it is a model capable of learning to distinguish between the frequency of terminal sequence motifs of nucleic acid fragments with normal chromosomal status and the frequency of terminal sequence motifs of nucleic acid fragments with chromosomal abnormalities. Preferably, it may be characterized as being any one selected from the group consisting of a multi-layer perceptron (MLP), a convolutional neural network (CNN), a deep neural network (DNN), a recurrent neural network (RNN), a transformer, and an autoencoder, but is not limited thereto.
[0089]
[0090] In the present invention, when the artificial intelligence model is an MLP, the loss function may be characterized by being expressed by the following Equation 1.
[0091]
[0092] In the present invention, when the artificial intelligence model is an MLP, the learning may be characterized by including the following steps:
[0093] i) A step of classifying the produced data into training, validation, and test data;
[0094] In this case, the training data is used to train the MLP model, the validation data is used to verify hyper-parameter tuning, and the test data is used for performance evaluation after producing the optimal model.
[0095] ii) A step of building an optimal MLP model through hyper-parameter tuning and training processes;
[0096] iii) A step of comparing the performance of various models obtained through hyper-parameter tuning using validation data, and determining the model with the best validation data performance as the optimal model;
[0097]
[0098] In the present invention, the hyper-parameter tuning process is a process of optimizing the values of various parameters (activation function, learning rate, number of dense layers, dropout rate, etc.) constituting the MLP model, and the hyper-parameter tuning process may be characterized by using Bayesian optimization and grid search techniques.
[0099] In the present invention, the learning process may be characterized by optimizing the internal parameters (weights) of the MLP model using predetermined hyper-parameters, determining that the model is overfitting when the validation loss relative to the training loss begins to increase, and stopping the model training before that.
[0100]
[0101] In the present invention, the result value of step (d) can be used without limitation as a specific score or a real number, and preferably may be characterized as a Deep Probability Index (DPI) value, but is not limited thereto.
[0102]
[0103] In the present invention, the Deep probability Index refers to a value expressed as a probability value by using a sigmoid function in the last layer of the artificial intelligence model to adjust the output of the artificial intelligence to a scale of 0 to 1.
[0104] In other words, the sigmoid function is used to train the model so that the DPI value becomes 1 when there is a chromosomal abnormality. For example, if a T21 sample and a normal sample are input, the model is trained so that the DPI value of the T21 sample is close to 1.
[0105]
[0106] In the present invention, when the artificial intelligence model is trained, if there is a chromosomal abnormality, the output result is trained to be close to 1, and if there is no chromosomal abnormality, the output result is trained to be close to 0. Based on 0.25, if the value is 0.25 or higher, it is determined that there is a chromosomal abnormality, and if the value is 0.25 or lower, it is determined that there is no chromosomal abnormality, and thus performance measurement is performed (Training, validation, test accuracy).
[0107] Here, it is obvious to a person of ordinary skill that the threshold value of 0.25 can be changed at any time. For example, if one wants to reduce false positives, one can set a threshold value higher than 0.25 to make the criteria for determining chromosomal abnormalities stricter, and if one wants to reduce false negatives, one can set a lower threshold value to make the criteria for determining chromosomal abnormalities slightly weaker.
[0108] Ideally, a reference value can be determined by applying unseen data (data with known answers that was not used for training) using a trained artificial intelligence model and checking the probability of the DPI value.
[0109]
[0110] In the present invention, the statistical analysis model may be characterized by including the following steps:
[0111] (i) a step of converting the obtained chromosome-specific nucleic acid fragment end sequence motif frequencies into Z-scores using the mean and standard deviation of normal samples; and
[0112] (ii) Step of calculating the combined Z-score for each chromosome using Equation 2 based on the transformed Z-scores for each chromosome:
[0113]
[0114] Here, w is a weight, a unit vector with all elements equal to 1, which can assign equal weight to all motif Z-scores, and
[0115] Z is the Z-score of individual motifs per chromosome, and
[0116] Σ (Sigma) is the covariance matrix between the chromosome-specific motif Z scores of normal individuals, and
[0117] w T z is the sum of all Z scores, and
[0118] sqrt(w T Σ w) is a normalization term that accounts for the correlated variability of the signals.
[0119]
[0120] In the present invention, the reference value of step (d) may be 0.25 when using an artificial intelligence model, and may be 3 when using a statistical analysis model.
[0121] In the present invention, the length of the nucleic acid fragment may be 150 bp or less.
[0122]
[0123] In another aspect, the present invention,
[0124] A decoding unit that extracts nucleic acids from a biological sample and decodes sequence information;
[0125] An alignment unit that aligns the decoded sequence to a standard chromosome sequence database; and
[0126] A nucleic acid fragment analysis unit that calculates the frequency of terminal sequence motifs of nucleic acid fragments based on aligned sequences on a chromosome-by-chromosome basis; and
[0127] The present invention relates to a chromosomal abnormality detection device comprising a chromosomal abnormality determination unit that inputs the calculated terminal sequence motif frequency of nucleic acid fragments by chromosome into a classification model for analysis and determines the presence or absence of chromosomal abnormalities by comparing it with a reference value.
[0128] In the present invention, the decoding unit may include a nucleic acid injection unit that injects nucleic acid extracted from an independent device; and a sequence information analysis unit that analyzes sequence information of the injected nucleic acid, and preferably may be an NGS analysis device, but is not limited thereto.
[0129] In the present invention, the alignment unit may be characterized by receiving sequence information data generated from an independent device and aligning it.
[0130]
[0131] In another aspect, the present invention,
[0132] A computer-readable storage medium comprising instructions configured to be executed by a processor for detecting chromosomal abnormalities,
[0133] (a) A step of extracting nucleic acids from a biological sample to obtain sequence information;
[0134] (b) A step of aligning the acquired sequence information (reads) to a standard chromosome sequence database (reference genome database);
[0135] (c) a step of calculating the frequency of terminal sequence motifs of nucleic acid fragments by chromosome using the above-mentioned aligned sequence information (reads); and
[0136] (d) a step of determining chromosomal abnormalities by comparing the output result value analyzed by inputting the frequency of the terminal sequence motif of the calculated nucleic acid fragment into a classification model with a cut-off value; thereby, the invention relates to a computer-readable storage medium comprising instructions configured to be executed by a processor for detecting chromosomal abnormalities.
[0137]
[0138] In other embodiments, the method according to the present invention may be implemented using a computer. In one embodiment, the computer includes one or more processors connected to a chipset. Additionally, memory, a storage device, a keyboard, a graphics adapter, a pointing device, and a network adapter are connected to the chipset. In one embodiment, the performance of the chipset is enabled by a memory controller hub and an I / O controller hub. In another embodiment, the memory may be used by being directly connected to the processor instead of the chipset. The storage device is any device capable of holding data, including a hard drive, a CD-ROM (Compact Disk Read-Only Memory), a DVD, or other memory device. The memory is involved in data and instructions used by the processor. The pointing device may be a mouse, a trackball, or other type of pointing device and is used to transmit input data to the computer system in combination with a keyboard. The graphics adapter displays images and other information on a display. The network adapter is connected to the computer system via a local or long-distance network. However, the computer used in this invention is not limited to the above configuration and may lack some configuration or include additional configuration, and may also be part of a Storage Area Network (SAN), and the computer of this invention may be configured to be suitable for the execution of a module in a program for performing the method according to this invention.
[0139]
[0140] In this invention, the term "module" may refer to a functional and structural combination of hardware for carrying out the technical concept according to this invention and software for driving said hardware. For example, said module may refer to a logical unit of a specific code and a hardware resource for executing said code, and it is obvious to those skilled in the art that it does not necessarily refer to physically connected code or to a single type of hardware.
[0141]
[0142] Examples
[0143] The present invention will be described in more detail below through examples. These examples are solely for illustrating the present invention, and it will be obvious to those skilled in the art that the scope of the present invention is not to be interpreted as being limited by these examples.
[0144]
[0145] Example 1. DNA was extracted from a sample, and next-generation sequencing was performed.
[0146] 10 mL of blood was collected from 3,044 normal subjects, 164 Trisomy 21 subjects, 54 Trisomy 18 subjects, and 17 Trisomy 13 subjects and stored in EDTA tubes. Within 2 hours of collection, the plasma portion was centrifuged once at 1200 g at 4°C for 15 minutes, and the plasma from the first centrifugation was centrifuged a second time at 16000 g at 4°C for 10 minutes to separate the plasma supernatant after excluding the precipitate. Cell-free DNA was extracted from the separated plasma using the Tiangenmicro DNA kit (Tiangen), and library preparation was performed using the Truseq Nano DNA HT library prep kit (Illumina). Subsequently, sequencing was conducted using a Nextseq500 instrument (Illumina) in 75 Single-end mode. As a result, it was confirmed that approximately 13 million reads were produced per sample.
[0147]
[0148] Example 2. Calculation of Nucleic Acid Fragment Terminal Sequence Motif Frequencies
[0149] 2-1. Machine Learning Models
[0150] The nucleic acid fragment end sequence motifs were set to 4 bases (A, T, G, C), and the total frequency of 256 types (4*4*4*4) of motifs was calculated for each of the 21 chromosomes excluding chromosome 19, X chromosome, and Y chromosome. Then, 5,376 (256*21) feature data per sample, generated by normalizing the frequency to relative frequency, were used as input data for a machine learning model.
[0151]
[0152] 2-2. Statistical Analysis Model
[0153] The nucleic acid fragment terminal sequence motifs were set to 4 bases (A, T, G, C), and the total frequency of 256 (4*4*4*4) types of motifs was calculated for all chromosomes (24), normalized to relative frequency, and then a frequency table for each sample was generated.
[0154]
[0155] Example 3. MLP Model Construction and Training Process
[0156] An MLP artificial intelligence model that distinguishes between normal and chromosomal abnormalities was trained using the data generated in Example 2-1 as input.
[0157] The entire sample was divided into Training, Validation, and Test datasets; the Training dataset was used for model training, the Validation dataset for hyper-parameter tuning, and the Test dataset for final model performance evaluation. The number of samples for each set is as follows.
[0158]
[0159]
[0160]
[0161] The hyper-parameter tuning process is a process of optimizing the values of various parameters (activation function, learning rate, number of dense layers, dropout rate, etc.) that make up the MLP model. Bayesian optimization techniques were used in the hyper-parameter tuning process, and when the validation loss started to increase relative to the training loss, it was determined that the model was overfitting and the model training was stopped.
[0162] The performance of various models obtained through hyperparameter tuning was compared using a validation dataset; the model with the best performance on the validation dataset was determined to be the optimal model, and a final performance evaluation was performed using the test dataset.
[0163] When the motif frequency calculation value of an arbitrary sample is input into the model created through the above process, the probability that the chromosome of the sample is 1) normal, T21, 2) normal, T18, or 3) normal, T13 is calculated through the softmax function, which is the last layer of the MLP model, and this probability value is defined as the Deep Probability Index (DPI).
[0164]
[0165] Example 4. Verification of the performance of the constructed MLP model
[0166] 4-1. T21 Model Performance Verification
[0167] As a result of verifying the performance of the T21 model constructed in Example 3, it was confirmed that when the reference value was set to 0.25 as described in Fig. 4, it showed 100% accuracy in the Test set.
[0168] 4-2. T18 Model Performance Verification
[0169] As a result of verifying the performance of the T18 model constructed in Example 3, it was confirmed that when the reference value was set to 0.25 as described in Fig. 5, it showed 100% accuracy in the Test set.
[0170] 4-3. T13 Model Performance Verification
[0171] As a result of verifying the performance of the T13 model constructed in Example 3, it was confirmed that when the reference value was set to 0.25 as described in Fig. 6, it showed 100% accuracy in the Test set.
[0172]
[0173] Example 5. Validation results of artificially generated positive samples of the constructed MLP model
[0174] Positive control samples prepared by shearing genomic DNA from Down syndrome (Trisomy 21) samples were analyzed using a Z-score-based method (RWK Chiu et al., Proc. Natl. Acad. Sci. USA Vol. 105(51), pp. 20458-20463, 2008) and an MLP model (EDM-NIPT model), respectively.
[0175] As a result, as shown in Figure 7, a strong signal exceeding 10 based on the Z-score of the positive sample was observed, whereas in the case of the EDM-NIPT model, a negative result with a DPI value of 0.002 was observed.
[0176] This means that the EDM-NIPT model reflects the signals of naturally occurring cell-free DNA, implying that more sophisticated analysis is possible.
[0177]
[0178] Example 6. Comparison of the constructed MLP model and the existing method
[0179] 6-1. Comparison with existing methods
[0180] Using 85 specimens identified as fetuses with Down syndrome and 2,152 samples identified as normal fetuses, the MLP model constructed by the method of Examples 1 to 3 was compared with the existing Z-score method (RWK Chiu et al., Proc. Natl. Acad. Sci. USA Vol. 105(51), pp. 20458-20463, 2008). At this time, the average number of fragments of the data used in the analysis was 11,325,171, and the distribution ranged from a maximum of 43,263,297 to a minimum of 159,558.
[0181] As a result, as shown in Figure 8, it was confirmed that the performance of the EDM-NIP model was superior in all indicators, such as ROC-AUC, Accuracy, Sensitivity, Specificity, and PPV.
[0182] 6-2. Performance Comparison During Downsampling
[0183] In order to compare performance based on the amount of data, in-silico dilution of the data was performed. Specifically, 50% data downsampling was performed on an average of 11M fragments to produce an average of 5.6M of data, and the performance was compared using the same method as in Example 6-1.
[0184] As a result, as shown in Figure 9, it was confirmed that the performance of the EDM-NIP model was superior in all indicators, such as ROC-AUC, Accuracy, Sensitivity, Specificity, and PPV.
[0185] In addition, 25% data downsampling was performed to produce an average of 2.8 million data points, and the performance was compared using the same method as in Example 6-1. As shown in Fig. 10, it was confirmed that the performance of the EDM-NIP model was superior in all indicators, including ROC-AUC, Accuracy, Sensitivity, Specificity, and PPV.
[0186] 6-3. Comprehensive Comparison
[0187] As shown in Table 6, it was confirmed that the EDM-NIPT model exhibited overall superior performance compared to existing Z-score-based algorithms.
[0188]
[0189] In other words, in the comparison between Z100 and EDM100, EDM100 recorded higher values in all indicators, including Accuracy (0.9991), Sensitivity (1), Specificity (0.99907), PPV (0.97701), and AUC (0.99999). Similarly, in the comparison between Z50 and EDM50, EDM50 demonstrated better performance in Accuracy (0.996), Specificity (0.99675), PPV (0.9222), and AUC (0.99925). While Z50 had a higher Sensitivity of 1, EDM50 also showed balanced results. Furthermore, in the comparison between Z25 and EDM25, EDM25 demonstrated overwhelmingly superior results in Accuracy (0.9839), Specificity (0.98513), PPV (0.71681), and AUC (0.99712). It was recorded that Z25 showed superiority in Sensitivity (1) but showed low values in other indicators.
[0190] Through this, it was confirmed that the EDM model provides higher accuracy and specificity than existing Z-score-based algorithms, and achieved significant improvements, particularly in positive predictive value (PPV).
[0191]
[0192] Example 7. End-motif frequency heatmap analysis
[0193] Using 723 normal individuals and 162 Down syndrome samples, the frequency of 256 End-Motifs corresponding to chromosome 21 was analyzed using a heatmap. As shown in Figure 11, it was confirmed that Down syndrome samples exhibited a different EDM distribution from normal samples.
[0194]
[0195] Example 8. Construction of a Statistical Analysis Model
[0196] The motif frequency table obtained in Example 2-2 was converted into a z-score using the mean and standard deviation of motif frequencies by chromosome of 150 reference subjects, as described in FIG. 12. The number of samples used to construct the analysis model is as shown in Table 7 below.
[0197]
[0198] For example, the Z-score value of the frequency of the AAAA motif on chromosome 1 of the sample was calculated using the following Equation 3.
[0199] Formula 3: z-score = (Sample chr1 AAAA motif frequency - Mean chr1 AAAA motif frequency of normal individuals) / (Standard deviation of chr1 AAAA motif frequency of normal individuals)
[0200] Subsequently, the combined Z-score per chromosome was calculated using Equation 2 from the Z-score vector of motif frequencies per chromosome.
[0201]
[0202] Here, w is a weight, a unit vector with all elements equal to 1, which can assign equal weight to all motif Z-scores, and
[0203] Z is the Z-score of individual motifs per chromosome, and
[0204] Σ (Sigma) is the covariance matrix between the chromosome-specific motif Z scores of normal individuals, and
[0205] w T z is the sum of all Z scores, and
[0206] sqrt(w T Σ w) is a normalization term that accounts for the correlated variability of the signals.
[0207]
[0208] Example 9. Verification of the performance of the statistical analysis model
[0209] As a result of comparing the integrated Z-score values of the constructed statistical analysis model by chromosome based on the reference value 3, it was confirmed that abnormalities in chromosome 21 in the T21 sample could be detected, as described in Figure 13.
[0210]
[0211] Example 10. Verification of influence in size selection mode
[0212] As fetal nucleic acid fragments are typically known to be less than 200 bp, the same method of Examples 1, 2-2, and 8 was performed by selecting a nucleic acid fragment length of 150 bp or less, and the performance was confirmed in two T21 samples, one T18 sample, and one T13 sample.
[0213] As a result, as shown in Table 8, it was confirmed that the chromosome-specific integrated Z score increased in size selection mode.
[0214]
[0215] FF is the fetal fraction, calculated using the adjY (Adjusted Y-chromosome fraction) method described in Hudecova I. et al., PLoS One. 2014 Feb 28;9(2):e88484.
[0216] In addition, as described in Figure 14, after size selection, the negative and positive groups are separated much more clearly, confirming that not only is the motif-based signal maintained during size selection, but the signal-to-noise ratio is further improved due to the fetal signal enrichment effect.
[0217]
[0218] Foregoing, specific parts of the present invention have been described in detail. It will be apparent to those skilled in the art that such specific descriptions are merely preferred embodiments and do not limit the scope of the invention. Accordingly, the actual scope of the invention is defined by the appended claims and their equivalents.
Claims
A method for detecting chromosomal abnormalities comprising the following steps: (a) A step of extracting nucleic acids from a biological sample to obtain sequence information; (b) A step of aligning the acquired sequence information (reads) to a standard chromosome sequence database (reference genome database); (c) a step of calculating the frequency of terminal sequence motifs of nucleic acid fragments by chromosome using the above-mentioned aligned sequence information (reads); and (d) A step of determining chromosomal abnormalities by comparing the output value analyzed by inputting the frequency of terminal sequence motifs of the calculated nucleic acid fragment into a classification model with a cut-off value. A method for detecting chromosomal abnormalities according to claim 1, wherein the above step (a) is performed by a method comprising the following steps: (ai) A step of obtaining nucleic acids from a biological sample; (a-ii) a step of removing proteins, lipids, and other residues from the collected nucleic acid using a salting-out method, a column chromatography method, or a beads method to obtain purified nucleic acid; (a-iii) a step of amplifying a purified nucleic acid or a nucleic acid randomly fragmented by enzymatic cutting, grinding, or a hydroshear method using the primer set of claim 1 or 2, and then producing a single-end sequencing or pair-end sequencing library; (a-iv) a step of reacting the produced library with a next-generation sequencer; and (av) A step of obtaining sequence information (reads) of nucleic acids from a next-generation gene sequencing machine. A method according to claim 1, characterized in that the terminal sequence motif of step (c) is a pattern of 2 to 30 base sequences at both ends of the nucleic acid fragment. A method according to claim 1, wherein the terminal sequence motif frequency of step (c) is the number of each motif detected in the entire nucleic acid fragment. A method according to claim 1, characterized in that the classification model is a learned artificial intelligence model or a statistical analysis model. A method according to claim 5, wherein the artificial intelligence model is trained to distinguish between the frequency of terminal sequence motifs of nucleic acid fragments with normal chromosomal status and the frequency of terminal sequence motifs of nucleic acid fragments with chromosomal abnormalities. A method according to claim 6, wherein the artificial intelligence model is selected from the group consisting of a multi-layer perceptron (MLP), a convolutional neural network (CNN), a deep neural network (DNN), a recurrent neural network (RNN), a transformer, and an autoencoder. A method according to claim 7, wherein when the artificial intelligence model is an MLP, the loss function is expressed by the following Equation 1: A method according to claim 1, characterized in that the statistical analysis model comprises the following steps: (i) a step of converting the obtained chromosome-specific nucleic acid fragment end sequence motif frequencies into Z-scores using the mean and standard deviation of normal samples; and (ii) Step of calculating the combined Z-score for each chromosome using Equation 2 based on the transformed Z-scores for each chromosome: Here, w is a weight, a unit vector with all elements being 1, which can assign equal weight to all motif Z-scores, and Z is the Z-score of individual motifs per chromosome, and Σ (Sigma) is the covariance matrix between the chromosome-specific motif Z scores of normal individuals, and w T z is the sum of all Z scores, and sqrt(w T Σ w) is a normalization term that accounts for the correlated variability of the signals. A method according to claim 1, wherein the reference value of step (d) is 0.25 or 3, and if the reference value is greater than or equal to the reference value, it is determined that there is a chromosomal abnormality. A method according to claim 1, characterized in that the length of the nucleic acid fragment is 150 bp or less. A decoding unit that extracts nucleic acids from a biological sample and decodes sequence information; An alignment unit that aligns the decoded sequence to a standard chromosome sequence database; and A nucleic acid fragment analysis unit that calculates the frequency of terminal sequence motifs of aligned sequence-based nucleic acid fragments by chromosome; and A chromosomal abnormality detection device comprising a chromosomal abnormality determination unit that inputs the calculated terminal sequence motif frequency of nucleic acid fragments per chromosome into a classification model for analysis and determines the presence or absence of chromosomal abnormalities by comparing with a reference value. A computer-readable storage medium comprising instructions configured to be executed by a processor for detecting chromosomal abnormalities, (a) A step of extracting nucleic acids from a biological sample to obtain sequence information; (b) A step of aligning the acquired sequence information (reads) to a standard chromosome sequence database (reference genome database); (c) a step of calculating the frequency of terminal sequence motifs of nucleic acid fragments by chromosome using the above-mentioned aligned sequence information (reads); and (d) a step of determining chromosomal abnormalities by inputting the frequency of terminal sequence motifs of the calculated nucleic acid fragments into a classification model and comparing the analyzed output value with a cut-off value; A computer-readable storage medium comprising instructions configured to be executed by a processor that detects chromosomal abnormalities through.