NGS library quality evaluation system
The system addresses inefficiencies in NGS library quality evaluation by using big data and logistic regression to predict library success, enhancing experimental efficiency and reducing costs in genomic research.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- DCGEN CO LTD
- Filing Date
- 2025-10-21
- Publication Date
- 2026-04-30
AI Technical Summary
Existing methods for evaluating NGS library quality are inefficient as they only identify issues during the experimental process, making it difficult to predict the success or failure of NGS analysis, leading to significant costs and time consumption before obtaining experimental results.
A system utilizing big data to analyze QC data values and evaluate high-quality NGS libraries using logistic regression, incorporating steps for nucleic acid extraction, quality control, and post-processing to predict library success before NGS analysis.
Enables the pre-selection of high-quality NGS libraries, reducing unnecessary analysis, saving time and costs, and increasing the success rate of experiments, particularly in large-scale genomic research.
Smart Images

Figure KR2025016735_30042026_PF_FP_ABST
Abstract
Description
NGS library quality evaluation system
[0001] The present invention relates to a system for selecting high-quality NGS libraries suitable for NGS data production by synthesizing QC information generated before performing NGS analysis. More specifically, the present invention relates to a technique for evaluating and predicting DNA or RNA libraries produced using logistic regression.
[0002] The process of acquiring genetic information begins with the extraction of DNA from various samples, such as tissues, formalin-fixed paraffin-embedded (FFPE) tissues, cells, and liquid biopsies. The extracted nucleic acids are utilized to construct libraries, and three stages of quality assessment are performed during this process. The first Quality Control (QC) takes place during the gDNA / RNA extraction process, where DNA quality is evaluated using equipment such as a Qubit 4 fluorometer, Bioanalyzer, and TapeStation used in gel electrophoretic methods. This step confirms the quality of the extracted nucleic acids. The second QC is performed during the library construction stage. Library quality is verified by checking the concentration, size (length), and the presence of adapter dimer contamination. The third QC is performed after the hybridization process with probes that bind to the target sequences of the constructed library. In this stage, the quality of the library is evaluated by analyzing the concentration and size of the selectively isolated libraries. This process is performed to verify the characteristics of the library during binding with specific nucleotide sequences and hybridization, and to guarantee the final data quality. Once the library is finalized, the nucleotide sequence is analyzed using NGS. Subsequently, the generated data is analyzed once again to evaluate the success of the experiment and to assess the performance of the library evaluation system.
[0003] The quality of DNA and RNA extracted from various biological samples possesses unique characteristics, which are evaluated using metrics such as DIN, RIN, and DV200. Meanwhile, the quality of the generated library is verified through quality control (QC) regarding concentration, size, and dimers. Although these indicators vary depending on the biological sample, they have the disadvantage of only identifying issues during the experimental process and making it difficult to verify the NGS results.
[0004] In order to achieve the purpose of predicting or determining the success or failure of an experiment prior to performing NGS analysis, the present invention was developed by utilizing big data to analyze QC data values and evaluate a high-quality NGS library as shown in FIG. 1.
[0005] Data quality for NGS analysis cannot be determined in advance solely based on QC data obtained from the extraction and library construction processes. Even if a sample is deemed problem-free during extraction and library construction, there is a possibility that unexpected low-quality data may be generated during NGS analysis, and vice versa. The success of an experiment can only be confirmed after the NGS analysis has been performed. This approach is inefficient as it results in significant costs and time consumption before experimental results can be obtained.
[0006] The present invention was devised to solve these problems and aims to provide a system capable of evaluating the quality of NGS data in advance by pre-selecting high-quality data from patient samples.
[0007]
[0008] Various embodiments of the present invention are described with reference to the drawings. In the following description, for a complete understanding of the present invention, various specific details, such as specific forms, compositions, and processes, are described. However, specific embodiments may be practiced without one or more of these specific details, or in combination with other known methods and forms. In other examples, known processes and manufacturing techniques are not described as specific details so as not to make the present invention unnecessary or obscure. Reference throughout this specification to one embodiment implies that a particular feature, form, composition, or characteristic described in association with the embodiment is included in one or more embodiments of the present invention. Accordingly, the circumstances of an embodiment expressed at various locations throughout this specification do not necessarily represent the same embodiment of the present invention. Additionally, a particular feature, form, composition, or characteristic may be combined in any suitable way in one or more embodiments.
[0009] Unless otherwise specifically defined in this invention, all scientific and technical terms used in this specification have the same meaning as commonly understood by those skilled in the art to which this invention pertains.
[0010]
[0011] According to one embodiment of the present invention, the invention relates to a method for obtaining quality control data for selecting a library.
[0012] In the present invention, the method may include a step of extracting nucleic acid from a biological sample of a target individual; and a first quality control step of confirming the quality of the nucleic acid.
[0013] In the present invention, the first quality control step may be performed by quantifying the nucleic acid using a fluorescence detector and verifying the quality of the extracted nucleic acid by evaluating the size distribution and purity of the sample before sequencing using a gel electrophoresis device. Additionally, the first quality control step is characterized by being performed during the nucleic acid extraction process.
[0014] In the present invention, the term "target individual" refers to an individual that has developed a disease or is highly likely to develop a disease, and may be a mammal including humans, and may be selected from the group consisting of, for example, humans, rats, mice, guinea pigs, hamsters, rabbits, monkeys, dogs, cats, cattle, horses, pigs, sheep, and goats, and specifically may be humans, but is not limited thereto.
[0015] In the present invention, the term "biological sample" includes, but is not limited to, solid tissue samples, tissue culture media, liquid tissue samples, cells, or cell fragments. In addition, as non-limiting examples of biological samples, whole blood, leukocytes, peripheral blood mononuclear cells, buffy coat, plasma, serum, sputum, tears, mucus, nasal washes, nasal aspirate, breath, urine, semen, saliva, peritoneal washings, ascites, cystic fluid, meningeal fluid, amniotic fluid, glandular fluid, pancreatic fluid, lymph fluid, pleural fluid, nipple aspirate, bronchial aspirate, synovial fluid, and joint aspirate It may include one or more selected from the group consisting of aspirate, organ secretions, cells, cell extracts, and cerebrospinal fluid, and specifically, as long as it corresponds to a sample from which nucleic acids can be extracted, it is not limited thereto.
[0016] In the present invention, the term "nucleic acid" generally refers to a single polyribonucleotide or polydeoxyribonucleotide. This may be unmodified RNA or DNA, or modified RNA or DNA. The nucleic acid may be at least one of single-stranded and double-stranded DNA, DNA containing single-stranded and double-stranded regions, single-stranded and double-stranded RNA, RNA containing single-stranded and double-stranded regions, deoxyribonucleic acid (DNA), or ribonucleic acid (RNA), but is not limited thereto.
[0017] In the present invention, the term "Library" refers to a product created by cutting nucleic acids into appropriate sizes for NGS analysis and then attaching adapters with specific base sequences to both ends. Specifically, a library created before undergoing the target capture step is referred to as a "Pre-library," and a library after target capture is completed is referred to as a "Post-library."
[0018] In the present invention, the term "library size" refers to the average fragment length of DNA fragments within a DNA or cDNA library produced for next-generation sequencing (NGS), and is generally expressed in units of base pairs (bp).
[0019] The above library size can be utilized as a key variable for evaluating library quality as a QC metric through pre-library, post-library, and size change analysis.
[0020] Additionally, "Library size" may include the meaning of "Library length" and can be interpreted as having the same meaning depending on the context.
[0021] In the present invention, the term "Quality Control (QC)" refers to a process of analyzing various items regarding data used in an experiment to obtain accurate and highly reliable experimental results, and based on this, selectively excluding data that may lower the reliability of the experimental results. The quality control process of the present invention may be performed in a first quality control step, a post-processing step, and a second quality control step. The quality control process of the present invention may be performed in pre-NGS sequencing steps, such as a nucleic acid extraction step, a library preparation step, and a hybridization process with a probe that binds to the target base sequence of the library, thereby enabling the acquisition of quality control data for the library. Specifically, the quality control process of the present invention comprises: a quality control process for the library obtained from quantitative analysis results; a quality control process for the first quality control data obtained by identifying and analyzing the size distribution and purity by comparing the obtained library with a database; and a post-processing step for deriving values for one or more items selected from a group consisting of library concentration, library size, and adapter dimer contamination level. It may include a quality control process for second quality control data in which values for one or more items selected from a group consisting of selectively separated library concentrations or selectively separated library sizes are derived. As a non-limiting example, a primary quality control process may be performed on a library obtained through a quantitative analysis experiment, and if the library is determined to be suitable, a secondary quality control process may be performed through a post-processing process.
[0022]
[0023] In the present invention, the method may additionally include a post-processing step after the first quality control step.
[0024] In the present invention, the post-processing step may be a step of deriving a value for one or more items selected from a group consisting of library concentration, library size, and adapter dimer contamination level. Additionally, the post-processing step is characterized by being performed in the step of producing the library.
[0025] In the present invention, the term "post-processing" refers to one of the steps involved in obtaining the final quality control data of the present invention, and means a step of removing low-quality libraries or contaminated libraries from among the libraries obtained through a database search after the first quality control data is determined to be normal, prior to the second quality control process mentioned in the present invention.
[0026] In the present invention, the method may further include a step of removing a low-quality library or a contaminated library after the post-processing step; and a second quality control step of verifying the quality of nucleic acids from a selectively separated high-quality library.
[0027] In the present invention, the term "cut-off" refers to a reference value for determining whether to proceed with NGS analysis based on the DQ score calculated by the NGS library quality evaluation algorithm.
[0028] In the present invention, the term "low-quality library" refers to a library that is removed to obtain experimental results of the desired reliability when applying the quality control method of the present invention. In the above, low quality is determined by a score value calculated for each library, and the score value is calculated as a positive or negative real or rational value, and is classified as 'FALSE' or 'TRUE' based on a threshold value (cut-off) of 0.756, specifically based on a threshold value of 0.756145, and most specifically based on a threshold value of 0.7561453178422982. According to one embodiment of the present invention, if 'FALSE' is assigned based on the calculated score value, it can be defined as a "low-quality library," and such a low-quality library can be removed in the post-processing step.
[0029] In the present invention, the term "high-quality library" refers to a library that is selectively separated after nucleic acids are extracted from a biological sample of a target individual, and after low-quality libraries and contaminated libraries are removed through a first quality control step and a post-processing step. Such a library may be the subject from which values for concentration or size are derived in the second quality control step.
[0030] In the present invention, the term "contaminated library" refers to a library containing foreign nucleic acids (contamination) that are clearly not of human origin when compared with a database, or nucleic acids that do not exist in the database. The contaminated library may refer to a library identified as containing sequences of contaminants (mostly non-human derived proteins, such as viruses) that can be expressed in a sample during the experimental process, which are artificially inserted into the sequence database used in the identification process. All contaminated libraries are removed during the pretreatment step.
[0031] In the present invention, the second quality control step may be a step of deriving a value for one or more items selected from a group consisting of selectively separated library concentrations or selectively separated library sizes. Additionally, the second quality control step is characterized by being performed after a hybridization process with a probe that binds to the target base sequence of the library.
[0032] In the present invention, the method relates to a method for obtaining quality control data for the selection of a library, wherein the library may be a library used for nucleic acid sequencing, and nucleic acid sequencing may be used interchangeably with base sequencing, sequence analysis, or sequencing. The nucleic acid sequencing may include, for example, NGS-based targeted sequencing, targeted deep sequencing, or panel sequencing.
[0033] In the present invention, the library used in the method is characterized by being used for next-generation sequencing (NGS). Specifically, the library may be a library used for next-generation sequencing (NGS), but is not limited thereto, as long as it corresponds to a technique for simultaneously sequencing a large number of nucleic acid fragments, which involves fragmenting the whole genome into chip-based and polymerase chain reaction (PCR)-based paired-end formats and performing ultra-high-speed sequencing of the fragments based on hybridization.
[0034] In the present invention, the term "Next-generation sequencing (NGS)" refers to a method of dividing a genome into countless fragments, decoding the genetic information of each fragment, combining them, and then analyzing the entire nucleotide sequence. It has the advantage of being able to analyze the nucleotide sequence of a genome at high speed and is also referred to as High-throughput sequencing, Massive parallel sequencing, or Second-generation sequencing. Compared to NGS, although Sanger sequencing can also read the entire human genome, it is limited in its scope of testing because it can only target known genes, and repeated experiments are required to examine multiple genes. For example, compared to Sanger sequencing, which requires approximately 3 million separate tests, next-generation sequencing is an analysis method that offers significant improvements in terms of time and cost. Massive parallel sequencing, made possible by next-generation sequencing (NGS) technology, is another method for approaching the counting of RNA transcripts in tissue samples, and RNA sequencing is a method that utilizes this. It is currently the most powerful analytical tool used for transcriptome analysis, including differences in gene expression levels between different physiological conditions or changes occurring during development or disease progression. Specifically, RNA sequencing can be used to study phenomena such as changes in gene expression, selective splicing events, allele-specific gene expression and gene fusion, novel transcripts, and chimeric transcripts, including RNA editing.For example, 454 platform (Roche), GS FLX Titanium, Illumina MiSeq, Illumina HiSeq, Illumina HiSeq 2500, Illumina Genome Analyzer, Solexa platform, SOLiD System (Applied Biosystems), Ion Proton (Life Technologies), Complete Genomics, Helicos Biosciences Heliscope, and Pacific Biosciences' Single Molecule Real-Time (SMRT). TM It may be performed by ) technology, or a combination thereof.
[0035]
[0036] According to another embodiment of the present invention, the invention relates to a method for evaluating a library using obtained quality control data.
[0037] In the present invention, the quality control data may be obtained during a quality control process for selecting a library.
[0038] In the present invention, the quality control process for library selection may include: a step of extracting nucleic acids from a biological sample of a target individual; a first quality control step of verifying the quality of the nucleic acids; a post-processing step; a step of removing low-quality libraries or contaminated libraries; and a second quality control step of verifying the quality of nucleic acids from a selectively separated high-quality library.
[0039] In the present invention, the method for evaluating the library may include the step of deriving a DCGen Quality score (DQ score) by applying it to the following Equation 1:
[0040] [Equation 1]
[0041]
[0042] In the above Equation 1,
[0043] The above a may be 0.5 or more and 2 or less, 0.6 or more and 1.9 or less, 0.7 or more and 1.8 or less, 0.8 or more and 1.7 or less, 0.9 or more and 1.6 or less, 1.0 or more and 1.5 or less, 1.01 or more and 1.4 or less, 1.02 or more and 1.3 or less, 1.1 or more and 1.2 or less, and specifically, 1.106.
[0044] The above b may be 0.1 or more and 0.5 or less, 0.15 or more and 0.4 or less, 0.2 or more and 0.3 or less, 0.25 or more and 0.3 or less, and specifically, 0.289.
[0045] The above c may be 0.5 or more and 1.5 or less, 0.6 or more and 1.4 or less, 0.7 or more and 1.3 or less, 0.8 or more and 1.0 or less, 0.83 or more and 0.87 or less, and specifically, 0.854.
[0046] The above d may be 1 or more and 2 or less, 1.1 or more and 1.9 or less, 1.2 or more and 1.8 or less, 1.3 or more and 1.7 or less, 1.4 or more and 1.6 or less, 1.41 or more and 1.5 or less, 1.42 or more and 1.49 or less, and specifically, 1.46.
[0047] In addition, the method for evaluating the library in the present invention may include the step of deriving a DCGen Quality score (DQ score) by applying it to the following Equation 2:
[0048] [Equation 2]
[0049]
[0050] In the above Equation 1 or Equation 2,
[0051] The above score1 is the result of evaluating the post-library size value, assigning '0' if it is less than 365, '1' if it is 365 or greater and 395 or less, and '2' if it exceeds 395.
[0052] The above score2 is the result of evaluating the size change value calculated by the difference between the pre-library size and the post-library size, and represents the value obtained by subtracting the pre-library size from the post-library size, and the above value is an integer. If the value is 1 or greater and 10 or less, '1' is assigned, and otherwise '0' is assigned.
[0053] The above score3 is an evaluation of the post-library concentration value, assigning '1' if it is 1 or more and 5 or less, and '0' otherwise.
[0054] In the present invention, the above-mentioned Equation 1 or Equation 2 may be a formula derived through the application of an algorithm from an embodiment of the present invention, and the DQ score value may be derived by subtracting the value of score2 multiplied by 0.289 from the value of score1 multiplied by 1.106, and adding the value of score3 multiplied by 0.854 and 1.46. Specifically, the DQ score value may be derived by subtracting the value of score2 multiplied by 0.2893 from the value of score1 multiplied by 1.1057, and adding the value of score3 multiplied by 0.8535 and 1.4595. Most specifically, the DQ score value may be derived by subtracting the value of score2 multiplied by 0.289318 from the value of score1 multiplied by 1.105715, and adding the value of score3 multiplied by 0.853536 and 1.45953209. In one embodiment of the present invention, if the DQ score value derived from Equation 1 or Equation 2 is 0.756 or less, which is the threshold value (cut-off), it is classified as 'FALSE', and it can be determined that library remanufacturing is required because the probability of failure during NGS analysis is high. Specifically, if the DQ score value derived from Equation 1 or Equation 2 is 0.7561453 or less, it is classified as 'FALSE', and it can be determined that library remanufacturing is required because the probability of failure during NGS analysis is high. Most specifically, if the DQ score value derived from Equation 1 or Equation 2 is 0.7561453178422982 or less, it is classified as 'FALSE', and it can be determined that library remanufacturing is required because the probability of failure during NGS analysis is high.
[0055] In another embodiment of the present invention, if the DQ score value derived from Equation 1 or Equation 2 exceeds the threshold value (cut-off) of 0.756, it is classified as 'TRUE', and it can be determined that it is a high-quality library with a high probability of success during NGS analysis. Specifically, if the DQ score value derived from Equation 1 or Equation 2 exceeds 0.7561453, it is classified as 'TRUE', and it can be determined that it is a high-quality library with a high probability of success during NGS analysis. Most specifically, if the DQ score value derived from Equation 1 or Equation 2 exceeds 0.7561453178422982, it is classified as 'TRUE', and it can be determined that it is a high-quality library with a high probability of success during NGS analysis.
[0056] In the present invention, the term "evaluation of a library" refers to the act of evaluating the quality of a library used for NGS analysis in advance to predict the results of the NGS analysis. More specifically, the evaluation of a library may be performed using indicators such as the concentration, integrity, and purity of the library; since the results of next-generation sequencing analysis may vary depending on the quality of the library, this can be interpreted to mean all acts of predicting such next-generation sequencing analysis results in advance.
[0057] In the present invention, the term "probability of failure / probability of success" may be determined by whether the results of next-generation sequencing analysis are reliable. A high probability of success specifically means that the reliability of the analysis results is 70% or higher, more specifically means that the reliability of the analysis results is 80% or higher, and most specifically means that the reliability of the analysis results is 90% or higher. The probability of failure may be interpreted as the opposite concept.
[0058] When using the above method of the present invention, all quality control data generated from gDNA / RNA to DNA library production are aggregated before NGS analysis is performed, thereby enabling the selection of high-quality NGS libraries suitable for NGS data production. Accordingly, by evaluating library quality in advance, the cost and time required for NGS experiments can be efficiently managed while increasing the success rate of the experiments, and it is expected that this can be usefully applied to large-scale genomic research requiring high-throughput analysis.
[0059]
[0060] In a first aspect of the present invention, a method for obtaining quality control data for selecting a library is provided, comprising: a step of extracting nucleic acid from a biological sample of a target individual; a first quality control step of confirming the quality of the nucleic acid; and a post-processing step after the first quality control step; wherein the first quality control step is a step of identifying and analyzing size distribution and purity by comparing a library with a database, and the post-processing step is a step of deriving values for one or more items selected from a group consisting of library concentration, library size, and adapter dimer contamination.
[0061]
[0062] In the first embodiment above, the second embodiment provides a method in which the nucleic acid is at least one of deoxyribonucleic acid (DNA) or ribonucleic acid (RNA).
[0063] In the first or second embodiment above, the third embodiment comprises whole blood, leukocytes, peripheral blood mononuclear cells, buffy coat, plasma, serum, sputum, tears, mucus, nasal washes, nasal aspirate, breath, urine, semen, saliva, peritoneal washings, ascites, cystic fluid, meningeal fluid, amniotic fluid, glandular fluid, pancreatic fluid, lymph fluid, pleural fluid, nipple aspirate, bronchial aspirate, The present invention provides a method comprising one or more selected from the group consisting of synovial fluid, joint aspirate, cerebrospinal fluid, organ secretions, cell, cell extract, and tissue.
[0064] In any one of the first to third embodiments above, the fourth embodiment provides a method in which the first quality control step quantifies the nucleic acid using a fluorescence detector and verifies the quality of the extracted nucleic acid by evaluating the size distribution and purity of the sample before sequencing using a gel electrophoresis device.
[0065] In any one of the first to fourth embodiments above, the fifth embodiment provides a method characterized in that the first quality control step is performed during the nucleic acid extraction process, and in any one of the first to fifth embodiments above, the sixth embodiment provides a method characterized in that the post-processing step is performed during the step of producing a library.
[0066] In any one of the first to sixth embodiments above, the seventh embodiment further comprises: a step of removing a low-quality library or a contaminated library after the post-processing step; and a second quality control step of verifying the quality of nucleic acids from a selectively separated high-quality library.
[0067] In any one of the first to seventh embodiments above, the eighth embodiment provides a method in which the second quality control step is a step of deriving a value for one or more items selected from a group consisting of selectively separated library concentrations or selectively separated library lengths.
[0068] In any one of the first to eighth embodiments above, the ninth embodiment provides a method characterized in that the second quality control step is performed after a hybridization process with a probe that binds to a target sequence of the library.
[0069] In any one of the first to ninth embodiments above, the tenth embodiment provides a method characterized in that the library is used for next-generation sequencing (NGS).
[0070] The eleventh embodiment provides a method for evaluating a library, comprising the step of applying quality control data obtained by the method of any one of the first to ten embodiments to the following Equation 1.
[0071] [Equation 1]
[0072]
[0073] In the above Equation 1, a is 0.5 or more and 2 or less, b is 0.1 or more and 0.5 or less, c is 0.5 or more and 1.5 or less, d is 1 or more and 2 or less, score1 is an evaluation of the post-library size value, score2 is an evaluation of the size change value obtained by subtracting the pre-library size from the post-library size, and score3 is an evaluation of the post-library concentration value.
[0074] In any one of the first to eleven embodiments above, the twelfth embodiment provides a method in which '0' is assigned to score 1 if it is less than 365, '1' if it is 365 or more and 395 or less, and '2' if it exceeds 395.
[0075] In any one of the first to twelfth embodiments above, the twelfth embodiment provides a method in which the score 2 is assigned '1' when it is 1 or more and 10 or less, and '0' otherwise.
[0076] In any one of the first to thirteenth embodiments above, the thirteenth embodiment provides a method in which the score 3 is assigned '1' when it is 1 or more and 5 or less, and '0' otherwise.
[0077] In any one of the first to fourth embodiments above, the fifth embodiment provides a method for determining that if the value derived from Equation 1 is below a threshold value (cut-off), it is classified as 'FALSE' and that library remanufacturing is required because the probability of failure during NGS analysis is high. Specifically, the value derived from Equation 1 is a DQ score value, and if the DQ score value is below the threshold value of 0.756, it is classified as 'FALSE' and that library remanufacturing is required because the probability of failure during NGS analysis is high. More specifically, if the DQ score value derived from Equation 1 is below 0.7561453, it is classified as 'FALSE' and that library remanufacturing is required because the probability of failure during NGS analysis is high. Most specifically, if the DQ score value derived from Equation 1 above is 0.7561453178422982 or less, it is classified as 'FALSE', and it can be determined that library reconstruction is required because the probability of failure during NGS analysis is high.
[0078] In any one of the first to fifteen embodiments above, the 16th embodiment provides a method for determining that a high-quality library is classified as 'TRUE' when the value derived from Equation 1 exceeds a cut-off value, and that the probability of success in NGS analysis is high. Specifically, the value derived from Equation 1 is a DQ score value, and when the DQ score value exceeds the cut-off value of 0.756, it is classified as 'TRUE', and it can be determined that a high-quality library is classified as having a high probability of success in NGS analysis. More specifically, when the DQ score value derived from Equation 1 exceeds 0.7561453, it is classified as 'TRUE', and it can be determined that a high-quality library is classified as having a high probability of success in NGS analysis. Most specifically, when the DQ score value derived from Equation 1 exceeds 0.7561453178422982, it is classified as 'TRUE', and it can be determined that a high-quality library is classified as having a high probability of success in NGS analysis.
[0079] In any one of the first to sixth embodiments above, the seventh embodiment provides a library evaluation device comprising: an extraction unit for extracting nucleic acids from a biological sample of a target individual; a first quality control unit for identifying and analyzing the size distribution and purity of the extracted nucleic acids by comparing them with a database; and a post-processing unit for deriving values for one or more items selected from a group consisting of library concentration, library length, and adapter dimer contamination.
[0080] In the present invention, the term "evaluation device" refers to a device equipped with the function of obtaining, analyzing, and evaluating quality control (QC) data for nucleic acids extracted from biological samples and libraries produced based thereon, comprising: an extraction unit that extracts nucleic acids from a target sample; a first quality control unit that identifies and analyzes the size distribution and purity of the extracted nucleic acids by comparing them with a database; a post-processing unit that calculates indicators such as library concentration, library length, and adapter dimer contamination; a removal unit that removes low-quality or contaminated libraries; a second quality control unit that performs additional quality verification on selected high-quality libraries; and an evaluation unit that combines the QC data, applies them to a DQ score formula, and determines the quality level of the library (failure probability, usability, and high quality), or is composed of a combination thereof. That is, it refers to a device that performs multi-stage quality verification (QC) and post-processing analysis on nucleic acids extracted from samples and libraries produced, calculates quantitative evaluation indicators (DQ scores), and determines the success probability of next-generation sequencing (NGS).
[0081] In any one of the first to seventh embodiments above, the eighth embodiment provides a library evaluation device further comprising: a removal unit for removing a low-quality library or a contaminated library after the post-processing unit; and a second quality control unit for verifying the quality of nucleic acids from a selectively separated high-quality library; wherein the second quality control step derives a value for one or more items selected from a group consisting of a selectively separated library concentration or a selectively separated library size.
[0082] In any one of the first to eighteen embodiments above, the 19th embodiment provides a library evaluation device that further includes an evaluation unit that applies quality control data obtained from a first quality control unit or a second quality control unit to the following [Equation 1].
[0083] [Equation 1]
[0084]
[0085] In the above Equation 1, a is 0.5 or more and 2 or less, b is 0.1 or more and 0.5 or less, c is 0.5 or more and 1.5 or less, d is 1 or more and 2 or less, score1 is an evaluation of the post-library size value, score2 is an evaluation of the size change value obtained by subtracting the pre-library size from the post-library size, and score3 is an evaluation of the post-library concentration value.
[0086] In any one of the above 1st to 19th embodiments, the 20th embodiment provides a library evaluation device in which the score 1 is assigned '0' when less than 365, '1' when 365 or more and 395 or less, and '2' when exceeding 395.
[0087] In any one of the first to 20 embodiments above, the 21st embodiment provides a library evaluation device in which the score 2 is assigned '1' when it is 1 or more and 10 or less, and '0' otherwise.
[0088] In any one of the above 1 to 21 embodiments, the 22nd embodiment provides a library evaluation device in which the score 3 is assigned '1' when it is 1 or more and 5 or less, and '0' otherwise.
[0089] In any one of the above 1 to 22 embodiments, the 23rd embodiment provides a library evaluation device that determines that if the value derived from Equation 1 is below a threshold (cut-off), it is classified as 'FALSE' and that library remanufacturing is required because the probability of failure during NGS analysis is high.
[0090] In any one of the first to 23 embodiments above, the 24th embodiment provides a library evaluation device that is classified as 'TRUE' when the value derived from Equation 1 exceeds a threshold (cut-off), and determines that it is a high-quality library with a high probability of success during NGS analysis.
[0091]
[0092] According to the present invention, by evaluating the quality of samples using QC data produced during the library production process, it is possible to pre-selectively classify library samples most suitable for NGS analysis. This allows for the production of high-quality NGS data while simultaneously reducing unnecessary NGS analysis and rapidly re-experimenting, thereby saving time and costs. It is expected that this will be useful for large-scale genomic research requiring high-throughput analysis.
[0093] Furthermore, the effects of the present invention are not limited to the effects described above, and should be understood to include all effects that can be inferred from the configuration of the invention described in the detailed description or claims of the present invention.
[0094]
[0095] Figure 1 is a diagram illustrating the library evaluation model development process according to the present invention.
[0096] Figure 2 is a schematic diagram of the process of QC data generation and NGS analysis progress determination according to the present invention.
[0097] Figure 3 is a figure showing the NGS data QC pass prediction and true value using the prediction method according to the present invention.
[0098] Figure 4 is a figure showing the NGS data QC pass prediction and true value using the prediction method of the present invention according to the present invention.
[0099] Figure 5 is the result of a performance analysis of NGS progress judgment for 54 samples according to the present invention.
[0100]
[0101] The present invention will be described in more detail below through examples. These examples are intended solely to explain the present invention more specifically, and it will be obvious to those skilled in the art that the scope of the present invention is not limited by these examples according to the gist of the invention.
[0102]
[0103] Examples
[0104] Although it is possible to evaluate anomalies occurring at each stage of the library production process by analyzing Quality Control (QC) data, there are no clear standards or guidelines for this. Furthermore, since there are no interpretation criteria for anomalies even if they occur, there is a problem in that NGS analysis proceeds entirely based on the researcher's judgment. As a result of diligent efforts to develop a method to provide more objective indicators, the inventors have derived a system that evaluates high-quality NGS libraries based on QC data of prepared libraries.
[0105]
[0106] Nucleic acid extraction from the sample
[0107] DNA and RNA extraction is performed using the automated Maxwell equipment. ® CSC 48 instrument IVD was used. DNA extraction was performed using Maxwell. ® DNA was extracted using the CSC Blood DNA Kit, and RNA was extracted using the Maxwell® CSC RNA Blood Kit. The experiment was performed according to the Promega manual.
[0108] Quality measurements after DNA and RNA extraction were performed by measuring gDNA and total RNA concentrations using an Invitrogen Qubit™ 4 Fluorometer, measuring integrity using an Agilent 4200 TapeStation, and measuring purity using a ThermoScientific NANODROP 8000 Spectrophotometer.
[0109]
[0110] Library Synthesis Kit and Method
[0111] An NGS library was synthesized using the xGen™ DNA Library Prep EZ Kit with the extracted total DNA. Additionally, the extracted total RNA was used to synthesize cDNA using the xGen™ RNA Library Prep Kit, and the process of synthesizing the cDNA library was carried out. During this process, the library synthesis followed the manufacturer's manual. The concentration of the synthesized pre-library was measured using an Invitrogen Qubit™ 4 Fluorometer, and the library size and concentration were verified using an Agilent 4200 TapeStation.
[0112]
[0113] Quality measurement of Total DNA and Total RNA
[0114] The indicators for evaluating the quality of total DNA and total RNA are different and are measured using the following three instruments. For total DNA, concentration was measured using an Invitrogen Qubit™ 4 Fluorometer, and integrity was measured using an Agilent 4200 TapeStation to determine the DNA Integrity Number (DIN). Purity was measured using a ThermoScientific NANODROP 8000 Spectrophotometer to determine absorbance, with A260 / 280 values set to 1.8 or higher and A260 / 230 values set to 2.0 to 2.2. For Total RNA, concentration was measured using an Invitrogen Qubit™ 4 Fluorometer. For integrity, only samples with an RNA Integrity Number (RIN) and a DV200 of 40% or higher were evaluated using an Automated Electrophoresis System (e.g., Agilent 4200 TapeStation or equivalent equipment) and this reagent was applied. Purity was measured by absorbance using a Spectrophotometer (e.g., ThermoScientific 'NANODROP 8000 Spectrophotometer' or equivalent equipment), with A260 / 280 values set at 1.8 or higher and A260 / 230 values between 2.0 and 2.2. The operation of these instruments was performed according to the manual.
[0115]
[0116] Pre-library quality measurement
[0117] The quality of each pre-library generated from DNA and RNA was measured using the same method. Pre-library concentration was measured using an Invitrogen Qubit™ 4 Fluorometer, and pre-library quality was measured using an Agilent 4200 TapeStation. The following results were verified using the software from the Automated Electrophoresis System manufacturer. 1) Electrophoresis diagram peaks: Three peaks were identified: the lower marker (25 bp), the NGS library main peak between 200 and 350 bp, and the upper marker (1500 bp). 2) Region selection and quantification: After designating the region between 180 bp and 700 bp, the concentration (ng / μL) value was read (Quantitative Precision: 15% CV). The following steps were performed based on the value obtained from the region.
[0118]
[0119] Post-library quality measurement
[0120] Measure the concentration and quality of the amplified captured DNA library using the Agilent 4200 TapeStation. Operation of the instrument follows the equipment manufacturer's manual.
[0121] Check the following results in the software of the Automated Electrophoresis System manufacturer. Electrophoresis diagram: Displays the purified amplified captured library in the interval (around 180 bp to 700 bp) for QC verification based on the standard marker (Low marker, 25 bp and Upper marker, 1500 bp). The concentration of target-captured library fragments in this interval is measured based on the standard marker (Lower marker, 25 bp and Upper marker, 1500 bp), and their sum is expressed as the 'concentration' of the final post-library and can be verified in ng / μL units. Similarly, the length of target-captured library fragments in the set interval (around 180 bp to 700 bp) is expressed as the 'average size' and can be verified in bp units.
[0122]
[0123] Sequencing Library Pooling
[0124] It is important to assign a unique index to each sample. Check to ensure that the indices of the samples to be pooled do not overlap. When pooling, consider the sample concentration, the amount of data required, the target size, etc., and perform pooling according to the NGS protocol.
[0125]
[0126] Sequencing
[0127] The final product from the above process is sequenced using Illumina’s NextSeq 550DX or MiseqDx - Illumina, Inc.
[0128]
[0129] Sequencing data analysis
[0130] DNA sequencing reads were aligned to the hg38 reference genome using BWA-0.7.17. Subsequently, Samtools was utilized to analyze sequencing coverage. Regions presumed to be putative duplications were marked using the GATK-4.1.8.0 module. Locations expected to contain small indels (insertions or deletions) were reordered using the GATK-4.1.8.0 module. For known variant locations, information from the 1000 Genomes Project Phase I, dbSNP-151, Mills, and 1000G gold standard indels was utilized.
[0131] The values for pre-library size, post-library size, size change, and post-library concentration were extracted from the QC database collected from the Qubit and agilent 4200 Tapestation of the extracted samples.
[0132] Data scaling was performed as a preprocessing step prior to building the prediction model. After excluding samples with missing values, it was confirmed that there were 177 usable data points. The post-library size was used to derive the variable 'score_1'. '0' is assigned if the value is less than 365, '1' if it is 365 or greater but 395 or less, and '2' if it exceeds 395. To derive the variable 'score_2', the size change value was calculated by subtracting the pre-library size from the post-library size. '1' is assigned if the size change value is between 1 and 10, and '0' otherwise. To derive the value of the variable 'score_3', the post-library concentration value was used to assign '1' if the value was 1 or greater and 5 or less, and '0' otherwise. For the selected variables 'score_1', 'score_2', and 'score_3' (hereinafter 'variables'), which were identified by selecting factors related to NGS success, StandardScaler was applied to scale each dataset. Through the above process, the mean of each variable was adjusted to 0 and the standard deviation to 1, thereby using normalized data values.
[0133] To develop a classification prediction model, Logistic Regression, one of the Supervised Learning methods, was applied. Before developing the model, the data was randomly split into a training set of 70% and a test set of 30%, and the performance of the prediction model was verified using the test set.
[0134] Logistic regression analysis is used when the dependent variable is categorical, and this model represents the probability that a predicted value obtained through regression analysis belongs to a specific category. Therefore, a cut-off threshold can be set to distinguish probability values as TRUE or FALSE.
[0135] [Equation 2]
[0136]
[0137] The above score1 is the result of evaluating the post-library size value, assigning '0' if it is less than 365, '1' if it is 365 or more and 395 or less, and '2' if it exceeds 395; the above score2 is the result of evaluating the size change value calculated by the difference between the pre-library size and the post-library size, assigning '1' if it is 1 or more and 10 or less, and '0' otherwise; and the above score3 is the result of evaluating the post-library concentration value, assigning '1' if it is 1 or more and 5 or less, and '0' otherwise.
[0138] The regression coefficients in Equation 2 above represent the influence of each variable on the prediction model. The larger the absolute value of the above values, the greater the influence the corresponding variable has on the judgment decision, and the smaller the absolute value, the relatively lower the influence of the corresponding variable. The inventors derived Equation 2 above by creating a model that finds the most suitable regression coefficients through logistic regression, and in the present invention, to optimize the accuracy of the prediction model, a process of finding an appropriate cut-off value where sensitivity and specificity are maximized was additionally performed.
[0139]
[0140] Experimental results
[0141] Table 1 below shows the quality control items generated during the process from DNA / RNA extraction from patient samples to library construction and NGS analysis.
[0142] Process QC Method QC Data Collection 1. DNA / RNA Extraction Qubit DNA Concentration Tape Station Size / DNA Concentration Nanodrop A 260 / 280 ratio, A 260 / 230 ratio 2. Library Synthesis (Pre-Library) Qubit Library Concentration Tape Station DNA Size / Library Concentration 3. Target Capture Step (Post-Library) Qubit Library Concentration Tape Station DNA Size / Library Concentration 4. NGS Results NGS Analysis Report On target read / target coverage
[0143]
[0144]
[0145] Among the quality control data above, pre-library size, post-library size, size change, and post-library concentration were used, and 'FALSE' or 'TRUE' was assigned through the optimal algorithm. Accordingly, whether to proceed with NGS analysis is distinguished as shown in Table 2 below.
[0146]
[0147] Predict results based on DQ score Decision on whether to proceed with NGS analysis FALSE: Re-experiment recommended TRUENGS: Proceed with analysis recommended
[0148]
[0149] As shown in Figure 2, research guidelines are established based on the score calculated through an algorithm utilizing various variables. As shown in Table 2 above, for samples receiving a 'FALSE' rating, researchers are advised to carefully consider whether to proceed with NGS analysis, taking into account the high probability of failure. Alternatively, the samples may be excluded from NGS analysis, and researchers are encouraged to re-prepare the library. For samples receiving a 'TRUE' rating, researchers are advised to proceed with NGS analysis, as the probability of failure is expected to be low and there is a high likelihood of generating high-quality NGS data. These guidelines save unnecessary costs and time for NGS analysis and enable efficient research through rapid response.
[0150]
[0151] Performance verification of NGS progress determination algorithm
[0152] The inventors developed an NGS library quality evaluation algorithm device, and to verify the performance of the NGS library evaluation algorithm device developed in this invention, a total of 177 nucleic acid samples were used in the experiment. The results of predicting the success or failure of NGS analysis using the quality evaluation device of this invention are shown in Figure 3. This shows the results for 123 samples used for model development out of the 177 samples. The model was derived with a cut-off value of 0.7561453178422982 through the algorithm, and the predicted NGS analysis values were classified as shown in Table 3 below by setting the above cut-off value.
[0153] DQ Score Cut-off Criteria Result Prediction DQ score ≤ 0.7561453178422982 FALSEDQ score > 0.7561453178422982 TRUE
[0154]
[0155] Samples requiring re-experimentation were marked as 'FALSE', while conversely, samples whose results passed the NGS data quality control test (passing both on target read and target coverage) were classified as 'TRUE', and samples that failed the NGS data quality control test (QC) were classified as 'FALSE'. As a result of the analysis in Figure 3, it was confirmed that the performance of the predicted values was 96% sensitive and 61% specific (see Figure 3).
[0156] Furthermore, among the 177 extracted samples, the NGS analysis success rate was to be predicted for 54 samples that were not used during algorithm development using the NGS analysis judgment algorithm device developed in the present invention. Figure 4 shows the results of predicting whether NGS analysis was successful using the quality evaluation device of the present invention, and is the result of classifying the predicted NGS analysis values by applying the previously derived threshold value (cut-off) 0.7561453178422982 (see Figure 4).
[0157] In the present invention, a process was performed to find an appropriate cut-off value that maximizes sensitivity and specificity in order to optimize the accuracy of the prediction model. As a result, it was confirmed that the optimal cut-off value that maximizes sensitivity and specificity is 0.7561453178422982. Table 4 below shows the results of a performance evaluation using 54 samples through the prediction algorithm system of the present invention. By confirming that each performance according to the cut-off value has a sensitivity of 88%, a specificity of 79%, a precision of 92%, and an accuracy of 85%, it is expected that the efficiency of NGS analysis can be improved by undergoing a library selection process suitable for NGS analysis through the simple application of a formula.
[0158] Item Percentage Sensitivity 88% Specificity 79% Precision 92% Accuracy 85%
[0159]
[0160] Foregoing, specific parts of the present invention have been described in detail. It is evident to those skilled in the art that such specific descriptions are merely preferred embodiments and do not limit the scope of the invention. Accordingly, the actual scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A step of extracting nucleic acids from a biological sample of a target organism; A first quality control step to verify the quality of nucleic acids; and It includes a post-processing step after the first quality control step mentioned above, and The above first quality control step is a step of obtaining size distribution and purity by comparing the library and the database to identify and analyze them, and A method for obtaining quality control data for library selection, wherein the above-mentioned post-processing step is a step of deriving values for one or more items selected from a group consisting of library concentration, library size, and adapter dimer contamination level.
2. In Paragraph 1, A method in which the nucleic acid is at least one of deoxyribonucleic acid (DNA) or ribonucleic acid (RNA).
3. In Paragraph 1, The above biological samples include whole blood, leukocytes, peripheral blood mononuclear cells, buffy coat, plasma, serum, sputum, tears, mucus, nasal washes, nasal aspirate, breath, urine, semen, saliva, peritoneal washings, ascites, cystic fluid, meningeal fluid, amniotic fluid, glandular fluid, pancreatic fluid, lymph fluid, pleural fluid, nipple aspirate, bronchial aspirate, synovial fluid, joint aspirate, A method comprising one or more selected from the group consisting of cerebrospinal fluid, organ secretions, cells, cell extracts, and tissues.
4. In Paragraph 1, The above first quality control step is a method of quantifying the nucleic acid using a fluorescence detector and verifying the quality of the extracted nucleic acid by evaluating the size distribution and purity of the sample before sequencing using a gel electrophoresis device.
5. In Paragraph 1, A method characterized in that the above-mentioned first quality control step is performed during the nucleic acid extraction process.
6. In Paragraph 1, A method characterized in that the above post-processing step is performed during the step of creating a library.
7. In Paragraph 1, A method further comprising: a step of removing a low-quality library or a contaminated library after the above post-processing step; and a second quality control step of verifying the quality of nucleic acids from a selectively separated high-quality library.
8. In Paragraph 7, The above second quality control step is a method for deriving a value for one or more items selected from a group consisting of selectively separated library concentrations or selectively separated library sizes.
9. In Paragraph 7, A method characterized in that the second quality control step is performed after a hybridization process with a probe that binds to the target base sequence of the library.
10. In Paragraph 1, A method characterized by the above library being used for next-generation sequencing (NGS).
11. A step of applying quality control data obtained by the method of any one of claims 1 to 10 to the following Formula 1, and [Equation 1] In the above Equation 1, a is 0.5 or more and 2 or less, b is 0.1 or more and 0.5 or less, c is 0.5 or more and 1.5 or less, and d is 1 or more and 2 or less, and A method for evaluating a library, wherein score1 is an evaluation of the post-library size value, score2 is an evaluation of the size change value obtained by subtracting the pre-library size from the post-library size, and score3 is an evaluation of the post-library concentration value.
12. In Paragraph 11, A method in which the above score 1 is assigned '0' if less than 365, '1' if 365 or more and 395 or less, and '2' if exceeding 395.
13. In Paragraph 11, A method in which the above score 2 is assigned '1' when it is 1 or more and 10 or less, and '0' otherwise.
14. In Paragraph 11, A method in which the above score 3 is assigned '1' when it is 1 or greater and 5 or less, and '0' otherwise.
15. In Paragraph 11, A method in which, if the value derived from the above Equation 1 is below the threshold value (cut-off), it is classified as 'FALSE', and a determination is made that library re-production is required because the probability of failure during NGS analysis is high.
16. In Paragraph 11, A method in which the value derived from the above Equation 1 is classified as 'TRUE' when it exceeds a threshold value (cut-off), and the probability of success during NGS analysis is high, thereby determining that it is a high-quality library.
17. An extraction unit for extracting nucleic acids from a biological sample of a target organism; A first quality control unit that identifies and analyzes the size distribution and purity of extracted nucleic acids by comparing them with a database; and A library evaluation device comprising: a post-processing unit that derives a value for one or more items selected from a group consisting of library concentration, library size, and adapter dimer contamination level.
18. In Paragraph 17, The above evaluation device includes a removal unit that removes low-quality libraries or contaminated libraries after the above post-processing unit; and It further includes a second quality control unit that verifies the quality of nucleic acids from a selectively separated high-quality library, and A library evaluation device in which the second quality control step derives a value for one or more items selected from a group consisting of selectively separated library concentrations or selectively separated library sizes.
19. In Paragraph 17, The above evaluation device further includes an evaluation unit that applies quality control data obtained from a first quality control unit or a second quality control unit to the following [Equation 1], and [Equation 1] In the above Equation 1, a is 0.5 or more and 2 or less, b is 0.1 or more and 0.5 or less, c is 0.5 or more and 1.5 or less, and d is 1 or more and 2 or less, and A library evaluation device, wherein score1 is an evaluation of the post-library size value, score2 is an evaluation of the size change value obtained by subtracting the pre-library size from the post-library size, and score3 is an evaluation of the post-library concentration value.
20. In Paragraph 19, A library evaluation device in which the above score 1 is assigned '0' if less than 365, '1' if 365 or more and 395 or less, and '2' if exceeding 395.
21. In Paragraph 19, A library evaluation device in which the above score 2 assigns '1' when the value is 1 or greater and 10 or less, and '0' otherwise.
22. In Paragraph 19, A library evaluation device in which the above score 3 assigns '1' when the value is 1 or greater and 5 or less, and '0' otherwise.
23. In Paragraph 19, A library evaluation device that classifies the value derived from the above Equation 1 as 'FALSE' when it is below a threshold value (cut-off), and determines that library reconstruction is required because the probability of failure during NGS analysis is high.
24. In Paragraph 19, A library evaluation device that classifies the value derived from the above Equation 1 as 'TRUE' when it exceeds a threshold value (cut-off), and determines that it is a high-quality library with a high probability of success during NGS analysis.