Nail-based NGS data analysis method and system for disease prediction

The nail-based NGS method addresses the inefficiencies and challenges of traditional genetic analysis by using nail samples and forensic DNA extraction, enabling accurate and cost-effective prediction and diagnosis of chronic diseases and cancers.

WO2025095682A1PCT designated stage expired Publication Date: 2025-05-08LEE EOJIN
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/017050
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-03
Filing Date
2024-11-01
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

Current methods for genetic analysis, such as SANGER-based sequence analysis, are inefficient and costly due to their limitations in analyzing long DNA sequences, and they face challenges in accurately detecting mutations, especially in samples like blood and tissue that are difficult to collect and transport.

Method used

A nail-based Next-Generation Sequencing (NGS) method and system that separates Genomic DNA (GDNA) from nail samples using a forensic DNA extraction method, avoiding phenol/chloroform, and then analyzes it using NGS technology to predict or diagnose chronic diseases and cancers.

Benefits of technology

This method allows for high-accuracy, cost-effective, and efficient analysis of genetic data from nail samples, which can be easily collected and transported, overcoming the limitations of traditional methods and improving the detection of mutations and prediction of diseases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024017050_08052025_PF_FP_ABST
    Figure KR2024017050_08052025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to a next generation sequencing (NGS) method for disease prediction through analysis of nail-derived germline mutation and somatic mutation, and a system therefor. Specifically, the present invention relates to a nail-based NGS analysis method for predicting or diagnosing diseases, comprising the steps of: isolating gDNA from nails; and analyzing the isolated gDNA using NGS.
Need to check novelty before this filing date? Find Prior Art

Description

Nail-based NGS data analysis method and system for disease prediction

[0001] The present invention relates to a Next Generation Sequencing (NGS) method and system for predicting diseases through germline mutation and somatic mutation analysis derived from nails. Specifically, the present invention relates to a nail-based Next Generation Sequencing (NGS) analysis method for predicting or diagnosing diseases, comprising the steps of isolating gDNA from nails and analyzing the isolated gDNA using NGS.

[0002]

[0003] Precision medicine involves proactively identifying each patient's unique genetic and environmental factors, disease history, and lifestyle habits, and then administering the appropriate medication at the right dose at the right time to provide an optimized treatment plan for each individual. In particular, as disease prevention is considered more effective than post-treatment, personal genome analysis serves as a key element in implementing future healthcare, including precision medicine and preventive care, by enabling the early identification and management of specific diseases and health risks.

[0004] To elucidate the genetic correlations between complex traits and diseases, large-scale whole genome sequencing (Whole Genome-seq), transcriptome-seq, and epigenome-seq are being conducted by numerous research groups. However, various limitations still exist, such as nonrandom sampling, disease complexity, lack of bio big data, and groundbreaking advancements in bioinformatics analysis methods. However, the necessity and utility of personal genome analysis are expected to continue to increase thanks to the achievements of clinical genomics research.

[0005] Since the COVID-19 crisis, awareness of the importance of regular health management has increased, as it has become clear that people with underlying conditions such as hypertension and diabetes are more prone to severe symptoms if infected with COVID-19. In particular, with the rising obesity rate, the incidence of chronic diseases accompanying hypertension is on the rise, and in Korea, a majority of diabetic patients also have dyslipidemia.

[0006] Hyperlipidemia refers to a condition in which blood cholesterol and blood triglyceride levels are abnormally high, and dyslipidemia is a broad concept that includes abnormalities in blood lipid levels. Currently, 1 in 4 adults in Korea suffers from hypercholesterolemia, and 2 in 5 adults suffer from dyslipidemia. Hypercholesterolemia is on the rise, with 23% of men and 25% of women reported to have hypercholesterolemia. While awareness is improving, 3 in 10 hypercholesterolemia patients are still unaware of their condition, and while treatment rates have improved significantly, only about half of hypercholesterolemia patients are taking lipid-lowering medication. Although a large number of diabetic patients in Korea suffer from hyperlipidemia, the awareness and treatment rates are only in the 20-30% range, indicating that improvement is urgent. 83.3% of adult diabetic patients had hyperlipidemia, and the prevalence was higher in women (88.3%) and men (78.1%). In particular, the prevalence rate in young people aged 19 to 39 was 88.5%, which was higher than in other age groups. This is a much higher figure than the prevalence rate of hyperlipidemia in people in their 20s and 30s in the general population (15-20%) reported in previous studies, suggesting that hyperlipidemia management is necessary from an earlier age in patients with diabetes (Seung Jae Kim, Oh Deog Kwon, Kyung-Soo Kim, Prevalence, awareness, treatment, and control of dyslipidemia among diabetes mellitus patients and predictors of optimal dyslipidemia control: results from the Korea National Health and Nutrition Examination Survey. Lipids Health Dis. 2021 Mar 26;20(1):29.).

[0007] Chronic diseases that are also common in cancer survivors include hypertension, hyperlipidemia, diabetes, osteoporosis, and anemia. Factors that can contribute to the development of cancer include genetic factors, lifestyle risk factors, cancer itself, cancer treatment, and existing chronic diseases. An analysis of the causes of death of approximately 240,000 long-term cancer survivors in Korea who were diagnosed with cancer between 1993 and 2000 and survived for more than five years showed that 24% of cancer survivors died from causes other than cancer, most of which were diabetes and cardiovascular disease (Dong Wook Shin, Young Ho Yun et al., Non-cancer mortality among long-term survivors of adult cancer in Korea: national cancer registry study. Cancer Causes Control 2010; 21: 919-29.). In addition, chronic diseases that occur more frequently in cancer survivors, such as hypertension, hyperlipidemia, osteoporosis, and anemia, need to be checked more frequently than in the general population.

[0008] Currently, there are no guidelines or guidelines for health screening items for cancer survivors, and general health management guidelines are provided only for the most common solid tumors. Compared to the general population, cancer survivors are at higher risk for chronic diseases such as obesity, hypertension, diabetes, hyperlipidemia, osteoporosis, and anemia, as well as secondary cancers. These risks vary somewhat by cancer type, and smoking cessation, alcohol abstinence, appropriate exercise, and nutritional management are crucial for this vulnerable group. Therefore, follow-up care and preventative monitoring are essential to predict the occurrence of secondary diseases.

[0009] Therefore, to fundamentally prevent and treat diseases, an approach tailored to individual genetic status and mutations is necessary. Sanger-based sequencing, a conventional method for analyzing base sequences, can typically read DNA fragments of 500-800 base pairs. Therefore, it is limited in the analysis of long DNA sequences. Therefore, its application in research fields requiring large amounts of DNA sequence information, such as personal genome analysis, is inefficient in terms of time, labor, and cost. For these reasons, there is a growing demand for new analytical techniques capable of obtaining large amounts of DNA sequence information efficiently and at low cost. Consequently, the popularization of next-generation sequencing (NGS) and the technologies utilizing it are advancing rapidly.

[0010] Genome mutation detection relies heavily on the performance of mutation detection software. While a variety of software is currently in use, their mutation detection capabilities vary significantly. To address this, extensive research is underway to identify optimal standard pipelines. However, the software that delivers the best performance may vary depending on the characteristics and quality of the input data.

[0011] In particular, Mutect2, a widely used variant detection software, is known to excel at detecting low-frequency variants. However, false negatives have been reported in clinical settings. While various additional options are available to overcome this issue, in some cases, they fail to detect clinically significant variants. Because software like Mutect2 utilizes a probabilistic model to calculate the probability of a variant, errors can occur during this calculation, potentially reducing the reliability of variant detection. Therefore, a new approach is needed to improve the accuracy of variant detection and address the issue of false negatives.

[0012] As multinational companies develop NGS equipment and improve reagents, and bioinformatics techniques advance, the cost of genotyping is gradually decreasing. NGS analysis offers several advantages over conventional Sanger-based methods, and is being utilized in numerous research, health screening, and diagnostic services.

[0013] Accordingly, tissue and blood samples have traditionally been used as the main source of DNA for genetic analysis, but tissue samples have limitations in obtaining them through surgery or other procedures, and blood samples have difficulties in collecting them due to factors such as arterial thrombosis, skin infection, anemia, hemophilia, people who have been taking drugs such as isotretinoin for less than 4 weeks, women who are pregnant or have given birth for a certain period of time, and patients with infectious diseases or leukemia.

[0014] Additionally, these tissues or blood samples require essential equipment and facilities for long-term storage due to the instability of the samples due to temperature changes during transport to the location for analysis or testing, the high cost of transportation for refrigeration or freezing, and the burden of freezing or ultra-low temperature equipment and facilities. In addition, there is a need for a system that can resolve the above shortcomings from alternative clinical samples and enable more periodic and smooth examination and diagnosis due to the burden of high transportation costs for refrigeration or freezing of samples due to temperature changes when transporting samples to the location for analysis or testing.

[0015] In order to solve this problem, the inventor of the present invention developed an analysis method and system that reduces the possibility of false negatives and has high accuracy through a nail specimen that can replace blood or tissue, thereby completing the present invention.

[0016]

[0017] (Prior art literature)

[0018] (Non-patent literature)

[0019] Seung Jae Kim, Oh Deog Kwon, Kyung-Soo Kim, Prevalence, awareness, treatment, and control of dyslipidemia among diabetes mellitus patients and predictors of optimal dyslipidemia control: results from the Korea National Health and Nutrition Examination Survey. Lipids Health Dis. 2021 Mar 26;20(1):29.

[0020] Dong Wook Shin, Young Ho Yun et al., non-cancer mortality among long-term survivors of adult cancer in Korea: national cancer registry study. Cancer Causes Control 2010; 21:919-29.

[0021] Barbitoff, Y. A., Abasov, R., Tvorogova, V. E. et al. Systematic benchmark of state-of-the-art variant calling pipelines identifies major factors affecting accuracy of coding sequence variant discovery. BMC Genomics 23, 155 (2022).

[0022]

[0023] The purpose of the present invention is to provide a method and system for analyzing genetic data based on NGS, which can predict the presence or absence of a disease based on individual characteristics using a nail (fingernail or toenail) specimen and can be easily applied to a screening center, hospital, or future DTC (Direct To Consumer) testing service.

[0024]

[0025] The present invention provides a nail-based NGS analysis method for predicting or diagnosing a disease and a system for predicting or diagnosing a disease using the same.

[0026] Specifically, the present invention provides a nail-based NGS analysis method capable of diagnosing or predicting chronic diseases such as obesity, hyperlipidemia, and diabetes, and the occurrence, recurrence, and metastasis of cancer, and a disease prediction or diagnosis system using the same.

[0027] The present invention provides a nail-based NGS analysis method for predicting or diagnosing the disease, comprising the steps of (a) isolating gDNA from a nail; and (b) analyzing the gDNA isolated in step (a) using NGS (Next Generation Sequencing).

[0028] The step of isolating gDNA from the nail (a) is a step of preparing a PCR or NGS analysis sample by extracting gDNA using a forensic method used in forensic medicine, rather than applying the conventional extraction method using phenol / chloroform, which is a common extraction method, in order to reduce DNA damage to the nail sample.

[0029] The above nail includes fingernails or toenails, or both fingernails and toenails. Furthermore, the nail is not limited to human nails, but includes all nails of animals (or individuals) with fingernails or toenails, such as dogs and cats.

[0030] Nails generally grow at a rate that can vary from person to person, but on average, they grow from 1.8mm to 4.5mm per month. The more stimulation they receive, the more blood flow to the nails increases, which promotes cell division in the nails, causing them to grow faster. Among other relatively non-invasive methods of testing, hair grows at an average of 0.3mm per day, which is about 1cm per month and about 12cm per year. This is because hair grows slower than nails, and in children or those without hair, or when a small amount of DNA is extracted from a hair sample, variables may occur in the NGS analysis results. Likewise, in the case of saliva, the results may be negatively affected by the use of mouthwash or the influence of foreign substances ingested at the time, excessive water intake, oral microorganisms, etc., and a disposable saliva collector and tube containing a separate preservative are required for sample collection, and the detection of trace amounts of DNA may limit specific genetic research or diagnosis. Therefore, the collection of nail specimens that can be transported and stored at room temperature and have no restrictions on age, gender, or specific diseases is most suitable for analysis.

[0031] The inventors confirmed, through the experimental results shown in Fig. 2, that nails, like other specimens, can be used for stable gDNA extraction. The sample was collected directly at room temperature by the client, stored in a paper container, and then subjected to testing. Nail specimens are environmentally friendly and convenient for clients, enabling them to be utilized for everything from screening to diagnosis at screening centers and hospitals, confirming their suitability as suitable samples for genetic testing.

[0032] Nail samples can be collected from individuals or patients and provided by local genetic testing institutions, health screening centers, clinics, and primary or secondary hospitals. By examining the gDNA (genomic DNA) contained in nail samples, researchers can predict the development of chronic diseases and related complications, such as aging, obesity, hyperlipidemia, and diabetes, and assess the risk of cancer development, metastasis, and recurrence.

[0033] In addition, the NGS-based data analysis method of the present invention may include a step of acquiring data by receiving existing clinical information from a screening institution or a primary or secondary hospital so that it can be compared with an existing blood sample of the analysis target.

[0034] Extraction of gDNA from nail samples can be performed for single nucleotide polymorphism (SNP) genotyping using phenol / chloroform extraction. This is a common method for extracting DNA from various samples, including nails. The phenol / chloroform extraction method disrupts the cell membranes of the cells in the nail sample, releasing the DNA. The DNA is then extracted from the solution using salt precipitation.

[0035] While the phenol / chloroform extraction method is a relatively simple and efficient way to extract DNA from nails, it is time-consuming and can damage DNA. Furthermore, the phenol / chloroform extraction method can also extract DNA from other sources, such as bacteria and fungi. This can lead to DNA sample contamination. While the phenol / chloroform extraction method can be used if nail-derived DNA is needed solely for SNP genotyping, a more efficient and less damaging method is needed if DNA is needed for other applications, such as next-generation sequencing (NGS) sequencing or polymerase chain reaction (PCR).

[0036] SNP genotyping is a technique used to identify SNPs in DNA samples. SNPs are variations of a single nucleotide (A, T, C, or G) in a DNA sequence, and can be used to identify individuals, track genetic variation in populations, and diagnose genetic diseases.

[0037] NGS can interfere with the results and lead to inaccurate conclusions if the sample is mixed with impurities such as salt, phenol, ethanol, and iron, or if heparin, EDTA, NaCl, and KCl are mixed during the DNA extraction step. Therefore, it is necessary to check for the cause of the error before testing.

[0038] The present invention aims to reduce such DNA damage by extracting gDNA using a forensic method instead of the phenol / chloroform extraction method. This eliminates time-consuming washing steps, such as precipitation with isopropanol or ethanol, thereby shortening extraction time and minimizing damage.

[0039] In the NGS analysis step of the above (b), NGS technology is a technology that fragments the DNA or RNA of a living organism into small pieces and reads the sequence by machine. NGS technology extracts DNA, fragments the DNA into short sequence fragments, and performs sequencing to analyze the bases contained in the sequence for each sequence fragment. In one embodiment of the present invention, the Illumina sequencing method may be used as the NGS method, but is not limited thereto.

[0040] The nail-based NGS analysis method of the present invention

[0041] (c) mapping genes in gDNA isolated from nails; and / or

[0042] (d) may additionally include a step of calculating the expression of a gene associated with the onset of a disease.

[0043] The mapping step of the above (c) means that after sequencing the gDNA separated from the nail, mapping is performed to align each sequence fragment with respect to a reference genome to confirm the location of each sequence fragment in the genome, thereby identifying the location of the sequence fragments within the genome. At this time, the reference genome is a virtual base sequence of an individual representing a species of organism and is a genome map that serves as a guide for mapping. After identifying the location of all sequence fragments, various analyses can be performed, such as analyzing whether there is a mutation in the DNA or measuring the amount of DNA transcribed into RNA.

[0044] The calculation step of the above (d) means calculating the expression level of a gene associated with disease onset by obtaining the probabilities that the analysis target data corresponds to each of the genotypes of the analysis target gene based on the mapping result of the above (c). Here, the "analysis target data" refers to genome sequence data obtained through an individual's nail-based NGS (Next-Generation Sequencing) analysis. The data includes information necessary to identify individual genetic characteristics, including genetic mutations associated with specific diseases or health conditions. In addition, the "analysis target genes" are genes known to be associated with the onset of a specific disease, and can be used to evaluate the influence on the possibility of disease onset, for example, by analyzing the genotype at a specific SNP (single nucleotide polymorphism) location.

[0045] The operation step of (d) above includes identifying the location and sequence of mutations or detecting and identifying somatic or germline mutations based on the mapping results of (c). The predicted results can be compared and analyzed with the results of whole genome sequencing (WGS), which analyzes the entire genome, whole exome sequencing (WES), which examines the protein-coding regions of the entire genome, and targeted panel sequencing (Targeted sequencing), which analyzes only specific gene regions.

[0046] In the present invention, genomic information extracted from nails is read using the Whole Exome Sequencing technique. The most important factor is to detect a small number of reads (DNA fragments whose base sequences are read by a sequencing device) without omission at low depth. Therefore, the inventors developed and applied a variant caller algorithm (Bullseye) that reads information from all reads to identify low-frequency mutations.

[0047] The operation step (d) according to the present invention includes aligning and mapping the sequenced reads to a reference genome, and then identifying the position and sequence of the mutation based on the information of the mapped reads using a variant caller algorithm. Since the variant caller algorithm (Bullseye) of the present invention can be used to detect and identify differences between the reference genome and the reads, particularly somatic mutations or germline mutations, Bullseye uses a method that has improved accuracy through comparative verification with the existing Mutect2, thereby reducing errors of false mutation detection (false positives) or missed mutations (false negatives).

[0048] In addition, the nail-based NGS analysis method of the present invention may additionally include a step of (e) constructing a genome database using the operation result of step (d).

[0049] The step of constructing a genome database of the above (e) may be a step of constructing a catalog of individual genome information based on a result report on disease prevention and prediction through confirmation of individual germline mutations and somatic mutations.

[0050] In addition, the present invention provides a method for providing information for the prevention, diagnosis, or prognosis of a disease based on a germline mutation and a somatic mutation of an individual.

[0051] The genome information catalog of the above (e) may include health information based on SNPs, which are individual-specific genotype characteristics, exercise and nutrition information, drug response information (Pharmacogenomics), information related to immunity, skin, mental health, digestion and metabolism-related hormones, metabolism, etc., and disease-related content including the occurrence, recurrence, and metastasis of chronic diseases and cancer.

[0052] Specifically, the above-mentioned genomic information catalog may include a genomic database such as SNV, INDEL, CNV, etc. through WGS (whole-genome sequencing), WES (whole-exome sequencing), and targeted sequencing based on nail-derived gDNA obtained from normal or patient samples, and a basic clinical data database of patients with chronic diseases or cancer.

[0053] The present invention relates to a system for predicting or diagnosing a disease using the above NGS analysis method, comprising the following steps:

[0054] (a1) loading a computer program to be executed by one or more processors;

[0055] (b1) A step of providing a memory for storing data on which genotypes of the target genome are determined;

[0056] (c1) a step of generating a personal genetic analysis result report based on the stored data; and

[0057] (d1) A system construction unit that builds a genetic database of the above analysis target and a report generation step that provides information from the system construction unit to the analysis target.

[0058] In the step (a1), one or more processors may include (i) an operation step of obtaining data by receiving a nail sample of the analysis target and general information and / or clinical information and storing the data, (ii) a gDNA sample analysis operation step of tracking disease-related genes such as the occurrence of chronic diseases and / or the occurrence, recurrence, and metastasis of cancer in a nail sample, preferably gDNA isolated from a nail, and mapping the analysis target data to each of base sequences having different genotypes for the analysis target genome, and (iii) a clinical information analysis operation step of calculating the expression of genotypes in the analysis target data based on the mapping result.

[0059] The general information in (i) above refers to information such as age, gender, family history, or lifestyle habits provided with the individual's consent, and clinical information refers to information such as blood pressure, blood sugar, smoking, drinking, disease history, vaccine and inoculation information, regular checkup items (blood pressure, blood sugar, cholesterol, cancer screening, etc.), type and dosage of prescribed medication, history of side effects, or allergic reactions, and can be provided from primary or secondary hospitals.

[0060] The disease prediction or diagnosis system according to the present invention may further include, after the report generation step (d1), the step of (e) building a big data database. The report may include information on an individual's constitution, customized nutrition and diet, exercise, stress relief and psychological mental health management, recommended activities, combinations of health functional foods and supplements, regular health checkup cycles and methods, or trackable health goals and progress.

[0061] The disease prediction or diagnosis system according to the present invention can also be applied as a veterinary genetic testing method in animals such as dogs and cats with nails.

[0062] In the present invention, diseases may include, but are not limited to, aging, obesity, hypertension, diabetes, hyperlipidemia, cancer, sarcopenia, osteoporosis, lung disease, cardiovascular disease, brain disease, liver disease, dementia, or Alzheimer's.

[0063]

[0064] By complementing the shortcomings of existing tissue and blood-based liquid biopsies, the present invention's nail-derived personal genome analysis can enable not only the identification of disease risks and health abnormalities and the provision of proactive measures, but also precision medicine for personalized treatment.

[0065] Disease prevention is crucial, as the incidence of diseases such as obesity, aging, hyperlipidemia, diabetes, and cancer is increasing due to the interaction between genetic markers and environmental factors. Diagnostic technologies utilizing easily collected, stored, and transportable specimens, such as nail (fingernail or toenail) samples, can facilitate proactive health management to prevent and improve human and animal disease, thereby extending healthy lifespans and achieving an effective quality of life (QOL).

[0066]

[0067] Figure 1 illustrates an NGS-based data analysis system for disease prevention and prediction based on an individual's genetic characteristics according to one embodiment of the present invention.

[0068] Figure 2 shows the results of a comparative analysis of gDNA derived from fingernails, gDNA derived from blood (whole blood), cfDNA derived from plasma, and gDNA derived from saliva.

[0069] Figure 3 shows the results of comparing the germline mutation detection rates of donors according to sample type.

[0070] Figure 4 shows the results of comparing the detection rate of somatic mutations in donors according to sample type.

[0071] Figure 5 shows the results of mutation status related to disease diagnosis and treatment decisions in a group of diabetic patients.

[0072] Figure 6 shows the results of mutation status related to diagnosis and treatment decisions in a group of patients with borderline diabetes.

[0073] Figure 7 shows the results of WES genetic mutation analysis related to sarcopenia when applying the developed in-house caller (Bullseye) and existing variant caller (Mutect2) algorithms.

[0074] Figure 8 shows the results of a comparative analysis of the non-small cell lung cancer standard material (Seracare Tri-Level Tumor Mutation DNA Mix v2) when applying the developed in-house caller (Bullseye) and existing variant caller (Mutect2) algorithms.

[0075] Figure 9 shows the results of analyzing somatic variants in three types of samples, namely fingernails, blood, and surgical tissue, from bladder cancer patients.

[0076]

[0077] Hereinafter, with reference to the attached drawings, embodiments and examples of the present invention will be described in detail so that those skilled in the art can easily implement the present invention. However, the present invention may be implemented in various forms and is not limited to the embodiments and examples described herein.

[0078] Throughout this specification, whenever a part is said to "include" a component, this means that it may include other components, but not to the exclusion of other components, unless otherwise stated.

[0079] The present invention will be described in more detail through the following examples; however, the following examples are for illustrative purposes only and are not intended to limit the scope of the present invention.

[0080]

[0081] [Example 1] Sample preparation

[0082] gDNA isolation from fingernail samples

[0083]

[0084] In the present invention, in order to minimize damage in nail DNA extraction, a method that does not use phenol / chloroform is applied based on a single nail weighing about 10 to 20 mg, and an extraction method utilizing Forensic DNA Kits produced by domestic and foreign companies (OMEGA BIO-TEK EZNA®, JinAll Biotechnology Exgene™ Forensic SV mini, etc.) is applied by adjusting some of the reaction solution volume and reaction time for DNA elution to suit the nail specimen.

[0085] Sample quality control was performed by checking DNA concentration, purity, fragment length, etc., and the quality assessment and analysis results for genomic DNA (gDNA) samples were confirmed according to the TapeStation gDNA Screen Tape method, one of the microelectrophoresis analysis technologies.

[0086] Sample quality assessment using TapeStation is a quick way to check the quality of gDNA samples, ensuring they are clean and intact and identifying issues due to contamination or degradation. It can also be used to measure gDNA concentration, allowing for a quantitative calculation of how much DNA is present in a sample. Additionally, the ability to visualize and analyze the size distribution of gDNA allows for the identification of the amount of DNA fragments within a specific size range or the detection of any anomalies in the size distribution, making it an ideal QC method for assessing the quality of extracted DNA.

[0087] The samples were compared and analyzed with fingernail-derived gDNA from the same individual who was diagnosed as at risk for borderline diabetes and hypertension through a general health checkup, and blood (whole blood)-derived gDNA, plasma-derived cfDNA, and saliva-derived gDNA, which are currently mainly used in liquid biopsy, as a control group for comparison. In addition, the fingernail-derived gDNA sample from a patient diagnosed with diabetes and a maternal donor was compared and analyzed to observe germline mutations and somatic mutations from actual fingernails, thereby deriving the results of DNA extraction suitable for NGS analysis for the prevention and prediction of diabetes. The results are shown in Table 1 and Fig. 2 below.

[0088]

[0089]

[0090]

[0091] [Example 2] NGS sequencing

[0092] NGS-based WES analysis

[0093]

[0094] Whole-exome sequencing (WES) is a method of analyzing only the base sequence of the exome region excluding non-coding regions. It can be seen as an example of target sequencing in that it targets a specific region, and its main purpose is to capture and amplify the target region and compare it with other samples to find specific variants. Compared to whole genome sequencing (WGS), which analyzes the entire genome sequence, the data capacity is small, making data analysis easy. In addition, since most variants known to be associated with diseases occur in the exon region, WES analysis was performed using NGS-based sequencing at a depth of 200X to 300X, which is an effective method in terms of time required and analysis cost.

[0095] Compared to panel-based NGS analysis with a depth of 1000x to 10000x or more that targets an average of 10 to 120 types of genetic mutations during cancer diagnosis, sequencing with a low depth of 200X to 300X may result in lower accuracy in selecting mutant genes, making it insufficient for use in the diagnostic field. Therefore, in order to secure improved mutation detection performance in terms of sensitivity and accuracy to reduce false positives and detect 100% of missed mutations while using WES analysis that can identify more than 20,000 mutations, we developed an important algorithm used to identify mutations by analyzing genomic data and conducted the analysis. The analysis results were derived through comparison and double validation with the Mutect2 variant caller, which is currently being utilized by the majority of domestic and foreign researchers and companies.

[0096] The developed variant caller algorithm is established through the following process, and the filtering step is improved based on a linear regression model obtained from simulation data to ensure high accuracy. In particular, it is an optimized method to maintain high sensitivity even in samples with low variant allele frequencies.

[0097]

[0098] 1. Importing Python packages

[0099] A. import argparse

[0100] ■ Package for receiving input arguments

[0101] B. import numpy as np

[0102] ■ Package required for depth calculation in the VCF writing stage

[0103] C. import pandas as pd

[0104] ■ A package that stores and processes mutation information as a data frame.

[0105] D. import pickle

[0106] ■ Packages required to load and use dictionaries (used for result filters)

[0107] E. import pysam

[0108] ■ Package used to load lead information from a bam file

[0109] F. import random

[0110] ■ Package used to encrypt mutation patterns

[0111] G. import re

[0112] ■ Package for manipulating strings (used to extract mutation information from cigarstring and md tags)

[0113] H. import scipy.stats as stats

[0114] ■ Scientific computing package (used to filter strand bias)

[0115] 2. Function

[0116] A. read_info

[0117] ■ The md tag obtained during the preprocessing process, the location information and base sequence of the reads contained in cigarstring and bamfile, and the corresponding reference genome are loaded.

[0118] ■ Extract snv and location information where deletion occurred using the re package. -cnt(list)

[0119] ■ Find out which types of bases are deleted from snv using the re package. - seq(list)

[0120] ■ Cut out the softclip from the read base sequence and modify the cigarstring accordingly. - cigar(string)

[0121] (1) A softclip is a portion that does not match (mainly both edges) when the read is aligned to the reference genome, and most of the time it corresponds to an adapter or molecular barcode sequence.

[0122] ■ Save the read sequence as a variable. - read_seq(string)

[0123] (1) When there is a softclip on the left, both sides, and right, only the purely mapped lead sequence is left based on the location to be deleted.

[0124] ■ Retrieves the reference genome sequence of the location where the read is mapped. - ref_seq(string)

[0125] ■ Save the genomic position of the lead as a variable. - block_start(int)

[0126] B. prep_indel

[0127] ■ Preparation process for detecting indels. Using cigar.

[0128] ■ Find the pattern of M(match), D(delete), and I(insert) in cigar. - indel_pattern(list)

[0129] ■ Find the location information of M (match), D (delete), and I (insert) in cigar. indel_iter(list)

[0130] ■ For example, if cigar is 52M2I28M2D63M, two insertions occur after a match of 52 bases, and two bases are deleted after a match of 28 bases, leaving 63 bases that match. In this case, indel_pattern becomes ['M', 'I', 'M', 'D', 'M'] and indel_iter becomes [52, 2, 28, 2, 63].

[0131] ■ Using a for loop, move the position of the read and the position of the reference according to the indel pattern, and find the relative position within the read of deletion and insertion, and the number of inserted / deleted bases. del_pos_list(list), ins_pos_list(list), del_size(list), ins_size(list)

[0132] ■ Explaining the example above

[0133] (1) Since the first indel element of indel_pattern is I (insertion), ref_pos is increased by the matching number (52) before the insertion. Since it is the first indel, alt_pos is also the same. ins_pos_list is based on alt_pos. The insertion position obtained in this way is added to ins_pos_list. That is, ref_pos: 52, alt_pos: 52, ins_pos_list:

[0052]

[0134] (2) The second indel element of indel_pattern is D(deletion), and ref_pos adds the matching number (28) after the existing ref_pos (step 1: 52). At this time, alt_pos (step 1: 52) adds the inserted 2 to the existing alt_pos since two bases were inserted in the previous step 1, and then adds the matching number (28). del_pos_list is based on ref_pos. That is, ref_pos: 80, alt_pos: 82, del_pos_list:

[0080]

[0135] (3) Summarizing the above results,

[0136] (A) del_pos_list =

[0080]

[0137] (B) del_size = [2]

[0138] (C) ins_pos_list =

[0052]

[0139] (D) ins_size = [2]

[0140]

[0141] C. call_indel

[0142] ■ The genomic position of the indel and the deleted or inserted base sequence are stored as dictionary variables using del_pos_list / del_size and ins_pos_list / ins_size obtained during the prep_indel process. - Key: Variant type (indel), Value: Genomic position, ref, alt (list)

[0143] (1) Insertion: The inserted base can be found in read_seq. The inserted base sequence is obtained using the insertion position and insertion size in read_seq.

[0144] (2) Deletion: The deleted base can be found in ref_seq. The deleted base sequence is obtained using the deletion position and deletion size in ref_seq.

[0145] D. call_snv

[0146] ■ The function that detects snv uses the cnt and seq lists obtained from the md tag.

[0147] ■ Since the location information of snv changes depending on the presence or absence of indel, the output variable of call_indel is entered as an input variable.

[0148] ■ Like prep_indel, move the lead and reference sequences from left to right to find the location corresponding to cnt. - Use for loop

[0149] ■ Set 0 as the initial value of ref_pos (reference) and alt_pos (lead). (Initial starting point)

[0150] ■ From cnt =

[0034] , seq = ['C'], we can see that the 34th C base was replaced with another base and that the total number of snvs is one.

[0151] ■ If you move 34 bases from the left in ref_seq, you get C, and if you move 34 bases in read_seq, you get G. In other words, the 35th C changed to G.

[0152] ■ If one of the cnt elements starts with ^, it means that it has been deleted, and since it is not snv, only ref_pos increases (because it has been deleted), and alt_pos does not change.

[0153] ■ On the other hand, when an insertion occurs, the information is not in cnt but in the ins_pos_list (alt_pos_list in the corresponding function), which is the output of prep_indel, and ref_pos does not change, but only alt_pos increases. alt_pos gets position information from cnt, but since cnt does not have insertion information, when calculating by referring to ins_pos_list, in order to prevent cases where insertion information is entered repeatedly, the relative positions of the substitution and insertion of the cnt must be considered in the calculation. In other words, alt_pos increases by the number of inserted bases only when the insertion appears before the substitution. At this time, the insertion position is saved in a variable (last_ins) to prevent alt_pos from increasing repeatedly.

[0154] ■ The information obtained from this is stored in the var_dict dictionary, just like call_indel. Key: variant type (snv), value: genomic position, ref, alt (list)

[0155] E. make_res

[0156] ■ The mutation information stored in the dictionary is converted into a list by dividing it into forward read / reverse read.

[0157] ■ Each element of this list contains a single string, which represents chromosome, genomic_position, ref, and alt, respectively.

[0158] ■ Convert this list into a data frame using pandas and count identical mutations.

[0159] ■ Calculate the depth of each mutation location using pysam's count function.

[0160] ■ The result is two data frames that store the mutations of each forward / reverse lead.

[0161] F. filter_output

[0162] ■ This is the process of filtering each mutation based on its own criteria.

[0163] ■ low_vaf

[0164] (1) Using the fastq simulator art_illumina, generate fastq data without mutation information. All mutations generated in this process are treated as errors. A linear regression model is created using the maximum error count for each depth.

[0165] (2) If the mutation from make_res is lower than the value of this model, it will be included in the error category and the low_vaf flag will be set in the subsequent make_vcf process.

[0166] ■ low_depth

[0167] (1) The average coverage of whole exome sequencing is 50, and variants with a depth of 40 or less are flagged with this flag. This value may vary depending on experimental results.

[0168] ■ strand_bias

[0169] (1) Strand bias refers to the phenomenon in which mutations are concentrated in the forward or reverse leads.

[0170] (2) Broad Institute recommends FisherStrand probability calculation, and this standard was applied to this algorithm.

[0171] (3) Get the depth and count information from the final data frame and save [[depth1, depth2], [read1, read2]] in variables.

[0172] (4) Calculate the p-value using the scipy.stats package.

[0173] (5) Calculate the Phred score.

[0174] (6) If the phred score is less than 60, it is considered that there is no strand bias.

[0175] ■ Germline / panel of normal filter

[0176] (1) The panel of normal VCF (variants appearing in the normal group) and germline mutations provided by the Broad Institute are excluded from the results.

[0177] G. find_pattern

[0178] ■ If multiple mutations are found in one read during the mutation detection process, they are saved as separate outputs.

[0179] ■ A unique ID (32-bit value) is assigned to each mutation using getrandbits from the random package. Identical mutations have identical IDs. The ID assigned to each mutation is stored in a dictionary called patterninfo.

[0180] ■ If there are three mutations in one lead, they are stored in pattern (list) in the form of [ID1, ID2, ID3].

[0181] ■ Once mutation detection is complete, the same patterns within the pattern are converted into a data frame and counted.

[0182] ■ Retrieve mutation information matching the ID from patterninfo.

[0183] H. make_vcf

[0184] ■ Process of creating vcf from the result data frame

[0185] ■ After loading the template vcf, find information in the data frame and fill in each column.

[0186] ■ It consists of '#CHROM', 'POS', 'ID', 'REF', 'ALT', 'QUAL', 'FILTER', 'INFO', 'FORMAT', and sample name.

[0187]

[0188] In order to compare the in house caller (Bullseye) developed in the above manner with the popular Variant caller (Mutect2), the preprocessed bam file was used to detect mutations and obtain a vcf file. Vcf files were annotated using ensembl's vep, and the Vep results and vcf information were integrated and processed. Subsequently, when filtering for specific disease-related genes and analyzing the results, as shown in Figure 7, for the WES genetic mutation analysis related to sarcopenia, the developed in-house caller algorithm detected approximately 40 times more disease-related mutations than the existing variant caller (Mutect2).

[0189] In addition, as a result of comparative analysis with the non-small cell lung cancer standard material (Seracare Tri-Level Tumor Mutation DNA Mix v2), as shown in Table 2 and Figure 8 below, the in house caller detected all mutations included in the standard material, whereas Mutect2 did not detect two mutations.

[0190]

[0191]

[0192]

[0193] There have been several reported cases of false negatives in Mutect2, which is the most representative and is known to have good performance at low frequencies (Krøigεrd, AB, Thomassen, M., Lζnkholm, A.-V., Kruse, TA, & Larsen, MJ (2016). Evaluation of Nine Somatic Variant Callers for Detection of Somatic Mutations in Exome and Targeted Deep Sequencing Data. PLOS ONE, 11(3), e0151664). In other words, there are cases where clinically significant variants are not detected, and this may be due to an error in the way the detection algorithm calculates the probability of a variant using its internal probability model.

[0194] The developed in-house caller (Bullseye) was able to find mutations even at low depth and frequency by reading all read information and finding differences from the reference genome sequence when comprehensively testing clinical samples and standard materials, and the performance evaluation results showed sensitivity and accuracy that were 1.3 to 39.7 times better than Mutect2.

[0195]

[0196] [Example 3] Genetic analysis test

[0197] Genetic analysis tests for patients with chronic diseases and cancer

[0198]

[0199] A genetic analysis test based on nail samples was conducted with patients with chronic diseases and cancer as the primary target. For primary verification, the results of a comparative analysis of genetic and non-genetic disease markers related to diabetes were derived. In order to confirm the concordance rate of the detected mutations based on whole blood in the results of observing genetic mutations related to diabetes, hypertension, hyperlipidemia, etc. based on germline mutations in sample donors with borderline diabetes and hypertension, when comparing the detection of nail, plasma, and saliva samples, similar or highest detection patterns were confirmed in nail samples, as shown in Figure 3. The results were calculated as a value expressed as a percentage by dividing the number of mutations common to those detected in whole blood for each sample by the total number of detections.

[0200] When comparing the detection of genetic mutations related to diabetes, hypertension, hyperlipidemia, etc. in fingernail, plasma, and saliva samples with whole blood based on somatic mutation criteria of sample donors with borderline diabetes and hypertension, similar or highest detection patterns were confirmed in fingernail samples, as shown in Fig. 4. The results were calculated as a value expressed as a percentage by dividing the number of mutations common to those detected in whole blood for each sample by the total number of detections.

[0201] In the case of the diabetic patient group, when only diabetes was filtered using the annotation tool, genes that were reported to be associated with existing diabetes with pathogenicity, such as ALAD, ANKRD17, ASH1L, ATXN3, AZIN1, CACNA1E, CASQ2, CDC25A, CEL, CEP104, CPZ, CRB1, CTNS, DMXL2, EPC1, ERBB3, FAM193A, FAM20C, GALC, INPP4B, IVD, MAF, MAP3K9, MBTD1, MECOM, MED13, MYH3, NCOA2, NRXN1, PACSIN1, PDE6B, PIAS2, PPP3CB, PRKRA, PUM2, QKI, ROCK1, RPL7, SPAG9, TFDP2, TNRC6A, TTK, VPS13A, YTHDF3, ZC3H15, By identifying mutations in genes such as ZNF326 and ZNRF3, we confirmed that gDNA from fingernails could potentially be used to predict and diagnose diabetes (Table 3). Common genetic mutations were detected in metabolic diseases such as obesity, hyperlipidemia (HL), type 2 diabetes mellitus (T2DM), and metabolic syndrome (MS), and the gray-colored group belongs to risk prediction markers based on type 1 diabetes mellitus (T1DM).

[0202]

[0203]

[0204]

[0205]

[0206]

[0207]

[0208]

[0209]

[0210]

[0211]

[0212]

[0213]

[0214]

[0215]

[0216]

[0217]

[0218]

[0219]

[0220]

[0221]

[0222]

[0223]

[0224]

[0225] Additionally, when we checked the markers related to pancreatic cancer, which is one of the most common cancers in diabetic patients, we were able to confirm the mutation status of pancreatic cancer-related genes marked in gray as in Table 4, and although there were no mutations corresponding to tier 1 and tier 2, which are mutation criteria related to diagnosis and treatment decisions of the disease, many cases corresponding to tier 3, which are mutations with unknown clinical significance that may be of interest in the future preventive medicine, were confirmed (Fig. 5), and it was confirmed that they were more than 10 times more common than the borderline diabetic sample group in Fig. 6 and Table 5.

[0226]

[0227]

[0228]

[0229]

[0230]

[0231]

[0232]

[0233] The above results indicate that Tier 3 is related to health sustainability and can be observed as a recurrence of the disease and a long-term outcome of treatment. Therefore, it can be said that this analysis result can positively contribute to the client and patient by enabling them to predict the changing status of the disease and prepare for prevention and management accordingly through periodic monitoring of this area.

[0234] In addition, by being able to observe the types of genetic mutations that specifically indicate a relationship with obesity and diabetes and a possibility of being a risk factor for pancreatic cancer in the future, we were able to derive results that can confirm the possibility of application to diagnosis related to prognosis or prediction, and by being able to confirm a similar result pattern of genetic mutations inherited through the maternal line, we showed the possibility of early diagnosis and preventive medicine utilization of diseases that can occur genetically among families in relation to chronic diseases such as diabetes.

[0235] Analysis of somatic mutations in three different samples—nails, blood, and surgical tissue—from bladder cancer patients revealed a higher prevalence of bladder cancer-related somatic mutations in nail samples compared to whole blood. As shown in Table 6 and Figure 9, when comparing each mutation, nail samples shared more mutations with tissue samples.

[0236]

[0237]

[0238]

[0239] Additionally, in the case of specific mutations detected by individual WES, the depth of coverage of the base sequence is relatively low compared to Targeted Sequencing, and since the number of genes analyzed is large, there is a possibility that variants of unclear significance (VOUS) may also be detected. Therefore, in cases where a specific genetic disease is suspected, a panel test (Targeted Sequencing, Panels) targeting a specific disease based on nail samples can be used to secondary screen the targeted genes and mutations, and then a panel kit can be produced to obtain high-intensity data of 1000X or more.

[0240] In the report generation stage, the results delivered from the nail sample analysis operation and the results delivered from the clinical information analysis operation can be collated and provided in the form of an HTML5-based responsive web that can be provided to terminals that can be checked by individuals in the case of frontline examination centers and local hospital DTC services. Each of these components can mean software or hardware such as FPGA (Field Programmable Gate Array) or ASIC (Application-Specific Integrated Circuit), but is not limited to software or hardware, and can be configured to be on an addressable storage medium, and can be configured to execute one or more, detailed components, or processors consisting of multiple components.

[0241] In addition, the data collected through the above method can be enhanced through machine learning for big data, and supervised learning or unsupervised learning can be utilized depending on the condition. In a state where a specific answer is available, the reliability can be increased through supervised learning that optimizes disease-related variables by adding clinical information (blood pressure, blood sugar, smoking, drinking, disease history, etc.), mutation information of chronic diseases or cancer genes, psychological questionnaire information, intestinal microbiome information, etc. As a large number of individual samples and clinical information are accumulated, the accuracy increases, and it can be utilized as a deep learning algorithm for the discovery of mutant genes with accurate prognosis prediction for monitoring and analysis of correlation with clinical findings.

[0242] Based on this analysis system, personalized screening information or diagnosis-based care can be provided to each client each month based on uniquely stored personal information. For example, online results reports can be provided to institutions, screening centers, hospitals, etc., where clients submitted nail samples. These institutions, screening centers, and hospitals can then assess the client's condition and, if additional testing is deemed necessary, recommend a more precise diagnosis. Furthermore, they can provide clients with superior wellness-related counseling services, genetic health screening counseling services, and other medical services.

Claims

1. A nail-based NGS (Next Generation Sequencing) analysis method for predicting or diagnosing a disease comprising the following steps: (a) a step of isolating gDNA from nails; and (b) A step of analyzing the gDNA separated in step (a) using NGS.

2. A nail-based NGS analysis method in claim 1, wherein the nail is at least one selected from the group consisting of fingernails and toenails.

3. A nail-based NGS analysis method in the first paragraph, wherein in step (a), NGS isolates gDNA from the nail using a forensic method.

4. A nail-based NGS analysis method, wherein in step (b), NGS analysis is performed using the Ilumina sequencing method.

5. A nail-based NGS analysis method according to claim 1, wherein the method further comprises the step of (c) mapping reads of a region corresponding to a gene in gDNA separated from the nail.

6. In the fifth paragraph, the mapping step of (c) is a nail-based NGS analysis method in which, after sequencing gDNA separated from the nail, each read is aligned and mapped based on a reference genome.

7. A nail-based NGS analysis method according to claim 5, further comprising the step of calculating the expression level of a gene associated with disease onset.

8. In the 7th paragraph, the calculation step of (d) is a nail-based NGS analysis method that calculates the expression of a gene associated with disease onset by confirming the probabilities that the analysis target data corresponds to each of the genotypes for the analysis target gene.

9. A nail-based NGS analysis method in paragraph 7, wherein the operation step of (d) includes identifying a position and sequence having a mutation or detecting and identifying a somatic mutation or germline mutation.

10. A nail-based NGS analysis method in the 7th paragraph, wherein the operation step of (d) is performed using a variant caller algorithm.

11. A nail-based NGS analysis method according to claim 7, wherein the method further comprises the step of (e) constructing a genome database.

12. In the 11th paragraph, the step of constructing a genome database of (e) is a nail-based NGS analysis method for constructing a catalog of individual genome information for disease prevention or prediction by confirming germline mutations and somatic mutations.

13. A system for predicting or diagnosing a disease using an NGS analysis method, comprising the following steps: (a1) loading a computer program to be executed by one or more processors; (b1) A step of providing a memory for storing data on which genotypes of the target genome are determined; (c1) a step of generating a personal genetic analysis result report based on the stored data; and (d1) A system construction unit that builds a genetic database of the above analysis target and a report generation step that provides information from the system construction unit to the analysis target.

14. In the 13th paragraph, at least one processor in the step (a1) (i) an operational step of obtaining data by receiving nail samples, general information, or clinical information of the subject of analysis and storing the data; (ii) gDNA sample analysis operation step of aligning and mapping DNA fragments obtained through NGS from nail samples to a reference genome sequence representing the species, and (iii) Clinical information analysis operation step for calculating the expression of genotypes from the above analysis target data. A system for predicting or diagnosing a disease, comprising:

15. A system for predicting or diagnosing a disease, wherein the nail is at least one selected from the group consisting of fingernails and toenails in the 13th paragraph.

16. A system for predicting or diagnosing a disease, wherein the general information in paragraph 14 includes age; gender; family history; or lifestyle habits.

17. A system for predicting or diagnosing a disease, wherein the clinical information in paragraph 14 includes blood pressure, blood sugar, smoking, drinking, disease history, vaccine and inoculation information, regular checkup items, types and dosages of prescribed medications, side effect history, or allergic reactions.

18. A disease prediction or diagnosis system according to paragraph 13, wherein the genetic analysis result report includes an individual's constitution, customized nutrition and diet, exercise, stress relief and psychological mental health management, recommended activities, combination of health functional foods, supplements, regular health check-up cycle and method, or trackable health goals and progress.

19. A nail-based NGS analysis method according to claim 1, wherein the disease is at least one selected from the group consisting of aging, obesity, hypertension, diabetes, hyperlipidemia, cancer, sarcopenia, osteoporosis, pulmonary disease, cardiovascular disease, brain disease, liver disease, dementia, and Alzheimer's.

20. A system for predicting or diagnosing a disease, wherein the disease is at least one selected from the group consisting of aging, obesity, hypertension, diabetes, hyperlipidemia, cancer, sarcopenia, osteoporosis, pulmonary disease, cardiovascular disease, cerebrovascular disease, liver disease, dementia, and Alzheimer's disease, in the 13th paragraph.

Citation Information

Patent Citations

  • User responsive kiosk using facial recognition technology

    KR1020230001817A

  • A cosmetic composition comprising a mixed extract of cotton, bitter melon, mugunghwa and hollyhock as an active ingredient, and a method for manufacturing the same

    KR102778142B1

  • KR20200135221A

  • KR20230037339A