Method for providing information on ATM gene function

The CRISPR/Cas9-based high-throughput method with deep learning models effectively evaluates ATM gene SNVs, addressing sequencing errors and improving genetic cancer risk assessment by correlating functional scores with clinical outcomes.

WO2026106315A1PCT designated stage Publication Date: 2026-05-21UI (UNIVERSITY IND FOUNDATION) YONSEI UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
UI (UNIVERSITY IND FOUNDATION) YONSEI UNIVERSITY
Filing Date
2025-11-12
Publication Date
2026-05-21

AI Technical Summary

Technical Problem

Existing methods struggle to accurately evaluate the phenotypic impact of single nucleotide variants (SNVs) in the ATM gene, particularly due to sequencing errors and the inefficiency of high-throughput functional evaluation, leading to uncertain variant classifications and challenges in precision medicine and genetic cancer risk assessment.

Method used

A high-throughput method using CRISPR/Cas9 technology and deep learning models to introduce single nucleotide variants and synonymous mutations in the ATM gene, combined with cell viability measurements and functional scoring, to assess the impact of SNVs on ATM gene function.

Benefits of technology

This approach provides precise functional evaluation of ATM gene variants, correlating with clinical significance in cancer incidence, prognosis, and drug selection, reducing sequencing errors and improving the accuracy of genetic risk assessments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025018639_21052026_PF_FP_ABST
    Figure KR2025018639_21052026_PF_FP_ABST
Patent Text Reader

Abstract

This specification discloses a technology associated with the function of the ATM gene including a single nucleotide variant (SNV). In particular, disclosed are a method for evaluating the function of a gene including a single nucleotide variant in an ATM gene and a method for using information derived through the method.
Need to check novelty before this filing date? Find Prior Art

Description

Method of providing ATM gene function information

[0001] This specification discloses technology related to ATM gene function, including single nucleotide variants (SNVs). In particular, it discloses technology related to precision medicine or genetic cancer risk assessment.

[0002]

[0003] single nucleotide variant (SNV)

[0004] The entire human genome sequence was determined by the Human Genome Project. Consequently, it is now possible to identify the reference sequence of a human reference genome that can be used as a reference. However, not every person's genome is identical to the reference genome. When examining an individual's genome sequence, it is very common for the nucleic acid sequence of a specific gene to differ from that of the reference gene. If a gene differs in nucleic acid sequence from the reference gene, it is referred to as a genetic variant. A Single Nucleotide Variant (SNV) refers to a variation in which a single base at a specific position in the reference sequence is substituted with another base, or the variant itself. In particular, a Single Nucleotide Variant possessed by more than 1% of individuals within a specific population is sometimes classified separately as a Single Nucleotide Polymorphism (SNP).

[0005] If a specific subject possesses a particular single nucleotide variant (SNU), it can have a significant impact on its phenotype. Conversely, if a subject possesses another specific SNU, it may have no effect on its phenotype. Therefore, it is important to identify the phenotypic effects of various SNUs. However, evaluating the phenotypic impact of SNUs has presented industrial challenges due to the difficulty of creating subjects with SNUs and the vast number of SNU variant types. Traditionally, statistical methods utilizing patient genome sequence information have been primarily employed. Consequently, there were limitations in data acquisition, and the majority of SNUs in almost all genes are Variants of Uncertain Significance (VUS), whose effects remain unknown.

[0006]

[0007] ATM gene (ATM gene, Ataxia-telangiectasia mutated gene)

[0008] During biological activities, DNA damage can occur for various reasons. If DNA damage is not properly repaired, it can have a significant impact on the organism's phenotype. Accordingly, most organisms possess means to repair DNA damage. The Ataxia-telangiectasia mutated (ATM) protein is one such factor for repairing DNA damage. The ATM protein has functions related to DNA damage repair, cell cycle regulation, and oxidative stress responses. The ATM protein is also referred to as ATM serine / threonine kinase, and the ATM gene, which encodes this protein, consists of a total of 63 exons. Furthermore, exons 2 through 63 of the ATM gene encode the ATM protein.

[0009] If there is a change in the function of the ATM gene, the ATM gene with the altered function (e.g., loss of function) can affect the phenotype of the individual carrying that ATM gene. Consequently, if a single nucleotide variant is introduced into the ATM gene and this introduction causes a loss of function in the ATM gene, the aforementioned single nucleotide variant can have a severe effect on the individual's phenotype.

[0010] Accordingly, if the ATM gene has a mutation, it is important to determine the effect of that mutation on the function of the ATM gene.

[0011]

[0012] The necessity of studying single nucleotide variants in the ATM gene

[0013] As mentioned above, the occurrence of single nucleotide mutations in the ATM gene can have a significant impact on its function. Therefore, it is necessary to study the function of ATM genes containing single nucleotide mutations. However, most single nucleotide mutations that may exist in the ATM gene are VUS variants, the effects of which have not been determined.

[0014]

[0015] Limitations of Statistical Studies on Single Nucleotide Variants

[0016] Due to the widespread nature of single nucleotide variants, databases regarding them have primarily been acquired through statistics. For example, single nucleotide variants in the ATM gene, which are frequently found in cancer patients, have been classified as oncogenic or pathogenic. However, such classifications are subject to statistical errors. In reality, single nucleotide variants that do not affect gene function may be classified as oncogenic due to sampling issues. In other words, it is difficult to regard this classification as a result derived from clearly establishing a causal relationship between the specific single nucleotide variant and the patient's cancer.

[0017] Accordingly, it is necessary to study single nucleotide variants by evaluating their actual impact on the function of specific genes. For example, it is necessary to study single nucleotide variants through indicators that vary depending on the function of a specific gene.

[0018] According to the present disclosure, a method utilizing a 'high-throughput library' is proposed to evaluate the effects of single nucleotide variants introduced into specific genes.

[0019]

[0020] Efficiency of the method of creating libraries for research in a high-throughput manner

[0021] The number of types of single nucleotide variants that can be introduced into a specific gene (e.g., ATM gene) is three times the number of base pairs in the specific gene sequence.

[0022] Accordingly, to study the effect of single nucleotide variants on the function of a gene in a sequence of 1,000 base pairs, a total of approximately 3,000 single nucleotide variants must be introduced and evaluated. However, this method presents the practical difficulty of having to create approximately 3,000 genetically engineered cell populations. In particular, the inefficiency of having to perform approximately 3,000 repetitions causes this difficulty.

[0023] Therefore, a high-throughput method that evaluates the effects simultaneously using a single nucleotide variant library containing cells with various types of SNVs is efficient.

[0024]

[0025] Difficulty in distinguishing single nucleotide variants due to sequencing errors

[0026] However, to identify various types of single nucleotide variants (SNUs) simultaneously using a high-throughput method, it is necessary to know the genomic sequence information of cells within a library containing various cells to which SNUs have been introduced. Since a very large number of SNUs must be distinguished by sequence, the sequence information of each cell must be read accurately. However, currently used sequencing technology is not perfect. It is known that the probability of misidentifying a single nucleotide during sequencing is approximately 0.1%. While 0.1% is a low figure in absolute terms, it is by no means insignificant when attempting to distinguish sequence information using a high-throughput method. The above content is explained in detail through the example below.

[0027] If we study the effect of a single nucleotide variant introduced into a 200 bp gene, there are 600 possible cases for the variant to occur (3 single nucleotide variants per 200 bases x bases). Furthermore, assuming that the library contains 60,000 cells and that a single nucleotide variant has been introduced into 20% of these cells, the following is explained. Among these, assuming a situation where a specific single nucleotide variant (described as 'Variant A' for convenience) is identified:

[0028] 1) Due to sequencing errors, approximately 60,000 x 80% (proportion of wild-type cells) x 0.1% (sequencing error rate) x 1 / 3 = 16 cells will actually have wild-type nucleic acid sequences but will be judged to have 'Variant A'. 2) Assuming all single nucleotide variants are produced in equal proportions, 60,000 x 20% (proportion of variant cells) x 1 / 600 = 20 cells will be included with 'Variant A'.

[0029] Therefore, performing sequencing will yield 36 genomes with 'variant A'. Of these, approximately 44% (16 out of 36) are false positive results due to sequencing errors. This is not a negligible error.

[0030] The problem with sequencing errors is that it is difficult to distinguish whether a sequencing result is noise caused by a false positive or a true positive. For example, the 36 genomes in the above example may be the result of 0.18% sequencing error in a library where the proportion of wild-type cells is 100%.

[0031] Accordingly, in order to consider a true positive result to exist in the sequencing result despite the influence of the aforementioned false positive result, a single nucleotide variant is treated as detected only when the frequency of the single nucleotide variant is above a certain level. Therefore, even if cells with the target single nucleotide variant introduced exist in the library, they may be treated as not having been detected.

[0032] In addition, in actual experimental settings, the proportion of cells into which single nucleotide mutations are introduced may be lower than 20% of the total. Therefore, there may be a higher number of false positive results attributable to sequencing errors. Consequently, the impact of sequencing errors may be greater in actual experimental settings.

[0033] Therefore, in order to use the 'high-throughput method,' the problem of these sequencing errors must be resolved.

[0034]

[0035] Necessity of functional evaluation for all possible single nucleotide variants in the ATM gene

[0036] The majority of single nucleotide mutations in the ATM gene are uncertain variants (VUS). The technical problem that this specification aims to solve is to evaluate the aforementioned uncertain variants as well. That is, this specification aims to evaluate the impact of all possible single nucleotide mutations on the function of the ATM gene. Hereinafter, the evaluation of the impact on ATM gene function is referred to as the functional evaluation. To this end, the inventors of this application designed a high-throughput functional evaluation experiment using CRISPR / Cas9 technology (prime editing).

[0037] However, due to various practical limitations, introducing 'all single nucleotide mutations' into a library can be very difficult in some cases. Accordingly, the inventors of the present application recognized the need to solve the problem of being unable to perform functional evaluations on single nucleotide mutations that were not introduced.

[0038] Therefore, a method is needed to evaluate the function of single nucleotide mutations located in areas with low prime editing efficiency.

[0039]

[0040] Simultaneous introduction of a single nucleotide variant and an intended synonymous mutation into the ATM gene

[0041] The inventors of the present application sought to minimize the impact of errors when sequencing genomic DNA by simultaneously introducing single nucleotide variants and intended synonymous mutations into the ATM gene.

[0042] In this case, synonymous mutations do not affect the translation outcome of ATM genes. Therefore, a genome with one or more intended synonymous mutations in addition to a single nucleotide polymorphism can be treated identically to a genome with only a single nucleotide polymorphism.

[0043] The probability of sequencing errors occurring at two or more bases is approximately 0.1% × 0.1% = 0.0001%. Therefore, if one or more intended synonymous mutations are added to a single nucleotide variant, the probability of false positives being included when reading the sequence becomes very low. Even when creating a library using a high-throughput method and attempting to detect each single nucleotide variant, the frequency of single nucleotide variants that can be said to be detected becomes very low.

[0044]

[0045] Functional evaluation of ATM genes with single nucleotide mutations

[0046] The inventors constructed a single nucleotide variant library using cells with single nucleotide variants and intended synonymous mutations in the ATM gene.

[0047] Furthermore, the inventors confirmed that the viability or survival rate of cells possessing a single nucleotide variant in the ATM gene can be used to evaluate the function of the ATM gene having a single nucleotide variant. This is because when ATM gene function is reduced, the survival rate (or viability) of cells possessing the corresponding gene decreases. In particular, in the presence of a PARP inhibitor, the survival rate (or viability) of cells with reduced ATM gene function decreases. Accordingly, the inventors confirmed that if the viability (or viability) of cells decreases after introducing a specific single nucleotide variant into the ATM gene, it can be treated as if ATM gene function has decreased.

[0048] The inventors evaluated the impact of most single nucleotide variants on ATM gene function through the aforementioned single nucleotide variant library and cell viability measurements. Additionally, they quantified the impact of each single nucleotide variant on ATM gene function to obtain a function score.

[0049]

[0050] Deep learning model for evaluating the function of ATM genes with single nucleotide polymorphisms

[0051] The inventors evaluated the impact of most single nucleotide variants that may exist in the ATM gene on ATM gene function using a high-throughput method. However, due to issues such as the efficiency of prime editing, there were difficulties in evaluating the function of some single nucleotide variants.

[0052] The inventors designed a deep learning model to evaluate the function of single nucleotide mutations located in regions with low prime editing efficiency. The deep learning model uses mutation information and function scores for each single nucleotide mutation obtained through experiments as input data. In addition, to consider the correlation between each amino acid residue, the deep learning model was designed with a transformer structure.

[0053] The inventors used the above model to obtain the function scores of all possible single nucleotide variants of the ATM gene. Accordingly, the inventors obtained the function scores for all possible single nucleotide variants of the ATM gene.

[0054]

[0055] Clinical Significance of Functional Evaluation Results of ATM Genes with Single Nucleotide Variants

[0056] To confirm the correlation between the function scores obtained through the method described above and actual clinical significance, the inventors compared the obtained function scores for each single nucleotide variant with a database of previously known single nucleotide variants.

[0057] The inventors confirmed, using existing databases, that the inclusion of single nucleotide variants with low function scores (i.e., reduced function of the ATM gene) increases the probability of developing cancer. Furthermore, the inventors confirmed, by comparing with existing databases, that the inclusion of single nucleotide variants with low function scores alters the prognosis for specific cancers. Additionally, the inventors confirmed that the inclusion of single nucleotide variants with low function scores reduces the frequency of those variants within the population.

[0058] The inventors confirmed that the function score obtained through the above method has clinical significance.

[0059]

[0060] By utilizing the results of the functional evaluation of ATM genes having single nucleotide variants disclosed in this specification, information regarding the function of ATM genes in individuals containing specific single nucleotide variants can be provided. In particular, information can be provided regardless of which single nucleotide variant the ATM gene of the subject contains.

[0061] In addition, it can provide information related to cancer incidence rates, cancer prognosis, selection of anticancer drugs, and diagnosis of ataxia telangiectasia syndrome.

[0062]

[0063] Figure 1 shows the structure of the ATM gene, indicating which ATM protein domains are encoded by each exon of the ATM gene. Each box in the ATM gene represents an exon. TAN stands for Tel1 / ATM N-terminal motif. FAT stands for FRAP-ATM-TRRAP domain.

[0064] Figure 2 shows the ClinVar classification for single nucleotide variants that may exist in each exon of the ATM gene. The upper graph shows the proportion of ClinVar classifications for each single nucleotide variant that are categorized as "benign / likely benign," "uncertain variant (VUS)," and "pathogenic / likely pathogenic." The lower graph shows the proportion of single nucleotide variants classified as pathogenic / likely pathogenic in ClinVar that exist in each exon.

[0065] Figure 3 shows the ratio of ATM-haploid-KO cells over time when ATM-haploid cells and ATM-haploid cells are cultured together.

[0066] Figure 4 shows the results of sequencing the ATM gene of ATM-haploid-KO cells.

[0067] Figure 5 shows the results of sequencing the ATM gene of HCT116 cells.

[0068] Figure 6 illustrates a method for producing and selecting ATM-haploid cells using HCT116 cells.

[0069] Figure 7 shows a schematic diagram of the process of creating ATM haploid cells using HCT116 cells.

[0070] Figure 8 shows the electrophoresis results for selecting cells from which the ATM gene containing c.3380C>T was removed during the process of creating ATM-haploid cells using HCT116 cells.

[0071] Figure 9 shows the sequencing results for selecting cells in which the ATM gene containing c.3380C>T was removed during the process of creating ATM-haploid cells using HCT116 cells.

[0072] Figure 10 shows the results of sequencing the ATM gene of ATM-haploid cells.

[0073] The upper graph in Fig. 11 represents the proportion of reads containing additional indels among reads with mutations identical to the target single nucleotide variant during the sequencing of a single nucleotide library. The lower graph represents the proportion of reads containing additional indels among reads without mutations identical to the target single nucleotide variant.

[0074] Figure 12 is a schematic diagram illustrating the process of conducting high-throughput experiments using a single nucleotide mutation library.

[0075] Figure 13 shows the correlation between the number of repetitions in the groups with and without olaparib treatment. The Pearson correlation coefficient (r) is also indicated.

[0076] Figure 14 shows the correlation between the olaparib-treated group and the DMSO-treated group. The Pearson correlation coefficient (r) is also indicated.

[0077] Figure 15 shows the ROC curves (Receiver-operating-characteristic curves) of sLFC values ​​determined by single nucleotide mutations in experiments using olaparib and DMSO. The graph on the left compares nonsense mutations (n ​​= 1,141) with synonymous mutations (n ​​= 4,837). Exons 62 and 63 were removed to exclude the effects of nonsense-mediated decay. The graph on the right compares pathogenic / pathogenically probable mutations (n ​​= 440) with benign / benignly probable mutations (n ​​= 1,163). The area under the curve (AUC) is also shown below.

[0078] Figure 16 compares the ROC curves for the results with and without olaparib for 17 variants that were annotated by experts and had nonsense variants removed.

[0079] Figure 17 shows the kernel density estimation plot of sLFCs of single nucleotide variants, classified by the type of single nucleotide variant and by the presence or absence of olaparib treatment. Additionally, each graph shows the number and proportion of single nucleotide variants lower than the values ​​obtained through the valid index (DMSO: -0.745 and olaparib: -0.912). In each graph, the upper values ​​represent the number and proportion of single nucleotide variants for DMSO, and the lower values ​​represent the number and proportion of single nucleotide variants for olaparib.

[0080] Figures 18 to 21 show the correlation of sLFCs between internal repeats introducing the same intended single nucleotide variant. Pearson correlation coefficients (r) are also indicated.

[0081] Figure 22 shows the correlation of sLFC between single nucleotide mutations that induce the same amino acid mutation. The Pearson correlation coefficient (r) is also indicated.

[0082] Figure 23 shows the results of Western blot experiments for ATM, phosphorylated ATM (p-ATM), and phosphorylated CHK2 (p-CHK2) in ATM haploid cells carrying the K331E or L969P mutation. Arrows indicate the molecular weight of the proteins. GAPDH was used as a control.

[0083] Figure 24 shows the results of verifying potential off-target effects. The numbers written above represent protospacer sequences from 1 to 19 and NGG PAM sequences from 20 to 22. Possible mismatches are indicated by the shade of the base.

[0084] Figure 25 shows the correlation between the function score and the score predicted by a known computational model. The Pearson correlation coefficient (r) is also indicated.

[0085] Figure 26 shows the correlation between the mean of the PhyloP score and the proportion of non-functional variants among missense variants by exon. The Pearson correlation coefficient (r) is also indicated.

[0086] Figure 27 shows the BLOSUM62 scores for each function score by classification using a violin plot.

[0087] Figure 28 shows the correlation between the BLOSUM62 score and the function score. Trends through linear regression are also indicated. In this case, the brightness of each dot was determined by neighboring dots. That is, it indicates the presence of dots located at a distance of 1.5 times the base radius.

[0088] Figure 29 shows the distribution of function scores for each type of variant. Introns represent +5, +4, and +3 in the 5' direction or -5, -4, and -3 in the 3' direction in the axon / intron adjacency region. Splice donor / recipient (splice AD) positions represent +2 and +1 in the 5' direction or -2 and -1 in the 3' direction in the axon / intron adjacency region. Boxes indicate the 1st, 2nd, and 3rd quartiles. Additionally, whiskers represent the top 10 or 90 percent.

[0089] Figure 30 shows the ratio of function scores by type of variation.

[0090] Figure 31 shows the ratio of functional scores by amino acid substitution.

[0091] Figure 32 shows the effect of the functional score distribution by type of amino acid substitution. NU (non-polar uncharged); NC (negative-charged); PC (positive-charged); PU (polar uncharged).

[0092] FIG. 33 shows a map of function scores for single nucleotide mutations in exons 2 to 63. The visibility of a dot is determined by neighboring dots, that is, the presence of dots located at a distance of 1.5 times the basic radius.

[0093] Figure 34 shows a function point map existing in the intron and splice AD ​​regions. The dotted lines represent exons existing between introns.

[0094] Figure 35 shows the ratio of functional score classifications for each exon. In addition, the solid line represents the average phyloP score for each exon.

[0095] Figure 36 shows the resistance to missense mutations in the ATM structure. The average functional score of the missense mutation at each position is indicated in bright light (from -5 to 1). In the enlarged box area, the p53 peptide and ANP (phosphoaminophosphate-adenylic acid ester, a synthetic analog of ATP) are indicated separately. Magnesium ions are also indicated by dots. The portions encoded in exons 59 and 60 of the amino acid residues are indicated by bars.

[0096] Figure 37 shows the results of classifying variants classified as positive / potentially positive (B / LB) and pathogenic / potentially pathogenic (P / LP) in ClinVar using functional scores.

[0097] Figure 38 shows a box plot of the functional scores of single nucleotide polymorphisms in the splice receptor and donor. The functional categories presented in the ACMG (American College of Medical Genetics and Genomics) guidelines are also indicated on the x-axis. PVS1 signifies "Pathogenic Very Strong Evidence," and N / A signifies "Not Applicable." The predicted functions of ATM decrease in the following order: PVS1 N / A > PVS1-Supporting > PVS1-Strong > PVS1.

[0098] Figure 39 shows the frequency and function score for each variant in the group in gnomAD v.4.1. In this case, the shape of the dot is displayed differently according to the ClinVar classification.

[0099] Figure 40 shows the cumulative cancer incidence rate according to the classification of mutation types and functions.

[0100] Figure 41 shows the hazard ratios for the values ​​and function scores of various computational models. The black bars represent the 95% confidence intervals.

[0101] In Figures 42 through 44, the left panel shows results including all single nucleotide variant types, while the right panel shows results including information only on missense variants and intact ATMs. P-values ​​are displayed comparing the non-functional group and the intact ATM group. Figure 42 shows the lifelong cancer incidence according to various mutation types among participants in the UK Biobank. Figure 43 shows the cumulative breast cancer incidence according to various mutation types among participants in the UK Biobank. Additionally, Figure 44 shows the lifetime breast cancer risk according to various mutation types among participants in the UK Biobank. The participants were classified based on their functional scores.

[0102] Figure 45 illustrates the functional subsets of missense variants and their association with occurrence as germline variants in breast cancer patients. The pathogenic variant subsets were determined using values ​​derived from computational model scores or functional scores in AlphaMissense, REVEL, and CADD. Odds ratios were calculated by comparing the frequency of each pathogenic variant subset in tumor samples with that of the benign variant subsets. Black bars represent 95% confidence intervals.

[0103] Figure 46 shows the kernel density estimation plot of the function scores of single nucleotide variants found in tumor sequencing data in the OncoKB dataset.

[0104] Figure 47 shows the correlation of sLFC with variants classified as carcinogenic or potentially carcinogenic in the GENIE data among repeats of the olaparib treatment group. In this case, the left graph includes only variants classified as carcinogenic or potentially carcinogenic. Additionally, the right graph includes all variants. Furthermore, in order to distinguish the variants indicated in the left graph, which confirmed the correlation, from the right graph for all variants, the variants indicated in the left graph are marked with triangular dots in both graphs.

[0105] Figure 48 shows the correlation of sLFC between repeats in the olaparib treatment group for variants classified as pathogenic or potentially pathogenic in the ClinVar data. In this case, the left graph includes only variants corresponding to pathogenicity or potentially pathogenicity. Additionally, the right graph includes all variants. Furthermore, in order to distinguish the variants indicated in the left graph, which confirmed the correlation, from the right graph for all variants, the variants indicated in the left graph are marked as triangular dots in both graphs.

[0106] Figure 49 shows the correlation of sLFCs for non-functionally classified variants among repeats in the olaparib treatment group. In this case, the left graph includes only variants corresponding to pathogenic or potentially pathogenic variants. Additionally, the right graph includes all variants. Furthermore, in order to distinguish the variants indicated in the left graph, which confirmed the correlation, from the right graph for all variants, the variants indicated in the left graph are marked as triangular dots in both graphs.

[0107] Figure 50 shows the number of occurrences and function scores found in tumor samples by mutation.

[0108] Figure 51 illustrates the association between missense variants and their occurrence in tumor samples. Pathogenic variant subsets were determined using values ​​derived from computational model scores or function scores in AlphaMissense, REVEL, and CADD. Odds ratios were calculated by comparing the frequency of each pathogenic variant subset in tumor samples with that of the benign variant subsets. Black bars represent 95% confidence intervals.

[0109] Figure 52 shows the odds ratios for non-functional single nucleotide variants. Variations in the ratios for each non-functional single nucleotide variant were induced by varying the values ​​used to distinguish non-functional single nucleotide variants. Each dotted line represents values ​​corresponding to 20% and 30% for non-functional single nucleotide variants. This is the same as the ratio of non-functional missense single nucleotide variants in this study and the GENIE tumor sequencing data.

[0110] Figures 53 and 54 show the prognosis of cancer patients with different types of single nucleotide mutations. Patients with chronic lymphocytic leukemia (CLL, n = 900) and stage 3 and 4 bladder cancer (n = 623) were classified into the following three categories: reducing mutations (non-functional, intermediate), functional mutations, and wild-type ATM. Survival analysis was performed using the Kaplan-Meier estimator. The p-values ​​for survival comparisons with the intact ATM group for functional or reducing mutations are shown. FFS (Failure-Free Survival), overall survival, and PFS (Progression-Free Survival)

[0111] Figures 55 and 56 show the number of A and T bases and the extent of the NGG PAM sequence for various genes.

[0112] Figure 57 shows an overall schematic diagram of the DeepATM model.

[0113] Figure 58 shows the Pearson correlation coefficients based on whether structural information is included (whether coordinate embedding is included).

[0114] Figure 59 shows the evaluation through a random forest model in addition to the Pearson correlation coefficient based on the inclusion of structural information.

[0115] Figure 60 shows the distribution of eDA scores for 23,092 single nucleotide variants for which values ​​were obtained experimentally and for the remaining 4,421 single nucleotide variants. In this case, a graph showing the eDA scores and function scores for experimentally obtained SNVs and SNVs not obtained experimentally is shown separately, and a graph showing them together is shown.

[0116] Figure 61 shows the correlation between eDA (predicted function score) and function score for 23,092 single nucleotide variants. -1.360 and -0.912, which correspond to the classification criteria for function scores, are indicated by dotted lines.

[0117] Figure 62 shows the kernel density estimation plot of eDA scores for unevaluated single nucleotide variants classified as positive / likely positive (B / LB) or pathogenic / likely pathogenic (P / LP) in ClinVar.

[0118] Figure 63 shows the ROC curves for 116 single nucleotide variants in the test dataset. The ROC curves for each test dataset are displayed separately, or a graph showing them together is displayed.

[0119] Figure 64 shows the ROC curves for 68 single nucleotide variants in the test dataset. In this case, the 68 single nucleotide variants are single nucleotide variants classified as positive / positive probability (B / LB) or pathogenic / pathogenic probability (P / LP) in ClinVar, and are classified with two or more stars. The ROC curves for each test dataset are displayed separately, or a graph showing them together is presented.

[0120] Figure 65 shows the ROC curves for 59 single nucleotide variants in the test dataset. In this case, the 59 single nucleotide variants are single nucleotide variants located in regions other than the kinase domain. ROC curves for each test dataset are displayed separately, or a graph showing them together is displayed.

[0121] Figure 66 shows the ROC curves for 240 unevaluated single nucleotide variants that are two or more stars in ClinVar. AUC values ​​are also indicated.

[0122] Figure 67 shows the ROC curves for 455 unevaluated single nucleotide variants that are one or more star single nucleotide variants in ClinVar. AUC values ​​are also indicated.

[0123] Figure 68 shows the distribution of eDA scores from exon 2 to exon 63.

[0124] Figure 69 shows the frequency of single nucleotide variants and eDA scores in gnomAD v.4.1 and the UK Biobank. ClinVar classification is represented through the form of dots.

[0125] Figure 70 shows the cumulative cancer incidence in the UBK data (n = 323,897) using only single nucleotide variants that were not experimentally evaluated. The left graph includes all variants. The right graph includes information only on missense single nucleotide variants and intact ATMs. The p-values ​​for comparison with the intact group are also indicated.

[0126] In Figures 71 through 73, the left panel shows the results including all single nucleotide variant types, while the right panel shows the results including information only on unevaluated missense variants and intact ATMs. P-values ​​are displayed comparing the non-functional and intermediate groups with the intact ATM group. Figure 71 shows the lifelong cancer incidence according to various variant types among participants in the UK Biobank. Figure 72 shows the cumulative breast cancer incidence according to various variant types among participants in the UK Biobank. Additionally, Figure 73 shows the lifetime breast cancer risk according to various variant types among participants in the UK Biobank. The participants were classified based on their eDA scores.

[0127] Figure 74 illustrates the association between missense variants and their occurrence in tumor samples. Pathogenic variant subsets were determined using computational model scores or eDA scores from AlphaMissense, REVEL, and CADD. Odds ratios were calculated by comparing the frequency of each pathogenic variant subset in tumor samples with that of the benign variant subsets. Black bars represent 95% confidence intervals.

[0128] Figure 75 shows the odds ratios for non-functional single nucleotide variants. Variations in the ratios for each non-functional single nucleotide variant were induced by varying the values ​​used to distinguish non-functional single nucleotide variants. Each dotted line represents values ​​corresponding to 20% and 30% for non-functional single nucleotide variants. This corresponds to the ratios of non-functional missense single nucleotide variants in this study and the GENIE tumor sequencing data.

[0129] Figure 76 shows the distribution of function scores by amino acid variant through a heatmap.

[0130] Figure 77 shows the kernel density estimation plot of the integrated score for all variants classified as positive / positive probability (B / LB) (n = 2,560) and pathogenic / pathogenic probability (P / LP) (n = 690) present in the ATM encryption sequence in ClinVar.

[0131] Figure 78 shows the combined score for the frequency of single nucleotide variants in gnomAD v.4.1.

[0132] Figure 79 shows the kernel density estimation plot of the combined scores of single nucleotide variants found in tumor sequencing data in the OncoKB dataset. -0.912 is indicated by the vertical line.

[0133] Figure 80 shows the combined score for the number of times it was found in tumor samples. The four variants most frequently observed in tumor samples and the variant strongly associated with breast cancer (c.7271T>G) are indicated by arrows.

[0134] Figure 81 illustrates the association between missense variants and their occurrence in tumor samples. Pathogenic variant subsets were determined using values ​​derived from computational model scores or integrated scores in AlphaMissense, REVEL, and CADD. Odds ratios were calculated by comparing the frequency of each pathogenic variant subset in tumor samples with that of the benign variant subsets. Black bars represent 95% confidence intervals.

[0135] Figure 82 shows the odds ratios for non-functional single nucleotide variants. Variations in the ratios for each non-functional single nucleotide variant were induced by varying the values ​​used to distinguish non-functional single nucleotide variants. Each dotted line represents values ​​corresponding to 20% and 30% for non-functional single nucleotide variants. This corresponds to the ratios of non-functional missense single nucleotide variants in this study and the GENIE tumor sequencing data.

[0136] Figure 83 shows the prognosis of cancer patients with different types of single nucleotide mutations. Patients with chronic lymphocytic leukemia (CLL, n = 900) and stage 3 and 4 bladder cancer (n = 623) were classified into the following three categories: reducing mutations (non-functional, intermediate), functional mutations, and wild-type ATM. Survival analysis was performed using the Kaplan-Meier estimator. The p-values ​​for survival comparisons with the intact ATM group for functional or reducing mutations are shown. FFS (Failure-Free Survival), overall survival, and PFS (Progression-Free Survival)

[0137] Figure 84 shows the results of examining the cumulative cancer incidence rate in ATM UKB participants (n = 458,524) determined by the integrated score. The left side includes participants for all single nucleotide variants. The middle side includes only participants for missense variants or intact ATM. The right side includes only participants for VUS or intact ATM. P-values ​​are displayed in comparison to the intact group.

[0138] Figure 85 shows the hazard ratios for the values ​​of various calculation models and the integrated score. The black bars represent the 95% confidence intervals.

[0139] In Figures 86 and 88, the left panel shows results including all single nucleotide variant types, while the right panel shows results including information only on unevaluated missense variants and intact ATMs. P-values ​​are displayed comparing the non-functional and intermediate groups with the normal ATM group. Figure 86 shows the lifelong cancer incidence according to various variant types among participants in the UK Biobank. Figure 87 shows the cumulative breast cancer incidence according to various variant types among participants in the UK Biobank. Additionally, Figure 88 shows the lifetime breast cancer risk according to various variant types among participants in the UK Biobank. The participants were classified based on their composite scores.

[0140] Figure 89 shows the risk ratios for various cancer types in participants including ATM variants in each functional class. The functional classes were determined through integrated scores. Additionally, the risk ratios were adjusted for sex and age. Each cancer was defined according to the ICD-10. The vertical bars represent 95% confidence intervals.

[0141] Figure 90 shows the results of sequencing the ATM gene of ATM-semi-KO cells.

[0142] Figure 91 shows the results of calculating the sLFC values ​​of single nucleotide mutations in exons 55 and 56 using ATM-haploid cells and ATM-semi-KO cells.

[0143] Figure 92 shows the proportion of cells with single nucleotide variants when wild-type cells and cells with single nucleotide variants are co-cultured. It was measured based on the presence or absence of olaparib. The p-values ​​compared to D0 for each experiment are also indicated. ns indicates not significant.

[0144] Figure 93 shows the proportion of cells with single nucleotide variants when wild-type cells and cells with single nucleotide variants are co-cultured. It was measured based on the presence or absence of olaparib. The p-values ​​compared to D0 for each experiment are also indicated. ns indicates not significant.

[0145] Figure 94 shows the correlation between LFC values ​​obtained using various PARP inhibitors.

[0146]

[0147] The best mode for carrying out the invention is disclosed below. However, the present application is not limited to the embodiments disclosed below, and the embodiments described in this paragraph are merely examples. Any variations that a person skilled in the art could conceive of regarding the examples described in this paragraph should be deemed to be included. Throughout the specification, when a part is described as "comprising" a certain component, this means that, unless specifically stated otherwise, it does not exclude other components but may include additional components.

[0148] In one embodiment, the present specification discloses a method for providing ATM gene function information comprising the following.

[0149] (a) Acquire ATM gene sequence information of the subject;

[0150] (b) determining whether the subject contains a specific single nucleotide variant; and

[0151] (c) Provide ATM gene function information based on the above mutation information.

[0152] In another embodiment, the present specification discloses a method for providing ATM gene function information comprising the following.

[0153] (a) Acquire ATM protein sequence information of the subject;

[0154] (b) determining whether the ATM protein of the subject contains a specific amino acid variation; and

[0155] (c) Provide ATM gene function information based on the above mutation information.

[0156] In this case, the specific single nucleotide variant may be a single nucleotide variant selected from among the single nucleotide variants listed in Table 1 of this specification. Additionally, the specific amino acid variant may be an amino acid variant selected from among the amino acid variants listed in Table 1 of this specification.

[0157] At this time, the ATM gene function information may be information regarding whether the function of the ATM gene is in a dysfunctional state.

[0158] In another embodiment, the present specification discloses a method for providing ATM gene function information comprising the following.

[0159] (a) Acquire ATM gene sequence information of the subject;

[0160] (b) determining whether the subject includes at least one of specific single nucleotide variants; and

[0161] (c) Provide ATM gene function information based on the above mutation information.

[0162] In another embodiment, the present specification discloses a method for providing ATM gene function information comprising the following.

[0163] (a) Acquire ATM protein sequence information of the subject;

[0164] (b) determining whether the ATM protein of the subject contains at least one of specific amino acid variants; and

[0165] (c) Provide ATM gene function information based on the above mutation information.

[0166] In this case, the specific single nucleotide variations may be the single nucleotide variations listed in Table 1 of this specification or parts thereof. Additionally, the specific amino acid variations may be selected amino acid variations among the amino acid variations listed in Table 1 of this specification or parts thereof.

[0167] At this time, the ATM gene function information may be information regarding whether the function of the ATM gene is in a dysfunctional state.

[0168]

[0169] Definition of Terms

[0170] In this specification, "mutation" and "variant" may be used interchangeably. However, "intended synonymous mutation" and "synonymous variant" are distinguished from each other in this specification. In this specification, an intended synonymous mutation refers to a mutation additionally introduced along with a single nucleotide variant to resolve sequencing error issues. Furthermore, in this specification, a synonymous variant refers to a single nucleotide variant that does not cause a change in protein sequence.

[0171] In this specification, "dysfunctional" and "non-functional" may be used interchangeably. In this case, "dysfunctional" refers to a state in which a gene fails to properly perform the function it is originally supposed to perform. Alternatively, "dysfunctional" refers to a state in which a gene performs a function different from its original function. In this specification, "functional" refers to a state in which a gene properly performs the function it is originally supposed to perform. In this specification, "intermediate" refers to an intermediate state that is not classified as "functional" or "dysfunctional."

[0172] In this specification, "reduction in gene function" means that a specific function of a gene has become unable to operate normally.

[0173] In this specification, "haploidize" means manipulating cells with a ploidy of 2 or more for a specific gene to have only one allele. It also means removing alleles other than one from cells with a ploidy of 2 or more for a specific gene.

[0174] In this specification, "pathogenic" means capable of causing disease in a cell or organism. It also means a property that causes death in a cell or organism. In this specification, "benign" means a property that does not cause disease in a cell or organism. It also means a property that does not cause death in a cell or organism.

[0175] In this specification, "coding" means that a specific nucleic acid or gene possesses information regarding a specific protein or a specific nucleic acid. Furthermore, it means that the coding nucleic acid or gene can produce the specific nucleic acid or a specific protein through translation. The nucleic acid and gene are not limited to specific sequences as long as they can produce the specific nucleic acid or a specific protein through translation.

[0176] In this specification, "subject" may refer to an individual containing a gene for which an evaluation of the gene's function is required. Alternatively, it may refer to a tissue or cell derived from the individual. For example, the subject may be a breast cancer patient, a prostate cancer patient, a bladder cancer patient, a leukemia patient, or a patient with ataxia telangiectasia syndrome. For another example, the subject may be a healthy person for whom an evaluation of the gene's function is to be performed. Alternatively, the subject may be a cell or tissue for which an evaluation of the gene's function is to be performed.

[0177]

[0178] Function of the ATM gene

[0179] The function of the ATM gene described herein refers to a function related to the repair of DNA double-strand breaks (DSBs). In this case, the function of the ATM gene is not limited to any function related to DSB repair. For example, it refers to a function in which the ATM protein, guided to the site of a DNA double-strand break by the MRN complex (a complex of MRE11, RAD50, and NBS1), operates to repair the DSB. For another example, it refers to a function of inducing homologous recombination repair by phosphorylating proteins such as BRCA1, BRCA2, and RAD51. For yet another example, it refers to a function of phosphorylating p53 to induce long-term cell cycle arrest or apoptosis.

[0180]

[0181] Definitions related to single nucleotide mutations

[0182] Expressions related to single nucleotide mutations

[0183] In this specification, a Single Nucleotide Variant (SNV) refers to a variation in which a single base at a specific position in a reference sequence is substituted with another base. Expressions related to single nucleotide variants are defined below.

[0184] When sequencing the gene sequence of a subject to verify the presence or absence of a single nucleotide variant, the nucleotides at specific positions may differ from the reference sequence. In this case, the gene may be described as containing (comprise) or having a specific single nucleotide variant. Additionally, the subject may be described as containing, having, or carrying a specific single nucleotide variant.

[0185] Furthermore, when sequencing a gene, the bases at multiple specific positions may differ from the reference sequence. In this case, the gene can be described as containing or possessing multiple specific single nucleotide variants. Additionally, the subject can be described as containing, possessing, or carrying multiple specific single nucleotide variants. Thus, if sequencing results show differences in multiple bases, the gene can be described as containing or possessing a combination of single nucleotide variants. The above explanation is intended to more clearly designate the types of variants possessed by the gene. Moreover, the above explanation does not apply to the concept of pseudo-single nucleotide variants. That is, if a gene differs from the reference sequence in the base corresponding to a single nucleotide variant and the base corresponding to a single intended synonymous mutation, the gene can be described as containing or possessing a single single nucleotide variant, but it is not described as containing a combination of single nucleotide variants. This must be interpreted appropriately depending on the type of data to be acquired.

[0186] The fact that a gene contains a specific single nucleotide variant means that the gene may contain other specific single nucleotide variants. In other words, even if a gene contains a combination of single nucleotide variants, it can also be expressed as containing a specific single nucleotide variant.

[0187] When a subject has polyploidy or greater polyploidy with respect to a specific gene, the sequences of the alleles contained in the subject may differ. In this case, if even one of the subject's alleles contains a specific single nucleotide variant, the subject may be described as containing a specific single nucleotide variant. Additionally, the subject's gene may be described as containing a specific single nucleotide variant.

[0188] If a subject exhibits polyploidy or greater polyploidy with respect to a specific gene, the sequences of the alleles contained within the subject may differ. In this case, if not all alleles of the subject contain a specific single nucleotide variant, the subject can be described as "carrying" a specific single nucleotide variant. Additionally, the subject's gene can be described as "carrying" a specific single nucleotide variant. If not all alleles of the subject contain a specific group of single nucleotide variants, the subject can be described as "carrying" a specific group of single nucleotide variants. Carrying a specific single nucleotide variant implies that it may contain additional specific single nucleotide variants.

[0189] In addition, if all alleles contained in the subject include at least one single nucleotide variant of a specific group, it can be expressed that each allele of the subject includes a single nucleotide variant of a specific group.

[0190] The following is explained in more detail through examples. First, assume a case where the subject contains only alleles A and B. Additionally, as an example, assume that the subject's allele A contains a single nucleotide variant a and the subject's allele B contains a single nucleotide variant b. In such a case, it can be expressed that the subject contains or carries a single nucleotide variant a. Additionally, it can be expressed that the subject contains or carries a single nucleotide variant b. Furthermore, if both a single nucleotide variant a and b are single nucleotide variants belonging to a specific group (e.g., single nucleotide variants classified as dysfunctional single nucleotide variants), it can be expressed that each allele of the subject contains a single nucleotide variant belonging to a specific group. As another example, assume a case where both allele A and allele B contain a single nucleotide variant a. In such a case, it can be expressed that all alleles of the subject contain a single nucleotide variant a.

[0191] A single nucleotide variant is expressed by its position and / or the type of base substitution. In this context, the position refers to the specific base location within a gene or genome where the variant exists. Therefore, even if the type of base substitution is the same, different positions constitute different types of single nucleotide variants. Furthermore, even if the positions are the same, different types of base substitutions constitute different types of single nucleotide variants.

[0192] The number of single base variants that can exist in a specific region refers to the number of all different types of single base variants that can exist in that region. For example, if the length of a specific region is 100 bp, the number of single base variants that can exist in that region is 300 (100 x 3 types of bases excluding the original base).

[0193]

[0194] Criteria for determining the presence of a single nucleotide mutation

[0195] A single nucleotide variant is defined based on a reference sequence. In this case, the reference sequence is not limited to any specific sequence, as long as it is a sequence comparable to the subject's gene sequence. The presence or absence of a single nucleotide variant may vary depending on the reference sequence. For example, assume that sequencing the ATM gene of subject A reveals a one-nucleotide difference from reference sequence 1. In this case, subject A can be said to contain a single nucleotide variant. Furthermore, assume that sequencing the ATM gene of subject A reveals a sequence identical to reference sequence 2. In this case, subject A does not contain a single nucleotide variant.

[0196]

[0197] Single nucleotide variation nomenclature

[0198] Single nucleotide variants may be named based on a reference sequence. Additionally, single nucleotide variants are named according to their type. Unless otherwise stated, the nomenclature described herein expresses single nucleotide variants according to the HGVS nomenclature. It will be obvious to a person skilled in the art that the HGVS nomenclature described herein can be appropriately interpreted. The expressions primarily used in this specification according to the HGVS nomenclature are as follows: "c" signifies the coding sequence of the reference DNA sequence. Additionally, the natural number following "c" signifies the position of the base. For example, "c.73" signifies the 73rd base sequence relative to the coding sequence of the reference DNA sequence. Furthermore, "+ natural number" and "- natural number" following the above natural number are used to indicate bases other than those in the coding sequence. "+ natural number" signifies the number of bases in the downstream sequence. "- natural number" signifies the number of bases in the upstream sequence. For example, assume that "c.72" is the last base of exon 2 (the base at the 3' end) and "c.73" is the first base of gene exon 3 (the base at the 5' end). In this case, "c.72+1" means the first base of intron 2 (the base at the 5' end). Also, "c.73-1" means the last base of intron 2 (the base at the 3' end). Additionally, a negative integer following "c" is used to denote an upstream sequence of the coding sequence. For example, "c.-1" means the first sequence among the upstream sequences of the coding sequence. Also, ">" is used to express the type of base substitution of a single nucleotide variant. For example, "c.73A>C" means that the 73rd base of the coding sequence of the reference sequence is A (adenine), while the 73rd base of the coding sequence of the target is C (cytosine).

[0199] In addition to this, single nucleotide variants can be named in various ways, provided that the base position in the referenced database can be appropriately expressed. It will be self-evident that an appropriate nomenclature can be properly interpreted by a person skilled in the art.

[0200]

[0201] Definitions related to pseudo-single nucleotide variants

[0202] Overview of Pseudo-Single Nucleotide Variants

[0203] Hereinafter, pseudo-single nucleotide variants, which are one of the key concepts used in this specification, will be described.

[0204] A pseudo-single nucleotide variant refers to a variant consisting of one single nucleotide variant and one or more intended synonymous mutations. In this case, the synonymous mutations do not affect the translation outcome of the ATM gene. Therefore, a genome containing one or more intended synonymous mutations in addition to a single nucleotide variant can be treated identically to a genome containing only a single nucleotide variant.

[0205] When classifying types of pseudo-single nucleotide variants (PSVs), intended synonymous mutations do not have a significant impact. In other words, PSVs are classified based on the location and type of the single nucleotide variant. For example, PSVs with the same type of single nucleotide variant but different intended synonymous mutations can be classified as the same type of PSV. This is because if a PSV affects gene function, it is likely due to the influence of the variant's location and type. A single type of single nucleotide variant can constitute the same PSV along with various types of intended single nucleotide variants. Accordingly, research on a specific single nucleotide variant can be conducted using multiple PSVs of the same type.

[0206]

[0207] Reason for introducing pseudo-single nucleotide mutations

[0208] As described above, the reason for inducing one or more intended synonymous mutations along with a single nucleotide variant is to address the effects of sequencing errors. In order to distinguish between false positive and true positive results caused by sequencing errors, in actual experiments, if the magnitude of the acquired signal does not exceed a certain level, it is classified as noise. By additionally inducing intended synonymous mutations, it becomes easier to acquire a signal greater than the noise level.

[0209]

[0210] Location of single nucleotide variants and intended synonymous mutations

[0211] In principle, intended synonymous mutations can only exist within the gene's coding sequence. That is, intended synonymous mutations cannot exist in introns or untranslated regions (UTRs). This is because, unlike coding sequences where the effects of mutations can be relatively predicted through codons, the effects on introns and similar regions cannot be predicted.

[0212] However, single nucleotide variants may exist not only in the coding sequence but also in introns or exon / intron adjacent regions. That is, there may exist pseudo-single nucleotide variants consisting of single nucleotide variants in introns or exon / intron adjacent regions and intended synonymous mutations within the coding sequence.

[0213] Single nucleotide variants and intended synonymous mutations are not significantly affected by their positional relationship to one another. In other words, the positional relationship between single nucleotide variants and intended synonymous mutations is not restricted. However, the positions of single nucleotide variants and intended synonymous mutations may be selected according to certain criteria. The detailed criteria may be the positional selection criteria for single nucleotide variants and intended synonymous mutations described in Chapter 2, "pegRNA Preparation."

[0214] Equivalence between pseudo-single nucleotide variants and single nucleotide variants

[0215] As mentioned above, synonymous mutations do not affect the translation outcome of the ATM gene. This can also be expressed as meaning that synonymous mutations have little impact on the function of the ATM gene. Therefore, a genome containing one or more intended synonymous mutations in addition to a single nucleotide variant can be treated identically to a genome containing only a single nucleotide variant. Consequently, when studying single nucleotide variants, pseudo-single nucleotide variants can be utilized. For example, to study the function of the ATM gene containing a single nucleotide variant, cells possessing pseudo-single nucleotide variants in the ATM gene can be used. The results of such a study would be close to or identical to those obtained using cells with a single nucleotide variant in the ATM gene.

[0216]

[0217] Definitions related to amino acid variations

[0218] The above amino acid variant refers to a variant in which an amino acid of a specific residue in the reference sequence is substituted with a different type of amino acid, or such a variant. Below, expressions related to amino acid variants are defined.

[0219] The aforementioned definition of a single nucleotide variant may also be applied to amino acid variants. Upon verification of the presence or absence of an amino acid variant, the amino acid of a specific residue may differ from that of the reference sequence. In this case, the protein may be described as containing or possessing a specific amino acid variant. Alternatively, the subject may be described as containing or possessing a specific amino acid variant.

[0220] The possible amino acid variations of the ATM protein encoded by the ATM gene having a single nucleotide mutation are as follows.

[0221] First, if the single nucleotide variant is a missense variant, it can result in a mutant ATM protein with a single amino acid variant (SAAV). Second, if the single nucleotide variant is a nonsense variant, it can result in a mutant ATM protein shorter than the reference sequence. Third, if the single nucleotide variant is a stop-loss mutation, it can result in a mutant ATM protein longer than the reference sequence.

[0222] Amino acid variations are described based on the reference sequence. Accordingly, the presence or absence of amino acid variations may vary depending on the reference sequence. The reference sequence may be an amino acid sequence represented by SEQ ID NO. 128 or 129, respectively.

[0223] Amino acid variations may be named based on a reference sequence. Unless otherwise stated, the nomenclature described herein names amino acid variations according to the following method. It will be obvious to a person skilled in the art that the nomenclature described herein can be appropriately interpreted. The expressions used to name amino acid variations in this specification are as follows. The first letter indicates the type of amino acid in the reference sequence. Additionally, the natural number following the first letter indicates the position of the residue. Additionally, the last letter indicates the changed amino acid. For example, "K25Q" indicates an amino acid variation in which lysine, the 25th residue of the reference sequence, is changed to glutamine.

[0224] Amino acid mutations caused by single nucleotide mutations that induce the aforementioned nonsense mutations and termination loss mutations may be named using "*" instead of an alphabet letter. In this case, "*" signifies termination. For example, "*3057W" signifies an amino acid mutation in which residue 3057 of the original reference sequence was absent but residue 3057 became Tryptophan. Additionally, "K24*" signifies an amino acid mutation in which residue 24 of the original reference sequence was lysine but was changed to termination.

[0225]

[0226] Background Knowledge_Prime Editing

[0227] Prime editing is a technology that modifies the CRISPR / Cas system to enable the editing of genes into intended sequences. Prime editing utilizes a method of recognizing a target site and rewriting the sequence of that site. To this end, prime editing uses a fusion protein of Cas nickase and reverse transcriptase and pegRNA (prime editing guide RNA). In this case, the pegRNA includes the sgRNA (single guide RNA), primer binding site, and reverse transcription template (RT-Template) of the CRISPR / Cas9 system.

[0228] Prime Editor

[0229] In this specification, "prime editor" refers to a fusion protein of Cas-nikase and reverse transcriptase for performing prime editing. Alternatively, if there are additional protein elements necessary or helpful for performing prime editing, it may mean including such elements. The prime editor can recognize and cleave a target nucleic acid together with pegRNA. The reverse transcriptase of the prime editor and the reverse transcript template of the pegRNA act together on the cleaved portion of the target nucleic acid to edit the gene as intended.

[0230] Prime Editor was developed by David R. Liu et al., and various embodiments of Prime Editor have been disclosed in numerous publications (Anzalone, Andrew V., et al. "Search-and-replace genome editing without double-strand breaks or donor DNA." Nature 576.7785 (2019): 149-157.; Chen, Peter J., et al. "Enhanced prime editing systems by manipulating cellular determinants of editing outcomes." Cell 184.22 (2021): 5635-5652.; and PCT applications [Application No. PCT / US2020 / 023712, Publication No. WO2020191233A1, etc.]).

[0231] pegRNA(prime editing guide RNA)

[0232] In this specification, pegRNA refers to an RNA sequence for performing prime editing, comprising a guide sequence, a scaffold, a reverse transcription template (RT-Template, RTT), and a primer binding primer in the direction from the 5' end to the 3' end.

[0233] The guide sequence is a sequence that recognizes the target sequence and enables the prime editor to induce a single-strand cut. The target sequence is selected according to the PAM (Protospacer Adjacent Motif) sequence. In one embodiment, when the prime editor uses Cas9 derived from Streptococcus pyogenes (spCas9), the PAM is NGG. The scaffold is a region that interacts with the prime editor. The scaffold can be appropriately designed for efficiency, etc. The reverse transcription template is a sequence containing a sequence complementary to the DNA sequence to be edited. Additionally, the primer binding site is a region where the DNA sequence with the 3' end exposed after a single-strand cut by the prime editor can bind. The reverse transcriptase of Prime Editor can polymerize a sequence complementary to the reverse transcript template at a target site using a DNA sequence with its 3' end exposed and a reverse transcript template of pegRNA. The reverse transcript template can be divided into the right homology arm (RHA) and the left homology arm (LHA) based on the region corresponding to the sequence to be modified through gene editing. In this case, the RHA refers to the reverse transcript template located at the 5' end of the region corresponding to the sequence to be modified through gene editing, based on the pegRNA. Additionally, the LHA refers to the reverse transcript template located at the 3' end of the region corresponding to the sequence to be modified through gene editing, based on the pegRNA. In this case, the DNA sequences complementary to the RHA and LHA may be referred to as the RHA and LHA, respectively, and should be interpreted according to the context.

[0234]

[0235] Background Knowledge_PARP Inhibitor

[0236] The meaning of PARP inhibitors

[0237] PARP inhibitors refer to inhibitors that inhibit PARP (Poly ADP-ribose polymerase). PARP is known to be involved in the repair of single-strand DNA breakage within the nucleus. Consequently, treatment with PARP inhibitors prevents the repair of single-strand DNA breakage, leading to the induction of double-strand DNA breakage.

[0238] PARP inhibitors can be used as anticancer agents against cancer cells with abnormalities in the repair mechanism for DNA double-strand damage. When PARP inhibitors are applied to cancer cells with abnormalities in the repair mechanism for DNA double-strand damage, DNA double-strand damage is induced and not repaired. Consequently, apoptosis is induced in the cancer cells. For this reason, PARP inhibitors are used as anticancer agents against some cancers with BRCA gene mutations.

[0239] PARP inhibitors may include any factor capable of inhibiting PARP. For example, PARP inhibitors may be olaparib, niraparib, rucaparib, or veliparib.

[0240] Reasons for using PARP inhibitors

[0241] The evaluation of ATM gene function in this application is characterized by evaluating ATM gene function through the viability (or viability) of cells having a single nucleotide mutation in the ATM gene. In this case, the inventors of this specification have discovered that while the function of the ATM gene can be evaluated by obtaining viability (or viability) through simple cell culture, the function of the ATM gene can be evaluated more accurately if treated with a PARP inhibitor.

[0242] This is because PARP inhibitors induce apoptosis in cells with abnormalities in the repair mechanism for DNA double-strand damage. As mentioned above, PARP inhibitors inhibit the repair of single-strand DNA damage. Accordingly, if cells with abnormal ATM gene function are treated with a PARP inhibitor, they will fail to repair DNA damage and will not survive. In contrast, cells with normal ATM gene function will be able to repair DNA double-strand damage even after treatment with a PARP inhibitor, and thus will be able to survive.

[0243] For the reasons mentioned above, the function of the ATM gene can be evaluated more accurately in the presence of a PARP inhibitor. Accordingly, if based on the survival rate (or viability) with respect to the PARP inhibitor, the effect of a single nucleotide variant on the function of the ATM gene can be evaluated more accurately.

[0244]

[0245] Chapter 1 Method for Evaluating Function of ATM Genes with Single Nucleotide Variants

[0246] The present application provides a method for providing information on the function of an ATM gene. The method for providing information on the function of an ATM gene is a method for providing information on the function of an ATM gene according to the ATM gene sequence for which the function is to be verified.

[0247] The above method for providing ATM gene function information is characterized by utilizing the results of the ATM gene function evaluation method of the present application. More specifically, the above method for providing ATM gene function information is characterized by utilizing information (function score or classification of single nucleotide variants) obtained through a single nucleotide variant library containing engineered cells.

[0248] Below, prior to describing the method for providing ATM gene function information of the present application, the method for evaluating ATM gene function having a single nucleotide variant of the present application is described.

[0249] Significance of the ATM gene function evaluation method with single nucleotide variants

[0250] The present application provides a method for evaluating the function of ATM genes having a single nucleotide mutation.

[0251] The above-mentioned functional evaluation method refers to a method for evaluating the effect of a single nucleotide variant on the function of the ATM gene. That is, the above-mentioned functional evaluation method evaluates changes in ATM gene function caused by the introduction of a single nucleotide variant. Alternatively, the above-mentioned functional evaluation method refers to a method for evaluating the function of an ATM gene containing a single nucleotide variant. The two meanings are substantially the same and may be appropriately selected and described depending on the context of the method to be described.

[0252] Characteristics of the ATM gene function evaluation method with single nucleotide mutations (1)

[0253] The method for evaluating ATM gene function having a single nucleotide variant in the present application can be classified into two types of methods.

[0254] First, the above-mentioned functional evaluation method may include a functional evaluation method through high-throughput experiments. The above-mentioned high-throughput experiments utilize a single nucleotide variant library. The above-mentioned single nucleotide variant library is characterized by containing cells having single nucleotide variants and intended synonymous mutations. That is, the above-mentioned single nucleotide variant library is characterized by containing cells having pseudo-single nucleotide variants. A detailed description is provided in "Method for Functional Evaluation of ATM Genes Having Single Nucleotide Variants_ High-Throughput Method".

[0255] Secondly, the above-mentioned function evaluation method may include a function evaluation method using a deep learning model. The function evaluation method using the deep learning model may be used to evaluate single nucleotide variants that could not be evaluated by a high-throughput method. The deep learning model is characterized by being trained using results and sequence information obtained through the high-throughput function evaluation method of the present application. A detailed description is provided in "Method for Function Evaluation of ATM Genes Having Single Nucleotide Variants_ Deep Learning Model".

[0256] Features of the ATM gene function evaluation method with single nucleotide mutations (2)

[0257] The method for evaluating the function of an ATM gene having a single nucleotide variant according to the present application evaluates the function of the ATM gene using cell viability or viability when the cell's single-strand damage repair function is impaired. Among the function evaluation methods of the present application, the high-throughput method evaluates the function of the ATM gene by directly obtaining cell viability information. Additionally, among the function evaluation methods of the present application, the method using a deep learning model evaluates the function of the ATM gene using a deep learning model trained on cell viability information. Both of the above methods evaluate the function of the ATM gene using cell viability (or viability). Therefore, the information derived by the function evaluation method of the present application can be considered a result obtained through cell viability information. Furthermore, the information derived by the function evaluation method of the present application can be considered a result obtained using cell viability (or viability).

[0258]

[0259] Chapter 2 Method for Evaluating Function of ATM Genes with Single Nucleotide Variants_ High-Throughput Method

[0260] Overview of the Method for Evaluating the Function of ATM Genes with Single Nucleotide Variants via a High-Throughput Approach

[0261] The following describes a method for evaluating the function of ATM genes containing single nucleotide variants using a high-throughput approach. Hereinafter, the method for evaluating the function of ATM genes containing single nucleotide variants using a high-throughput approach may be abbreviated as the high-throughput function evaluation method.

[0262] The high-throughput function evaluation method of the present application utilizes cells having a single nucleotide variant in the ATM gene. More specifically, it utilizes cells having a pseudo-single nucleotide variant in the ATM gene. The high-throughput function evaluation method of the present application may utilize a single nucleotide variant library comprising cells having a pseudo-single nucleotide variant in the ATM gene. In this case, the library may be a saturated library for single nucleotide variants.

[0263] The high-throughput function evaluation method of the present application may include preparing cells having a single nucleotide variant and a synonymous mutation in the ATM gene simultaneously. Alternatively, it may include preparing a library containing cells having a pseudo-single nucleotide variant in the ATM gene. In this case, the library may be saturated with the single nucleotide variant.

[0264] The high-throughput functional evaluation method of the present application evaluates the function of an ATM gene by utilizing the viability or viability of cells having pseudo-single nucleotide mutations in the ATM gene. In this case, the method is not limited as long as the cell viability (or viability) can be utilized. As an example, the high-throughput functional evaluation method of the present application may be characterized by obtaining viability information by comparing the results of sequencing a library containing cells having pseudo-single nucleotide mutations in the ATM gene at different time points.

[0265] The high-throughput functional evaluation method of the present application can evaluate the function of the ATM gene by utilizing the survival rate (or viability) of cells having pseudo-single nucleotide variants in the ATM gene against a PARP inhibitor. For example, the high-throughput functional evaluation method of the present application may be characterized by obtaining survival information by comparing sequencing results of a library at time points before and after treatment with a PARP inhibitor.

[0266] The high-throughput function evaluation method of the present application comprises obtaining survival information of cells having pseudo-single nucleotide mutations in the ATM gene. As an example, it may include a) sampling at a first time point, b) sampling at a second time point, and c) comparing the sequencing result for the sampling at the first time point with the sequencing result for the sampling at the second time point. As another example, it may include a) dividing a library into a first library and a second library, b) sampling the first library at a first time point, c) sampling library 2 at a second time point, and d) comparing the sequencing result for the sampling at the first time point with the sequencing result for the sampling at the second time point. As another example, it may include a) dividing a library into a first library and a second library, b) sampling the first library at a first time point, c) sampling the second library treated with a PARP inhibitor at a second time point, and d) comparing the sequencing result for the sampling at the first time point with the sequencing result for the sampling at the second time point.

[0267] The high-throughput function evaluation method of the present application may include obtaining functional information of an ATM gene having a specific single nucleotide variant using acquired survival information. To this end, the high-throughput function evaluation method of the present application may include quantifying survival information for the specific single nucleotide variant. In this case, the method is not limited to obtaining a function score using the survival information. Alternatively, the high-throughput function evaluation method of the present application may include obtaining qualitative information using the survival information.

[0268] Hereinafter, the ATM high-throughput function evaluation method of the present application will be described in detail.

[0269] Prepare the library

[0270] The high-throughput function evaluation method of the present application comprises preparing a library. In this case, the library preparation is not limited to methods as long as a single nucleotide mutation library can be prepared. For example, the library preparation may be subdivided into a) preparing cells and pegRNA and b) introducing pegRNA. This is described in a separate section. For another example, the library preparation may be to prepare a library containing various types of engineered cells.

[0271] The above library contains engineered cells. In this case, engineered cells refer to cells having one single nucleotide variant and one intended synonymous mutation in the ATM gene. Additionally, engineered cells refer to cells in which the sequence of the ATM gene is identical to the reference sequence except for the one single nucleotide variant and one intended synonymous mutation.

[0272] The engineered cell has a pseudo-single nucleotide variant in the ATM gene. Alternatively, the engineered cell has a pseudo-single nucleotide variant in the genomic region of interest. In this case, the location and type of the single nucleotide variant are not limited. Additionally, the location and type of the intended synonymous mutation are not limited. Furthermore, in one embodiment, the intended synonymous mutation may be present in the genomic region of interest. In another embodiment, the type and location of the intended synonymous mutation may be the same as the type and location of the intended synonymous mutation described in "PegRNA Design Method_Introduction of Synonymous Mutation" below. In this case, it is not considered whether the engineered cell additionally contains a single nucleotide variant in a region other than the ATM gene or the genomic region of interest.

[0273] The above library contains various types of engineered cells. In this case, the types of engineered cells are distinguished based on the single nucleotide mutations possessed by the engineered cells. That is, engineered cells of the same type have the same type of single nucleotide mutations. Additionally, engineered cells of the same type have the same or different types of intended synonymous mutations. Furthermore, engineered cells of different types have different types of single nucleotide mutations. Additionally, engineered cells having different types of single nucleotide mutations are different types of engineered cells.

[0274] The number of types of engineered cells included in the library varies. In one embodiment, if the library of the high-throughput function evaluation method of the present application is a saturated library for single nucleotide variants, the library may contain a number of types of engineered cells equal to or nearly equal to the number of single nucleotide variants that may exist in the genomic region of interest. In another embodiment, the number of types of engineered cells may be a specific ratio to the number of single nucleotide variants that may exist in the genomic region of interest. In one specific embodiment, the number of types of engineered cells may be 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% or more of the number of single nucleotide variants that may exist in the genetic region of interest. Preferably, the number of types of engineered cells may be 80% or more of the number of single nucleotide variants that may exist in the genetic region of interest.

[0275] The above genetic region of interest may vary. For example, the above genetic region of interest may be a region corresponding to a specific exon of the ATM gene. For another example, the above genetic region of interest may be an ATM gene region corresponding to a sequence selected from SEQ ID NOs 4 to 65. For another example, the above genetic region of interest may be a region corresponding to -5 bp bases of the upstream sequence of a specific exon of the ATM gene up to +5 bp bases of the downstream sequence of the specific exon. For another example, the above genetic region of interest may be a region corresponding to 5 bp of a specific exon and the upstream and downstream regions adjacent to the specific exon. In this case, the reason for including the 5 bp sequences of the upstream and downstream sequences of the exon is that the sequences may be related to splicing. As another example, the genetic region of interest may be an ATM gene region corresponding to one sequence selected from SEQ ID NOs 66 to 127. As another example, the genetic region of interest may be a region corresponding to specific exons of the ATM gene. As another example, for specific exons of the ATM gene, the genetic region of interest may be a region corresponding to a combination of -5 bp bases in the upstream sequence of each exon and +5 bp bases in the downstream region. As another example, the genetic region of interest may be a region corresponding to specific exons and 5 bp in the upstream and downstream regions adjacent to each exon. As another example, the genetic region of interest may be a region corresponding to exons 2 to 63 of the ATM gene. That is, the genetic region of interest may be a region corresponding to SEQ ID NOs 4 to 65 of the ATM gene.As another example, the genetic region of interest may be a region corresponding to -5 bp of the upstream sequence of each exon and +5 bp of the downstream sequence for exons 2 through 63 of the ATM gene. As another example, the genetic region of interest may be a region corresponding to 5 bp of the upstream and downstream sequences adjacent to exons 2 through 63, respectively. That is, the genetic region of interest may be a region corresponding to sequence numbers 66 through 127 of the ATM gene.

[0276] The ploidy of the engineered cell with respect to the ATM gene may vary. In one embodiment, the engineered cell may be haploid with respect to the ATM gene. In another embodiment, the engineered cell may be a haploidized cell with respect to the ATM gene. In another embodiment, the engineered cell may be a semi-KO cell with respect to the AMT gene. In another embodiment, the engineered cell may be diploid with respect to the ATM gene.

[0277] Preparing the library (1)_Cell preparation

[0278] Preparing the library of the present application may include cell preparation. The cells prepared in the cell preparation refer to cells into which single nucleotide mutations and intended synonymous mutations have not been introduced. The cell preparation is not limited to preparing cells capable of introducing single nucleotide mutations and intended synonymous mutations.

[0279] The above cell may have originated from a single type of cell.

[0280] The above cell may be derived from one type of cell line. The type of cell line is not particularly limited. For example, it may be a cell line having the characteristics of one or more cells described below. For other examples, it may be a cell line derived from cancer cells, immortalized cells, or stem cells. For other examples, it may be a breast-derived cell line, lung-derived cell line, colon-derived cell line, liver-derived cell line, stomach-derived cell line, pancreas-derived cell line, skin-derived cell line, kidney-derived cell line, or hematopoietic stem cell-derived cell line.

[0281] Below, the characteristics that the cells prepared in the above cell preparation may possess are described.

[0282] 1) Characteristics of Cells 1_ATM Gene Sequence

[0283] The cells prepared in the cell preparation of the present application are used to produce cells having a single nucleotide variant and an intended synonymous mutation in the ATM gene. Accordingly, the sequence of the ATM gene in said cells is used as a reference sequence for the single nucleotide variant and the intended synonymous mutation. It is important to determine the sequence of the ATM gene in said cells.

[0284] For example, the cell may be a sequence containing SEQ ID NO. 1 as the sequence of the ATM gene. For another example, the cell may be a sequence containing SEQ ID NO. 2 as the sequence of the ATM gene. For another example, the cell may be a sequence containing SEQ ID NO. 3 as the sequence of the ATM gene.

[0285] The single nucleotide variants and synonymous mutations described below are explained based on one of the above reference sequences.

[0286] 2) Characteristics of Cells 2_Ploydity

[0287] The ploidy of ATM genes in the cells prepared in the cell preparation of the present application may vary. The number of ATM genes contained in the cells prepared in the cell preparation of the present application may vary.

[0288] In one embodiment, the cell may be diploid with respect to the ATM gene. Accordingly, the cell may contain two ATM genes within its gene genome. In this case, the two ATM genes included in the gene genome may have identical or different sequences. In one specific embodiment, the cell may contain two ATM genes having sequences selected from SEQ ID NOs 1 to 3 within its gene genome. In another specific embodiment, the cell may contain two ATM genes having different sequences selected from SEQ ID NOs 1 to 3 within its gene genome.

[0289] In one embodiment, the cell may be haploid with respect to the ATM gene. That is, the cell may contain one ATM gene within its gene genome. In one specific embodiment, the cell may contain one ATM gene having a sequence selected from SEQ ID NOs 1 to 3 within its gene genome.

[0290] In one embodiment, the cell may be a cell that is haploidized with respect to the ATM gene. That is, the cell may have originally been polyploid or more polyploid with respect to the ATM gene, but with all but one ATM gene removed. Accordingly, the cell may have the same genetic composition as a cell that is haploid with respect to the ATM gene.

[0291] In one embodiment, the cell may be a cell in which all ATM genes except one are knocked out with respect to the ATM gene. That is, the cell may be polyploid or polyploid with respect to the ATM gene, but all ATM genes except one are no longer functional. The cell may be referred to as a semi-KO cell. In one embodiment, the cell may contain one ATM gene having a sequence selected from SEQ ID NOs 1 to 3 and a knocked-out ATM gene within its gene genome.

[0292] Using haploid cells allows for a more accurate evaluation of the function of ATM genes with single nucleotide mutations compared to using cells with 2 or more ploidy. Cells with 2 or more ploidy exhibit a phenotype influenced by multiple ATM genes. Accordingly, even if a cell with 2 or more ploidy has a single nucleotide mutation in one ATM gene, it may exhibit a phenotype influenced by other ATM genes. In contrast, since haploid cells possess only one ATM gene, the phenotype can be determined solely by the ATM gene present within the cell. Therefore, it is preferable that the cells prepared in the cell preparation of the present application be haploid or haploidized cells with respect to ATM genes.

[0293] 3) Cell Characteristics 3_Functional BRCA1, BRCA2, and TP53 Expression

[0294] The method for evaluating the function of an ATM gene having a single nucleotide variant according to the present specification is characterized by evaluating the function of an ATM gene related to DNA double-strand break repair function.

[0295] In this case, if the cell has abnormalities in genes other than the ATM gene involved in DNA double-strand repair function, it becomes impossible to be certain whether the impact on DNA double-strand repair function is solely due to single nucleotide mutations in the ATM gene. Therefore, genes other than the ATM gene involved in DNA double-strand repair function must be normal for the function of the ATM gene to be evaluated more accurately.

[0296] Accordingly, it is preferable that the cells prepared in the cell preparation of the present application are normal cells with genes related to DNA double-strand break repair function other than the ATM gene. For example, it is preferable that the cells prepared in the cell preparation of the present application express functional BRCA1, BRCA2, and TP53. For another example, it is preferable that the cells prepared in the cell preparation of the present application have at least one functional BRCA1, BRCA2, and TP53 gene on their genome.

[0297] 4) Cell Characteristics 4_Prime Editor Expression.

[0298] The cells prepared in the cell preparation of the present application may be cells that express a prime editor. The method of causing the cells to express the prime editor is not limited. In one embodiment, the cells may be cells into which nucleic acids capable of expressing the prime editor have been introduced. Or, the cells may be cells into which nucleic acids encoding prime data have been transduced into the genome. Or, the cells may be cells into which nucleic acids encoding prime data have been knocked in the genome.

[0299] The Prime Editor expressed by the cell is not limited to any specific type. For example, the cell may express Prime Editor 1 (PE1), Prime Editor 2 (PE2), Prime Editor 2 max (PE2max), Prime Editor 3 (PE3), Prime Editor 5 (PE4), Prime Editor 6 (PE6), or Prime Editor 7 (PE7). In a preferred embodiment, the cell may express Prime Editor 2 max.

[0300] For cells that do not express Prime Editor, an additional process of processing Prime Editor is required for library preparation. A detailed explanation of this is provided in "Introduction of pegRNA" below.

[0301] Preparing the library (2)_pegRNA preparation

[0302] Preparing the library of the present application may include preparing pegRNA (prime editing guide RNA). That is, preparing the library of the present application may involve treating cells with the pegRNA library to prepare the library. In this case, as long as pegRNA can be prepared, the method is not limited. In one embodiment, pegRNA may be synthesized according to a method known to a person skilled in the art.

[0303] The above pegRNA library refers to a library capable of introducing various types of pegRNA into cells. Therefore, the above pegRNA library may refer to a library containing various pegRNAs. Additionally, the above pegRNA library may refer to a library containing individual nucleic acids that encode each of various types of pegRNAs. Furthermore, the above pegRNA library includes a library of viruses capable of introducing nucleic acids encoding pegRNAs into the cell genome. In other words, the above pegRNA library means including various types of libraries designed to administer various pegRNAs to cells. The following description is based on a library containing pegRNA. It will be obvious to a person skilled in the art that such a description can be applied to libraries designed to introduce various pegRNAs.

[0304] Below, the pegRNA and pegRNA library prepared in pegRNA preparation are described.

[0305] 1) pegRNA design

[0306] Hereinafter, the guide sequence, scaffold, reverse transcription template, and primer binding site of the pegRNA used in preparing the library of the present application will be described.

[0307] The target sequence of the above pegRNA is selected as a location where a PAM sequence exists. However, if a PAM sequence does not exist, flexibility regarding the PAM sequence may be utilized. For example, if the PAM sequence is NGG and the NGG sequence is absent, the target sequence may be selected as a location where an NGA or NAG sequence exists.

[0308] The scaffold of the above pegRNA may be a sequence appropriately modified to increase prime editing efficiency. For example, it may be an sgRNA scaffold for SpCas9 appropriately modified.

[0309] The above reverse transcription template is designed to introduce single nucleotide variants and intended synonymous mutations. Additionally, the length of the above reverse transcription template is not limited. Preferably, the length of the reverse transcription template may be 40 bp or less. The length of the RHA of the above reverse transcription template is not limited. Preferably, the length of the RHA may be 4 bp or more. The length of the LHA of the above reverse transcription template is not limited.

[0310] The sequence of the above reverse transcription template is designed in two processes.

[0311] Primarily, the sequence of the reverse transcription template is designed according to the single nucleotide variant to be introduced. At this time, a guide sequence is selected first according to the single nucleotide variant to be introduced. At this time, the efficiency of prime editing may vary depending on the selection of the single nucleotide variant and the guide sequence. Therefore, the sequence of the reverse transcription template is designed by appropriately selecting the single nucleotide variant and the guide sequence. In one embodiment, the sequence of the reverse transcription template may be designed by the DeepPrime-FT model. For a description of the DeepPrime-FT model, the entire contents of the specification of application KR 2023-0111504 are incorporated by reference into this specification. Additionally, to increase accuracy and efficiency, the leading nucleotide of the guide sequence may be guanine (G).

[0312] Secondly, the sequence of the above reverse transcription template is designed by modifying the sequence of the primarily selected reverse transcription template to enable the introduction of the intended synonymous mutation. That is, the sequence of the above reverse transcription template is designed according to the location and type of the intended synonymous mutation. A detailed explanation of the location and type of the synonymous mutation is provided in "PegRNA Design Method_Introduction of Synonymous Mutation" below.

[0313] The length of the primer binding site is not limited. Preferably, the length of the primer binding site may be 17 bp or less.

[0314] At this time, the pegRNA may include an additional sequence at the 3' end position. In one embodiment, it may include a linker sequence and a sequence to prevent degradation of the pegRNA. In one specific embodiment, the linker sequence may be an 8 bp long linker sequence designed with pegLIT tools. In one specific embodiment, the sequence to prevent degradation of the pegRNA at the 3' end position may include a 37 bp tevopreQ1 sequence. In one specific embodiment, the additional sequence at the 3' end position may include a polyT(U) at the 3' end.

[0315] 2) pegRNA Design Method_Introduction of Intended Synonymous Mutations

[0316] Below, the types and locations of intended synonymous mutations introduced by pegRNA are described.

[0317] The aforementioned intended synonymous mutation may be introduced into a codon that is the same as or different from the single nucleotide variant to be introduced. Preferably, the aforementioned intended synonymous mutation may be introduced into a codon different from the single nucleotide variant to be introduced.

[0318] The aforementioned intended synonymous mutation is introduced into the exon region. That is, the aforementioned intended synonymous mutation is not introduced into the intron region. In particular, the aforementioned intended synonymous mutation is not introduced into the exon / intron adjacent region. Additionally, the aforementioned intended synonymous mutation is not introduced into the 2bp sequence (region) of the exon adjacent to the exon / intron junction. This is because the regions described above may be related to splicing and may affect the protein in which the intended synonymous mutation is expressed.

[0319] The above-mentioned intended synonymous mutation can be introduced at a location that disrupts or eliminates the PAM sequence recognized by the Prime Editor. That is, through the introduction of the above-mentioned intended synonymous mutation, the PAM sequence recognized by the Prime Editor can be changed to a sequence that the Prime Editor does not recognize. For example, when using a spCas9-derived Prime Editor, a synonymous mutation intended to change the PAM sequence NGG to the NTG sequence can be introduced. In this case, if the above-mentioned intended synonymous mutation is introduced at a location that disrupts or eliminates the PAM sequence recognized by the Prime Editor, there is an advantage that the target sequence is not recognized by the Prime Editor. Therefore, it is prioritized to introduce the above-mentioned intended synonymous mutation at a location that disrupts or eliminates the PAM sequence recognized by the Prime Editor.

[0320] If it is impossible to introduce the intended synonymous mutation into a location that destroys or eliminates the PAM sequence recognized by the prime editor, the intended synonymous mutation is introduced into a location as close as possible to the PAM sequence in the LHA.

[0321] If it is impossible to introduce the intended equivalent mutation into the LHA, the intended equivalent mutation is introduced into the RHA at a position as close as possible to the single nucleotide variant to be introduced.

[0322] 3) Design of pegRNA library

[0323] The preparation of the library for the present application may be performed by treating cells with the pegRNA library. Below, the types of pegRNA included in the pegRNA library are described. Each pegRNA included in the pegRNA library is designed according to the criteria described above.

[0324] The types of pegRNAs included in the above pegRNA library can be defined in various ways. For example, the types of pegRNAs can be classified according to the type of single nucleotide variant to be introduced. For another example, the types of pegRNAs can be classified according to the intended synonymous mutation to be introduced. For yet another example, the types of pegRNAs can be classified according to the RTT and guide sequence.

[0325] The following description is based on the types of pegRNA classified according to the type of single nucleotide variant to be introduced. The types of pegRNA included in the above pegRNA library may be classified according to the type of single nucleotide variant to be introduced. In this case, the number of types of pegRNA included in the above pegRNA library may vary. In one embodiment, it is assumed that the above pegRNA library is intended to create a saturated library for single nucleotide variants for the ATM gene. In this case, the number of types of pegRNA included in the above pegRNA library may be equal to or nearly equal to the number of single nucleotide variants that may exist in the genetic region of interest. In another embodiment, the number of types of pegRNA included in the above pegRNA library may be a specific ratio to the number of single nucleotide variants that may exist in the genetic region of interest. In one specific embodiment, the number of types of pegRNA included in the pegRNA library may be 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% or more of the number of single nucleotide variants that may exist in the genetic region of interest. In this case, the greater the number of types of pegRNA, the more information on as many types of single nucleotide variants as possible can be obtained in a single high-throughput experiment. Therefore, preferably, the number of types of pegRNA included in the pegRNA library may be equal to or nearly equal to the number of single nucleotide variants that may exist in the genetic region of interest. The description of the genetic region of interest is the same as described above in "Preparing the Library".

[0326] The above pegRNA library may include various types of pegRNAs to introduce a specific type of single nucleotide variant. That is, the pegRNAs included in the above pegRNA library may include various sequences even when introducing a specific type of single nucleotide variant. In one embodiment, the above pegRNA library may include pegRNAs with different sequences expected to show high introduction efficiency for each single nucleotide variant to be introduced. In one specific embodiment, the above pegRNA library may include pegRNAs with different sequences expected to show high introduction efficiency by the DeepPrime-FT model. In one embodiment, the above pegRNA library may include pegRNAs with different RTT sequences or guide sequences for each single nucleotide variant to be introduced. In one embodiment, the above pegRNA library may include pegRNAs with different intended synonymous mutations to be introduced for each single nucleotide variant to be introduced.

[0327] Prepare library (3)_introduce pegRNA

[0328] Preparing the library of the present application may include introducing pegRNA into prepared cells.

[0329] The method of introducing the above pegRNA is not limited as long as it allows for the introduction of pegRNA. In one embodiment, the pegRNA itself may be introduced. Alternatively, a nucleic acid encoding pegRNA may be introduced. In another embodiment, the pegRNA or the nucleic acid encoding pegRNA may be introduced via electroporation, microinjection, introduction via LNP, introduction via a lentiviral vector, or via AAV (Adeno-Associated Virus).

[0330] The above introduction method may be selected so that a nucleic acid encoding one pegRNA is introduced into one cell. In one embodiment, the concentration of the nucleic acid encoding the pegRNA may be controlled so that a nucleic acid encoding one pegRNA is introduced into one cell. In one specific embodiment, when a lentivirus is used for the introduction of the nucleic acid encoding the pegRNA, the titer of the lentivirus may be controlled so that one lentivirus is introduced into one cell.

[0331] The introduction of the above pegRNA may proceed through different processes depending on the cells prepared for the introduction of pegRNA. For example, if the cells are cells that express the Prime Editor, the objective can be achieved solely through the introduction of pegRNA. As another example, if the cells are cells that do not express the Prime Editor, the process of introducing the Prime Editor or nucleic acid encoding the Prime Editor is additionally included. In this case, the introduction method is not limited as long as the Prime Editor can be introduced.

[0332] Acquired survival information

[0333] The high-throughput function evaluation method of the present application includes obtaining survival information. The obtaining of survival information means obtaining information related to the survival rate (or viability) of cells having a single nucleotide mutation in the ATM gene by using the single nucleotide mutation library prepared in "preparing the library".

[0334] The above-mentioned survival data refers to information that can be used to infer the survival rate (or viability) of cells. Additionally, the above-mentioned survival data may refer to information regarding changes in the proportion of cells having a specific single nucleotide variant within a library. The said change in proportion may simply mean that the proportion changes over time. Alternatively, the said change in proportion may mean that the proportion changes over time in the presence of a PARP inhibitor. For example, the method for obtaining the above-mentioned survival data is not limited to any method as long as it is possible to obtain survival data of cells containing a single nucleotide variant in the ATM gene.

[0335] The above-mentioned acquisition of survival information may be obtained by verifying changes in the proportion of cells having a specific single nucleotide variant within a library. This utilizes the fact that if the viability (or viability) of a specific cell is low, the proportion of said specific cell within the library decreases over time. An example is described below assuming that a library containing cells with low viability (or viability) and cells with high viability (or viability) is cultured. As culture time elapses, the proportion of cells with low viability (or viability) in the library will decrease.

[0336] The above-mentioned survival information may be obtained by utilizing sequencing results for a single nucleotide variant library. For example, the above-mentioned survival information may be obtained by comparing the single nucleotide variant library sequencing results at different time points. As another example, the above-mentioned survival information may be obtained by comparing the single nucleotide variant library sequencing results at time points before and after PARP inhibitor treatment.

[0337] In one embodiment, the acquisition of survival information may include a) sequencing at a first time point, b) sequencing at a second time point, and c) comparing the sequencing results. In another embodiment, the acquisition of survival information may include a) sampling at a first time point, b) sampling at a second time point, and c) comparing the sequencing results for the sampling at the first time point with the sequencing results for the sampling at the second time point. In another embodiment, the acquisition of survival information may include a) sampling at a first time point without treatment with a PARP inhibitor, b) sampling at a second time point with treatment with a PARP inhibitor, and c) comparing the sequencing results for the sampling at the first time point with the sequencing results for the sampling at the second time point.

[0338] The above-mentioned survival information acquisition may additionally include dividing the library to sequence the library at different times. In one specific embodiment, the above-mentioned survival information acquisition may include a) dividing the library into a first library and a second library, b) sampling the first library at a first time point, c) sampling the second library at a second time point, and d) comparing the sequencing result for the sampling at the first time point with the sequencing result for the sampling at the second time point. In another specific embodiment, the above-mentioned survival information acquisition may include a) dividing the library into a first library and a second library, b) sampling the first library at a first time point without treating it with a PARP inhibitor, c) sampling the second library treated with a PARP inhibitor at a second time point, and d) comparing the sequencing result for the sampling at the first time point with the sequencing result for the sampling at the second time point.

[0339] The above sampling refers to extracting part or all of the library to obtain information about the state of the library at each time point. At this time, the sampled extract may be used immediately for sequencing, but it may also undergo preservation treatment to preserve the state at each time point. For example, the extract may be cryopreserved and then thawed for use in sequencing.

[0340] The above comparison of sequencing results refers to comparing the frequencies of reads containing specific single nucleotide variants and intended synonymous mutations in the sequencing results. Through this, survival information of cells having pseudo-single nucleotide variants in the ATM gene is obtained.

[0341] Function points can be obtained through the acquisition of the above survival information. A detailed explanation of this is provided in "Acquiring Function Points" below.

[0342] Below, the timing, sampling, and sequencing described above will be explained in more detail.

[0343] Acquired survival information_First Time Point and Second Time Point

[0344] The acquisition of survival information in this application is characterized by obtaining survival information by comparing the sequencing results of a library for a first time point and a second time point. This is intended to obtain survival information by utilizing the fact that the proportion of specific cells included in the library changes over time.

[0345] The above-mentioned first and second time points are concepts intended to represent the temporal difference for the proportion of specific cells included in the library to change. Alternatively, the above time points are concepts intended to represent the temporal difference for the sampling or sequencing of the library. Therefore, as long as there is a temporal difference between the first and second time points, they are not limited to a specific time point.

[0346] In one embodiment, the first time point may be a time point after a certain period has passed since the introduction of pegRNA. In one specific embodiment, the first time point may be a time point after a day selected from 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, and 20 days has passed since the introduction of pegRNA, or a time point between two selected different days. Preferably, the first time point may be a time point after 13 days have passed since the introduction of pegRNA. This is to ensure that the introduction of pseudo-single nucleotide mutations via pegRNA is properly carried out. In one embodiment, the second time point may be a time point after a certain period has passed compared to the first time point. In one embodiment, the second time point may be a time point selected from days 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, and 20 that have elapsed compared to the first time point. Preferably, the second time point may be a time point 10 days after the first time point. In one embodiment, the first time point may be a time point before treatment with a PARP inhibitor. In one embodiment, the first time point may be a time point 13 days after pegRNA introduction and before treatment with a PARP inhibitor. In one embodiment, the second time point may be a time point after treatment with a PARP inhibitor. In one specific example, the second time point may be a time point 10 days after the first time point, and a time point 10 days after the PARP inhibitor was administered.

[0347] Acquiring survival information_sequencing

[0348] As described above, the high-throughput functional evaluation method of the present application can utilize viability information of various types of engineered cells in a single nucleotide mutation library. In this case, for a high-throughput method, a method is required to acquire viability information of various types of engineered cells included in the single nucleotide mutation library all at once. To this end, the acquisition of viability information of the present application utilizes the results of sequencing a library containing various types of engineered cells.

[0349] The library of the present application contains various types of engineered cells. Accordingly, sequencing results for all or part of the cells constituting the library include sequencing results for various types of engineered cells. In this case, the frequency of a specific single nucleotide variant (or pseudo-single nucleotide variant) obtained by the sequencing can be treated as a specific type of engineered cell. This is because an engineered cell is a cell containing a single type of single nucleotide variant (or pseudo-single nucleotide variant). Therefore, the frequency of engineered cells included in the library can be obtained as a result of the sequencing. Ultimately, viability information can be obtained by comparing the sequencing results.

[0350] In order to perform the above sequencing, genomic DNA extraction may be performed first. At this time, the method of extracting genomic DNA is not limited as long as genomic DNA from cells included in the library can be extracted.

[0351] To perform the above sequencing more effectively, a Polymerase Chain Reaction (PCR) may be performed beforehand. At this time, the method and conditions for performing the PCR are not limited. For example, PCR may be performed using exon-specific primers of the ATM gene.

[0352] The above sequencing may be performed by a method known to a person skilled in the art. For example, the above sequencing may be performed through Sanger sequencing, Next-Generation Sequencing (NGS), Illumina sequencing, Nanopore sequencing, or PacBio sequencing.

[0353] The high-throughput function evaluation method of the present application is characterized by using engineered cells that simultaneously possess one single nucleotide variant and one intended synonymous mutation. The reason for using engineered cells that simultaneously possess one single nucleotide variant and one intended synonymous mutation is to reduce the impact of sequencing errors. Accordingly, only reads that simultaneously possess one single nucleotide variant and one intended synonymous mutation without other discrepancies are evaluated as having originated from engineered cells that simultaneously possess one single nucleotide variant and one intended synonymous mutation. Therefore, the sequencing results can be processed to include only reads identical to the reads and reference sequences. That is, the sequencing results can be processed to filter out anything other than reads that simultaneously possess one single nucleotide variant and one intended synonymous mutation without other discrepancies or reads identical to the reference sequence.

[0354] Acquiring functional information_Information processing

[0355] Functional evaluation of a specific single nucleotide variant can be performed using the aforementioned survival information. The reason why the effect of a specific single nucleotide variant introduced into ATM on the function of ATM can be evaluated based on survival information is that, as previously mentioned, the cell viability (or viability) may vary depending on whether the ATM gene functions normally when cells are treated with a PARP inhibitor. Therefore, the high-throughput functional evaluation method of the present application includes obtaining functional information.

[0356] The above functional information refers to information regarding the function of ATM genes having specific single nucleotide mutations. In this case, as previously mentioned, since the function and survival of the ATM gene in a cell are related, functional information and survival information may be used interchangeably depending on the context. The two terms are selected to appropriately indicate the meaning of the information to be expressed. Furthermore, the above functional information may be quantified information (i.e., quantitative information) or information regarding the functional state (i.e., qualitative information).

[0357] If the above-mentioned functional information is quantified information, the functional information is obtained by quantifying the comparison of sequencing results for the libraries of the aforementioned first and second time points. In this case, it is more preferable that the quantification be based on a common standard. Accordingly, in one embodiment, the functional evaluation method of the present application may include quantifying functional information based on a common standard. Hereinafter, a functional score, which is an example of quantification, is described. This is an example, and even if quantification is performed in a different way, it is not limited as long as it is quantified according to a certain standard. The value obtained by quantifying the functional evaluation result for a specific single nucleotide variant using a specific method is referred to as the functional score of the specific single nucleotide variant. Alternatively, the above-mentioned functional score refers to a value obtained by numerically expressing information regarding ATM gene function using a specific method.

[0358] The specific method for obtaining the above function score is as follows. Prior to this, with respect to one specific single nucleotide variant (single nucleotide variant A), the sequencing results may include the following information: 1) a read containing single nucleotide variant A and the intended synonymous mutation 1, 2) a read containing single nucleotide variant A and the intended synonymous mutation 2, 3) a read containing single nucleotide variant A and the intended synonymous mutation 3, and 4) the results of repeated experiments for 1), 2), and 3).

[0359] First, the frequency of reads containing a specific single nucleotide variant and an intended synonymous mutation at the first time point and the frequency of reads containing a specific single nucleotide variant and an intended synonymous mutation at the second time point are expressed as LFC (log-fold change). In this case, the value calculated through the above process is referred to as the LFC value of the specific single nucleotide variant and the intended synonymous mutation.

[0360] To correct for differences that may exist between exons, the LFC values ​​are standardized based on the LFC values ​​of synonymous single nucleotide variants that do not cause amino acid mutations. At this time, a regressed LFC is used to obtain the LFC values ​​of synonymous single nucleotide variants for each position (i.e., to obtain them even for positions without synonymous single nucleotide variants). The regressed LFC is obtained through LOWESS (Locally Weighted Scatterplot Smoothing) regression. Additionally, a standardized LFC value is obtained using the interquartile range values ​​of the LFC values ​​of synonymous single nucleotide variants for each exon. That is, the standardized LFC (Standardized LFC) is obtained by subtracting the regressed LFC value of the synonymous single nucleotide variant for each position from the LFC values ​​of the specific single nucleotide variant and the intended synonymous mutation for each position, and dividing by the interquartile range values ​​of the LFC values ​​of the synonymous single nucleotide variant for each exon.

[0361] At this time, the normalization is performed separately for each exon. That is, it is normalized and displayed based on the LFC value of the synonymous single nucleotide variant in the exon containing the specific single nucleotide variant and the intended synonymous mutation. In addition, the normalized LFC refers to the normalized LFC value of the specific single nucleotide variant and the specific intended synonymous mutation.

[0362] To express the integrated LFC values ​​of sequences that have the same single nucleotide variant but different intended synonymous mutations, the normalized LFC values ​​are weighted averaged. That is, through the weighted average, the LFC value is represented as a single single nucleotide variant. This is because synonymous mutations do not affect the translation outcome of the ATM gene, and most of the impact on function is due to the single nucleotide variant. The weighted average is calculated using the frequency at the first time point and the normalized LFC values ​​for sequences that have the same single nucleotide variant but different intended synonymous mutations.

[0363] Function scores can be obtained by averaging the values ​​obtained through repeated experiments with respect to the above-mentioned weighted average LFC value. A more detailed explanation of the classification and meaning of the above-mentioned function scores is provided in "Function scores of ATM genes with single nucleotide polymorphisms" below. The function scores obtained by the above-mentioned high-throughput method may also be referred to as function evaluation scores to distinguish them from function scores obtained using deep learning models.

[0364] Alternatively, if the above functional information is qualitative information, the functional information may qualitatively represent a comparison of sequencing results for the libraries at the aforementioned first and second time points. For example, the functional information may be information indicating that a specific single nucleotide variant is in a functional state, an intermediate state, or a state of dysfunction. Descriptions of each state are provided in Chapter 5.

[0365] Chapter 3 Method for Evaluating Function of ATM Genes with Single Nucleotide Variants_ Deep Learning Model

[0366] Reasons why deep learning model feature evaluation methods are necessary

[0367] According to the disclosure of the high-throughput functional evaluation method of this specification, functional evaluation of single nucleotide mutations is possible. However, since the high-throughput functional evaluation method uses CRISPR / Cas9, functional evaluation of all possible single nucleotide mutations of the ATM gene is not possible due to limitations of CRISPR / Cas9 technology. For example, functional evaluation of some single nucleotide mutations may not be possible due to the absence of a PAM (Protospacer Adjacent Motif) sequence or issues with the efficiency of pegRNA. In particular, when using Cas9 derived from Streptococcus pyogenes (spCas9), functional evaluation of single nucleotide mutations present in AT-rich regions (NGG sequence-deficient regions) may not be possible.

[0368] As described above, there is a need for a method to perform functional evaluation even for single nucleotide mutations that are impossible or difficult to evaluate using high-throughput functional evaluation methods. Below, a method for evaluating the functionality of a deep learning model to solve the aforementioned problem is described.

[0369] Overview of Deep Learning Model Function Evaluation Methods

[0370] The method for evaluating the function of ATM genes having single nucleotide variations in the present application may include a method for evaluating the function of ATM genes having single nucleotide variations through a deep learning model. The method for evaluating the function of ATM genes having single nucleotide variations using a deep learning model will be described below. In the following, the method for evaluating the function of ATM genes having single nucleotide variations using a deep learning model may be referred to as the deep learning function evaluation method.

[0371] The deep learning function evaluation method described above is characterized by designing or / and using a deep learning model. The deep learning model is a model that uses the results of a high-throughput function evaluation method and information on variant sequences as training data. Furthermore, the deep learning model refers to a model that predicts the function score of an ATM gene or ATM protein having a corresponding sequence when sequence information of an ATM gene having a single nucleotide variant or an ATM protein having an amino acid variant is input.

[0372] The above method for evaluating deep learning functions may include designing or / and using a deep learning model. For example, the above method for evaluating deep learning functions may include a) designing a deep learning model, b) training a deep learning model, and c) obtaining information on the function of ATM genes having single nucleotide mutations using the deep learning model. For another example, the above method for evaluating deep learning functions may include a) preparing a deep learning model and b) obtaining information on ATM genes having single nucleotide mutations through the deep learning model. For yet another example, the above method for evaluating deep learning functions may include obtaining information on ATM genes having single nucleotide mutations through the deep learning model. In this case, a detailed explanation of the deep learning model, the design of the deep learning model, and the training of the deep learning model will be provided in a separate section below.

[0373] The deep learning model of the deep learning function evaluation method described above uses the results of the high-throughput function evaluation method as training data. Accordingly, the deep learning function evaluation method may include obtaining results through the high-throughput function evaluation method. For example, the deep learning function evaluation method may include obtaining survival information or function scores. In this case, the process of obtaining results through the high-throughput function evaluation method is not limited to the order of execution as long as it is performed only before training the deep learning model.

[0374] The deep learning model of the deep learning function evaluation method above outputs a function score based on amino acid variation. Accordingly, obtaining information about the ATM gene having the above single nucleotide variation may include converting the output value of the deep learning model into a function score based on single nucleotide variation.

[0375] The method for designing and training the above-mentioned deep learning model is not limited as long as it has the structure described below. Below, the training data, structure, and training method of the above-mentioned deep learning model are described.

[0376] Deep learning model training data

[0377] The deep learning model of the present application is characterized by using 1) sequence information of ATM proteins and 2) function scores obtained from the aforementioned high-throughput function evaluation method as training data. The deep learning model uses ATM protein sequences, rather than ATM genes, as training data. This is because it is the ATM protein that directly influences the phenotype.

[0378] The function scores obtained from the aforementioned high-throughput function evaluation method are based on single nucleotide variants. Therefore, to use them as training data, the function scores obtained from the aforementioned high-throughput function evaluation method must be converted to amino acid variants. A specific single nucleotide variant is associated with a specific amino acid variant. For example, assume that a specific single nucleotide variant is a missense variant. In this case, the protein encoded by the gene containing the corresponding single nucleotide variant possesses an amino acid variant. Therefore, the function score for the single nucleotide variant causing the missense mutation can also be viewed as the function score for the aforementioned amino acid variant. Alternatively, assume that a specific single nucleotide variant causes a synonymous variant. In this case, the protein encoded by the gene containing the corresponding single nucleotide variant has the same sequence as the reference protein sequence. Therefore, the function score for the single nucleotide variant causing the synonymous variant can also be viewed as the function score for the reference protein sequence.

[0379] Consequently, the single nucleotide variant-based function scores obtained from the aforementioned high-throughput function evaluation method can be converted into protein-based function scores. If there are multiple single nucleotide variants inducing a specific amino acid variation, the single nucleotide variant-based function scores are averaged and used.

[0380] The deep learning model above converts the feature scores obtained from the high-throughput feature evaluation method into training data using the above method.

[0381] The above function points may be processed to prevent skewness. In one embodiment, the function points may be used after performing an inverse hyperbolic sine transformation (arcsinh transformation).

[0382] The above deep learning model uses domain information of ATM proteins as training data. This is to include structural information of ATM proteins.

[0383] The deep learning model described above uses information on the X, Y, and Z axes of the alpha-carbon atoms for each residue of the ATM protein or the ATM protein having amino acid variants as training data. This is to include structural information of the ATM protein. At this time, the structural information can be obtained through a known computational model. The known computational model is not limited. For example, it could be AlphaFold 3.

[0384] The sequence information, domain information, and structural information of the ATM protein described above are embedded and utilized. A detailed explanation of this is provided in the "Deep Learning Model Structure" below.

[0385] The deep learning model above uses scores obtained through a known computational model of ATM proteins with amino acid variations as training data. The use of information from the other computational model is intended to produce more accurate results. The scores obtained through the known computational model are not limited to scores of ATM proteins with amino acid variations obtained using the model. In one embodiment, it may be one or more of the scores derived from SIFT, FATHMM, MutationTaster, LRT, DANN, PolyPhen-2 HVAR, PROVEAN, REVEL, CADD, phyloP100, GERP, ESM1b, EVE, AlphaMissense, BoostDM, and SpliceAI. In one embodiment, it may be all of the scores derived from SIFT, FATHMM, MutationTaster, LRT, DANN, PolyPhen-2 HVAR, PROVEAN, REVEL, CADD, phyloP100, GERP, ESM1b, EVE, AlphaMissense, BoostDM, and SpliceAI. If a score cannot be derived from an already known computational model for an ATM protein with an amino acid variation, an appropriately set value may be used. For example, if a score can be derived for several computational models for an ATM protein with a single amino acid variation, the average of these scores may be used as the score for other models. As another example, for an ATM protein identical to a reference sequence or an ATM protein that has been prematurely terminated, a score of 0 may be entered. The method by which the score derived from the above already known computational model is used in the model is explained in detail in "Deep Learning Model Structure".

[0386] Deep learning model structure

[0387] The deep learning model described above is a transformer-based regression model that utilizes a transformer model and a neural network regression model. In particular, it is a model that uses one of the vectors produced through the transformer model as the input value of the input layer of the neural network regression model.

[0388] Below, after explaining the structure of the Transformer model, the relationship with the neural network regression model and the structure of the neural network regression model are explained.

[0389] Hereinafter, the length of the input sequence in the above Transformer model is set according to the length of the ATM protein. The following description is based on the assumption that the sequence length is 3056. It will be obvious to a person skilled in the art that the length of the sequence input into the Transformer model can be appropriately changed.

[0390] The above transformer model utilizes the sequence information, domain information, and structural information of the ATM protein by embedding them. That is, the input embedding vector of the above transformer model includes the sequence information, domain information, and structural information of the ATM protein. Furthermore, the input embedding vector of the above transformer model is a vector created by summing the amino acid embedding vector, the domain embedding vector, and the coordinate embedding vector. Each embedding vector is described below.

[0391] The sequence information of the ATM protein is encoded in an amino acid embedding vector. In this case, the sequence information of the ATM protein refers to the sequence information of the ATM protein encoded by the ATM gene having a single nucleotide mutation. The sequence information of the ATM protein is embedded by numerically representing the amino acids at each position. The method of embedding is not limited as long as the amino acids at each position can be numerically represented. In one embodiment, the amino acids at each position are represented using a 64-dimensional embedding vector. In this case, the values ​​input into the 64-dimensional embedding vector are initially set randomly. Additionally, the values ​​input into the 64-dimensional embedding vector change during the learning process.

[0392] The domain information of the ATM protein is encoded in a domain embedding vector. In this case, the domain information is intended to reflect that the ATM protein contains multiple domains. The domain embedding vector can be designed to reflect that the domain information of the ATM protein corresponds to a specific domain region depending on the sequence position in the sequence input to the transformer. In one embodiment, the fact that a specific domain region corresponds to a sequence position in the sequence input to the transformer can be specified from the beginning. For example, it can be specified from the beginning that the 1st to 166th positions of the sequence input to the transformer correspond to the TAN (Tel1 / ATM N-terminal motif) domain, the 1940th to 2566th positions correspond to the FAT (FRAP-ATM-TRRAP) domain, the 2686th to 2998th positions correspond to the PI3 / 4 kinase domain, and the 3024th to 3056th positions correspond to the FATC domain. The method of embedding the domain information of the ATM protein is not limited as long as the domain information can be expressed numerically. In one embodiment, information for each domain is represented using a 64-dimensional embedding vector. In this case, the input value to the 64-dimensional embedding vector is initially set randomly. Additionally, the input value to the 64-dimensional embedding vector changes during the learning process.

[0393] The above ATM protein structure information is encoded in a coordinate embedding vector. The above ATM protein structure information refers to the X, Y, and Z axis information of each alpha-carbon atom of the ATM protein obtained through a previously known computational model. The method of embedding the ATM protein structure prediction information obtained through the above computational model is not limited as long as the domain information can be expressed numerically. In one embodiment, the X, Y, and Z axis information can be embedded by transforming a 64-dimensional embedding vector using a two-layer multi-layer perceptron.

[0394] The above transformer model may be a transformer model that includes only transformer encoder layers or encoder layers. The number of transformer encoder layers in the above transformer model is not limited. In one embodiment, the number of transformer encoder layers in the above transformer model may be two. Accordingly, in the above transformer model, the embedding vector may be learned by passing through a total of two transformer encoder layers. In addition, the number of attention heads included in each transformer encoder layer is not limited. In one embodiment, the transformer encoder layer may include a total of eight attention heads.

[0395] The above Transformer model ultimately produces output embedding vectors. At this time, one of the output embedding vectors is subsequently input into the input layer of a neural network regression model. The output embedding vector corresponds to the position of the amino acid variation based on the training data.

[0396] Furthermore, the neural network regression model utilizes the score of the ATM protein obtained from an existing computational model by inputting it into the input layer of the neural network regression model. That is, the neural network regression model utilizes a vector formed by concatenating the score of the ATM protein obtained from an existing computational model with the output embedding vector corresponding to the location of the amino acid variation by inputting it into the input layer.

[0397] The above neural network regression model may be a model composed of a fully connected layer. In this case, the number of hidden units included in the neural network regression model is not limited. In one embodiment, the neural network regression model may be a model including 128 hidden units and a ReLU activation function.

[0398] The above neural network regression model is a model that ultimately outputs a single output value. In this case, the output value may be referred to as the Deep_ATM score. The output value is information regarding the function of the ATM gene or ATM protein. The output value is used to calculate the function score. The function score is described in detail below in "Function Score Using Deep Learning Models".

[0399] Deep learning model dataset

[0400] The dataset for training and testing the deep learning model is not limited as long as it can reflect the characteristics of the deep learning model.

[0401] The above dataset consists of function scores based on single nucleotide variants obtained by the aforementioned high-throughput function evaluation method. Additionally, as previously mentioned, the function scores based on single nucleotide variants are converted into information about proteins and used in a deep learning model.

[0402] The criteria for classifying the above dataset as a test dataset are not limited. In one embodiment, the test dataset may include information on single nucleotide variants disclosed in a database regarding their impact on the ATM gene. In one specific embodiment, ClinVar may identify pathogenicity and classify information on 116 single nucleotide variants as the test dataset.

[0403] The criteria for dividing the above dataset into a training dataset are not limited. In one embodiment, the training dataset may be obtained by excluding information included in the test dataset from function score information based on single nucleotide variants obtained by a high-throughput function evaluation method. In one embodiment, the training dataset may be obtained by excluding information regarding single nucleotide variants that have the same effect on ATM gene translation as the single nucleotide variants included in the test dataset from function score information based on single nucleotide variants obtained by a high-throughput function evaluation method. In one embodiment, the training dataset may be obtained by excluding information regarding single nucleotide variants present in the stop codon of the ATM gene from function score information based on single nucleotide variants obtained by a high-throughput function evaluation method.

[0404] Deep learning model training methods

[0405] The method of training the deep learning model of the present application is not limited.

[0406] In one embodiment, training can be performed by setting an appropriate loss function. In one embodiment, training can be performed by setting a weighted loss function. In one embodiment, training can be performed by setting an appropriate overfitting prevention method. In one embodiment, performance improvement can be tracked, and early stopping can be performed under certain criteria. In another embodiment, a cross-validation technique can be used.

[0407] The number of epochs for training the deep learning model of the present application is not limited. In one embodiment, the number of epochs may be 150. Preferably, the number of epochs may be 150 or less by early termination under certain criteria.

[0408] The number of batches for training the deep learning model of the present application is not limited. In one embodiment, the number of batches may be 20. Additionally, the batches may be allocated according to the type of single nucleotide variant. In one embodiment, 90% of the batches may be allocated to single nucleotide variants causing missense variants, 5% of the batches to single nucleotide variants causing synonymous variants, and 5% of the batches to single nucleotide variants causing nonsense variants.

[0409] Feature points using deep learning models

[0410] Below, feature scores using deep learning models are explained.

[0411] The function score obtained using the deep learning model above is a function score based on amino acid mutations. This can be converted into a function score based on single nucleotide mutations according to the method described above. In this case, if there are multiple single nucleotide mutations that induce a specific amino acid mutation, the multiple single nucleotide mutations may have the same function score.

[0412] The function score obtained using the deep learning model described above can be obtained using the output value of the deep learning model described above. The method of obtaining the function score is not limited as long as it has a higher correlation with the function score of the high-throughput function evaluation method. Preferably, the output value of the deep learning model can be converted into a function score through rank-based conversion and regression using the function score of the high-throughput function evaluation method.

[0413] The function score obtained using the deep learning model described above may be referred to as a function score along with the function score of the high-throughput function evaluation method. The function score obtained using the deep learning model described above is a value predicted using the deep learning model. Therefore, the function score obtained using the deep learning model described above may also be referred to as a function prediction score. This may be used to distinguish between the function score of the high-throughput function evaluation method and the function score obtained using the deep learning model. This should be interpreted according to the context.

[0414] The above-mentioned function prediction score may be a score that has verified its association with the function score. In one embodiment, verification can be performed through the association of the distributions between the function prediction score and the function score.

[0415] The above-mentioned function prediction score may be verified to show higher accuracy than scores obtained by other known prediction models (e.g., AlphaMissense, ESM1b, phyloP, and PROVEAN). In one embodiment, this can be verified by examining the area under the receiver operating characteristic curve for some single nucleotide variants. The said single nucleotide variants may be single nucleotide variants that disclose their effects on ATM genes in known databases.

[0416]

[0417] Chapter 4 Method for Evaluating the Function of ATM Proteins with Amino Acid Variants

[0418] This specification aims to provide a method for evaluating the function of ATM proteins having amino acid variants. Hereinafter, a method for evaluating the function of ATM proteins having amino acid variants will be described.

[0419] The method for evaluating the function of an ATM protein having an amino acid variation according to the present application is characterized by including each process or part of each process of the method for evaluating the function of an ATM gene having a single nucleotide variation according to the present application. More specifically, the method for evaluating the function of an ATM protein having an amino acid variation according to the present application utilizes the result (function score) of the method for evaluating the function of an ATM gene having a single nucleotide variation according to the present application.

[0420] The method for evaluating the function of ATM genes having single nucleotide variants through a high-throughput method of the present application obtains a function score based on the single nucleotide variant. At this time, the function score based on the single nucleotide variant can be converted to obtain a function score based on the amino acid variant. This is because, as previously stated, the function score based on the single nucleotide variant is associated with the function score based on the amino acid variant.

[0421] In one embodiment, the method for evaluating the function of an ATM protein having the amino acid variation may precede the method for evaluating the function of an ATM gene having a single nucleotide variation through a high-throughput method of the present application. In one specific embodiment, the method for evaluating the function of an ATM gene having the amino acid variation may include a) obtaining information on the function of an ATM gene having a single nucleotide variation and b) obtaining a function score based on the amino acid variation using the information on the function.

[0422] In one embodiment, the method for evaluating the function of an ATM protein having the amino acid variation may include converting the result of the method for evaluating the function of an ATM gene having a single nucleotide variation through the high-throughput method of the present application (function score based on single nucleotide variation) into a function score based on amino acid variation.

[0423] Converting to a function score based on the above amino acid variation criteria is not limited to the method provided that the function score for a single nucleotide variation can be converted to a function score for an amino acid variation.

[0424] It is obvious that a person skilled in the art can easily perform the conversion of the above single nucleotide variant-based function points to the amino acid variant-based function points using a codon table. Furthermore, if there are multiple single nucleotide variants that induce a specific amino acid variant during the conversion process, the single nucleotide variant-based function points can be averaged and used as the amino acid variant-based function points.

[0425] Unlike high-throughput methods, the deep learning model of the method for evaluating the function of ATM genes having single nucleotide variations using the deep learning model of the present application outputs a function score based on amino acid variation.

[0426] Accordingly, in one embodiment, the method for evaluating the function of an ATM protein having the amino acid variation may include a part of the method for evaluating the function of an ATM gene having a single nucleotide variation using the deep learning model of the present application. In one specific embodiment, it may include obtaining functional information of an ATM protein having the amino acid variation using the deep learning model of the method for evaluating the function of an ATM gene having a single nucleotide variation using the deep learning model of the present application.

[0427] Chapter 5 Functional Scores of ATM Genes with Single Nucleotide Variants

[0428] Meaning of Function Points

[0429] The function score of the present application is a numerical representation of how the function of the ATM gene will change when a specific single nucleotide variant is introduced into the ATM gene. Alternatively, the function score is a numerical representation of information regarding the function of the ATM gene having a specific single nucleotide variant.

[0430] In relation to the above-mentioned function scores, a single nucleotide variant may also be referred to as a variant. That is, the function score for a specific single nucleotide variant described below may also be referred to as the function score for a specific variant. Furthermore, the classification of single nucleotide variants described below may be referred to as a classification of variants.

[0431] The above function score is a numerical representation of how the subject's phenotype will change when a specific single nucleotide variant is introduced into the ATM gene. In one embodiment, the above function score is a numerical representation of information regarding the phenotype of a subject containing a specific single nucleotide variant in the ATM gene. In another embodiment, it is a numerical representation of information regarding the phenotype of a subject carrying a specific single nucleotide variant in the ATM gene. As described above, the correlation between the function score of the present application and the phenotype related to the ATM gene of an individual containing a single nucleotide variant in the ATM gene can be expressed as clinical significance. The clinical significance can be confirmed through a comparison between the function score of the present application and a known database. In this case, the above function score may be a function evaluation score or a function prediction score.

[0432] The above function score is assigned to each single nucleotide variant. As previously mentioned, the reference sequence used to define the types of single nucleotide variants is SEQ ID NO. 1. Therefore, below, the types of single nucleotide variants can be defined based on the reference sequence of SEQ ID NO. 1, and the function score can be calculated individually for each type of single nucleotide variant.

[0433] Based on the above function scores, a function class can be determined (assigned) for ATM genes having specific single nucleotide variants. For example, if the function score of a specific single nucleotide variant is below the first criterion, the function class of the specific single nucleotide variant may be determined as "dysfunctional." Additionally, if the function score of a specific single nucleotide variant is above the second criterion, the function class of the specific single nucleotide variant may be determined as "functional." In this case, the second criterion refers to a value higher than the first criterion. If the function score of a specific single nucleotide variant is between the second and first criterions, the function class of the specific single nucleotide variant may be determined as "intermediate." A detailed explanation of the above criterion is provided in the "Method for Determining Criteria" below.

[0434] A lower function score indicates that the ATM gene is not functioning normally. Conversely, a higher function score indicates that the ATM gene is functioning normally.

[0435] The function score of the present application collectively refers to scores derived using a high-throughput method and scores derived through a deep learning model. If you wish to distinguish between the two, the former may be referred to as a function evaluation score and the latter as a function prediction score.

[0436] The functional prediction score may be used only for single nucleotide variants for which there is no functional evaluation score. That is, if a functional evaluation score exists for a specific single nucleotide variant, the functional evaluation score is used, and if a functional evaluation score does not exist for another specific single nucleotide variant, the functional prediction score can be used. The scores derived for each single nucleotide variant through the above method may be referred to as the functional combined score.

[0437] Method for determining classification criteria

[0438] Single nucleotide variants can be classified under certain criteria based on the function score of the present application.

[0439] In one embodiment, the function scores of the present application can be classified by percentile. In one specific embodiment, the function scores of the present application can be classified based on values ​​corresponding to the bottom 5%, bottom 10%, bottom 15%, bottom 20%, bottom 25%, bottom 30%, bottom 40%, bottom 50%, bottom 60%, bottom 70%, bottom 80%, bottom 85%, bottom 90%, or bottom 95%. In this case, the function scores of the present application may be function scores of a specific type of single nucleotide variant.

[0440] In another embodiment, the function score of the present application may be distinguished by a certain statistical calculation. In one specific embodiment, let us assume a case where a single nucleotide variant having a function score higher than a specific function score is distinguished as the gene functioning properly. In this case, the specific function score may be a value that enables distinction between a single nucleotide variant causing a nonsense variant that is highly likely to cause the ATM gene to not function properly and a single nucleotide variant causing a synonymous variant that is highly likely to cause the ATM gene to function properly. At this time, the specific function score can be obtained using Youden's Index.

[0441] The function score of the present application was classified through two classification criteria.

[0442] First, in order to distinguish that a single nucleotide variant having a function score lower than a specific function score will not function properly, the first distinction criterion of the present application is set to distinguish function scores based on the lower 5% of the function scores of the single nucleotide variants of the present application. Accordingly, the first distinction criterion of the present application may be -1.360.

[0443] Secondly, in order to distinguish whether a single nucleotide variant having a function score higher than a specific function score will function properly, the second distinction criterion of the present application was established to distinguish function scores based on a standard obtained using the aforementioned Youden's Index. Accordingly, the second distinction criterion of the present application may be -0.912.

[0444] If the function score of a specific single nucleotide variant is lower than -1.360, the specific single nucleotide variant is described as a dysfunctional single nucleotide variant. Additionally, ATM genes containing specific single nucleotide variants are described as being in a dysfunctional state.

[0445] If the function score of a specific single nucleotide variant is higher than -0.912, the specific single nucleotide variant is expressed as a functional single nucleotide variant. Additionally, ATM genes containing the specific single nucleotide variant are expressed as being in a functional state.

[0446] If the function score of a specific single nucleotide variant is between -1.360 and -0.912, the specific single nucleotide variant is represented as an intermediate single nucleotide variant. Additionally, ATM genes containing specific single nucleotide variants are represented as being in an intermediate state.

[0447] Single nucleotide mutation of dysfunction

[0448] The functional integration score of the present application obtained by the method for evaluating the function of ATM genes having single nucleotide mutations of the present application is as shown in Table 1. The functional integration score of the present application in Table 1 is described only for single nucleotide mutations of functional impairment.

[0449] [Table 1]

[0450]

[0451]

[0452]

[0453]

[0454]

[0455]

[0456]

[0457]

[0458]

[0459]

[0460]

[0461]

[0462]

[0463]

[0464]

[0465]

[0466]

[0467]

[0468]

[0469]

[0470]

[0471]

[0472]

[0473]

[0474]

[0475]

[0476]

[0477]

[0478]

[0479]

[0480]

[0481]

[0482]

[0483]

[0484]

[0485]

[0486]

[0487]

[0488]

[0489]

[0490]

[0491]

[0492]

[0493]

[0494]

[0495]

[0496]

[0497]

[0498]

[0499]

[0500]

[0501]

[0502]

[0503]

[0504]

[0505]

[0506]

[0507]

[0508]

[0509]

[0510]

[0511]

[0512]

[0513]

[0514]

[0515]

[0516]

[0517]

[0518]

[0519]

[0520]

[0521]

[0522]

[0523]

[0524]

[0525]

[0526]

[0527]

[0528]

[0529]

[0530]

[0531]

[0532]

[0533]

[0534]

[0535]

[0536]

[0537]

[0538]

[0539]

[0540]

[0541]

[0542]

[0543]

[0544]

[0545]

[0546]

[0547]

[0548]

[0549]

[0550]

[0551]

[0552]

[0553]

[0554]

[0555]

[0556]

[0557]

[0558]

[0559]

[0560]

[0561]

[0562]

[0563]

[0564]

[0565]

[0566]

[0567]

[0568]

[0569]

[0570]

[0571]

[0572]

[0573]

[0574]

[0575]

[0576]

[0577]

[0578]

[0579]

[0580]

[0581]

[0582]

[0583]

[0584]

[0585]

[0586]

[0587]

[0588]

[0589]

[0590]

[0591]

[0592]

[0593]

[0594]

[0595]

[0596]

[0597]

[0598]

[0599]

[0600]

[0601]

[0602]

[0603]

[0604]

[0605]

[0606]

[0607]

[0608]

[0609]

[0610]

[0611]

[0612]

[0613]

[0614]

[0615]

[0616]

[0617]

[0618]

[0619]

[0620]

[0621]

[0622]

[0623]

[0624]

[0625]

[0626]

[0627]

[0628]

[0629]

[0630]

[0631]

[0632]

[0633]

[0634]

[0635]

[0636]

[0637]

[0638]

[0639]

[0640]

[0641]

[0642]

[0643]

[0644]

[0645]

[0646]

[0647]

[0648]

[0649]

[0650]

[0651]

[0652]

[0653]

[0654]

[0655]

[0656]

[0657]

[0658]

[0659]

[0660]

[0661]

[0662]

[0663]

[0664]

[0665]

[0666]

[0667]

[0668]

[0669]

[0670]

[0671]

[0672]

[0673]

[0674]

[0675]

[0676]

[0677]

[0678]

[0679] "NO." indicates the position of the single nucleotide variant relative to the reference sequence (Sequence No. 1). "Type" indicates the position and type of the single nucleotide variant. For example, I signifies an intron, P signifies a splice acceptor / donor site, M signifies a missense variant, N signifies a nonsense variant, and S signifies a synonymous variant. "Single nucleotide variant" refers to the type of single nucleotide variant according to HGVS nomenclature. "Function score" refers to the function score of the single nucleotide variant. "AA variant" refers to the type of amino acid variation caused by the single nucleotide variant. "AA_Function score" refers to the function score based on the amino acid variation. "Prediction" indicates whether the integrated function score is a function evaluation score or a function prediction score. If indicated as N, it is a function evaluation score, and if indicated as Y, it is a function prediction score.

[0680] Moderate single nucleotide mutation

[0681] The list of intermediate single nucleotide variants obtained by the ATM gene function evaluation method having single nucleotide variants of the present application is as follows.

[0682] c.-5G>A, c.-3A>G, c.4A>C, c.14T>G, c.20A>T, c.34T>G, c.37C>T, c.48A>T, c.49C>G, c.50A>C, c.52G>C, c.55A>C, c.56G>T, c.57A>C, c.68G>T, c.70A>G, c.71A>C, c.72G>A, c.72+2T>G, c.73-2A>G, c.73-2A>T, c.73-1G>A, c.73-1G>C, c.73A>T, c.76G>T, c.82G>T, c.85A>T, c.89T>C, c.89T>G, c.95G>C, c.101T>G, c.109C>A, c.110C>T, c.113A>T, c.115A>T, c.116C>A, c.131A>C, c.134G>C, c.140C>G, c.147C>A, c.157A>G, c.158A>G, c.167A>T, c.168T>A, c.168T>G, c.174T>G, c.174T>C, c.177T>G, c.183T>A, c.185+2T>A, c.187T>C, c.189T>G, c.189T>A, c.192A>T, c.192A>C, c.199T>A, c.203T>C, c.207G>A, c.236A>T, c.245T>G, c.246A>G, c.247T>A, c.250G>C, c.251C>T, c.253T>A, c.253T>C, c.254C>T, c.255A>C, c.255A>G, c.256A>T, c.256A>G, c.259C>G, c.260A>C, c.264C>T, c.269G>T, c.270G>T, c.277A>C, c.290T>G, c.295A>C, c.296G>C, c.296G>T, c.299T>C, c.305A>C, c.305A>T, c.311T>G, c.315C>T, c.317A>T, c.318A>T, c.320G>A, c.321T>A, c.325A>C, c.333A>C, c.338C>A, c.340A>G, c.355G>C, c.359T>G, c.364A>T, c.365A>T, c.367T>G, c.371T>G, c.371T>A, c.373A>C, c.375G>A, c.386A>C, c.387A>T, c.389A>T, c.389A>G, c.390T>C, c.391T>C, c.413G>T, c.416C>G, c.427A>T, c.427A>G, c.436C>A, c.441A>G, c.444C>T, c.445A>T, c.451T>C, c.452C>A, c.455T>G, c.455T>A, c.462A>T, c.462A>C, c.464A>T, c.469T>A, c.469T>C, c.470G>C, c.472G>A, c.473A>C, c.474A>T, c.478T>C, c.479C>G, c.481C>G, c.482A>G, c.488A>C, c.497-5T>C, c.497-2A>G, c.497A>T, c.497A>C, c.499T>G, c.500T>C, c.504C>A, c.525T>G, c.536C>G, c.541G>A, c.545T>C, c.558A>T, c.565A>G, c.567A>C, c.572T>C, c.580G>T, c.584C>A, c.621C>A, c.637T>G, c.639T>G, c.645G>A, c.647C>A, c.651T>G, c.654G>C, c.655T>C, c.662+1G>C, c.663-3C>T, c.663-2A>T, c.663-1G>A, c.667G>C, c.682G>C, c.703G>C, c.704C>G, c.722A>C, c.723G>A, c.728T>A, c.730G>C, c.735C>T, c.736A>C, c.737A>C, c.744A>T, c.749G>A, c.751G>T, c.752T>G, c.754T>C, c.757G>T, c.763G>C, c.766G>T, c.769G>T, c.772A>T, c.776T>C, c.776T>A, c.780C>A, c.788T>G, c.793A>C, c.803A>C, c.810G>T, c.832G>A, c.838A>G, c.847T>A, c.869A>T, c.871C>G, c.874C>G, c.884C>G, c.897A>T, c.898A>T, c.900A>G, c.902G>T, c.905C>A, c.922T>C, c.924G>C, c.924G>T, c.927A>T, c.946T>A, c.963T>C, c.967A>C, c.969A>T, c.972T>A, c.980G>T, c.987A>C, c.992A>C, c.992A>G, c.1000T>C, c.1009C>A, c.1010G>C, c.1010G>T, c.1014T>A, c.1020C>A, c.1034T>C, c.1042T>C, c.1051G>T, c.1052A>T, c.1052A>G, c.1059T>G, c.1061A>C, c.1065+5A>T, c.1066-5T>A, c.1066-3C>A, c.1066-2A>C, c.1067T>A, c.1075G>A, c.1088C>A, c.1088C>T, c.1095G>C, c.1096A>T, c.1110C>A, c.1115C>A, c.1122A>G, c.1153A>G, c.1154A>G, c.1162A>G, c.1169A>G, c.1175G>A, c.1178G>A, c.1179G>C, c.1179G>T, c.1187T>C, c.1202A>C, c.1216G>T, c.1219T>A, c.1229T>A, c.1232C>G, c.1234T>C, c.1235+4A>G, c.1236-5T>A, c.1236-4T>G, c.1236-2A>T, c.1236G>T, c.1244T>G, c.1268A>C, c.1276A>C, c.1294C>T, c.1318T>C, c.1333C>G, c.1343A>C, c.1351C>G, c.1372T>G, c.1387G>A, c.1397A>G, c.1399G>A, c.1405A>C, c.1412A>T, c.1415T>A, c.1415T>C, c.1428A>G, c.1430A>G, c.1436A>T, c.1458A>T, c.1463G>C, c.1464G>C, c.1467T>G, c.1471A>C, c.1473C>T, c.1477C>G, c.1484T>A, c.1493A>C, c.1494G>A, c.1494G>T, c.1500A>G, c.1503A>G, c.1512C>T, c.1513T>G, c.1516G>A, c.1523T>A, c.1528G>C, c.1546T>G, c.1550T>G, c.1558G>T, c.1559A>T, c.1559A>C, c.1560C>G, c.1571G>T, c.1577T>C, c.1580T>G, c.1585G>C, c.1598G>A, c.1601C>A, c.1606T>A, c.1607G>C, c.1607+3A>T, c.1608-4A>T, c.1615G>T, c.1640C>T, c.1671G>A, c.1675A>G, c.1687A>T, c.1721A>T, c.1721A>G, c.1723T>C, c.1723T>G, c.1724C>G, c.1730T>G, c.1730T>A, c.1733A>C, c.1742T>C, c.1757A>T, c.1757A>G, c.1758G>T, c.1759G>A, c.1759G>T, c.1787C>T, c.1793T>A, c.1796T>C, c.1802+3A>G, c.1802+4A>G, c.1823T>C, c.1832T>G, c.1833T>A, c.1847C>A, c.1850T>A, c.1851G>T, c.1851G>C, c.1851G>A, c.1852A>G, c.1857C>G, c.1857C>A, c.1858T>C, c.1859G>A, c.1862A>T, c.1864G>C, c.1871T>A, c.1872G>T, c.1872G>A, c.1905C>G, c.1925A>C, c.1933T>A, c.1934T>A, c.1939G>A, c.1947A>C, c.1949A>T, c.1951C>G, c.1955T>A, c.1957C>T, c.1961A>C, c.1961A>T, c.1964C>A, c.1965A>G, c.1966A>C, c.1968T>A, c.1989A>C, c.2024A>C, c.2048A>C, c.2093C>T, c.2116T>G, c.2118A>T, c.2123A>C, c.2124+3G>A, c.2124+4A>C, c.2125-5T>A, c.2125-5T>G, c.2155T>C, c.2156C>T, c.2156C>G, c.2165T>A, c.2169G>T, c.2170G>A, c.2171G>T, c.2171G>C, c.2172T>A, c.2172T>C, c.2178T>A, c.2179G>C, c.2181C>G, c.2181C>T, c.2187C>A, c.2209G>A, c.2249A>T, c.2258T>G, c.2268A>T, c.2270G>C, c.2271A>T, c.2272G>T, c.2276G>T, c.2278A>G, c.2279T>G, c.2283T>C, c.2286G>T, c.2300C>G, c.2302A>G, c.2310A>G, c.2313C>A, c.2336T>G, c.2341C>A, c.2342A>C, c.2350A>G, c.2352A>C, c.2356T>C, c.2358C>A, c.2362A>C, c.2362A>G, c.2364C>G, c.2368T>A, c.2376+1G>A, c.2376+1G>C, c.2377-5T>A, c.2377-1G>A, c.2377-1G>C, c.2377-1G>T, c.2419T>C, c.2420T>A, c.2428A>T, c.2431C>T, c.2433A>T, c.2434A>T, c.2455T>C, c.2467-4G>T, c.2467-1G>C, c.2497G>A, c.2499A>T, c.2507A>C, c.2509T>A, c.2509T>C, c.2511A>C, c.2515G>A, c.2518G>C, c.2533A>C, c.2547G>T, c.2570T>C, c.2606C>G, c.2607A>G, c.2608A>T, c.2612A>T, c.2613A>T, c.2614C>A, c.2615C>T, c.2615C>A, c.2638+4A>G, c.2638+4A>T, c.2639-2A>G, c.2639-1G>T, c.2639G>A, c.2654T>G, c.2669T>A, c.2670G>A, c.2675A>C, c.2680G>C, c.2681A>T, c.2681A>C, c.2687T>A, c.2696A>C, c.2699T>G, c.2702T>C, c.2706G>T, c.2714G>T, c.2726C>A, c.2736G>A, c.2751C>T, c.2775G>C, c.2828A>T, c.2831T>A, c.2835T>A, c.2841T>C, c.2847G>A, c.2858A>G, c.2859G>A, c.2864C>A, c.2878C>T, c.2891A>T, c.2893G>T, c.2894A>G, c.2897T>A, c.2915C>T, c.2922-1G>T, c.2928G>A, c.2929T>G, c.2931T>G, c.2937G>C, c.2938T>A, c.2939A>G, c.2941C>G, c.2945G>C, c.2947G>C, c.2948A>T, c.2950C>A, c.2950C>G, c.2951A>T, c.2955T>A, c.2992G>T, c.3002T>G, c.3007C>T, c.3048A>T, c.3056T>A, c.3068G>T, c.3070G>A, c.3073T>A, c.3076T>G, c.3077+3A>T, c.3078-3C>G, c.3078-2A>C, c.3078-2A>G, c.3078-1G>T, c.3078G>C, c.3079C>G, c.3083T>G, c.3089A>G, c.3090G>T, c.3092A>T, c.3101A>G, c.3102T>G, c.3106T>A, c.3111T>A, c.3146T>G, c.3152A>T, c.3153+3G>C, c.3157G>T, c.3158A>C, c.3171A>C, c.3176C>A, c.3188T>A, c.3192G>A, c.3211A>G, c.3214G>C, c.3215A>T, c.3215A>G, c.3216A>G, c.3219A>C, c.3219A>T, c.3246T>A, c.3256C>G, c.3260T>C, c.3283A>G, c.3334C>A, c.3355G>C, c.3358T>C, c.3359T>C, c.3367G>C, c.3371A>T, c.3379G>C, c.3384G>C, c.3391A>C, c.3399A>T, c.3403-2A>G, c.3403-1G>C, c.3424G>T, c.3434A>G, c.3434A>C, c.3436G>C, c.3436G>T, c.3438A>G, c.3439A>C, c.3441T>A, c.3442T>A, c.3444T>G, c.3447T>A, c.3449G>C, c.3454T>A, c.3462A>T, c.3470T>A, c.3476C>T, c.3481G>T, c.3482T>A, c.3485T>G, c.3502T>C, c.3504C>A, c.3510A>T, c.3514G>C, c.3518T>A, c.3527T>A, c.3527T>C, c.3529T>C, c.3531T>A, c.3538G>C, c.3554T>A, c.3566T>C, c.3568G>A, c.3582A>C, c.3586A>C, c.3590T>G, c.3593C>A, c.3593C>T, c.3608A>C, c.3609T>G, c.3619G>C, c.3621A>T, c.3623A>T, c.3624C>A, c.3626T>C, c.3660A>C, c.3661T>A, c.3665T>A, c.3673C>G, c.3674A>G, c.3689A>G, c.3690C>T, c.3693A>G, c.3700T>C, c.3700T>A, c.3703C>T, c.3713T>C, c.3716T>C, c.3718A>T, c.3718A>G, c.3730A>T, c.3741C>G, c.3741C>A, c.3746+3A>T, c.3747A>T, c.3754T>G, c.3769C>T, c.3769C>G, c.3773A>T, c.3774T>A, c.3791A>G, c.3794T>G, c.3807G>A, c.3811A>T, c.3812T>C, c.3815C>T, c.3821A>C, c.3824T>G, c.3824T>A, c.3830A>G, c.3833A>T, c.3833A>G, c.3836G>T, c.3844C>A, c.3845T>G, c.3848T>G, c.3852A>C, c.3852A>G, c.3856T>C, c.3869T>G, c.3873T>A, c.3880A>T, c.3882T>C, c.3894T>C, c.3896C>T, c.3896C>A, c.3901G>A, c.3902A>G, c.3908C>A, c.3917G>T, c.3952G>T, c.3975A>C, c.3983T>C, c.3985G>A, c.3985G>C, c.3986G>A, c.3992A>C, c.3993+2T>C, c.4010T>A, c.4016A>C, c.4021C>T, c.4022C>A, c.4022C>T, c.4028T>G, c.4037A>G, c.4041A>C, c.4044G>C, c.4044G>T, c.4045A>C, c.4047G>A, c.4053A>C, c.4054C>A, c.4056T>A, c.4057G>A, c.4058A>G, c.4061C>T, c.4073G>T, c.4080T>G, c.4081C>A, c.4082A>G, c.4084A>C, c.4085G>C, c.4085G>A, c.4085G>T, c.4097G>T, c.4100A>T, c.4124C>A, c.4127C>G, c.4130A>T, c.4135C>A, c.4135C>G, c.4136C>T, c.4141T>A, c.4142T>C, c.4147T>G, c.4151A>G, c.4163C>T, c.4169T>G, c.4171G>A, c.4178T>C, c.4182C>G, c.4182C>A, c.4186T>A, c.4186T>G, c.4187G>T, c.4201T>G, c.4208G>C, c.4217A>C, c.4223T>G, c.4223T>A, c.4236+5G>C, c.4237-1G>A, c.4252A>T, c.4253T>G, c.4256T>A, c.4265T>G, c.4273C>T, c.4282G>A, c.4285A>C, c.4293T>G, c.4324T>C, c.4325A>C, c.4328A>G, c.4338T>A, c.4349T>A, c.4357A>T, c.4394T>A, c.4399G>A, c.4400A>C, c.4410T>G, c.4411A>T, c.4418T>A, c.4419T>G, c.4435A>G, c.4436G>C, c.4438C>T, c.4439C>T, c.4449C>T, c.4452G>C, c.4453G>A, c.4454A>G, c.4478T>C, c.4484G>C, c.4486G>C, c.4491A>T, c.4502T>C, c.4505G>A, c.4510A>T, c.4514C>T, c.4517T>A, c.4518G>T, c.4525T>A, c.4528A>C, c.4529A>T, c.4530G>T, c.4533T>A, c.4534G>C, c.4537C>G, c.4538T>G, c.4543A>G, c.4547A>C, c.4551T>A, c.4552C>T, c.4556T>A, c.4564G>A, c.4568C>G, c.4576C>T, c.4576C>G, c.4583T>A, c.4592A>G, c.4599G>C, c.4600G>C, c.4600G>A, c.4601T>A, c.4604A>T, c.4610A>C, c.4611+4A>T, c.4613T>G, c.4629A>G, c.4634T>A, c.4639A>G, c.4669A>T, c.4670T>C, c.4676T>C, c.4676T>G, c.4679A>T, c.4680G>C, c.4733A>C, c.4738A>T, c.4749C>G, c.4753A>T, c.4754G>T, c.4774G>T, c.4775A>C, c.4776G>C, c.4776+1G>C, c.4776+2T>A, c.4776+2T>C, c.4776+2T>G, c.4777G>C, c.4781T>A, c.4782T>G, c.4783A>T, c.4784A>G, c.4787A>C, c.4789T>G, c.4789T>C, c.4790T>C, c.4790T>G, c.4791T>C, c.4794C>G, c.4797A>C, c.4797A>T, c.4798G>T, c.4800A>C, c.4801A>C, c.4804G>C, c.4804G>A, c.4810G>T, c.4811A>T, c.4814C>A, c.4819C>G, c.4837G>A, c.4841T>A, c.4849C>A, c.4849C>T, c.4850T>G, c.4852C>T, c.4870C>G, c.4874A>T, c.4880A>G, c.4880A>C, c.4884G>T, c.4891A>C, c.4896G>A, c.4897A>G, c.4901C>G, c.4902T>A, c.4902T>G, c.4903T>G, c.4910-4C>A, c.4910-4C>T, c.4915C>A, c.4915C>G, c.4924G>C, c.4928T>G, c.4928T>A, c.4929T>G, c.4937A>C, c.4937A>G, c.4943T>G, c.4947C>G, c.4948A>T, c.4949A>T, c.4951T>G, c.4953G>C, c.4973C>T, c.4979A>T, c.4988G>T, c.4994A>T, c.4999G>T, c.5002C>G, c.5003T>C, c.5004A>G, c.5005+3A>G, c.5005+4A>T, c.5009C>T, c.5013T>G, c.5019C>A, c.5021G>T, c.5026G>A, c.5031A>C, c.5039C>T, c.5056A>G, c.5069A>C, c.5070T>G, c.5075A>T, c.5083T>C, c.5089A>T, c.5089A>C, c.5093A>G, c.5107T>A, c.5162C>G, c.5162C>A, c.5162C>T, c.5170G>C, c.5171A>C, c.5195C>G, c.5200G>T, c.5204C>G, c.5206T>A, c.5207G>C, c.5208T>A, c.5210T>A, c.5218A>C, c.5221T>A, c.5225C>T, c.5232G>T, c.5239C>A, c.5245T>A, c.5246T>A, c.5265G>C, c.5266A>G, c.5269A>G, c.5272G>A, c.5275C>T, c.5276C>G, c.5276C>A, c.5282T>C, c.5288A>T, c.5290C>G, c.5294A>C, c.5296C>T, c.5303G>T, c.5304A>C, c.5309C>A, c.5310A>T, c.5313A>C, c.5318A>C, c.5322T>A, c.5354C>T, c.5360A>T, c.5360A>C, c.5366T>G, c.5366T>A, c.5371G>T, c.5378A>T, c.5395A>C, c.5404C>T, c.5405A>G, c.5407G>T, c.5407G>C, c.5408A>T, c.5411T>G, c.5412T>G, c.5415G>A, c.5419A>G, c.5430T>A, c.5432G>A, c.5435C>T, c.5437T>G, c.5438T>A, c.5438T>C, c.5439T>A, c.5442G>C, c.5443G>T, c.5443G>A, c.5446A>C, c.5448T>A, c.5450G>T, c.5461T>A, c.5468T>A, c.5473C>A, c.5473C>G, c.5477T>C, c.5484G>T, c.5485C>A, c.5489T>C, c.5492G>C, c.5497-5T>C, c.5499G>C, c.5501A>C, c.5504C>A, c.5507A>T, c.5508C>A, c.5509T>A, c.5512T>A, c.5516A>C, c.5518A>C, c.5522T>G, c.5527C>A, c.5527C>G, c.5528C>G, c.5531A>C, c.5532C>G, c.5539C>T, c.5540A>T, c.5543A>C, c.5548T>G, c.5558A>T, c.5564A>T, c.5574G>T, c.5575A>G, c.5575A>T, c.5576G>T, c.5577A>T, c.5582T>G, c.5582T>C, c.5585T>G, c.5585T>C, c.5587T>G, c.5594A>T, c.5597T>G, c.5600A>T, c.5603G>T, c.5603G>C, c.5606T>A, c.5606T>G, c.5607T>A, c.5611A>C, c.5612C>T, c.5657C>T, c.5661A>G, c.5661A>T, c.5662A>G, c.5674+3A>C, c.5674+3A>T, c.5675A>T, c.5675A>C, c.5690T>C, c.5702T>A, c.5703G>T, c.5707A>C, c.5709A>T, c.5710A>C, c.5712A>C, c.5714C>T, c.5717A>C, c.5723C>G, c.5737G>A, c.5752A>G, c.5753G>T, c.5754A>T, c.5755C>A, c.5762G>T, c.5763A>C, c.5776A>C, c.5777C>G, c.5779A>T, c.5782T>G, c.5782T>A, c.5783T>C, c.5786A>T, c.5788G>T, c.5791G>C, c.5794T>C, c.5797T>A, c.5804A>C, c.5809A>C, c.5819A>G, c.5820A>C, c.5821G>C, c.5822T>C, c.5825C>G, c.5829G>T, c.5829G>C, c.5830G>A, c.5833G>T, c.5836C>A, c.5842T>C, c.5843G>T, c.5845G>C, c.5856T>A, c.5857A>T, c.5865A>T, c.5866C>A, c.5869T>C, c.5873C>T, c.5878A>G, c.5880C>G, c.5882A>T, c.5885C>T, c.5885C>G, c.5887G>A, c.5888A>C, c.5891A>C, c.5891A>T, c.5903A>T, c.5911G>T, c.5918+1G>C, c.5918+2T>C, c.5919-3C>T, c.5951C>T, c.5953A>G, c.5957T>A, c.5963G>C, c.5965T>G, c.5966T>G, c.5967G>C, c.5967G>T, c.5969G>A, c.5972A>T, c.5972A>G, c.5973A>T, c.5976A>T, c.5976A>C, c.5977A>G, c.5977A>T, c.5979T>A, c.5979T>G, c.5983G>C, c.5983G>A, c.5984A>G, c.5985A>T, c.5985A>C, c.5986G>C, c.5986G>A, c.5987A>G, c.5989A>C, c.5989A>G, c.5990C>T, c.5992G>A, c.5993G>A, c.5995A>G, c.5997A>G, c.5999G>A, c.5999G>T, c.6001T>G, c.6001T>A, c.6007-5T>G, c.6007G>A, c.6009T>G, c.6016T>C, c.6018A>T, c.6022A>G, c.6024C>T, c.6025T>C, c.6026A>C, c.6028A>G, c.6031A>C, c.6034A>T, c.6043C>G, c.6050G>T, c.6056A>T, c.6058G>A, c.6062G>C, c.6062G>T, c.6064G>A, c.6064G>C, c.6065G>A, c.6069A>G, c.6070G>A, c.6072G>T, c.6086C>A, c.6092C>A, c.6104C>G, c.6117A>C, c.6118G>A, c.6123G>C, c.6128G>T, c.6142A>T, c.6146A>C, c.6151C>T, c.6157A>C, c.6158C>T, c.6166C>T, c.6173C>T, c.6175A>C, c.6184G>T, c.6185C>A, c.6190A>G, c.6195T>C, c.6199-5T>A, c.6199G>A, c.6209A>C, c.6211T>G, c.6213G>C, c.6215G>C, c.6223C>T, c.6223C>A, c.6229C>A, c.6233C>T, c.6236T>G, c.6247G>A, c.6249A>T, c.6252G>A, c.6254A>G, c.6293T>G, c.6295C>G, c.6305C>G, c.6305C>A, c.6307G>A, c.6308C>G, c.6323A>T, c.6329A>T, c.6336C>G, c.6348-4A>T, c.6371A>C, c.6378A>T, c.6379T>C, c.6384G>T, c.6384G>C, c.6390T>A, c.6401C>T, c.6402T>C, c.6405A>C, c.6407G>C, c.6407G>T, c.6408A>T, c.6409G>T, c.6414A>C, c.6416A>T, c.6416A>G, c.6416A>C, c.6423T>A, c.6426A>G, c.6427T>A, c.6428T>A, c.6439C>G, c.6447T>C, c.6448G>T, c.6449C>A, c.6452+1G>C, c.6452+5T>A, c.6453A>T, c.6455T>C, c.6458A>G, c.6460G>T, c.6476G>T, c.6479A>G, c.6481C>A, c.6481C>T, c.6482G>C, c.6483C>G, c.6483C>T, c.6485G>T, c.6486C>A, c.6487C>G, c.6493T>A, c.6494C>T, c.6502T>C, c.6508T>G, c.6511C>T, c.6512C>A, c.6521G>T, c.6522C>G, c.6524G>T, c.6525G>T, c.6526T>A, c.6528G>C, c.6528G>T, c.6529C>A, c.6530A>T, c.6543G>T, c.6548A>C, c.6553A>C, c.6554T>C, c.6554T>G, c.6555T>C, c.6556G>T, c.6556G>A, c.6557G>T, c.6557G>C, c.6558G>A, c.6558G>C, c.6558G>T, c.6560A>T, c.6561G>A, c.6561G>C, c.6563T>C, c.6564T>G, c.6573-5T>G, c.6573-3C>A, c.6573-2A>T, c.6573A>C, c.6574T>A, c.6575C>T, c.6582A>C, c.6587G>C, c.6590A>T, c.6594C>G, c.6597T>A, c.6602T>C, c.6605A>T, c.6615G>T, c.6622C>T, c.6638A>C, c.6646G>T, c.6649T>G, c.6656T>G, c.6664C>G, c.6668T>A, c.6672G>C, c.6672G>T, c.6673G>A, c.6674C>G, c.6689T>G, c.6692T>C, c.6704T>G, c.6713A>T, c.6717G>A, c.6723C>A, c.6724T>G, c.6736T>A, c.6752T>G, c.6770A>T, c.6773T>A, c.6801C>G, c.6801C>A, c.6804T>A, c.6807+3A>C, c.6808-2A>G, c.6811C>T, c.6821C>G, c.6825A>C, c.6826T>G, c.6849A>T, c.6873G>C, c.6884A>C, c.6886G>T, c.6897C>G, c.6905A>C, c.6910G>A, c.6925C>G, c.6926T>G, c.6929G>T, c.6944T>C, c.6944T>A, c.6947T>G, c.6947T>A, c.6972A>G, c.6975+4T>C, c.6983C>A, c.6988C>T, c.6991A>G, c.6992A>C, c.7000T>A, c.7001A>C, c.7017G>C, c.7026C>T, c.7027A>C, c.7027A>G, c.7028A>T, c.7052A>T, c.7052A>C, c.7061C>A, c.7065C>G, c.7066A>G, c.7070T>G, c.7071G>A, c.7073A>C, c.7081C>A, c.7089+5G>C, c.7089+5G>T, c.7103C>T, c.7130A>C, c.7131T>C, c.7134G>C, c.7137A>C, c.7137A>T, c.7148A>T, c.7157C>A, c.7158A>C, c.7166C>T, c.7170A>T, c.7171G>A, c.7178T>A, c.7187C>A, c.7193A>C, c.7196A>T, c.7197A>G, c.7198A>G, c.7202T>C, c.7207A>T, c.7209C>G, c.7209C>A, c.7211A>G, c.7214T>C, c.7215G>A, c.7215G>T, c.7215G>C, c.7218A>C, c.7220C>G, c.7223C>T, c.7226A>C, c.7228T>A, c.7238A>C, c.7239G>C, c.7239G>T, c.7253A>T, c.7259C>T, c.7262A>T, c.7263A>T, c.7282A>G, c.7284G>C, c.7284G>T, c.7286A>T, c.7286A>G, c.7287A>G, c.7289A>G, c.7290T>A, c.7291A>C, c.7291A>G, c.7295T>C, c.7298A>T, c.7300A>C, c.7304A>T, c.7307+2T>C, c.7307+4A>G, c.7316T>G, c.7326G>A, c.7328G>A, c.7334T>A, c.7335G>C, c.7340T>C, c.7344T>A, c.7344T>C, c.7351G>A, c.7361C>G, c.7361C>A, c.7364T>A, c.7364T>C, c.7371G>T, c.7372G>A, c.7389A>C, c.7390T>A, c.7395A>T, c.7395A>C, c.7396G>A, c.7397C>A, c.7399G>T, c.7403A>T, c.7404A>T, c.7406A>C, c.7409A>T, c.7417T>A, c.7430G>T, c.7432G>A, c.7433A>G, c.7436A>G, c.7437A>T, c.7447T>G, c.7448G>C, c.7449G>C, c.7450G>T, c.7451T>A, c.7451T>C, c.7454T>G, c.7459C>A, c.7466C>G, c.7474C>G, c.7479A>T, c.7480A>C, c.7493C>A, c.7498G>T, c.7503T>G, c.7508T>A, c.7510A>T, c.7513A>G, c.7515+3A>G, c.7517G>A, c.7519G>A, c.7521C>A, c.7532T>C, c.7535C>A, c.7537A>T, c.7539A>T, c.7543A>C, c.7545A>C, c.7545A>T, c.7546T>G, c.7546T>C, c.7549T>A, c.7552C>A, c.7556T>C, c.7559T>C, c.7559T>A, c.7562A>T, c.7567T>G, c.7571C>A, c.7574C>T, c.7576A>G, c.7577G>A, c.7578A>T, c.7581G>C, c.7583G>A, c.7584G>A, c.7585A>C, c.7586C>A, c.7587C>G, c.7589A>G, c.7590G>A, c.7597G>C, c.7598G>T, c.7610T>A, c.7610T>C, c.7625A>G, c.7627A>C, c.7627A>G, c.7628A>C, c.7629+3A>T, c.7629+4A>T, c.7634T>C, c.7636T>G, c.7651G>T, c.7663C>G, c.7664A>G, c.7670T>A, c.7671G>A, c.7675A>C, c.7682T>C, c.7688T>G, c.7693A>C, c.7693A>G, c.7696G>T, c.7700A>T, c.7704A>G, c.7708G>T, c.7720A>T, c.7726G>T, c.7735A>T, c.7765A>T, c.7769A>G, c.7773C>A, c.7774T>C, c.7776T>A, c.7781T>C, c.7783G>T, c.7783G>A, c.7783G>C, c.7788G>T, c.7788+1G>T, c.7788+2T>G, c.7788+3A>G, c.7789G>C, c.7795A>C, c.7802C>T, c.7803T>A, c.7804G>A, c.7805C>T, c.7825A>C, c.7830A>G, c.7831A>T, c.7838G>T, c.7847T>C, c.7850T>C, c.7852A>G, c.7855A>T, c.7863G>T, c.7874A>G, c.7883T>C, c.7891G>T, c.7894A>T, c.7899A>G, c.7904C>T, c.7907C>A, c.7919C>T, c.7927+5G>C, c.7928-2A>T, c.7928-1G>T, c.7934T>A, c.7943C>A, c.7943C>T, c.7951C>G, c.7954C>G, c.7955C>T, c.7957A>T, c.7970A>T, c.7973A>T, c.8006T>G, c.8014G>C, c.8015A>T, c.8042T>C, c.8045C>A, c.8059A>G, c.8080G>C, c.8081G>C, c.8089A>T, c.8095C>A, c.8102T>G, c.8119T>G, c.8120C>G, c.8137A>G, c.8140C>G, c.8146G>C, c.8148T>A, c.8151+5G>T, c.8152G>C, c.8153G>T, c.8162A>G, c.8164C>G, c.8169A>C, c.8170C>G, c.8171A>G, c.8177C>A, c.8182A>G, c.8184G>T, c.8184G>A, c.8184G>C, c.8186A>G, c.8186A>T, c.8187A>C, c.8190G>C, c.8191G>A, c.8194T>G, c.8199G>A, c.8199G>C, c.8200A>C, c.8200A>T, c.8201T>G, c.8203T>C, c.8204G>A, c.8205T>G, c.8205T>A, c.8210C>G, c.8218C>A, c.8221A>T, c.8225A>G, c.8228C>G, c.8234C>A, c.8234C>T, c.8238G>T, c.8240A>G, c.8240A>T, c.8241G>C, c.8242A>C, c.8245A>G, c.8254A>T, c.8264A>C, c.8264A>G, c.8265T>C, c.8269-5T>A, c.8269-5T>C, c.8269-4C>T, c.8269-4C>G, c.8269-3C>A, c.8271G>C, c.8272G>A, c.8274T>C, c.8274T>A, c.8275C>A, c.8277C>T, c.8277C>A, c.8287C>G, c.8308T>A, c.8312C>G, c.8315G>A, c.8317A>T, c.8321T>G, c.8321T>A, c.8323C>A, c.8323C>T, c.8324C>G, c.8324C>T, c.8329G>A, c.8333A>G, c.8333A>C, c.8345A>C, c.8354A>G, c.8369G>C, c.8374A>T, c.8374A>G, c.8376G>C, c.8377C>T, c.8378C>G, c.8378C>T, c.8381A>G, c.8390G>A, c.8393C>T, c.8394C>A, c.8395T>C, c.8395T>A, c.8396T>G, c.8398C>A, c.8399A>T, c.8400G>A, c.8401T>C, c.8404C>A, c.8404C>T, c.8405A>T, c.8409G>T, c.8409G>C, c.8412A>C, c.8412A>T, c.8413A>G, c.8413A>T, c.8413A>C, c.8419-5T>C, c.8419-2A>G, c.8419-2A>C, c.8425C>A, c.8428A>C, c.8435C>A, c.8438T>G, c.8439T>A, c.8439T>G, c.8440G>C, c.8441A>T, c.8442A>T, c.8442A>C, c.8445G>C, c.8445G>T, c.8449T>A, c.8450A>G, c.8452G>C, c.8452G>A, c.8456T>A, c.8464G>T, c.8476A>G, c.8478T>G, c.8503T>A, c.8506A>C, c.8507T>C, c.8522A>G, c.8524C>T, c.8530A>T, c.8540A>G, c.8545C>A, c.8549T>C, c.8557A>T, c.8558C>T, c.8558C>A, c.8564G>A, c.8566G>A, c.8568A>G, c.8569G>A, c.8575T>G, c.8576C>G, c.8584+3A>G, c.8584+4A>G, c.8584+4A>C, c.8584+5T>C, c.8584+5T>A, c.8585-5T>C, c.8585-3C>A, c.8585T>C, c.8586T>C, c.8591A>T, c.8594T>C, c.8618T>C, c.8632A>G, c.8644T>C, c.8647G>T, c.8651A>C, c.8652A>T, c.8652A>C, c.8656G>C, c.8656G>T, c.8662A>G, c.8665G>C, c.8673T>C, c.8674G>C, c.8681T>A, c.8688G>A, c.8691C>A, c.8691C>T, c.8694A>G, c.8697C>A, c.8701C>A, c.8702C>T, c.8703T>A, c.8703T>G, c.8704A>T, c.8706T>G, c.8706T>A, c.8712G>C, c.8716G>A, c.8728C>A, c.8740A>C, c.8741T>C, c.8749G>T, c.8751C>T, c.8757C>T, c.8764G>C, c.8768T>C, c.8776G>A, c.8779T>C, c.8781C>A, c.8781C>G, c.8786+3A>T, c.8794G>C, c.8796G>T, c.8800A>G, c.8804T>A, c.8822C>G, c.8826G>C, c.8832T>A, c.8836T>G, c.8841C>A, c.8841C>T, c.8842A>G, c.8850+3A>G, c.8855T>G, c.8860T>C, c.8861A>G, c.8863G>A, c.8866C>T, c.8869C>T, c.8870T>A, c.8878T>G, c.8885T>C, c.8891C>A, c.8894T>G, c.8895G>T, c.8897A>G, c.8898A>T, c.8898A>G, c.8900C>A, c.8901T>A, c.8905T>G, c.8909T>C, c.8931A>C, c.8933C>G, c.8934T>C, c.8934T>A, c.8938C>T, c.8940T>C, c.8943C>G, c.8943C>A, c.8943C>T, c.8944C>A, c.8944C>G, c.8947A>C, c.8961T>A, c.8984T>A, c.8987G>A, c.8987G>T, c.8996A>C, c.9007A>C, c.9008A>C, c.9019G>A, c.9026T>A, c.9026T>C, c.9030A>T, c.9032T>A, c.9039A>C, c.9043G>A, c.9047A>C, c.9047A>T, c.9050T>C, c.9053A>C, c.9054A>C, c.9061G>C, c.9067G>A, c.9068G>A, c.9072T>A, c.9077T>A, c.9079A>C, c.9081T>A, c.9082G>C, c.9082G>T, c.9083T>G, c.9083T>A, c.9092A>T, c.9093A>G, c.9094G>A, c.9095T>C, c.9104T>A, c.9121G>T, c.9128A>C, c.9133C>T, and c.9168G>A.

[0683] Single nucleotide variants not listed in "dysfunctional single nucleotide variants" and "moderate single nucleotide variants" are functional single nucleotide variants.

[0684] Characteristics of function points

[0685] The function score of the present application is a score derived using cells having pseudo-single nucleotide variants in the ATM gene. The function score of the present application may be equivalent to a score derived using cells having single nucleotide variants in the ATM gene. The function score of the present application may be a score verified to be equivalent to a score derived using cells containing single nucleotide variants. In one embodiment, for cells to which pseudo-single nucleotide variants have been introduced, verification can be performed by changing the intended synonymous mutation and comparing them with each other. That is, verification can be performed by confirming that the function scores are similar if the single nucleotide variants remain the same, even if the intended synonymous mutations are changed. In another embodiment, verification can be performed by comparing with a database already known regarding the pathogenicity of the single nucleotide variant.

[0686] The function score of the present application may have characteristics depending on the location of the single nucleotide mutation or the type of amino acid mutation that induces it.

[0687] For example, single nucleotide mutations that cause missense mutations located at species-conserved positions tend to show low function scores. As another example, among missense single nucleotide mutations, single nucleotide mutations that change tryptophan to another amino acid tend to show low function scores. As another example, among missense single nucleotide mutations, single nucleotide mutations that change a nonpolar amino acid to a positively charged amino acid or a negatively charged amino acid tend to show low function scores. As another example, among missense single nucleotide mutations, single nucleotide mutations located in exons 57 to 60 tend to show low function scores. As another example, among missense single nucleotide mutations, single nucleotide mutations located in the activation loop (residues 2888 through 2911) and catalytic loop (residues 2867 through 2875) of the kinase site tend to show low function scores. As another example, among missense single nucleotide mutations, single nucleotide mutations located in exon 17 tend to show high function scores. As another example, nonsense single nucleotide mutations or single nucleotide mutations located at splicing receptor / donor sites tend to show low function scores. As another example, among synonymous single nucleotide mutations, single nucleotide mutations located in exon sites close to introns tend to show low function scores.

[0688] Chapter 6 Functional Scores of ATM Proteins with Amino Acid Variations

[0689] Meaning of ATM protein function scores

[0690] If an ATM gene having a specific single nucleotide variant encodes an ATM protein having a specific amino acid variant, the function score of the specific single nucleotide variant can be considered the function score of the specific amino acid variant. Therefore, the function score of the ATM protein of the present application relates to the function score of Chapter 5.

[0691] The function score of the ATM protein of the present application refers to a value obtained by converting the function score of Chapter 5 above based on amino acid variations. In this case, if there are multiple single nucleotide mutations that induce a specific amino acid variation during the conversion process, the function scores based on single nucleotide mutations may be averaged and used as the function score of the ATM protein.

[0692] The meaning of the function score of the ATM protein of the present application is identical to the meaning of the function score in Chapter 5 above, except that it is converted based on amino acid variation. Additionally, the function score of the ATM protein of the present application may also be referred to as a function score. This should be interpreted according to whether the relevant subject is a gene or a protein.

[0693] The function score of the ATM protein in this application is a numerical representation of how the function of the ATM protein changes when a specific amino acid variant is introduced into the ATM protein. Alternatively, the function score of the ATM protein is a numerical representation of information regarding the function of the ATM protein having a specific amino acid variant.

[0694] The function score of the ATM protein above is a numerical representation of how the subject's phenotype will change when a specific amino acid mutation is introduced into the ATM protein. Alternatively, the function score of the ATM protein above is a numerical representation of information regarding the phenotype of a subject having a specific amino acid mutation in the ATM protein.

[0695] The function score of the above ATM protein is assigned to each amino acid variant. The reference sequence used to define the types of amino acid variants is SEQ ID NO. 128. Therefore, below, the types of amino acid variants can be defined based on the reference sequence of SEQ ID NO. 128, and the function score can be calculated individually for each type of amino acid variant.

[0696] Based on the function score of the ATM protein mentioned above, a function class may be determined (assigned) for an ATM protein having a specific amino acid variant. For example, if the function score of a specific amino acid variant is below the first classification criterion, the function class of the specific amino acid variant may be determined as "dysfunctional." Additionally, if the function score of a specific amino acid variant is above the second classification criterion, the function class of the specific amino acid variant may be determined as "functional." In this case, the second classification criterion refers to a value higher than the first classification criterion. If the function score of a specific amino acid variant is between the second classification criterion and the first classification criterion, the function class of the specific amino acid variant may be classified as "moderate." A detailed explanation of the above classification criteria is provided in "Classification of Function Score of ATM Protein" below.

[0697] The lower the function score value of the above ATM protein, the more it means that the ATM protein is not functioning normally. Alternatively, the higher the function score value of the above ATM protein, the more it means that the ATM protein is functioning normally.

[0698] Classification of amino acid variations

[0699] Amino acid variations can be classified under certain criteria based on the functional score of the ATM protein of the present application.

[0700] For example, the criteria for classifying function scores in Chapter 5 may also be used as criteria for classifying function scores of ATM proteins. In one specific example, -1.360 may be used as the first criterion and -0.912 as the second criterion.

[0701] If the function score of an ATM protein with a specific amino acid variant is lower than -1.360, the specific amino acid variant is represented as a dysfunctional amino acid variant. Additionally, an ATM protein with a specific amino acid variant is represented as being in a dysfunctional state.

[0702] If the function score of an ATM protein with a specific amino acid variant is higher than -0.912, the specific amino acid variant is expressed as a functional amino acid variant. Additionally, an ATM protein with a specific amino acid variant is expressed as being in a functional state.

[0703] If the function score of an ATM protein with a specific amino acid variant has a value between -1.360 and -0.912, the specific amino acid variant is represented as an intermediate amino acid variant. Additionally, an ATM protein with a specific amino acid variant is represented as being in an intermediate state.

[0704]

[0705] Chapter 7 Method for Providing Functional Information_Results of ATM Gene / Protein Function Evaluation Method

[0706] This specification aims to provide a method using the results of an ATM gene / protein function evaluation method. As an example, it aims to provide a method for providing functional information using the results of an ATM gene / protein function evaluation method.

[0707] Through the aforementioned method for evaluating the function of ATM genes / proteins in this application, integrated functional scores were derived for 27,513 single nucleotide variants that may exist in the coding sequence of the ATM gene. Additionally, integrated functional scores were derived for single nucleotide variants present in some introns. In particular, as an example, integrated functional scores for "single nucleotide variants of dysfunction" were described in this specification in Chapter 5. Furthermore, the functional class of a single nucleotide variant was derived through the functional scores in this application. As an example, single nucleotide variants for which the functional class was determined as "single nucleotide variants of dysfunction" and "single nucleotide variants of moderate severity" were described in this specification in Chapter 5. Additionally, as described in Chapter 6, functional scores and functional classes for amino acid variants were derived. By utilizing the derived information, for example, if a specific ATM gene contains a specific single nucleotide variant, information can be provided that the specific ATM gene is in a state of dysfunction.

[0708] Hereinafter, a method for providing ATM gene function information and a method for providing ATM protein function information using the information obtained as described above will be described.

[0709] The method for providing information on the function of the present application is characterized by using information obtained by the method for evaluating the function of the present application. Accordingly, while information that has already been obtained may be used, a process for obtaining information may be optionally included. That is, the method for providing information on the function of the present application may include obtaining information by the method for evaluating the function of the present application.

[0710] In one embodiment, the present application’s ATM gene function information providing method may further include a) preparing a cell, b) preparing a pegRNA library introducing a pseudo-single nucleotide variant, c) introducing the pegRNA library into the cell, d) sequencing at time point 1, e) sequencing at time point 2, and f) obtaining information on a function score or a specific single nucleotide variant / specific amino acid variant through comparison of sequencing results.

[0711] In one embodiment, the present application’s ATM gene function information providing method may further include a) obtaining a function score, b) preparing a deep learning model, and c) obtaining information about the function score or a specific single nucleotide variant / specific amino acid variant using the deep learning model.

[0712] In one embodiment, the present application’s ATM gene function information providing method may further include a) preparing a library containing engineered cells, b) obtaining viability information, c) preparing a deep learning model using the viability information, and d) obtaining information on function scores or specific single nucleotide variants / specific amino acid variants using the viability information and the deep learning model using the viability information.

[0713] In one embodiment, the present application’s ATM gene function information providing method may further comprise a) preparing a library containing engineered cells, b) obtaining viability information, and c) using the viability information to obtain information on function scores or specific single nucleotide variants / specific amino acid variants.

[0714] In one embodiment, the present application’s ATM gene function information providing method may further include a) preparing a library containing engineered cells and b) obtaining viability information, wherein information regarding function scores or specific single nucleotide variants / specific amino acid variants is obtained.

[0715] As previously mentioned, information regarding single nucleotide variations and information regarding amino acid variations can be converted into one another. Therefore, the information regarding a specific single nucleotide variation / specific amino acid variation refers to the information regarding the specific single nucleotide variation and / or specific amino acid variation.

[0716] The information regarding the specific single nucleotide variant mentioned above may be, but is not limited to, information that the specific single nucleotide variant is a dysfunctional single nucleotide variant, an intermediate single nucleotide variant, or a functional single nucleotide variant.

[0717] The information regarding the specific amino acid variation mentioned above may be, but is not limited to, information that the specific amino acid variation is a dysfunctional amino acid variation, an intermediate amino acid variation, or a functional amino acid variation.

[0718] Chapters 2 through 4 are included for reference as a detailed explanation of the above-mentioned optional addition process. Additionally, Chapter 5 is included for a detailed explanation of information regarding single nucleotide mutations, and Chapter 6 is included for reference as a detailed explanation of information regarding amino acid mutations.

[0719] Chapter 8 Method of Providing ATM Gene Function Information

[0720] Characteristics of the ATM gene function information provision method

[0721] This specification aims to provide a method for providing information regarding ATM gene functions. Below, the features of the method for providing ATM gene function information are described.

[0722] The above method for providing ATM gene function information is a method for providing information on the function of an ATM gene including a mutation. Alternatively, the above method for providing ATM gene function information is a method for providing information on the change in function when a mutation is introduced into an ATM gene. In this case, the mutation may be a single nucleotide mutation or a combination of single nucleotide mutations.

[0723] The above method for providing ATM gene function information may provide functional information of an ATM gene according to the ATM gene sequence for which the function is to be verified. Accordingly, the above method for providing ATM gene function information may include obtaining the ATM gene sequence for which the function is to be verified. Alternatively, the above method for providing ATM gene function information may include obtaining the ATM gene sequence of a target. In this case, the target refers to a target containing the ATM gene for which the function is to be verified.

[0724] The above method for providing ATM gene function information may provide information based on the variation contained in the ATM gene. Accordingly, the above method for providing ATM gene function information may include obtaining information regarding the type and location of the variation contained in the ATM gene or the ATM gene of the target object whose function is to be verified. That is, it may include obtaining variation information. In this case, the variation may be a single nucleotide variation or a combination of single nucleotide variations. Alternatively, the variation means that the ATM gene differs from the reference sequence by one or more nucleotides.

[0725] The above method for providing ATM gene function information can provide function information of the ATM gene using a function score or a single nucleotide variant distinguished by a function score. Accordingly, the above method for providing ATM gene function information may include providing function information of the ATM gene.

[0726] The present application's method for providing ATM gene function information is described in more detail below.

[0727] ATM gene sequence obtained

[0728] The present application's method for providing ATM gene function information includes obtaining an ATM gene sequence.

[0729] The above ATM gene sequence means that it includes not only information about the nucleotide sequence of the ATM gene, but also information capable of determining the presence, type, and location of specific nucleotides contained in the ATM gene. Therefore, the above ATM gene sequence can also be expressed as ATM gene information. For example, the above ATM gene sequence may be the full-length sequence of the ATM gene. For another example, the above ATM gene sequence may be a partial sequence of the ATM gene. For another example, the above ATM gene sequence may be a specific exon sequence of the ATM gene. For another example, the above ATM gene sequence may be information about the presence, type, or location of specific nucleotides contained in the ATM gene.

[0730] Acquiring the above ATM gene sequence means acquiring the sequence of the ATM gene whose function is to be verified. Therefore, when seeking to verify the function of the ATM gene possessed by a subject, acquiring the above ATM gene sequence means acquiring the ATM gene sequence of the subject.

[0731] The means for obtaining the ATM gene sequence mentioned above are not limited to any specific method, as long as the ATM gene sequence can be obtained. For example, the acquisition of the ATM gene sequence may involve obtaining the ATM gene sequence through sequencing of the ATM gene. For another example, the acquisition of the ATM gene sequence may involve obtaining information regarding whether a specific sequence or specific base is present in the ATM gene through PCR. For yet another example, the acquisition of the ATM gene sequence may involve obtaining information regarding whether a specific sequence or specific base is present in the ATM gene using a microarray. For yet another example, the acquisition of the ATM gene sequence may involve obtaining sequence information regarding an ATM gene that has already been analyzed.

[0732] The acquisition of the ATM gene sequence described above is not limited to the subject of acquisition, as long as it involves acquiring a sequence for the ATM gene. For example, the sequence for the ATM gene may be acquired from the ATM gene or a target. As another example, the sequence for the ATM gene may be acquired from a third party in the form of data.

[0733] When attempting to verify the function of the ATM gene possessed by a subject, the ploidy of the subject's ATM gene may be polyploid or higher. Additionally, the sequences of the subject's ATM alleles may differ from one another. In this case, if the sequence of the subject's ATM gene is to be obtained, there is no restriction on which specific ATM allele information is acquired. For example, information on one of the subject's ATM alleles may be obtained. As another example, information on all of the subject's ATM alleles may be obtained. As yet another example, ATM gene information from which specific ATM allele of the subject is unknown may be obtained.

[0734] Acquired mutation information

[0735] The method for providing ATM gene function information of the present application includes obtaining mutation information. The mutation may be a single nucleotide mutation or a combination of single nucleotide mutations.

[0736] The acquisition of the above mutation information refers to obtaining information regarding the presence, location, or type of mutation contained in the ATM gene whose function is to be verified. The ATM gene whose function is to be verified may be the ATM gene of a subject.

[0737] When comparing the sequence of the ATM gene whose function is to be verified, or the ATM gene of a target, with the reference sequence, a single nucleotide may differ. In such cases, the ATM gene can be said to contain a specific single nucleotide variant. Additionally, when comparing the sequence of the ATM gene with the reference sequence, two or more nucleotides may differ. In such cases, it may be expressed that the ATM gene contains a specific variant or specific variants. Alternatively, it may be expressed that the ATM gene contains a specific single nucleotide variant or specific single nucleotide variants. Or, it may be expressed that the ATM gene contains a combination of single nucleotide variants. Expressing it in this way is intended to more accurately designate the variant contained in the ATM gene. That is, to more accurately designate the nucleotides and locations that differ from the reference sequence, even when nucleotides differ at two or more locations, the different nucleotides at each location can be distinguished and referred to as single nucleotide variants. Furthermore, even if nucleotides differ at other locations, if the difference of a specific nucleotide at a specific location within the ATM gene sequence is significant, the ATM gene can be expressed as containing a specific single nucleotide variant. Alternatively, it can be expressed that the subject contains a single nucleotide variant. Expressing it in this way is also intended to more accurately designate the variant state of the ATM gene.

[0738] The method of obtaining the above mutation information is not limited to that method, provided that information regarding the presence, location, or type of mutation contained in the ATM gene whose function is to be verified can be obtained.

[0739] In one embodiment, obtaining the variant information may include a) comparing the obtained ATM gene sequence with a reference sequence, and b) obtaining information regarding the presence, location, or type of variant contained in the ATM gene whose function is to be verified. In another embodiment, obtaining the variant information may include a) comparing the obtained ATM gene sequence with a reference sequence, and b) identifying or determining whether the ATM gene whose function is to be verified contains a specific variant. In another embodiment, obtaining the variant information may include a) comparing the obtained ATM gene sequence with a reference sequence, and b) obtaining information regarding whether the ATM gene whose function is to be verified contains a specific variant.

[0740] In another embodiment, obtaining the mutation information may include confirming or determining whether the ATM gene whose function is to be verified includes a specific mutation. In another embodiment, obtaining the mutation information may include obtaining information regarding whether the ATM gene whose function is to be verified includes a specific mutation.

[0741] In another embodiment, obtaining the variant information may include confirming or determining whether the ATM gene whose function is to be verified includes one or more variants from a specific list of variants. In another embodiment, obtaining the variant information may include obtaining information regarding whether the ATM gene whose function is to be verified includes one or more variants from a specific list of variants.

[0742] In another embodiment, obtaining the mutation information may include comparing the obtained ATM gene sequence with Table 1 of the present application or a part of Table 1.

[0743] Checking or determining whether the above includes may be checking or determining that it includes. Or, checking or determining whether the above includes may be checking or determining whether it does not include.

[0744] Obtaining information regarding whether the above includes may mean obtaining information that it includes. Or, obtaining information regarding whether the above includes may mean obtaining information that it does not include.

[0745] The above specific variant is not limited to variants with specified location and type. For example, the above specific variant may be a single nucleotide variant classified as a "single nucleotide variant of dysfunction," a "single nucleotide variant of moderate severity," or a "functional single nucleotide variant" in Chapter 5 of this specification. Alternatively, the above specific variant may be a variant including a "single nucleotide variant of dysfunction," a "single nucleotide variant of moderate severity," or a "functional single nucleotide variant" in Chapter 5 of this specification.

[0746] Variants in the aforementioned specific list are not limited to a list or group consisting of two or more variations with specified locations and types. For example, the variations in the aforementioned specific list may be single nucleotide variations classified as "single nucleotide variations of functional impairment" in Chapter 5 of this specification. For example, the single nucleotide variations in the aforementioned specific list may be variations classified as "single nucleotide variations of moderate severity" in Chapter 5 of this specification. For example, the single nucleotide variations in the aforementioned specific list may be variations classified as "functional single nucleotide variations" in Chapter 5 of this specification. In one embodiment, they may be some of the variations listed in Table 1 of this specification.

[0747] The reference sequence for determining the above mutation may be SEQ ID NO. 1 or a part thereof. The above mutation may be a mutation named through the ATM gene composed of SEQ ID NO. 1. Meanwhile, if another specific sequence differs from SEQ ID NO. 1 only in that it contains a functional or intermediate single nucleotide variation, said other specific sequence may be used as the reference sequence for determining the mutation. This is because, depending on the information to be provided, the presence of a functional or intermediate single nucleotide variation may not have a significant impact. In one specific example, the sequence corresponding to SEQ ID NO. 3 may be used as the reference sequence.

[0748] The above list of specific variants may be a plurality of lists of specific variants. That is, obtaining the variant information may involve obtaining information regarding which list the ATM gene sequence whose function is to be verified includes the variants. In one specific embodiment, obtaining the variant information may involve verifying or determining which list of single nucleotide variants the ATM gene whose function is to be verified includes among a plurality of lists of specific single nucleotide variants.

[0749] Provides functional information

[0750] The present application's ATM gene function information provision method includes providing function information.

[0751] The provision of the above-mentioned functional information is characterized by providing information related to the variant contained in the ATM gene whose function is to be verified. In this case, if the variant of the ATM gene whose function is to be verified differs by one base compared to the reference sequence, information may be provided depending on whether it contains a specific single nucleotide variant. Additionally, even if the variant of the ATM gene whose function is to be verified differs by two or more bases compared to the reference sequence, information may be provided depending on whether it contains a specific single nucleotide variant. That is, the provision of the above-mentioned functional information is characterized by providing information depending on which single nucleotide variant the ATM gene whose function is to be verified contains.

[0752] The above functional information is not limited to any specific type of information as long as it is information regarding the function of the ATM gene to be verified.

[0753] In one embodiment, the function information may be information regarding whether the function of the ATM gene to be checked is in a functional state, an intermediate state, or a dysfunctional state.

[0754] In one embodiment, if functional information of the ATM gene of a subject having polyploidy or greater polyploidy is provided, said functional information may be information regarding the function of a specific allele of the subject. In another embodiment, said functional information may be information that the subject contains an allele in a specific state. In another embodiment, said functional information may be information regarding the function of all alleles of the subject. That is, said functional information may vary depending on the acquired ATM gene sequence and may vary depending on the purpose to be provided.

[0755] In one embodiment, the function information may be quantified information regarding the function of the ATM gene to be verified. Additionally, the quantified information may be a function score. In one specific embodiment, if the ATM gene to be verified contains multiple single nucleotide variants, the lowest score among the function scores corresponding to each single nucleotide variant may be provided.

[0756] The information provided regarding the above-mentioned function information may vary depending on what the acquired mutation information is.

[0757] In one embodiment, if the acquired variant information is information regarding whether the ATM gene whose function is to be verified includes a variant of dysfunction, providing the function information may be providing information that the ATM gene whose function is to be verified is in a state of dysfunction.

[0758] In another embodiment, if the acquired mutation information is information that the ATM gene whose function is to be verified does not contain a mutation of dysfunction, providing the function information may be information that the ATM gene whose function is to be verified is in a normal or intermediate state.

[0759] In another embodiment, where the variant information is information regarding which list of specific variant lists the ATM gene whose function is to be verified contains a single nucleotide variant, providing the function information may involve providing information corresponding to each list. In this case, the information corresponding to each list may indicate that the ATM gene whose function is to be verified is in a functional, intermediate, or impaired state, but is not limited thereto.

[0760] In one embodiment, if the acquired variant information is information regarding the presence, location, or type of variant contained in the ATM gene whose function is to be verified, providing the function information may involve providing ATM gene function information using the acquired variant information. In one specific embodiment, providing the function information may involve providing information by comparing the acquired variant information with Table 1 of this specification or a part of Table 1.

[0761]

[0762] Chapter 9 Method for Providing ATM Protein Function Information

[0763] Characteristics of the ATM protein function information provision method

[0764] This specification aims to provide a method for providing functional information regarding ATM proteins. The features of the method for providing ATM protein functional information are described below.

[0765] The above method for providing information on the function of an ATM protein can provide information on the function of an ATM protein having an amino acid variation. In this case, the information on the function of the ATM protein having the amino acid variation can be equated with the information on the function of the ATM gene encoding the ATM protein having the amino acid variation.

[0766] The above method for providing ATM protein function information may provide ATM protein function information according to the sequence of the ATM protein whose function is to be verified. Accordingly, the above method for providing ATM protein function information may include obtaining the sequence of the ATM protein whose function is to be verified. Alternatively, the above method for providing ATM protein function information may include obtaining the ATM gene sequence encoding the ATM protein whose function is to be verified. Alternatively, the above method for providing ATM protein function information may include obtaining the ATM protein sequence of a target or the ATM gene sequence encoding the ATM protein of a target. In this case, the target refers to a target containing the ATM gene whose function is to be verified.

[0767] The above method for providing ATM protein function information can provide functional information of the ATM protein according to the type and location of amino acid variations contained in the ATM protein. Accordingly, the above method for providing ATM protein function information may include obtaining information regarding the type and location of amino acid variations contained in the ATM protein or the ATM protein of the target body whose function is to be verified. That is, it may include obtaining amino acid variation information.

[0768] The above method for providing ATM protein function information may include providing function information of an ATM protein. Alternatively, the above method for providing ATM protein function information may include providing information on the function of an ATM gene encoding an ATM protein having an amino acid variation.

[0769] The method for providing ATM protein function information in this application is described in more detail below.

[0770] ATM protein sequence obtained

[0771] The method for providing ATM protein function information of the present application includes obtaining an ATM protein sequence.

[0772] The above ATM protein sequence means that it includes not only information about the sequence of the ATM protein, but also information capable of determining the presence, type, and location of specific amino acid residues contained in the ATM protein.

[0773] Acquiring the above ATM protein sequence means acquiring the sequence of the ATM protein whose function is to be verified. Therefore, if the function of the ATM protein possessed by the subject is to be verified, acquiring the above ATM protein sequence may mean acquiring the ATM protein sequence of the subject.

[0774] The means of obtaining the above ATM protein sequence are not limited to any means that allow the ATM protein sequence to be obtained. For example, information may be obtained by analyzing the ATM protein sequence. For another example, information on an ATM protein sequence that has already been analyzed may be obtained. For yet another example, it may be obtained through the ATM gene sequence encoding the ATM protein. In this case, the contents of Chapter 8 are included as a detailed explanation of obtaining the above ATM gene sequence.

[0775] The above ATM protein sequence is not limited to any specific type of information, as long as it contains information capable of determining the presence, type, or location of amino acid variations included in the ATM protein sequence. For example, the above ATM protein sequence may be an ATM protein sequence. For example, the above ATM protein sequence may be the whole or a part of an ATM protein sequence. As another example, the above ATM protein sequence may be information regarding an ATM gene sequence encoding an ATM protein. In this case, the contents of Chapter 8 are included as a detailed description of the above ATM gene sequence.

[0776] The subject from which the sequence for the ATM protein is obtained is not limited. For example, the sequence for the ATM protein may be obtained from the ATM protein, the ATM gene, or a subject. For another example, the sequence for the ATM protein may be obtained from a third party.

[0777] When attempting to verify the function of an ATM protein possessed by a subject, the ploidy of the subject's ATM gene may be polyploid or higher. Additionally, the sequences of the subject's ATM alleles may differ. Consequently, the sequences of the ATM proteins contained within the subject may vary. In the above cases, as long as the sequence of the subject's ATM protein is to be obtained, there is no restriction on which specific ATM protein information is acquired. For example, information regarding the sequence of a specific ATM protein contained in the subject may be obtained. For example, information regarding one of the ATM alleles encoding the subject's ATM protein may be obtained. For another example, information regarding all ATM alleles encoding the subject's ATM protein may be obtained. For yet another example, an ATM gene sequence encoding an ATM protein whose origin from which ATM allele of the subject is unknown may be obtained.

[0778] Acquired amino acid variation information

[0779] The present application’s method for providing ATM protein function information includes obtaining amino acid variation information.

[0780] Acquiring the above amino acid variation information means obtaining information regarding the presence, location, or type of amino acid variation contained in the ATM protein whose function is to be verified.

[0781] When comparing the sequence of an ATM protein whose function is to be verified or an ATM protein of a target with a reference sequence, there may be differences of two or more amino acids. In this case, even if amino acids differ at other positions, if the difference of a specific amino acid at a specific position in the ATM protein sequence is significant, it may be said that the ATM protein contains a specific amino acid variation. This expression is intended to more accurately describe the state of the amino acid variation possessed by the ATM protein.

[0782] Acquiring the above amino acid variation information is not limited to the method provided that information regarding the presence, location, or type of amino acid variation contained in the ATM protein whose function is to be verified can be obtained.

[0783] In one embodiment, obtaining the amino acid variation information may include a) comparing the obtained ATM protein sequence with a reference sequence, and b) obtaining information regarding the presence, location, or type of amino acid variation contained in the ATM protein whose function is to be verified. In another embodiment, obtaining the amino acid variation information may include a) comparing the obtained ATM protein sequence with a reference sequence, and b) verifying or determining whether the ATM protein whose function is to be verified contains a specific amino acid variation. In another embodiment, obtaining the amino acid variation information may include a) comparing the obtained ATM protein sequence with a reference sequence, and b) obtaining information regarding whether the ATM protein whose function is to be verified contains a specific amino acid variation.

[0784] In another embodiment, obtaining the amino acid variation information may include confirming or determining whether the ATM protein whose function is to be verified contains a specific amino acid variation. In another embodiment, obtaining the amino acid variation information may include obtaining information regarding whether the ATM protein whose function is to be verified contains a specific amino acid variation.

[0785] In another embodiment, obtaining the amino acid variation information may include confirming or determining whether the ATM protein whose function is to be verified contains one or more amino acid variations from a specific list of amino acid variations. In another embodiment, obtaining the amino acid variation information may include obtaining information regarding whether the ATM protein whose function is to be verified contains one or more amino acid variations from a specific list of amino acid variations.

[0786] In another embodiment, obtaining the amino acid variation information may include comparing the obtained ATM protein sequence with Table 1 of the present application or a part of Table 1. In another embodiment, obtaining the amino acid variation information may include comparing with amino acid variations listed in Table 1 of the present application or a part of Table 1, wherein the AA_function score is -1.360 or lower.

[0787] Checking or determining whether the above includes may be checking or determining that it includes. Or, checking or determining whether the above includes may be checking or determining whether it does not include.

[0788] Obtaining information regarding whether the above includes may mean obtaining information that it includes. Or, obtaining information regarding whether the above includes may mean obtaining information that it does not include.

[0789] The specific amino acid variant mentioned above is not limited to amino acid variants with specified location and type. For example, the specific amino acid variant mentioned above may be a single nucleotide variant classified as a "dysfunctional amino acid variant," a "moderate amino acid variant," or a "functional amino acid variant" in Chapter 6 of this specification.

[0790] The amino acid variations in the above specific list are not limited to a list or group consisting of two or more amino acid variations with specified positions and types. For example, the amino acid variations in the above specific list may be the amino acid variations listed in Table 1 of this specification. For another example, the amino acid variations in the above specific list may be the amino acid variations classified as "dysfunctional amino acid variations" in Chapter 6 of this specification. That is, the amino acid variations in the above specific list may be amino acid variations among the amino acid variations listed in Table 1 or part of Table 1 of this specification that have an AA_function score of -1.360 or less.

[0791] The reference sequence for determining the above amino acid variation may be SEQ ID NO. 128 or a part thereof. The above amino acid variation may be an amino acid variation named through the ATM protein composed of SEQ ID NO. 128. Meanwhile, if another specific sequence differs from SEQ ID NO. 128 only in the degree of having a functional or intermediate variation, said specific sequence may be used as the reference sequence for determining the variation. This is because, depending on the information to be provided, the presence of a functional or intermediate variation may not have a significant impact. In one embodiment, the sequence corresponding to SEQ ID NO. 129 may be used as the reference sequence.

[0792] The above list of specific amino acid variants may be a list of multiple specific amino acid variants. That is, obtaining the amino acid variant information may involve obtaining information regarding which list the amino acid variants corresponding to the ATM protein sequence whose function is to be verified are included. In one specific embodiment, obtaining the amino acid variant information may involve identifying or determining which list of amino acid variants the ATM protein whose function is to be verified includes among the multiple lists of specific amino acid variants.

[0793] Provides functional information

[0794] The method for providing functional information of an ATM protein according to the present application includes providing functional information. The provision of said functional information is characterized by providing information related to amino acid variations contained in the ATM protein whose function is to be verified. That is, the provision of said functional information is characterized by providing information depending on which amino acid variation the ATM protein whose function is to be verified contains.

[0795] The above functional information is not limited to any specific type, provided it is information regarding the function of the ATM protein whose function is to be verified. In this case, the information regarding the function of the ATM protein having the amino acid variation may be equated with the information regarding the function of the ATM gene encoding the ATM protein having the amino acid variation. Therefore, in the following description, the functional information of the ATM protein whose function is to be verified may also be the functional information of the ATM gene encoding the ATM protein whose function is to be verified.

[0796] In one embodiment, the functional information may be information regarding whether the function of the ATM protein to be verified is in a functional state, an intermediate state, or a dysfunctional state.

[0797] In another embodiment, if functional information of an ATM protein of a subject having polyploidy or greater ploidy with respect to the ATM gene is provided, said functional information may be information regarding the function of a specific ATM protein contained in the subject. That is, said functional information may be the extent that the subject contains an ATM protein in a specific functional state. Alternatively, said functional information may be information that the subject contains only an ATM protein in a specific functional state.

[0798] In another embodiment, the function information may be quantified information regarding the function of the ATM protein to be verified. The quantified information may be a function score. In one specific embodiment, if the ATM protein to be verified contains multiple amino acid variations, the lowest score among the function scores corresponding to each amino acid variation may be provided.

[0799] The information provided regarding the above-mentioned function information may vary depending on what the amino acid variation information is.

[0800] In one embodiment, if the amino acid variation information is information regarding whether the ATM protein whose function is to be verified contains an amino acid variation of dysfunction, providing the function information may be providing information that the ATM protein whose function is to be verified is in a dysfunctional state.

[0801] In another embodiment, if the amino acid variation information is information that the ATM protein whose function is to be verified does not contain an amino acid variation of dysfunction, providing the function information may be information that the ATM protein whose function is to be verified is in a normal or intermediate state.

[0802] In another embodiment, where the amino acid variation information is information regarding which of a list of amino acid variations the ATM protein whose function is to be verified contains, providing the function information may involve providing information corresponding to each list. In this case, the information corresponding to each list may indicate that the ATM protein whose function is to be verified is functional, moderately functional, or in a dysfunctional state, but is not limited thereto.

[0803] In one embodiment, if the amino acid variation information is information regarding the presence, location, or type of amino acid variation contained in the ATM protein whose function is to be verified, the provision of the function information may be the provision of function information for the ATM protein whose function is to be verified using the information of the acquired amino acid variation. In one specific embodiment, the provision of the function information may be the provision of information by comparing the information of the acquired amino acid variation with Table 1 or a part of Table 1 of this specification. In another specific embodiment, the provision of the function information may be the provision of information by comparing the information of the acquired amino acid variation with amino acid variations listed in Table 1 or a part of Table 1 of this specification, wherein the AA_function score is -1.360 or lower.

[0804]

[0805] Chapter 10: Utilization of Function Information Provision Methods

[0806] The function of the ATM gene is known to be associated with specific phenotypes, characteristics, or symptoms. For example, the function of the ATM gene is associated with cancer incidence. As another example, depending on the type of cancer, the function of the ATM gene is associated with cancer prognosis. As yet another example, the function of the ATM gene is associated with sensitivity to anticancer drugs that are PARP inhibitors. As yet another example, the function of the ATM gene is associated with ataxia-telangiectasia syndrome. As yet another example, the function of the ATM gene is associated with sensitivity to PARP inhibitors.

[0807] In the context of the above, the function score of the present application is a score associated with a phenotype, characteristic, or symptom related to the function of the ATM gene. In one embodiment, the function score of the present application is a score confirmed to be associated with the incidence of cancer. In another embodiment, the function score of the present application is a score confirmed to be associated with the prognosis of cancer. In another embodiment, the function score of the present application is a score obtained using the survival rate (or viability) against a PARP inhibitor. Accordingly, the function score of the present application or single nucleotide variants classified according to the function score of the present application can be used to provide information regarding a phenotype, characteristic, or symptom associated with the function of the gene.

[0808] Accordingly, through the functional information provided by the method of providing functional information of the present application, information regarding phenotypes, characteristics, or symptoms associated with the function of a gene can be provided. In one embodiment, through the method of providing functional information of the present application, information can be provided for predicting cancer incidence rates, predicting cancer prognosis, diagnosing ataxia telangiectasia syndrome, and selecting PARP inhibitors. That is, the method of providing functional information of the present application can be utilized as a method for predicting cancer incidence rates, a method for predicting cancer prognosis, a method for providing information for assisting in the selection of anticancer drugs (assistance in selecting PARP inhibitors), or a method for providing information for diagnosing ataxia telangiectasia syndrome.

[0809] This specification aims to provide a method for providing ATM gene function information or ATM protein function information for the present application. As previously mentioned, information regarding ATM gene function including single nucleotide mutations is related to information regarding ATM proteins including amino acid mutations. Therefore, the utilization of the ATM gene function information and ATM protein function information for the present application will be abbreviated as "utilization of the function information provision method" below. Furthermore, although the following description is based on the utilization of the ATM gene function information provision method, it is also applicable to the utilization of the ATM protein function information provision method.

[0810] The above utilization means that the functional information provided by the method for providing functional information of the present application has been utilized. The utilization of the above method for providing functional information may include all or part of the method for providing functional information of the present application. However, the information provided in the utilization of the above method for providing functional information is not functional information, but information regarding specific phenotypes, characteristics, or symptoms related to ATM gene function. That is, the utilization of the above method for providing functional information is a method of providing information regarding specific phenotypes, characteristics, or symptoms related to ATM gene function according to the functional score and single nucleotide variant classification of the present application.

[0811] The utilization of the above method for providing functional information may be preceded by an optional additional process. In one embodiment, the utilization of the above method for providing functional information may include obtaining information as a method for evaluating the function of an ATM gene having a single nucleotide variant according to the present application.

[0812] Below, the utilization of each of the methods for providing functional information of the present application is described.

[0813] Cancer Prediction Methods

[0814] One of the applications of the functional information provision method of the present application is a method for predicting the likelihood of cancer development or the cancer incidence rate.

[0815] The above-described cancer incidence rate prediction method predicts the likelihood of cancer development based on functional information obtained from the functional information provision method of the present application. That is, the cancer incidence rate prediction method of the present application includes providing cancer incidence rate information while including all or part of the functional information provision method of the present application.

[0816] The functional information in the method for providing functional information of the present application includes information regarding the incidence of cancer. In one embodiment, it includes information that the incidence of cancer is high when the functional score of a single nucleotide variant contained in the ATM gene of a subject is low. In one embodiment, it includes information that the incidence of cancer is high when the function of the ATM gene of a subject is moderate. In one embodiment, it includes information that the incidence of cancer is high when the function of the ATM gene of a subject is impaired. At this time, the criteria for determining that the incidence of cancer is high may vary. For example, it may be determined that the incidence of cancer is high in an individual containing a specific single nucleotide variant by comparing it with the incidence of cancer in an individual that does not contain a specific single nucleotide variant. As another example, it may be determined that the incidence of cancer is high in an individual containing a specific single nucleotide variant when checking the distribution of cancer incidence rates among individuals containing various individual single nucleotide variants. As yet another example, it may be determined that the incidence of cancer is high in an individual containing a specific single nucleotide variant by comparing it with the incidence of cancer in an individual containing only functional single nucleotide variants. As another example, it may be determined that the cancer incidence rate of an individual with a specific single nucleotide mutation is high when compared to an absolute standard established for cancer incidence rates. In this case, the cancer incidence prediction method of the present application provides information based on the functional integration score of the present application, and accordingly, the standard for determining a high cancer incidence rate is also based on the functional integration score. Furthermore, the functional integration score has been classified as a standard that corresponds to the above-mentioned judgment criteria. Therefore, the above-mentioned judgment criteria are interpreted as mutually complementary and should be interpreted according to the context.

[0817] The above-described method for predicting cancer incidence is characterized by providing information related to the cancer incidence rate of a subject. Additionally, the above-described method for predicting cancer incidence is characterized by providing information related to the cancer incidence rate based on the ATM gene function or ATM protein function of the subject. At this time, the type of cancer is not limited. In one embodiment, the cancer may be breast cancer, lung cancer, colorectal cancer, stomach cancer, liver cancer, pancreatic cancer, prostate cancer, bladder cancer, ovarian cancer, cervical cancer, esophageal cancer, kidney cancer, lymphoma, blood cancer, leukemia, skin cancer, brain tumor, osteosarcoma, thyroid cancer, oral cancer, biliary tract cancer, or melanoma. In one specific embodiment, the cancer may be breast cancer or prostate cancer.

[0818] In the above method for predicting cancer incidence rates, the number of alleles of the ATM gene containing a single nucleotide variant in the subject is not limited. For example, the subject may be a subject having a specific single nucleotide variant. For another example, the subject may be a subject carrying a specific single nucleotide variant. This is because the cancer incidence rate is related even if the function of just one of the alleles contained in the subject is impaired.

[0819] Even if the subject contains a combination of single nucleotide variants, the cancer incidence prediction method described above may provide cancer incidence information based on whether it contains a specific single nucleotide variant. This is because if a specific single nucleotide variant affects the function of the ATM gene to a state of dysfunction, the ATM gene will remain in a state of dysfunction even if it contains other single nucleotide variants. For example, if the subject contains an ATM allele containing a specific A single nucleotide variant and a specific B single nucleotide variant, and an ATM allele identical to the wild type (or reference sequence), information may be provided based on each of the specific A single nucleotide variant and the specific B single nucleotide variant. Alternatively, information may be provided based on the specific A single nucleotide variant or the specific B single nucleotide variant. For example, if the subject contains an ATM allele containing a specific A single nucleotide variant and an ATM allele containing a specific B single nucleotide variant, information may be provided based on the specific A single nucleotide variant and / or the specific B single nucleotide variant.

[0820] The above-mentioned method for predicting cancer incidence may include all or part of the present application’s method for providing ATM gene function information or the present application’s method for providing ATM protein function information. In this regard, the contents of Chapters 8 and 9 of this specification are included. For example, the above-mentioned method for predicting cancer incidence may include a) obtaining the ATM gene sequence of a subject and b) obtaining mutation information and c) providing cancer incidence information. For another example, the above-mentioned method for predicting cancer incidence may include a) obtaining mutation information of a subject and b) providing cancer incidence information. For another example, the above-mentioned method for predicting cancer incidence may include a) obtaining the ATM protein sequence of a subject, b) obtaining amino acid mutation information, and c) providing cancer incidence information. For another example, the above-mentioned method for predicting cancer incidence may include a) obtaining amino acid mutation information of a subject and b) providing cancer incidence information. In this case, the provision of cancer incidence information may be provided according to the present application’s function score or classified single nucleotide variant.

[0821] The above-described method for predicting cancer incidence includes providing information related to cancer incidence. In one embodiment, if a single nucleotide variant in the subject's ATM gene or an amino acid variant in the subject's ATM protein is a dysfunctional variant, information is provided that the cancer incidence rate is high. In one embodiment, if a single nucleotide variant in the subject's ATM gene or an amino acid variant in the subject's ATM protein is an intermediate variant, information is provided that the cancer incidence rate is high.

[0822] The above method for predicting cancer incidence rates may include optional additional processes. In one embodiment, the method for predicting cancer incidence rates may include performing cancer diagnostic tests frequently when the subject's cancer incidence rate is high. In another embodiment, the method for predicting cancer incidence rates may include performing cancer diagnostic tests frequently when the subject carries a specific single nucleotide mutation. In another embodiment, it may include providing information, diagnosis, or recommendation to perform cancer tests frequently.

[0823] In one embodiment, the cancer incidence prediction method is characterized by providing information, particularly regarding breast cancer or prostate cancer. That is, the cancer incidence prediction method may additionally include frequently performing diagnostic tests for breast cancer or prostate cancer when the cancer incidence rate of the subject is high. In another embodiment, the cancer incidence prediction method may include frequently performing diagnostic tests for breast cancer or prostate cancer when the subject carries or contains a specific single nucleotide mutation. Alternatively, it may include providing information, diagnosis, or recommendation to frequently perform cancer diagnostic tests. In one embodiment, the breast cancer or prostate cancer diagnostic test may be an ultrasound examination.

[0824] The above method for predicting the cancer incidence rate may additionally include removing the breast or the prostate if the subject's cancer incidence rate is high.

[0825] Cancer prognosis prediction methods

[0826] One of the applications of the functional information provision method of the present application is a cancer prognosis prediction method.

[0827] This specification discloses a method for predicting cancer prognosis. In this case, the method for predicting cancer prognosis may also be referred to as a method for providing information related to the prognosis of cancer in a cancer patient. The method for predicting cancer prognosis is characterized by providing information regarding cancer prognosis using the method for providing functional information of the present application. That is, the method for predicting cancer prognosis includes providing cancer prognosis information while including all or part of the method for providing functional information of the present application.

[0828] The functional information in the method for providing functional information of the present application includes information regarding cancer prognosis. In one embodiment, it includes cancer prognosis information regarding a case where the function score of a single nucleotide variant contained in the ATM gene of a cancer patient is low. In one embodiment, it includes cancer prognosis information regarding a case where the function of the ATM gene of a cancer patient is moderate. In another embodiment, it includes cancer prognosis information regarding a case where the function of the ATM gene of a cancer patient is impaired. In one specific embodiment, it includes information that the cancer prognosis will be poor if the function of the ATM gene of a patient with chronic lymphocytic leukemia is impaired. In one specific embodiment, it includes information that the cancer prognosis will be good if the function of the ATM gene of a patient with bladder cancer is impaired.

[0829] In this case, the cancer prognosis may refer to the degree of cancer progression. Alternatively, the cancer prognosis may refer to symptoms or results resulting from cancer progression. In one embodiment, the cancer prognosis may be survival time after cancer diagnosis, overall survival, failure-free survival, or progression-free survival. The survival time after cancer diagnosis refers to the period from cancer diagnosis until the subject's death. The overall survival refers to the period from the start of cancer treatment until the subject's death. The failure-free survival refers to the period from the start of treatment until cancer progression, recurrence, or death. The progression-free survival refers to the period during which cancer progression stops and there is no further progression. The cancer prognosis may be expressed as "good cancer prognosis" or "bad cancer prognosis." In this case, a good cancer prognosis means that the survival period is long. Furthermore, a bad cancer prognosis means that the survival period is short.

[0830] In this case, the criteria for determining cancer prognosis can vary. For example, the prognosis of a cancer patient with a specific single nucleotide mutation may be determined by comparing it to the prognosis of cancer patients who do not have the specific single nucleotide mutation. Another example may be determining the prognosis of a cancer patient with a specific single nucleotide mutation by comparing it to the distribution of cancer prognoses among cancer patients with the mutation. Yet another example may be determining the prognosis of a cancer patient with a specific single nucleotide mutation by comparing it to the prognosis of cancer patients with functional single nucleotide mutations. A third example may be determining the prognosis of an individual with a specific single nucleotide mutation by comparing it to an absolute standard established for cancer prognosis.

[0831] The above-described cancer prognosis prediction method is characterized by providing information related to the cancer prognosis of a subject. Additionally, the above-described cancer prognosis prediction method is characterized by providing information related to the cancer prognosis based on the function of the subject's ATM gene or the function of the subject's ATM protein. At this time, the type of cancer is not limited. In one embodiment, the cancer may be breast cancer, lung cancer, colorectal cancer, stomach cancer, liver cancer, pancreatic cancer, prostate cancer, bladder cancer, ovarian cancer, cervical cancer, esophageal cancer, kidney cancer, lymphoma, blood cancer, leukemia, skin cancer, brain tumor, osteosarcoma, thyroid cancer, oral cancer, biliary tract cancer, or melanoma. In one specific embodiment, the cancer may be chronic lymphocytic leukemia or bladder cancer.

[0832] In the above-described cancer prognosis prediction method, the number of alleles of the ATM gene containing a single nucleotide mutation in the subject is not limited. For example, the subject may be a subject having a specific single nucleotide mutation. For another example, the subject may be a subject carrying a specific single nucleotide mutation. This is because even if the function of just one of the alleles contained in the subject is impaired, it is related to the cancer prognosis.

[0833] In the above-described cancer prognosis prediction method, the number of single nucleotide variants included in the subject is not limited. For example, if the subject includes an ATM allele containing a specific A single nucleotide variant and a specific B single nucleotide variant, and an ATM allele identical to the wild type (or reference sequence), information may be provided based on each of the specific A single nucleotide variant and the specific B single nucleotide variant. Alternatively, information may be provided based on the specific A single nucleotide variant or the specific B single nucleotide variant. For example, if the subject includes an ATM allele containing a specific A single nucleotide variant and an ATM allele containing a specific B single nucleotide variant, information may be provided based on the specific A single nucleotide variant and / or the specific B single nucleotide variant.

[0834] The above-described cancer prognosis prediction method may include all or part of the present application’s ATM gene function information provision method or the present application’s ATM protein function information provision method. In this regard, the contents of Chapters 8 and 9 of this specification are included. For example, the above-described cancer prognosis prediction method may include a) obtaining the ATM gene sequence of a subject, b) obtaining mutation information, and c) providing cancer prognosis information. For another example, the above-described cancer prognosis prediction method may include a) obtaining mutation information of a subject and b) providing cancer prognosis information. For another example, the above-described cancer prognosis prediction method may include a) obtaining the ATM protein sequence of a subject, b) obtaining amino acid mutation information, and c) providing cancer prognosis information. For another example, the above-described cancer prognosis prediction method may include a) obtaining amino acid mutation information of a subject and b) providing cancer prognosis information. In this case, the provision of cancer prognosis information may provide information based on the present application’s function score or classified single nucleotide variants. Additionally, the above-described subject may be a cancer patient.

[0835] The above cancer prognosis prediction method includes providing cancer prognosis information.

[0836] In one embodiment, the cancer prognosis prediction method provides information that the prognosis of a specific cancer will be poor if a single nucleotide mutation contained in the subject's ATM gene is a dysfunctional mutation or an intermediate mutation. Alternatively, the cancer prognosis prediction method provides information that the prognosis of a specific cancer will be poor if an amino acid mutation contained in the subject's ATM protein is a dysfunctional mutation or an intermediate mutation. In one embodiment, the specific cancer may be leukemia. In another embodiment, the specific cancer may be chronic lymphocytic leukemia.

[0837] In one embodiment, the cancer prognosis prediction method provides information that the prognosis of a specific cancer will be good if a single nucleotide mutation contained in the ATM gene of a subject is a dysfunctional mutation or an intermediate mutation. Alternatively, the cancer prognosis prediction method provides information that the prognosis of a specific cancer will be good if an amino acid mutation contained in the ATM protein of a subject is a dysfunctional mutation or an intermediate mutation. In one specific embodiment, the specific cancer may be bladder cancer.

[0838] The above-described cancer prognosis prediction method may include optional additional processes. For example, the cancer prognosis prediction method may additionally include performing tests for cancer progression or diagnosis more frequently. In this case, the cancer may be leukemia. For another example, the cancer prognosis prediction method may additionally include performing tests for cancer progression or diagnosis less frequently. In this case, the cancer may be bladder cancer. For yet another example, the cancer incidence prediction method may include administering treatment of an intensity appropriate to the expected prognosis.

[0839] Method for providing information to assist in anticancer drug selection (assistance in PARP inhibitor selection)

[0840] One application of the method for providing functional information in this application may be to assist in the selection of anticancer drugs. Based on the functional information of this application, information regarding sensitivity or resistance to anticancer drugs can be obtained and used.

[0841] A representative example of the above anticancer agent is a 'PARP inhibitor'. In one embodiment, the ATM function score of the present application can be obtained using the survival rate (or viability) for a PARP inhibitor, and in this case, the function score of the present application may also be referred to as a sensitivity score for a PARP inhibitor.

[0842] In the present disclosure, the method for providing information to assist in the selection of an anticancer drug is characterized by utilizing the method for providing functional information of the present application. That is, the method for providing information to assist in the selection of an anticancer drug is characterized by providing information to assist in the selection of an anticancer drug while including all or part of the method for providing functional information of the present application. In this case, as a representative example, the anticancer drug may be a PARP inhibitor, and thus, the present method may be provided as a method for providing information to assist in the selection of a PARP inhibitor.

[0843] The above method for providing information to assist in the selection of an anticancer drug may include all or part of the method for providing ATM gene function information or the method for providing ATM protein function information of the present application. In this regard, the contents of Chapters 8 and 9 of this specification are included. For example, the above method for providing information to assist in the selection of an anticancer drug may include a) obtaining the ATM gene sequence of a subject, b) obtaining mutation information, and c) providing information to assist in the selection of an anticancer drug. For another example, the above method for providing information to assist in the selection of an anticancer drug may include a) obtaining mutation information of a subject, and b) providing information to assist in the selection of an anticancer drug. For another example, the above method for providing information to assist in the selection of an anticancer drug may include a) obtaining the ATM protein sequence of a subject, b) obtaining amino acid mutation information, and c) providing information to assist in the selection of an anticancer drug. For another example, the above method for providing information to assist in the selection of an anticancer drug may include a) obtaining amino acid mutation information of a subject, and b) providing information to assist in the selection of an anticancer drug. At this time, the information provided to assist in the selection of the anticancer drug may be provided according to the function score of the present application or classified single nucleotide mutations. In addition, the subject may be a patient with breast cancer, lung cancer, colorectal cancer, stomach cancer, liver cancer, pancreatic cancer, prostate cancer, bladder cancer, ovarian cancer, cervical cancer, esophageal cancer, kidney cancer, lymphoma, blood cancer, leukemia, skin cancer, brain tumor, osteosarcoma, thyroid cancer, oral cancer, biliary tract cancer, or melanoma. In addition, the anticancer drug may be olaparib, niraparib, rucaparib, or veliparib.

[0844] In the above method for providing information to assist in the selection of anticancer drugs, the number of alleles of the ATM gene containing a single nucleotide mutation in the subject is not limited. For example, the subject may be a subject having a specific single nucleotide mutation. For another example, the subject may be a subject carrying or containing a specific single nucleotide mutation. This is because the sensitivity to PARP inhibitors is related even when the function of a single allele contained in the subject is impaired.

[0845] In the method for providing information to assist in the selection of anticancer drugs described above, the number of single nucleotide variants included in the subject is not limited. For example, if the subject includes an ATM allele containing a specific A single nucleotide variant and a specific B single nucleotide variant, and an ATM allele identical to the wild type (or reference sequence), information may be provided based on each of the specific A single nucleotide variant and the specific B single nucleotide variant. Alternatively, information may be provided based on the specific A single nucleotide variant or the specific B single nucleotide variant. For example, if the subject includes an ATM allele containing a specific A single nucleotide variant and an ATM allele containing a specific B single nucleotide variant, information may be provided based on the specific A single nucleotide variant and / or the specific B single nucleotide variant.

[0846] The above method for providing information to assist in the selection of anticancer drugs includes providing information related to the selection of anticancer drugs.

[0847] As stated above, the above anticancer agent may be a PARP inhibitor, and the above disclosure may be interpreted as 'PARP inhibitor' instead of 'anticancer agent'.

[0848] In one embodiment, the method for providing information to assist in the selection of an anticancer drug provides information that a PARP inhibitor will be effective for the subject if a single nucleotide variant contained in the subject's ATM gene is a dysfunctional variant or an intermediate variant. In another embodiment, the method for providing information to assist in the selection of an anticancer drug provides information that the subject will be sensitive to a PARP inhibitor if a single nucleotide variant contained in the subject's ATM gene is a dysfunctional variant or an intermediate variant. Alternatively, the method for providing information to assist in the selection of an anticancer drug provides information that a PARP inhibitor will not be effective if an amino acid variant contained in the subject's ATM protein is a dysfunctional variant or an intermediate variant. In another embodiment, the method for providing information to assist in the selection of an anticancer drug provides information that the subject will be sensitive to a PARP inhibitor if an amino acid variant contained in the subject's ATM protein is a dysfunctional variant or an intermediate variant.

[0849] In another embodiment, the method for providing information to assist in the selection of an anticancer drug may provide information that a PARP inhibitor can be selected as an anticancer drug if the single nucleotide mutation contained in the ATM gene of the subject is a dysfunctional mutation or an intermediate mutation.

[0850] The above method for providing information to assist in the selection of an anticancer drug is characterized by being able to provide information regarding whether the anticancer drug is effective. Accordingly, the above method for providing information to assist in the selection of an anticancer drug may further include treating cancer patients carrying or containing a specific single nucleotide mutation using a PARP inhibitor. In another embodiment, the above method for providing information to assist in the selection of an anticancer drug may further include treating the cancer of cancer patients carrying or containing a specific single nucleotide mutation using a PARP inhibitor. In another embodiment, the above method for providing information to assist in the selection of an anticancer drug may further include prescribing a PARP inhibitor to cancer patients carrying or containing a specific single nucleotide mutation.

[0851] In another embodiment, the method for providing information to assist in the selection of an anticancer drug may be used in a method for treating cancer by administering a PARP inhibitor to a cancer patient. Alternatively, the method for providing information to assist in the selection of an anticancer drug may be used in a method for treating cancer by administering a PARP inhibitor to a cancer patient. That is, the present application specification discloses a method for treating cancer using a PARP inhibitor based on the aforementioned functional information.

[0852] In this case, the cancer treatment method is a method of treating a cancer patient using a PARP inhibitor by identifying which mutation is present in the ATM gene. The cancer treatment method may include a) obtaining ATM gene information of the cancer patient, b) obtaining mutation information of the cancer patient, and c) treating the cancer by administering a PARP inhibitor. In this case, the ATM gene information of the cancer patient may be information derived from cancer tissue derived from the cancer patient. In another embodiment, the cancer treatment method may include a) obtaining ATM gene information of the cancer patient, and b) treating the cancer by administering a PARP inhibitor when it is determined that the cancer patient contains a single nucleotide mutation causing dysfunction. The treatment of the cancer may be referred to in various ways. In another embodiment, the cancer treatment method may include a) obtaining ATM gene information of the cancer patient, and b) treating the cancer by administering a PARP inhibitor when it is determined that the cancer patient contains a moderate single nucleotide mutation and / or a single nucleotide mutation causing dysfunction. The treatment of the cancer may be referred to in various ways. For example, treating the cancer may be referred to as treating a cancer patient. For another example, treating the cancer may be referred to alleviating the cancer. For another example, treating the cancer may be referred to removing the cancer from a cancer patient.

[0853] The above PARP inhibitor may include all factors capable of inhibiting PARP. The above PARP inhibitor may be olaparib, niraparib, rucaparib, or veliparib.

[0854] Alternatively, in the same context as above, the present application specification discloses a pharmaceutical composition comprising a PARP inhibitor for treating cancer in a subject that does not contain a single nucleotide mutation of dysfunction in the ATM gene. In this case, the foregoing details regarding the single nucleotide mutation of dysfunction, the cancer, and the PARP inhibitor are incorporated by reference.

[0855] Method of providing information for the diagnosis of ataxia telangiectasia syndrome

[0856] One application of the functional information provision method of the present application is a method for diagnosing ataxia-telangiectasia syndrome (AT). Alternatively, one application of the functional information provision method of the present application is a method for providing information for diagnosing ataxia-telangiectasia syndrome. Although the following description is based on the diagnostic method, the following content may also be applied to a method for providing information for diagnosing ataxia-telangiectasia syndrome.

[0857] The above diagnostic method is characterized by utilizing the method for providing functional information of the present application. That is, it is characterized by diagnosing an ATM or providing information related to an ATM while including all or part of the method for providing functional information of the present application.

[0858] The functional information in the method for providing functional information in the present application includes information regarding the incidence of cancer. In one embodiment, if both ATM genes possessed by a subject contain dysfunctional single nucleotide mutations, the information may include that the subject has ataxia telangiectasia syndrome or that the subject is highly likely to be diagnosed with ataxia telangiectasia syndrome.

[0859] The above diagnostic method may include all or part of the present application’s ATM gene function information provision method or the present application’s ATM protein function information provision method. In this regard, the contents of Chapters 8 and 9 of this specification are included. For example, the above diagnostic method may include a) obtaining the ATM gene sequence of a subject, b) obtaining mutation information, and c) diagnosing ataxia telangiectasia syndrome. For another example, the above diagnostic method may include a) obtaining mutation information of a subject, and b) diagnosing that the subject has ataxia telangiectasia syndrome. For another example, the above diagnostic method may include a) obtaining the ATM protein sequence of a subject, b) obtaining amino acid mutation information, and c) diagnosing ataxia telangiectasia syndrome. For another example, the above diagnostic method may include a) obtaining amino acid mutation information of a subject, and b) diagnosing that the subject has ataxia telangiectasia syndrome. If the method is for providing information for the diagnosis of ataxia telangiectasia syndrome, the diagnosis may be providing information about ataxia telangiectasia syndrome.

[0860] The above diagnosis may vary depending on the number and condition of ATM genes possessed by the subject. In one embodiment, the diagnosis may be that the subject is diagnosed with ataxia telangiectasia syndrome only when the function of all ATM genes possessed by the subject is impaired or in an intermediate state. That is, the diagnosis may be that the subject is diagnosed with ataxia telangiectasia syndrome when they possess ATM genes in an impaired or intermediate state. In another embodiment, the diagnosis may be that the subject is diagnosed as a carrier of ataxia telangiectasia syndrome when some ATM genes possessed by the subject are impaired or in an intermediate state. That is, the diagnosis may be that the subject is diagnosed as a carrier of ataxia telangiectasia syndrome when they carry ATM genes in an impaired or intermediate state.

[0861] In the above diagnostic method, the number of single nucleotide variants included in the subject is not limited. For example, if the subject has an ATM gene containing a specific A single nucleotide variant and a specific B single nucleotide variant, and an ATM gene containing a specific C single nucleotide variant and a specific D single nucleotide variant, information can be provided by determining whether the chromosome based on the specific A single nucleotide variant and the specific B single nucleotide variant is dysfunctional, and determining whether the chromosome based on the specific C single nucleotide variant and the specific D single nucleotide variant is dysfunctional.

[0862] The above diagnostic method includes diagnosing that the subject has telangiectasia ataxia syndrome.

[0863] In one embodiment, if a single nucleotide mutation in the ATM gene of a subject or an amino acid mutation in the ATM protein of a subject is dysfunctional, the subject may be diagnosed with ataxia telangiectasia syndrome. Alternatively, the method for providing information for diagnosing ataxia telangiectasia syndrome may include providing information related to ataxia telangiectasia syndrome.

[0864] In one embodiment, if a single nucleotide variant in the subject's ATM gene or an amino acid variant in the subject's ATM protein is dysfunctional, information may be provided that the subject has ataxia telangiectasia syndrome. Alternatively, information may be provided that the subject is likely to have ataxia telangiectasia syndrome.

[0865] The above diagnostic method or the method for providing information for diagnosing ataxia telangiectasia syndrome may include additional processes. In one embodiment, it may include treating ataxia telangiectasia syndrome. In one specific embodiment, it may further include treating by correcting a single nucleotide mutation in the subject. In this case, the correction may involve replacing the base at the single nucleotide mutation site of the subject for which the diagnosis or information is provided with another base. Furthermore, the substitution with another base may be a substitution that does not result in a single nucleotide mutation that is dysfunctional.

[0866] The above method for providing information for the diagnosis of ataxia telangiectasia syndrome may provide different information depending on the subject acquiring the single nucleotide variant information. For example, if the subject is a fetus or a fertilized egg, the above method for providing information for the diagnosis of ataxia telangiectasia syndrome may provide information that the fetus or fertilized egg will be diagnosed with ataxia telangiectasia syndrome when born as an individual.

[0867]

[0868] Possible embodiments of the invention

[0869] Hereinafter, embodiments of the present application will be described. The following embodiments are intended to explain more specifically what is disclosed in this specification, and it will be obvious to those skilled in the art that the scope of this specification is not limited by the following embodiments.

[0870] Library containing engineered cells

[0871] [Example 1] Library 1

[0872] As a library containing engineered cells,

[0873] Each of the above-mentioned engineered cells contains only one single nucleotide variant in the genetic region of interest and one intended synonymous mutation,

[0874] The above-mentioned engineered cells are classified into various types, and

[0875] Different types of engineered cells contain different single nucleotide mutations,

[0876] The number of types of the above-mentioned engineered cells is more than 80% of the number of single nucleotide variants that may exist in the genetic region of interest.

[0877] [Example 2] Library 2

[0878] As a library containing engineered cells,

[0879] Each of the above-mentioned engineered cells contains only one single nucleotide variant in the genetic region of interest and one intended synonymous mutation,

[0880] The types of the above-mentioned engineered cells are classified by the single nucleotide mutations they contain, and

[0881] The number of types of the above-mentioned engineered cells is more than 80% of the number of single nucleotide variants that may exist in the genetic region of interest.

[0882] [Example 3] Limitation of genetic region of interest

[0883] In Example 1 or 2, the genetic region of interest is a library selected from the following:

[0884] (a) A specific exon region of the ATM gene;

[0885] (b) a region corresponding to -5 bp bases of the upstream sequence of a specific exon of the ATM gene to +5 bp bases of the downstream sequence of a specific exon;

[0886] (c) a specific exon of the ATM gene and a region corresponding to 5 bp in the upstream and downstream regions adjacent to the specific exon;

[0887] (d) Specific exon regions of the ATM gene;

[0888] (e) For specific exons of the ATM gene, regions corresponding to combinations of -5 bp bases in the upstream sequence of each exon to +5 bp bases in the downstream region;

[0889] (f) specific exons of the ATM gene and regions corresponding to 5 bp upstream and downstream adjacent to each exon; and

[0890] (g) In (d) to (f) above, the specific multiple exons are regions selected from 2 or more of exons 2 to 63.

[0891] [Example 4] Limited to synonymous mutations

[0892] In any one of Examples 1 to 3, the synonymous mutation is present in the coding sequence, in a library.

[0893] [Example 5] Reference sequence limitation

[0894] A library in any one of Examples 1 to 4, wherein the single nucleotide variant and the synonymous mutation are defined based on one sequence selected from SEQ ID NOs 1 to 3.

[0895] [Example 6] Cell-limited

[0896] In any one of Examples 1 to 5, the engineered cell is a library of cells comprising one or more features selected from the following:

[0897] (a) The ploidy of the ATM gene is haploid or polyploid, or the ATM gene is haploidized;

[0898] (b) functional BRCA1, BRCA2, and TP53 expression; and

[0899] (c) Prime Editor manifestation.

[0900] Method for evaluating the function of ATM genes with single nucleotide variants_High-throughput method

[0901] [Example 7] High-throughput method 1

[0902] A method for evaluating ATM gene function having a single nucleotide variant comprising the following.

[0903] (a) Prepare a cell library of any one of Examples 1 to 6;

[0904] (b) acquiring survival information of engineered cells; and

[0905] (c) Using the above survival information, obtain functional information.

[0906] [Example 8] High-throughput method 2

[0907] A method for evaluating ATM gene function having a single nucleotide variant comprising the following.

[0908] (a) Prepare a library of any one of Examples 1 to 6; and

[0909] (b) Acquiring survival information of engineered cells, thereby acquiring information on specific single nucleotide mutations.

[0910] [Example 9] Limited library preparation

[0911] In Example 8 or 9, the preparation of the library comprises a method for evaluating ATM gene function having a single nucleotide variant, comprising:

[0912] (a-1) Prepare cells and a pegRNA library; and

[0913] (a-2) Introduce a pegRNA library into cells.

[0914] [Example 10] pegRNA limitation

[0915] In Example 9, the pegRNA library comprises various types of pegRNA, and

[0916] Each of the above pegRNAs is designed to introduce only one single nucleotide variant and one intended synonymous mutation into the genetic region of interest, and

[0917] The types of the above pegRNA are classified according to the type of single nucleotide mutation that the pegRNA intends to introduce, and

[0918] A method for evaluating the function of an ATM gene having single nucleotide variants, wherein the number of types of pegRNAs is 90% or more or 100% of the number of single nucleotide variants that may exist in the genetic region of interest.

[0919] [Example 11] Cell-limited

[0920] In Example 9, a method for evaluating ATM gene function having a single nucleotide variant, wherein the cell is a cell comprising one or more features selected from the following:

[0921] (a) The genomic ATM gene contains the sequence of SEQ ID NO. 1, 2 or 3;

[0922] (b) the ploidy of the ATM gene is haploid or polyploid, or the ATM gene is haploidized;

[0923] (c) functional BRCA1, BRCA2, and TP53 expression; and

[0924] (d) Prime Editor manifestation.

[0925] [Example 12] Method for introducing pegRNA 1

[0926] In Example 11, when the cell does not express prime data, introducing the pegRNA library comprises a method for evaluating ATM gene function having a single nucleotide variant, comprising:

[0927] (a-2-1) Introduce a vector containing a prime editor protein or a nucleic acid encoding a prime editor into the cell; and

[0928] (a-2-2) Introduced a pegRNA library.

[0929] [Example 13] Method for introducing pegRNA 2

[0930] In any one of Examples 9 to 12, the pegRNA is a method for evaluating the function of an ATM gene having a single nucleotide variant introduced via a lentivirus.

[0931] [Example 14] Limited to obtaining survival information 1

[0932] In any one of Examples 7 to 13, obtaining the survival information comprises a method for evaluating the function of an ATM gene having a single nucleotide variant, comprising:

[0933] (a) Divide the library into a first library and a second library;

[0934] (b) Sample the first library at the first time point;

[0935] (c) sampling the second library at the second time point; and

[0936] (d) Compare the sequencing results of the sampling at time 1 with the sequencing results of the sampling at time 2.

[0937] [Example 15] Limited to obtaining survival information 2

[0938] In any one of Examples 7 to 13, obtaining the survival information comprises a method for evaluating the function of an ATM gene having a single nucleotide variant, comprising:

[0939] (a) Divide the library into a first library and a second library;

[0940] (b) Sampling at time 1 without treating the first library with a PARP inhibitor;

[0941] (c) sampling the second library treated with a PARP inhibitor at a second time point; and

[0942] (d) Compare the sequencing results of the sampling at time 1 with the sequencing results of the sampling at time 2.

[0943] [Example 16] Limited to obtaining survival information 3

[0944] A functional evaluation method in any one of Examples 7 to 15, wherein obtaining the survival information is obtaining survival information for each engineered cell type.

[0945] [Example 17] Function Information Limitation

[0946] A method for evaluating the function of an ATM gene having a single nucleotide variant, wherein in any one of Examples 7 to 16, the functional information or information regarding the single nucleotide variant is a numerical value of survival information or qualitative information regarding the single nucleotide variant.

[0947] Deep learning model for evaluating the function of ATM genes with single nucleotide polymorphisms

[0948] [Example 18] DeepATM learning method

[0949] A model training method for evaluating the function of an ATM gene having a single nucleotide variation or an ATM protein having an amino acid variation, comprising the following.

[0950] (a) Prepare amino acid variant information of the ATM protein encoded by the ATM gene having a single nucleotide variant, or amino acid variant information of the ATM protein, and survival information;

[0951] (b) Pass amino acid mutation information and survival information through the transformer layer;

[0952] (c) Obtain an output value by passing the value corresponding to the amino acid mutation position among the output values ​​of the transformer layer through the neural network layer; and

[0953] (d) Compare the above output value with the above survival information,

[0954] [Example 19] DeepATM data acquisition method

[0955] In Example 17, the survival information is obtained through a learning method using any one of the methods in Examples 7 to 17.

[0956] [Example 20] Embedding limitation

[0957] In Example 18 or 19, the amino acid variation information is embedded in an amino acid embedding vector, a domain embedding vector, and a coordinate embedding, and

[0958] In the above amino acid embedding vector, the value for the amino acid at each position is randomly set at the start of learning, and

[0959] In the above domain embedding vector, values ​​for each domain are randomly set at the start of training, and

[0960] A learning method in which the coordinate information of the alpha carbon of the amino acid at each position is set in the above coordinate embedding vector by a certain method.

[0961] [Example 21] Coordinate Embedding Limitation

[0962] In Example 20, the certain method is a learning method comprising the following:

[0963] (a) Using a protein structure model prediction model, derive the coordinates of the alpha carbon of the amino acid at each position; and

[0964] (b) The value derived by passing the above coordinates through a multilayer perceptron is set as the embedding value.

[0965] [Example 22] Neural network layer limited

[0966] A learning method in any one of Examples 18 to 21, wherein the output value corresponding to the amino acid variation position is concatenated with the score of the ATM protein using an existing computational model and passes through a neural network layer.

[0967] Method for Evaluating Function of ATM Genes with Single Nucleotide Variants_Deep Learning

[0968] [Example 23] Deep Learning 1

[0969] A method for evaluating ATM gene function with a single nucleotide variant comprising the following:

[0970] (a) Prepare a model trained by any one of Examples 18 to 22 and ATM gene variant information having a single nucleotide variant;

[0971] (b) converting ATM gene variant information with single nucleotide variations into ATM protein variant information with amino acid variations; and

[0972] (c) Functional information of the ATM gene having a single nucleotide mutation was obtained using the above model and ATM protein mutation information.

[0973] [Example 24] Deep Learning 2

[0974] A method for evaluating ATM gene function with a single nucleotide variant comprising the following:

[0975] (a) preparing ATM protein variant information encoded by an ATM gene having a single nucleotide variant and a model trained by any one of Examples 18 to 22; and

[0976] (b) Using the above model and ATM protein mutation information, functional information of the ATM gene having a single nucleotide mutation is obtained.

[0977] Method for evaluating the function of ATM proteins with amino acid variations

[0978] [Example 25] Amino Acid Variation Criteria High Throughput Method 1

[0979] Method for evaluating ATM protein function with amino acid variations including the following:

[0980] (a) Prepare a cell library of any one of Examples 1 to 6;

[0981] (b) acquiring survival information of engineered cells; and

[0982] (c) Using the above survival information, functional information for ATM proteins having amino acid mutations is obtained.

[0983] [Example 26] Amino Acid Variation Criteria High Throughput Method 2

[0984] Method for evaluating ATM protein function with amino acid variations including the following:

[0985] (a) Prepare a cell library of any one of Examples 1 to 6; and

[0986] (b) Acquiring survival information of engineered cells, thereby acquiring information on specific amino acid mutations.

[0987] [Example 27] Deep learning method based on amino acid variation

[0988] Method for evaluating ATM protein function with amino acid variations including the following:

[0989] (a) preparing a model and ATM protein variant information trained by any one of Examples 18 to 22; and

[0990] (b) Using the above model and ATM protein mutation information, functional information of the ATM protein having amino acid mutations is obtained.

[0991] Method for evaluating the function of ATM genes having single nucleotide variations or methods for evaluating the function of ATM proteins having amino acid variations

[0992] [Example 28] Function evaluation method (integration) 1

[0993] Function evaluation method including the following:

[0994] (a) Prepare a library of any one of Examples 1 to 6;

[0995] (b) Acquire survival information of engineered cells;

[0996] (c) preparing a model trained by any one of Examples 17 to 22 using survival information; and

[0997] (d) Functional information of ATM genes with single nucleotide mutations or ATM proteins with amino acid mutations is obtained using survival information and a deep learning model utilizing survival information.

[0998] [Example 29] Function evaluation method (integration) 2

[0999] Function evaluation method including the following:

[1000] (a) Prepare a library of any one of Examples 1 to 6;

[1001] (b) Acquire survival information of engineered cells;

[1002] (c) preparing a model trained by any one of Examples 17 to 22 using survival information; and

[1003] (d) Using survival information, functional information of the ATM gene with a single nucleotide mutation or the ATM protein with an amino acid mutation is obtained.

[1004] [Example 30] Function evaluation method (integration) 3

[1005] Function evaluation method including the following:

[1006] (a) Prepare a library of any one of Examples 1 to 6.

[1007] (b) Acquiring survival information of engineered cells, thereby acquiring functional information of ATM genes having single nucleotide mutations or ATM proteins having amino acid mutations.

[1008] Method of providing ATM gene function information

[1009] [Example 31] Method for providing functional information 1

[1010] Method for providing ATM gene function information including the following:

[1011] (a) Acquired ATM gene sequence;

[1012] (b) obtaining mutation information; and

[1013] (c) Provide ATM gene function information based on the above mutation information.

[1014] [Example 32] Method for Providing Function Information 2_Addition of Object

[1015] Method for providing ATM gene function information including the following:

[1016] (a) Obtain the ATM gene sequence of the subject;

[1017] (b) determining or confirming whether the subject comprises at least one of specific single nucleotide variants; and

[1018] (c) Provide ATM gene function information based on the above mutation information.

[1019] [Example 33] Method for Providing Functional Information 3_Addition of Protein-Related Content

[1020] Method for providing ATM gene function information including the following:

[1021] (a) Obtain the ATM protein sequence encoded by the ATM gene of the subject;

[1022] (b) determining or confirming whether the ATM protein of the subject contains at least one of specific amino acid variants; and

[1023] (c) Provide ATM gene function information based on the above mutation information.

[1024] [Example 34] Limited provision of functional information

[1025] A method in any one of Examples 31 to 33, wherein providing functional information according to the mutation information is providing information that the ATM gene is in a dysfunctional state, information that the ATM gene is in an intermediate state, or information that the ATM gene is in a functional state.

[1026] [Example 35] Method for providing function information_Addition of function information options

[1027] Method for providing ATM gene function information including the following:

[1028] (a) Obtain the ATM gene sequence of the subject;

[1029] (b) determining or confirming whether the subject comprises at least one of specific single nucleotide variants; and

[1030] (c-1) If the subject contains at least one of specific single nucleotide variants, the subject’s ATM gene provides information that it is in a dysfunctional state,

[1031] (c-2) If the subject does not contain at least one of specific single nucleotide variants, it provides information that the subject's ATM gene is not in a dysfunctional state.

[1032] [Example 36] Single nucleotide mutation limitation

[1033] A method in Example 32, 34, or 35, wherein the specific single nucleotide variants are the single nucleotide variants described in Table 1 of the present application specification or a part thereof.

[1034] [Example 37] Amino acid variation limitation

[1035] A method in Example 33, wherein the specific amino acid variants are amino acid variants or parts thereof having a functional score of -1.360 or less among the functional scores of amino acid variants listed in Table 1 of the present application.

[1036] [Example 38] Addition of information acquisition process

[1037] A method in which, in any one of Examples 31 to 37, functional information is obtained through any one of Examples 7 to 17 and Examples 23 to 30.

[1038] Method of providing ATM protein function information

[1039] [Example 39] Method for providing protein function information 1

[1040] Method for providing ATM protein function information including the following:

[1041] (a) Acquired ATM protein sequence;

[1042] (b) obtaining mutation information; and

[1043] (c) Provides ATM protein function information based on the above mutation information.

[1044] [Example 40] Method for Providing Protein Function Information 2_Addition of Target

[1045] Method for providing ATM protein function information including the following:

[1046] (a) Obtain the ATM protein sequence of the subject;

[1047] (b) determining whether the ATM protein contains at least one of a specific amino acid variant; and

[1048] (c) Provides ATM protein function information based on the above mutation information.

[1049] [Example 41] Limited provision of functional information

[1050] A method according to Example 39 or 40, wherein providing functional information according to the mutation information is providing information that the ATM protein is in a dysfunctional state, information that the ATM protein is in an intermediate state, or information that the ATM protein is in a functional state.

[1051] [Example 42] Method for Providing Protein Function Information 3_Granting Function Information Options

[1052] Method for providing ATM protein function information including the following:

[1053] (a) Obtain the ATM protein sequence of the subject;

[1054] (b) determining or confirming whether the ATM protein contains at least one of a specific amino acid variant; and

[1055] (c-1) If the ATM protein contains at least one of specific amino acid variants, the ATM protein of the subject provides information that it is in a state of dysfunction.

[1056] (c-2) If the ATM protein does not contain at least one of the specific amino acid variants, it provides information that the ATM protein of the subject is not in a state of dysfunction.

[1057] [Example 43] Amino acid variation limitation

[1058] A method in any one of Examples 40 to 42, wherein the specific amino acid variants are amino acid variants or parts thereof having a functional score of -1.360 or less among the functional scores of amino acid variants listed in Table 1 of the present application.

[1059] [Example 44] Addition of information acquisition process

[1060] A method in which, in any one of Examples 39 to 43, functional information is obtained through any one of Examples 7 to 17 and Examples 23 to 30.

[1061] Utilization of information provision methods

[1062] [Example 45] Utilization of Information Provision Method (Gene) 1

[1063] Method of providing information including the following:

[1064] (a) Obtain the ATM gene sequence of the subject;

[1065] (b) determining or confirming whether the subject comprises at least one of specific single nucleotide variants; and

[1066] (c) Provide information based on the above mutation information.

[1067] [Example 46] Utilization of Information Provision Method (Protein) 1

[1068] Method of providing information including the following:

[1069] (a) Obtain the ATM protein sequence of the subject;

[1070] (b) determining or confirming whether the ATM protein contains at least one of specific amino acid variants; and

[1071] (c) Provide information based on the above mutation information.

[1072] [Example 47] Utilization of Information Provision Method (Gene) 2_Addition of Option

[1073] Method of providing information including the following:

[1074] (a) Obtain the ATM gene sequence of the subject;

[1075] (b) determining or confirming whether the subject comprises at least one of specific single nucleotide variants; and

[1076] (c-1) If the subject comprises at least one of specific single nucleotide variants, the first information is provided,

[1077] (c-2) If the subject does not contain at least one of specific single nucleotide variants, provide second information.

[1078] [Example 48] Utilization of Information Provision Method (Gene) 2_Addition of Option (Protein)

[1079] Method of providing information including the following:

[1080] (a) Obtain the ATM protein sequence of the subject;

[1081] (b) determining or confirming whether the ATM protein sequence comprises at least one of specific amino acid variants; and

[1082] (c-1) If the ATM protein sequence comprises at least one of specific amino acid variants, the first information is provided.

[1083] (c-2) If the above ATM protein sequence does not contain at least one of specific amino acid variants, provide second information.

[1084] [Example 49] Addition of information acquisition ...

Claims

1. A method for providing functional information on the ATM gene (Ataxia-telangiectasia mutated gene) of a subject including the following: (a) Prepare a cell library, At this time, the cell library includes engineered cells, and Each of the above-mentioned engineered cells contains one single nucleotide variant and one intended synonymous mutation in the gene region of interest, and The above-mentioned engineered cells are classified into various types according to single nucleotide mutations in the gene region of interest, and The number of types of the above-mentioned engineered cells is more than 80% of the number of single nucleotide variants that may exist in the gene region of interest, and The above-mentioned gene region of interest consists of ATM gene exons 2 to 63 and regions corresponding to 5 bp upstream and downstream adjacent to each exon; (b) by obtaining survival information of each type of the above-mentioned engineered cell, if the ATM gene contains at least one single nucleotide variant selected from the list of single nucleotide variants below, information is obtained that the ATM gene is dysfunctional; At this time, the above single nucleotide variant is named using the HGVS nomenclature with the sequence of SEQ ID NO. 1 as the reference sequence, and At this time, the sequence of SEQ ID NO. 1 is a reference sequence corresponding to the region from -5bp of exon 2 to 5bp of exon 63 of the ATM gene; (c) Acquire the ATM gene sequence of the subject; (d) determining, based on the result of (c) above, whether the ATM gene of the subject contains at least one single nucleotide variant selected from the following list of single nucleotide variants; and (e) If it is determined that the ATM gene of the subject contains at least one single nucleotide variant selected from the list of single nucleotide variants below, information is provided that the ATM gene of the subject is in a dysfunctional state. <List of single nucleotide variants> Single nucleotide mutations in Table 1 2. A method for providing functional information on the ATM gene (Ataxia-telangiectasia mutated gene) of a subject including the following: (a) Prepare a cell library, At this time, the cell library includes engineered cells, and Each of the above-mentioned engineered cells contains one single nucleotide variant and one intended synonymous mutation in the gene region of interest, and The above-mentioned engineered cells are classified into various types according to single nucleotide mutations in the gene region of interest, and The number of types of the above-mentioned engineered cells is more than 80% of the number of single nucleotide variants that may exist in the gene region of interest, and The above-mentioned gene region of interest consists of ATM gene exons 2 to 63 and regions corresponding to 5 bp upstream and downstream adjacent to each exon; (b) by obtaining survival information of each type of the above-mentioned engineered cell, if the ATM gene encodes an ATM protein containing at least one amino acid variant selected from the list of amino acid variants below, information is obtained that the ATM gene is dysfunctional; At this time, the above amino acid variation is named using the sequence of SEQ ID NO. 128 as a reference sequence, and At this time, the sequence of SEQ ID NO. 128 is a reference sequence corresponding to the reference ATM protein; (c) Obtain the ATM protein sequence encoded by the ATM gene of the subject; (d) Based on the result of (c) above, determine whether the ATM protein of the subject comprises at least one amino acid variant selected from the following list of amino acid variants; and (e) If it is determined that the ATM protein of the subject contains at least one amino acid variant selected from the following list of amino acid variants, information is provided that the ATM gene of the subject is in a dysfunctional state. <List of Amino Acid Variants> Amino acid variations of dysfunction in Table 1 3. A method for providing functional information on the ATM gene (Ataxia-telangiectasia mutated gene) of a subject including the following: (a) Obtain the ATM gene sequence of the subject; (b) determining whether the ATM gene of the subject contains at least one single nucleotide variant selected from the list of single nucleotide variants below; and (c) If it is determined that the ATM gene of the subject contains at least one single nucleotide variant selected from the list of single nucleotide variants below, information is provided that the ATM gene of the subject is in a dysfunctional state. <List of single nucleotide variants> Single nucleotide variants in Table 1, At this time, the single nucleotide variants in Table 1 are named using the HGVS nomenclature with the sequence of SEQ ID No. 1 as the reference sequence, and The sequence of SEQ ID NO. 1 above is a reference sequence corresponding to the region from -5 bp of exon 2 to 5 bp of exon 63 of the ATM gene.

4. A method for providing functional information on the ATM gene (Ataxia-telangiectasia mutated gene) of a subject including the following: (a) Obtain the ATM protein sequence encoded by the ATM gene of the subject; (b) determining whether the ATM protein of the subject comprises at least one amino acid variant selected from the following list of amino acid variants; and (c) If it is determined that the ATM protein of the subject contains at least one amino acid variant selected from the following list of amino acid variants, information is provided that the ATM gene of the subject is in a dysfunctional state. <List of Amino Acid Variants> Amino acid variations of dysfunction in Table 1 At this time, the above amino acid variation is named using the sequence of SEQ ID NO. 128 as a reference sequence, and The sequence of SEQ ID NO 128 above is a reference sequence corresponding to the reference ATM protein.

5. Methods for treating cancer in cancer patients, including the following: (a) Obtain the ATM gene sequence of the subject; (b) If the subject’s ATM gene contains at least one single nucleotide variant selected from the list of single nucleotide variants below, treat the cancer by administering a PARP inhibitor, <List of single nucleotide variants> Single nucleotide variants in Table 1, At this time, the single nucleotide variants in Table 1 are named using the HGVS nomenclature with the sequence of SEQ ID No. 1 as the reference sequence, and The sequence of SEQ ID NO. 1 above is a reference sequence corresponding to the region from -5 bp of exon 2 to 5 bp of exon 63 of the ATM gene.

6. Methods for treating cancer in cancer patients, including the following: (a) Obtain the ATM protein sequence encoded by the ATM gene of the subject; (b) If the ATM protein of the subject comprises at least one amino acid variant selected from the following list of amino acid variants, treat the cancer by treating with a PARP inhibitor. <List of Amino Acid Variants> Amino acid variations of dysfunction in Table 1 At this time, the above amino acid variation is named using the sequence of SEQ ID NO. 128 as a reference sequence, and The sequence of SEQ ID NO 128 above is a reference sequence corresponding to the reference ATM protein.

7. A method according to any one of claims 5 to 6, wherein the cancer is breast cancer, lung cancer, colorectal cancer, stomach cancer, liver cancer, pancreatic cancer, prostate cancer, bladder cancer, ovarian cancer, cervical cancer, esophageal cancer, kidney cancer, lymphoma, blood cancer, leukemia, skin cancer, brain tumor, osteosarcoma, thyroid cancer, oral cancer, bile duct cancer, or melanoma.

8. A method according to any one of claims 5 to 6, wherein the PARP inhibitor is olaparib, niraparib, rucaparib, or veliparib.

9. A pharmaceutical composition for treating cancer comprising a PARP inhibitor for treating cancer in a cancer patient having at least one single nucleotide mutation selected from the list of single nucleotide mutations below in the ATM gene: <List of single nucleotide variants> Single nucleotide variants in Table 1, At this time, the single nucleotide variants in Table 1 are named using the HGVS nomenclature with the sequence of SEQ ID No. 1 as the reference sequence, and The sequence of SEQ ID NO. 1 above is a reference sequence corresponding to the region from -5 bp of exon 2 to 5 bp of exon 63 of the ATM gene.

10. A pharmaceutical composition for treating cancer comprising a PARP inhibitor for treating cancer in a cancer patient in which the ATM protein encoded by the ATM gene comprises at least one amino acid mutation selected from the list of amino acid mutations below: <List of Amino Acid Variants> Amino acid variations of dysfunction in Table 1 At this time, the above amino acid variation is named using the sequence of SEQ ID NO. 128 as a reference sequence, and The sequence of SEQ ID NO 128 above is a reference sequence corresponding to the reference ATM protein.

11. The method of claim 9 or 10, wherein the cancer is breast cancer, lung cancer, colorectal cancer, stomach cancer, liver cancer, pancreatic cancer, prostate cancer, bladder cancer, ovarian cancer, cervical cancer, esophageal cancer, kidney cancer, lymphoma, blood cancer, leukemia, skin cancer, brain tumor, osteosarcoma, thyroid cancer, oral cancer, bile duct cancer, or melanoma.

12. The method of claim 9 or 10, wherein the PARP inhibitor is olaparib, niraparib, rucaparib, or veliparib.