Systems and methods for identifying the association between mutations and phenotypes
The Candidate Explorer system uses machine learning to integrate gene mapping data into a single score, addressing the limitations of existing methods by accurately predicting mutation-phenotype associations, enhancing the identification of causative mutations and their role in phenotypes.
Patent Information
- Application Number
- JP2025500187
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-01
- Filing Date
- 2023-06-21
- Publication Date
- 2025-07-23
AI Technical Summary
Existing methods for identifying causative mutations in forward genetics are limited by false positives due to co-segregation of multiple mutations, lack of homozygotes, and accidental recognition, leading to challenges in accurately associating phenotypes with specific mutations.
A Candidate Explorer (CE) system using machine learning algorithms integrates gene mapping data into a single numerical score to predict the probability of a mutation-phenotype association, incorporating features like damage scores, essentiality scores, and linkage data to improve the identification of causative mutations.
The CE system enhances the accuracy of identifying causative mutations by providing a CE score and verification probability, accurately distinguishing between true and false associations, thereby facilitating rapid identification of mutations causing phenotypes in mice and potentially human diseases.
Smart Images

Figure 2025523638000001_ABST
Abstract
Description
Technical Field
[0001] 1. Technical Field Aspects of the concepts of the present invention generally relate to systems and methods for mutagenesis, and more particularly, to systems and methods for identifying the relationship between phenotype and mutation.
[0002] [Cross - reference to related applications] This application claims priority to U.S. Provisional Patent Application No. 63 / 357,803, filed on July 1, 2022, entitled "SYSTEMS AND METHODS TO IDENTIFY MUTATION AND PHENOTYPE ASSOCIATION", which is hereby incorporated by reference in its entirety.
[0003] [Acknowledgment of government support] This invention was made with government support under grant numbers AI125581 and AI100627 awarded by the National Institutes of Health. The government has certain rights in this invention.
Background Art
[0004] 2. Consideration of related technologies Phenotype refers to a set of observable characteristics resulting from the interaction between genotype and environment. In some cases, gene mutations can cause phenotypes. Mutations generally refer to changes in deoxyribonucleic acid (DNA) sequences. Mutations can occur due to DNA copying during cell division, ionizing radiation, mutagens, or viral infections.
Summary of the Invention
[0005] Certain aspects of the disclosed technology can provide a method for mutagenesis. The method generally includes receiving one or more input features including phenotypic data and mutation data, and generating, via a machine learning model, a candidate explorer (CE) score indicative of the probability of an association between a phenotype and a mutation based on the one or more input features, and outputting an indicator of the association between the phenotype and the mutation based on the CE score.
[0006] In some examples, the indicator of the association includes a candidate status of the association based on the CE score and an algorithm score indicating the likelihood that a mutation is the cause. The mutation data can include a damage score indicating the likelihood that a protein associated with the mutation is functionally impaired. Additionally, the method can include generating the damage score via another machine learning model trained using known deleterious and neutral mutations. Also, the one or more input features can further include an essentiality score indicating the likelihood of pre-weaning lethality in mice homozygous for a robust knockout allele of a gene associated with the mutation. Additionally, or alternatively, the method can include generating the essentiality score via another machine learning model trained using genes known not to be essential for survival and genes known to be essential for survival. The one or more input features can further include features associated with an algorithm score indicating the likelihood that a mutation is the cause.
[0007] Furthermore, in some examples, one or more input features can further include linkage data generated using automatic meiotic mapping (AMM). The method also includes determining which of two or more mutations is a more robust candidate cause of a phenotype by omitting instances of co-segregation of two or more mutations, and generating a CE score based on that determination.Furthermore, one or more input functions are the number of phenotypes having an algorithm score for a mutation that meets a threshold, where the algorithm score is the number of phenotypes indicating the likelihood of being caused by the mutation, the average number of AMM operations resulting in a p-value that meets the threshold for each allele of the gene associated with the mutation, the algorithm score of the mutation or phenotype, the number of AMM operations resulting in a p-value that meets the threshold of the gene associated with the mutation, the damage score of the mutation indicating the likelihood that the protein associated with the mutation is functionally impaired, the number of families within the superpedigree associated with the gene, and whether the p-value of the result of the AMM operation for the superpedigree meets the threshold, the number of phenotypes having a p-value of the superpedigree that meets the threshold, the number of families contributing to the p-value of the superpedigree that meets the threshold and the number of families of the superpedigree, the percentage of fluorescence-activated cell sorting (FACS) screens having a p-value that meets the threshold of the mutation, the minimum value of the p-value from the AMM operation, the percentage of mutant allele (VAR) mice whose screening results overlap with the screening results of B6 mice, whether the result of the AMM operation for the superpedigree meets the thresholds of the null allele and missense allele, whether the result of the AMM operation for the superpedigree meets the threshold of the null allele, the percentage of VAR mice whose screening results overlap with the screening results of reference allele (REF) mice, the difference in the results of the AMM operation for heterozygous (HET) mice and VAR mice, the number of female REF mice used in the AMM operation, the percentage of body weight screens having a p-value that meets the threshold of the mutation, the number of female HET mice used in the AMM operation, and / or at least one of the differences in the results of the AMM operation for REF mice and VAR mice.
[0008] Furthermore, certain aspects of the disclosed technology can provide an apparatus for mutagenesis processing. The apparatus generally includes a memory and one or more processors coupled to the memory and configured to receive one or more input features including phenotypic data and mutation data, generate a CE score indicating the probability of the relevance between a phenotype and a mutation based on the one or more input features via a machine learning model, and output an indicator of the relevance between the phenotype and the mutation based on the CE score.
[0009] In some examples, the relevance indicator includes a relevance candidate status based on the CE score and an algorithm score indicating the likelihood that a mutation is the cause. The mutation data can also include a damage score indicating the likelihood that a protein associated with the mutation is functionally impaired. Further, the one or more processors can be further configured to generate the damage score via another machine learning model trained using known deleterious and neutral mutations. Also, the one or more input features can further include an essentiality score indicating the likelihood of pre-weaning lethality in homozygous mice for a robust knockout allele of a gene associated with the mutation.
[0010] In some examples, one or more processors can be further configured to generate an essentiality score via another machine learning model trained using genes known not to be essential for survival and genes known to be essential for survival. One or more input features can also include features associated with an algorithm score indicating the likelihood of being caused by a mutation. Additionally, one or more input functions can include linkage data generated using automated meiotic mapping (AMM). Further, if two or more mutations are co-segregated, one or more processors can be further configured to determine which of the two or more mutations is a more robust candidate cause of the phenotype by omitting instances of the shared homozygosity of the two or more mutations, and one or more processors can be configured to generate a CE score based on that determination.
[0011] Certain aspects of the disclosed technology provide a non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to receive one or more input features including phenotypic data and mutation data, generate a CE score indicating the probability of a relationship between a phenotype and a mutation based on the one or more input features via a machine learning model, and output an indicator of the relationship between the phenotype and the mutation based on the CE score.
[0012] Other implementations are also described and cited herein. Further, while multiple implementations are disclosed, other implementations of the presently disclosed technology will be apparent to those skilled in the art from the following detailed description, which illustrates and describes exemplary implementations of the presently disclosed technology. As will be understood, the presently disclosed technology is capable of modification in various aspects without departing from the spirit and scope of the presently disclosed technology. Accordingly, the drawings and detailed description are to be regarded as illustrative in nature and not restrictive.
Brief Description of the Drawings
[0013]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5A
Figure 5B
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10A
Figure 10B
Figure 10C
Figure 11
DETAILED DESCRIPTION OF THE INVENTION
[0014] After considering the entire disclosure, it will be apparent to those skilled in the art that the steps shown in the above drawings may be performed in an order other than the described order, and that one or more of the steps shown in these drawings may be optional.
[0015] Particular aspects of the concepts of the present invention are directed to methods and systems for using machine learning algorithms to identify chemically induced mutations that cause screened phenotypes. For example, a candidate explorer (CE) system can determine the probability that a mutation will be verified as the cause of a phenotype when a gene is independently targeted for knockout or recapitulation of the mutation. The CE system (also abbreviated as "CE") uses a number of parameters (e.g., 67 parameters) from mapping data including gene, mutation, genotype, allele, and phenotype information to determine a CE score and a verification probability.
[0016] Forward genetic studies typically use meiotic mapping to provide evidence that a particular mutation, induced by a germline mutagen, is the cause of a particular phenotype. In small families in particular, co-segregation of multiple mutations, accidental lack of recognition of mutations, and lack of homozygotes can lead to false declarations of cause and effect. Certain aspects of the concepts of the present invention provide a system for improving the identification of mutations that cause immunophenotypes that can be identified in mice. The CE system can use machine learning to integrate the features of gene mapping data into a single numerical score that can be mathematically transformed into a probability of verifying any putative mutation-phenotype association.
[0017] The CE system can be used to assess the association between putative mutations and phenotypes that arise by screening for deleterious mutations in mouse genes (e.g., about 55%) and examining their effect on flow cytometric measurements of immune cells in the blood. The CE system can identify more than half of the genes containing mutations that may cause flow cytometric phenotype changes in Mus musculus (e.g., laboratory mice). Most of these genes may not have been previously known to support immune function or homeostasis. Mouse geneticists can use CE data to identify causative mutations within quantitative trait loci, which are regions of DNA associated with particular phenotypic traits. Clinical geneticists can use CE to help associate causative mutant forms with rare genetic diseases of the immune system, even in the absence of linkage information. CE displays integrated mutation, phenotype, and linkage data.
[0018] FIG. 1 shows an exemplary computing device 100 according to a particular aspect of the concepts of the present invention. The computing device 100 can include a processor 103 for controlling the overall operation of the computing device 100 and related components of the computing device 100, including input / output device 109, communication interface 111, and / or memory 115. A data bus can interconnect the processor 103, memory 115, I / O device 109, and / or communication interface 111.
[0019] The input / output (I / O) device 109 can include a microphone, keypad, touch screen, and / or stylus through which a user of the computing device 100 can provide input, and can also include one or more of a speaker for providing audio output and a video display device for providing text, audiovisual, and / or graphic output. Software can be stored in the memory 115 to provide instructions to the processor 103 and enable the computing device 100 to perform various actions. For example, the memory 115 can store software used by the computing device 100, such as an operating system 117, application programs 119, and / or associated internal databases 121. The various hardware memory units of the memory 115 can include volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storing information such as computer readable instructions, data structures, program modules, or other data. The memory 115 can include one or more physical persistent memory devices and / or one or more non-persistent memory devices. The memory 115 can include, but is not limited to, random access memory (RAM), read only memory (ROM), electrically erasable programmable read only memory (EEPROM), flash memory or other memory technology, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other media that can be used to store the required information and can be accessed by the processor 103.
[0020] As described herein, the communication interface 111 can include one or more transceivers, digital signal processors, and / or additional circuitry and software for communicating via any wired or wireless network using any protocol. The processor 103 can include a single central processing unit (CPU) that can be a single-core or multi-core processor (e.g., dual-core, quad-core, etc.), or can include multiple CPUs. The processor 103 and associated components enable the computing device 100 to execute a series of computer-readable instructions to perform some or all of the processes described herein. Although not shown in FIG. 1, various elements within the memory 115 or other components within the computing device 100 can include one or more caches, such as a CPU cache used by the processor 103, a page cache used by the operating system 117, a disk cache of a hard drive, and / or a database cache used to cache content from the database 121. In implementations that include a CPU cache, the CPU cache can be used by one or more processors 103 to reduce the memory latency and access time. Since the processor 103 can obtain data from or write data to the CPU cache rather than reading / writing to the memory 115, the speed of these operations can be improved. In some examples, a database cache can be created in which specific data from the database 121 is cached in another small database within a memory separate from the database, such as RAM or another computing device. For example, in a multi-tier application, using a database cache on an application server can eliminate the need to communicate with the backend database server via the network, thus reducing the time for data acquisition and data manipulation.These types of caches and other caches can be included in various implementations and can provide potential benefits such as reduced response time and reduced dependence on the network state during data transmission and reception in certain implementations of a software deployment system.
[0021] Forward genetics often starts with phenotypes induced by random germline mutagens and ends with the discovery of the causative mutations. Certain aspects provide techniques for rapidly identifying causative mutations in mice with N-ethyl-N-nitrosourea (ENU)-induced germline mutations. Certain aspects provide techniques for mutating male inbred lines of mice (e.g., C57BL / 6J) and breeding them on a C57BL / 6J background to create founder males of the first generation (G1), daughters of the second generation (G2), and both sexes of the third generation (G3) of mice. To achieve greater than 10-fold coverage of over 99% of the target exome, the exomes of all G1 founders of the pedigree can be sequenced. Identified mutant types (e.g., with respect to the C57BL / 6J reference genome) have their genotypes determined in G2 and G3 mice prior to phenotypic screening. Using various phenotypic screens, G3 mice can be tested for phenotypic variance compared to C57BL / 6J mice or a control population of G3 mice. Proof of linkage between a mutant phenotype detected in the screen and a particular mutation is achieved by automated meiotic mapping (AMM) performed by a linkage analyzer algorithm (or program or software), which tests the null hypothesis for all mutations within the pedigree (e.g., "Mutation A is unrelated to the phenotypic performance in screen α"). In contrast, mutations that are associated with a mutant phenotype at a higher frequency than predicted by chance alone are likely to confer a phenotype. With Bonferroni correction for multiple comparisons, a causal relationship can be considered to be suggested if the p-value is less than or equal to 0.05 and the null hypothesis is rejected. Verification by independently generated alleles can be used to confirm the association.
[0022] From the experience of the association of thousands of mutations and phenotypes identified by AMM and verified or excluded by testing CRISPR / Cas9 target alleles, it has been shown that the p-value determined by AMM is not the only indicator of causality. The p-value is a measure of the probability that the observed difference occurred by chance alone. The lower the p-value, the greater the statistical significance of the observed difference.
[0023] Mutations linked to phenotypes with a p-value less than 0.05 may not be the causative mutations. Many other factors, such as the nature of the mutation (benign, harmful, null), the essentiality of the gene for pre-weaning survival, the size of the pedigree, the number of homozygotes tested, the magnitude of the phenotypic effect, the data dispersion characteristics of the screening of the problem, the number of different phenotypes caused by the mutation, the presence or absence of co-segregating mutations, and the observation of other alleles with similar effects, affect the correct selection of the true causative mutation. The CE system described herein can estimate the likelihood of validating the putative mutation and phenotype associations suggested by AMM.
[0024] Immune cell populations, specifically B cells, T cells, conventional and plasmacytoid dendritic cells (DCs), macrophages, neutrophils, natural killer (NK) cells, and NK1.1 + Changes in T cells can be analyzed. Flow cytometric analysis of peripheral blood leukocytes from G3 mutant mice with ENU-induced mutations can detect and measure cell populations and subpopulations. In some cases, the CE system is used to evaluate the association of 87,795 mutations and phenotypes (e.g., having P < 0.05), from which the CE system identified over 1,270 genes highly defined as having a probability of having a verifiable importance in leukocyte development or maintenance. Many of these genes were not previously known to be important in immune function.
[0025] The CE system may be useful when researchers predict whether a mutation associated with a phenotype by AMM is actually the causative mutation. The CE system evaluates the association between mutations and phenotypes that pass a specific basic filter for conventional good candidates. The default filter for the data that can be used includes a p-value of less than 0.05 (Bonferroni corrected), more than 10 mice in the tested pedigree, and more than 2 screened homozygous reference mice, but the user can set more stringent criteria. The core of the CE system is a supervised machine learning algorithm that outputs a numerical score (CE score), a categorical evaluation (candidate status), and a verification probability for each association between a mutation and a phenotype based on the input phenotype data (e.g., from screening), mutation data, gene data, and meiotic mapping data.
[0026] Referring to FIG. 1, the processor 103 and / or the memory 115 may be used to implement the CE system. For example, the processor 103 may include a circuit 120 for receiving one or more input features (e.g., receiving at least one of a phenotype feature, a linkage data feature, a mutation feature, a gene feature, or an algorithm score). The processor 103 may also include a circuit 122 for generating a CE score based on one or more input features. For example, the circuit 122 may be a trained machine learning model. The machine learning model may be trained based on the phenotypic evaluation of mice having a target null allele or substitution allele of a candidate gene. The processor 103 may also include a circuit 124 for outputting an indicator of the association between the phenotype and the mutation based on the CE score.
[0027] Memory 115 may be coupled to processor 103 and, when executed by processor 103, may store code for performing the operations described herein. For example, memory 115 may include code 130 for receiving one or more input features, code 132 for generating a CE score, or code 134 for outputting an indicator of relatedness.
[0028] As described, the CE system may include a machine learning model. The machine learning model may be trained using an objective function. For example, candidate solutions may be provided to the model and evaluated against a training data set. An error score (also referred to as the loss of the model) may be calculated by comparing the solution to the training data set. The machine learning model may be trained to minimize the error score. For example, the machine learning model may be trained to implement a CE system that includes memory 115 and processor 103. The CE system may be trained based on the phenotypic evaluation of mice having a target null allele or substitution allele of a candidate gene. Due to the dynamic status of the database, during predictions run four times a day, CE uses all defined features of the original pedigree screening data to estimate the probability of candidate validation. CE may be used to query the association between mutations identified in flow cytometry screens and phenotypic associations in bone radiation screens (dual energy X-ray absorptiometry (DEXA) scans).
[0029] Figure 2 shows the input and output functions of a CE system according to a particular aspect of the concepts of the present invention. The CE machine learning system 212 can be a supervised machine learning algorithm that outputs, based on various input features, a numerical score (e.g., CE score 214), a categorical evaluation (e.g., candidate status 218), and a verification probability 216 for the relevance of each mutation to a phenotype. The input features can include input phenotype data (e.g., phenotype feature 202), mutation data (e.g., mutation feature 206), gene data (e.g., gene feature 208), meiotic mapping data (e.g., linkage data feature 204), and algorithm score 210. The mutation feature can include a damage score indicating the likelihood that the protein associated with the mutation is functionally impaired. The damage score can be generated using an ML system 230 that may use known harmful and neutral mutations, as described in more detail herein. The gene feature 208 can include an essentiality score indicating the likelihood of pre-weaning lethality in homozygous mice for a robust knockout allele of the gene associated with the mutation. The essentiality score can be generated using an ML system 240 that can be trained using genes known to be non-essential for survival and genes known to be essential for survival. The algorithm score can be a score generated based on a set of rules 260 associated with empirical observations, as described in more detail herein. The meiotic mapping data can be generated using automated meiotic mapping (AMM), as described herein. The generated CE score 214 can be used to determine the verification probability 216. The CE score, along with the algorithm score, can be used to generate the candidate status 218 (e.g., whether the relevance of the mutation to the phenotype is an excellent candidate, a good candidate, a potential candidate, or not a good candidate).
[0030] As described herein, a CE system can be trained using a CE training set. A CE training set (e.g., used to train a machine learning model of CE system 212) can include verified (e.g., 1,903 verified) and excluded (e.g., 3,013 excluded) mutation-phenotype associations (a total of 4,916 evaluations) based on germline retargeting of genes (e.g., 514 genes). To generate knockout alleles of candidate genes in mice of a pure reference background (C57BL / 6J or C57BL / 6N), germline retargeting can be performed using CRISPR / Cas9. Alternatively, if there is evidence of homozygous lethality of the null allele (e.g., using the essentiality scores described herein), or if an N-ethyl-N-nitrosourea (ENU) mutation is suspected of causing a hypermorphic, neomorphic, or antimorphic effect, the original ENU allele can be recreated by CRISPR / Cas9 targeting (designated as a "substitution" allele). Mice having a targeted germline knockout or substitution allele can be expanded to form a pedigree including homozygous mice for the reference allele (REF), heterozygous (HET) mice, and homozygous mice for the mutant allele (VAR). Compound heterozygous mice having two or more mutant alleles of a gene can be generated. Fresh pedigrees of mice having CRISPR-targeted alleles can be subjected to a phenotypic screening in which the original ENU mutation was recorded as a hit. In some embodiments, a CRISPR-targeted mutation can be considered verified according to criteria including: (1) the same phenotype with the same direction of change as observed for the original ENU allele is observed with a p-value better than 0.01; (2) the same phenotype with the opposite direction of change as observed for the original ENU allele is observed with a p-value better than 0.001; or (3) a phenotype (e.g., not seen in the original screening) is newly observed with a p-value better than 0.001.
[0031] Figure 3 is a graph 300 showing a polynomial regression analysis of the CE score and the average percentage of the verified mutation-phenotype association according to a particular aspect of the concept of the present invention. Each data point represents a group of mutation-phenotype associations. The percentage of verified associations (e.g., on the y-axis of graph 300) is plotted against the CE score range (e.g., on the x-axis of graph 300) in bins of 0.01 (e.g., from 0.35 to 0.36, from 0.37 to 0.38, etc.), with n = 4,916 mutation-phenotype associations and 514 CRISPR / Cas9 target genes. The CE score (e.g., in the range from 0 to 1) is a class probability associated by a polynomial function with the actual probability of verification by the CRISPR-target allele as determined by regression analysis. This is used, in combination with the algorithm score, by the CE system to assign one of four candidate statuses (excellent, good, potential, or not good) for each mutation-phenotype association. In some aspects, an excellent candidate corresponds to a CE score ≧ 0.39 and an algorithm score ≧ -0.5, a good candidate corresponds to a CE score ≧ 0.39 and -4.5 ≦ algorithm score < -0.5, a potential candidate corresponds to a CE score ≧ 0.39 and an algorithm score < -4.5 or a CE score < 0.39 and an algorithm score ≧ -0.5, and a not good candidate corresponds to a CE score < 0.39 and an algorithm score < -0.5.
[0032] In some aspects, good or excellent candidates may be selected for CRISPR / Cas9 targeting and further study. However, as shown in Figure 3, the CE score is not strictly proportional to the probability of verification, and some "good" or "excellent" candidates may fail verification. Conversely, "potential" and "not good" candidates may also be verified as true positive associations. As more alleles are obtained and tested (e.g., as approaching saturation), true candidates may achieve a high CE score and thus may ultimately be verified.
[0033] Figure 4 is a graph 400 showing the receiver operating characteristic (ROC) curve of the CE score according to a particular aspect of the concept of the present invention. The performance of the CE prediction model established using the training set can be evaluated using the repeated 10-fold cross-validation method. The area under the curve (AUC) of the ROC curve is 0.943, and the cut-off may be set at 0.39, which corresponds to the point where the distance to the upper left corner of the ROC curve is minimized.
[0034] Figure 5A is a table 500 showing the CE performance of the flow cytometry phenotype according to a particular aspect of the concept of the present invention. As shown, when the CE ranking is good or better, the accuracy is about 80% (e.g., correctly calling the verified candidates "true" and the false detection rate is 20%), and the recall can correspond to 87% (e.g., true positive rate).
[0035] Figure 5B is a table 501 showing the CE performance in the scoring of co-localized mutations according to a particular aspect of the concept of the present invention. The CE system can identify which mutation is the cause when two or more mutations co-segregate (e.g., determined by being driven by software as described herein). Among such 961 cases, CE can on average identify 76.5% of the causative mutations as the top CE score, and as shown in Figure 5B, the performance generally improves as the number of co-segregated mutations decreases. As further training is performed, the CE performance continues to improve as the total amount of screening data increases (e.g., the number of alleles and the overall density of the allele sequences increase concomitantly).
[0036] Multiple alleles of a given gene may be subjected to a given phenotypic screening, and as a result, there may be multiple mutations associated with the same gene and phenotype. The association between each mutation and the phenotype may be independently assigned an allele verification probability (AVP) estimate for the mutation in question, extrapolated from the polynomial regression analysis of the CE score and the average percentage of the verified mutation-phenotype associations (e.g., as shown in FIG. 3). Further, a composite estimate (e.g., gene verification probability (GVP)) that one or more mutations (e.g., N mutations) within a gene can be verified as the cause of a phenotype is expressed as follows. GVP = 1 - (1 - AVP1)(1 - AVP2)(1 - AVP3)...(1 - AVP N ) The AVP of alleles that cause phenotypic changes in the same direction in a given screen is included in the calculation.
[0037] FIG. 6 is Table 600 showing an example of input features to the machine learning model of the CE system according to a particular aspect of the concept of the present invention. The CE prediction model can incorporate 67 features of the input data, including 34 phenotypic features (e.g., phenotypic feature 202 in FIG. 2), 20 linkage analysis features (e.g., linkage analysis feature 204 in FIG. 2), 9 mutation features (e.g., mutation feature 206 in FIG. 2), 2 gene features (e.g., gene feature 208 in FIG. 2), and 2 other features (e.g., algorithm score 210 in FIG. 2).
[0038] Table 600 shows examples of input functions. For example, phenotypic characteristics include the percentage of VAR mice whose screening results overlap with those of B6 mice, the percentage of VAR mice whose screening results overlap with those of REF mice, the difference between HET and VAR results, the direction of the results (whether the average of the VAR screening results is greater than or less than the average of the REF screening results), the difference between REF and VAR results, the number of female HET mice, the number of female REF mice, the number of male REF mice, the number of male HET mice, the number of male VAR mice, the number of female VAR mice, phenotypic identification (e.g., fluorescence-activated cell sorting (FACS) T cells), phenotypic group identity (e.g., FACS screen or bone screen), the number of outliers in REF mice, the number of outliers in HET mice, the number of outliers in VAR mice, the difference between REF and B6 results, the difference between REF and HET results, whether the variance of REF is large (e.g., whether it exceeds a threshold), whether the variance of HET is large (e.g., whether it exceeds a threshold), whether the variance of VAR is large (e.g., whether it exceeds a threshold), whether the average age of this mutant / phenotypic mouse is higher than the average age of all mice tested for this phenotype, whether the average age of VAR mice is younger than the average age of REF mice, the number of families with this gene / phenotype, the direction of the position superpedigree results of this mutant / phenotype, the number of significant single families in the significant position superpedigree of this mutant / phenotype (e.g., a significant family has a p-value of the association between the mutation and the phenotype < 0.(referring to the linkage analysis of pedigrees or super-pedigrees by AMM at 05), the number of pedigrees included in the significant position super-pedigree results of this mutation / phenotype, the direction of the gene super-pedigree results (null alleles) for this phenotype, the direction of the gene super-pedigree results (null + missense alleles) for this phenotype, whether there is a trimmed result corresponding to the untrimmed data if the trimmed result is raw data normalized against cell viability (e.g., only if the VAR result is greater than the REF result), how similar the VAR result is to the B6 result, how similar the HET result is to the B6 result, how similar the REF result is to the B6 result, or whether the REF and B6 results are different, which may include at least one of these. The characteristics of the linkage are the average number of executions of the linkage analyzer with p-value < 0.00005 for each allele of the gene, the number of phenotypes having significant selective gene super-pedigree results for this gene, the number of executions of the linkage analyzer with p-value < 0.00005 for this gene, the number of pedigrees within the selective gene super-pedigree, and whether the result is significant for this gene / phenotype, the number of pedigrees contributing to the significant gene super-pedigree result (null alleles), the number of pedigrees of the significant gene super-pedigree result (null alleles), the minimum p-value of a single linkage analyzer result for this mutation / phenotype, the proportion of body weight screening with p-value < 0.0001 for this mutation, the proportion of FACS screening with p-value < 0.0001 for this mutation, whether the gene super-pedigree result is significant for this phenotype (null + missense), whether the p-value is significant in both the raw assay and the normalized assay for this mutation / phenotype, whether the minimum p-value is for a recessive genetic model (not dominant or additive), whether this phenotype is driven by another mutation, the proportion of DSS screening with p-value < 0.0001 for this mutation, the number of FACS phenotypes with p-value < 0.0001 for this mutation, the number of Dejerine-Sottas syndrome (DSS) phenotypes with p-value < 0.0001 for this mutation, the p-value of this mutation < 0.The number of weight phenotypes that are 0001, whether the positional superfamily results are significant for this mutation / phenotype, whether the gene superfamily results are significant for this phenotype (null allele), or whether the gene superfamily results are significant for this phenotype (missense allele), may include at least one of them.
[0039] The mutation feature (e.g., mutation feature 206) may include at least one of the damage score of the mutation, the number of alleles the gene has, whether the mutation is autosomal, whether the mutation co-localizes with another mutation of this phenotype, whether the mutation co-localizes with a mutation verified for this phenotype, whether the mutation co-localizes with an excluded mutation of this phenotype, whether the mutation co-localizes with a mutation with a higher damage score, the number of splice mutant forms of the gene containing this mutation, or the ratio of the number of named mutations of this amino acid change to the number of incidental mutations. The gene feature may include at least one of the p-value of the lethal phenotype or the probability that the gene is an essential gene (e.g., based on the E-score calculated as described herein). Other features (e.g., algorithm score features) may include at least one of the number of phenotypes for which the algorithm score of this mutation is -0.5 or greater, or the algorithm score of this mutation / phenotype. Although various features are shown in Table 600, only a subset of features, such as the features shown in bold, may be used.
[0040] In some embodiments, the damage score and the essentiality score (E-score) are obtained from independent machine learning programs. The rule-based algorithm score is obtained from the computational execution of a fixed algorithm.
[0041] A damage score (e.g., in the range from 0 to 1), which is a mutation feature, has biologically significant relevance. The damage score indicates the likelihood that a protein is functionally impaired and is determined by a machine learning algorithm that integrates an independent prediction score (e.g., a 37 - score) from the human database of Nonsynonymous Functional Prediction (dbNSFP) and the probability of protein damage against the phenotypic variance caused by mouse mutations. The higher the score, the higher the likelihood that the mutation is harmful and thus the higher the likelihood that it is causative (although this is not always the case). The damage score prediction model can be implemented using a machine learning model trained based on known harmful mutations (e.g., 871 mutations) and known neutral mutations (e.g., 1,797 mutations). To test the performance of the established model, mutations with known effects (e.g., 666 mutations) may be used, which can generate an ROC curve with an AUC of 0.852. Harmful mutations refer to genetic changes that increase susceptibility or predisposition to a particular disease or disorder. Neutral mutations refer to mutations that are neither beneficial nor harmful to an organism's survival and reproductive ability.
[0042] The E score (e.g., in the range from 0 to 1) is a gene feature that indicates the likelihood of lethality before weaning age (e.g., 4 weeks after birth) in homozygous mice for a robust knockout allele of a gene. The E score is calculated using a machine learning algorithm that incorporates various independent features of the gene, including gene conservation, protein - protein interaction networks, expression stages, and the viability / proliferation ability of human cell lines in which the gene has been mutated. A machine learning model (also called an E score prediction model, for example) for generating the E score can be trained on lethal / survivable mutations. The E score prediction model can be trained at monthly intervals. The training dataset is determined based on the annotations of the Mouse Genome Informatics (MGI) database and the observed effects of CRISPR - targeted null mutations generated in C57BL / 6J mice, and may include known non - essential genes (E score = 0) (e.g., 3,538 non - essential genes) and known essential genes (E score = 1) (e.g., 2,070 essential genes). The cut - off value may be set greater than 0.5 for essential genes and less than 0.5 for non - essential genes, and is used to inform gene - targeting efforts where knockout alleles or replacement genes identical to the original ENU alleles are created for phenotypic validation. To test the performance of the established model, genes known to have an impact on viability (e.g., 1,041 genes) may be used, resulting in the generation of an ROC curve with an AUC of 0.894.
[0043] The assessment of the association between a mutation and a phenotype can be performed using a human - developed algorithm that outputs a point - based score called an algorithm score (e.g., having a range from - 13.5 to 3.5). The algorithm score appears twice among the important features contributing to the CE algorithm and provides an overall assessment of the likelihood that a mutation is causative.
[0044] FIG. 7 is a table 700 showing rules for algorithm score determination according to a particular aspect of the concept of the present invention. The algorithm includes a set of rules based on empirical observations. The score of the algorithm increases or decreases for each feature that either supports or opposes the credibility of the association between the mutation and the phenotype. The features used in the score calculation by the algorithm are similar to those used in the CE machine learning algorithm, but are static (e.g., not affected by exposure to new training data, etc.), and the performance of the rule-based algorithm itself does not reach that of the CE prediction model. The association between each mutation and the phenotype starts from an algorithm score 0 that is adjusted according to the rules described herein with respect to FIG. 7.
[0045] FIG. 8 is a graph 800 showing the ROC curve 802 of the algorithm score. As shown, the AUC of the ROC curve 802 is 0.733, which is lower than the performance of the CE prediction model with an AUC of 0.943. Other input features to the CE algorithm (e.g., the linked data feature 204) can be generated by an algorithm called a driven by algorithm that evaluates linked candidate mutations and unlinked candidate mutations to determine the optimal candidate. Since clusters of linked mutations may fail to segregate during meiosis, multiple mutations may be candidates for the cause of the phenotype. In other cases, by chance coincidence, a homozygote of an unlinked mutation that is not causal may also be a homozygote of the causal mutation. Usually, this occurs when the number of homozygotes of non-causal mutations is small. The driving algorithm omits all instances of shared homozygosity of both mutations, recalculates the p-value testing for deviation from the null hypothesis in recessive, additive, and dominant transmission models, and determines which mutation is a more robust candidate cause. This mutation is assigned a "driver" status. Based on the driver status and other factors (e.g., which mutation is the most deleterious, which mutation is the most important for survival to weaning age, and which mutation has evidence of other alleles with a similar phenotype), CE may be able to identify the causal mutation from a set of co-localized mutations and assign a significantly better CE score.
[0046] Finally, the allelic series investigated in phenotypic screening provide important clues to causality and are considered in CE assessment. When multiple alleles of the same gene are associated with the same phenotype, this strongly suggests that mutations in this gene caused the observed phenotype. There are three types of superfamilies (complexes of multiple families analyzed in the same screen). Gene superfamilies pool different alleles of a given gene that have undergone the same screening. Positional superfamilies pool only identical alleles. Identical alleles can arise from 1) accidental mutations of the same nucleotide, 2) transmission of a single mutation to multiple G1 progeny of a single G0 mouse, and 3) background mutations present in a mutagenized stock and shared by multiple G0 mice. Selective gene superfamilies show an intentionally biased view of the mutation effect, as they incorporate only alleles with a common direction of effect and a p-value < 0.05 in a given phenotypic screening. Since many (but not all) ENU-induced mutations are functionally hypomorphic, selective gene superfamilies of sets of mutations in a particular gene suggest that the gene may be strongly involved in the phenotypes investigated by the screening in question. The number of families (and alleles) tested is also important. For very large genes, hundreds of alleles may be tested, and the result that two or three alleles score in a particular screening may be due to chance. The CE system takes this into account when calculating the probability of causality.
[0047] Figure 9 is Table 900 showing flow cytometric screening parameters according to a particular aspect of the concepts of the present invention. The flow cytometric screen examines 42 parameters of peripheral blood cells and measures the frequencies of various immune cell populations and the expression levels of several cell surface markers, as illustrated. Of the 7,109,669 mutation-phenotype associations tested by AMM in the flow cytometric screen, 87,795 passed the default initial filter and became available for analysis by CE. These putative mutation-phenotype associations arose from 39,685 mutations in 14,809 genes present in 142,653 G3 mice of 3,987 pedigrees. By restricting to good or excellent candidates, the number of mutation-phenotype associations decreased to 7,676, arising from 2,336 mutations in 1,279 genes present in 1,634 pedigrees.
[0048] Figures 10A, 10B, and 10C show the gene-phenotype association characteristics of genes having at least one good / excellent mutation-phenotype association according to a particular aspect of the concepts of the present invention, according to a particular aspect of the concepts of the present invention. Figure 10A is graph 1000 showing the number of good / excellent phenotype associations plotted against the number of genes. Figure 10B shows the number of good / excellent gene associations plotted against flow cytometric parameters. Figure 10C shows the number and percentage of essential and non-essential genes.
[0049] Various observations may have been made regarding the association between genes and phenotypes. First, mutations in the majority (872 genes, 68.2%) of 1,279 genes may result in associations with three or fewer favorable / excellent phenotypes. As shown in Figure 10A, 533 genes (41.7%) have an association with a single favorable / excellent phenotype. In contrast, only 30 genes (2.3%) have associations with at least 20 favorable / excellent phenotypes, 26 of which are well-known immunomodulatory genes. Second, the number of favorable / excellent gene associations may vary greatly depending on the type of cells affected. As shown in Figure 10B, B cell and T cell phenotypes are associated with the most genes, while conventional and plasmacytoid DC phenotypes are associated with very few genes. Finally, 449 genes (35.1%) that are known or predicted to be essential for survival (in this case, E-score > 0.55) may be associated with at least one flow cytometry phenotype, indicating that, as shown in Figure 10C, many developmentally important genes are likely to also have postnatal functions in leukocytes.
[0050] A total of 1,354 mutations in 667 genes that are evaluated as favorable / excellent by CE and suspected or proven to be the cause of the flow cytometry phenotype can be allele-named and annotated as phenotypic mutations in the Mutagenetix database, regardless of their current candidate status. Named alleles are likely to be causal, but it is unclear whether unnamed alleles are also non-causal. In fact, 27% of the named alleles have an AVP ≤ 0.5. Among the unnamed alleles, some are designated as "linked to" or "driven by" another mutation within the same pedigree. This may indicate that they are not causal, but it does not always guarantee this, and in some cases, it suggests that two named alleles may be linked and both mutations may be declared causal (e.g., even if they may co-segregate). Evidence of such double causation can be presented by CRISPR / Cas9 targeting.
[0051] Highly expressed gene ontology (GO) annotations can be identified that are associated with 667 genes with named alleles. Annotations of biological processes are most highly enriched for terms related to immune system processes (211 genes, P = 9.82e-42), lymphocyte activation (113 genes, P = 5.21e-39), immune system development (117 genes, P = 9.73e-36), and other immune development / regulation processes, which is consistent with a manual evaluation that identified 281 (42.1%) of the 667 genes as previously known immune regulators. By manual evaluation, 386 genes represent "new" immunologically important genes, each of which is used in a normal flow cytometry profile. For many of these genes, mutant alleles may not have been available in mice to date, and primary immunological data or other phenotypic data are not available. This may be due in part to the known or predicted lethality (E-score > 0.5) caused by null alleles of 146 of these 386 genes. Enriched GO terms associated with the 386 new immunologically important genes are occupied by metabolic process terms, including cellular metabolic process (232 genes, P = 3.73e-12), organic substance metabolic process (240 genes, P = 8.50e-12), cellular macromolecule metabolic process (178 genes, P = 3.59e-7), and protein metabolic process (127 genes, P = 0.000264). The 386 genes may be assigned to a defined set of broad GO annotations for biological processes regardless of enrichment. Based on fine-grained GO annotations, each gene may be assigned to one of 70 related parent GO terms.In particular, 31 out of 386 genes may be associated with the term "immune system process" based on genetic interactions, association with the immune system of ancestral genes, sequence homology with other genes associated with immune system processes, or association with immune system processes of orthologous human genes. Further, 300 out of 386 genes have been detected to be expressed at moderate levels (11 to 1,000 transcripts per million) or high levels (more than 1,000 transcripts per million) in the spleen and / or thymus by RNA sequencing.
[0052] Currently, among the genes related to flow cytometry phenotypes, there are a total of 603 genes with GVP > 0.5, 332 genes with GVP > 0.8, and 222 genes with GVP > 0.95. Among the genes with GVP > 0.95 where the flow cytometry phenotype is almost certain to occur, 121 (55%) are known to affect flow cytometry measurements and 101 (45%) are novel.
[0053] Using the CE system enables rapid examination of mutations and genes that are strongly predicted to affect (or not affect) the phenotypes of interest measured in forward genetics screening. Generally, CE has the ability to integrate parameters that are intuitively unfavorable or not harmful for linkage analysis, and can perform this evaluation more rapidly and on a larger scale, thus providing advantages to human researchers when assessing the association between mutations and phenotypes. Using the numerical CE scores and categorical evaluations provided by the CE system, mutations can be easily ranked in a priority list for further detailed study. Furthermore, the causative mutations can often be identified among multiple co-localized mutations. As millions of coding / splicing mutations are introduced into the mouse genome for each family, a wider range of allelic series may occur, and it may be possible to identify with high confidence almost all genes that may harbor causative loss-of-function mutations.
[0054] The CE system is used not only as a tool for rapidly identifying mutations that cause ENU-induced phenotypes, but is also useful to mouse geneticists studying complex traits (e.g., Collaborative Cross). In meiotic mapping, phenotypes can be limited to relatively broad genomic intervals that contain multiple candidate genes with mutational differences. When the phenotype is immunological, knowledge of all genes in which the flow cytometry phenotype occurs is an important starting point for causal studies that can target these genes.
[0055] CE is also valuable to clinical geneticists attempting to identify the causes of human diseases. In cases where patients have immunopathological and flow cytometry abnormalities but no mutations in “classical” causative genes, CE can be used to evaluate mutant forms of other genes. Mouse gene symbols corresponding to all loci mutated within a patient (e.g., those identified by whole genome or whole exome sequencing) can be entered into CE and searched en masse. Those that cause flow cytometry abnormalities in mice and evoke the abnormalities in patients can be considered prime candidates. When gene mapping has been performed in a human family and a particular chromosomal region has been identified, using CE that also accepts human chromosomal coordinates as search inputs can identify candidate genes with a higher degree of confidence. By using CE in combination with the analysis of large-scale human genome / phenotype datasets, CE can also facilitate and accelerate the identification of causative mutant forms within disease-associated loci discovered by genome-wide association studies (GWAS). CE can query the association between related mutations and phenotypes for each candidate gene within loci identified by GWAS, and mutant forms of genes in mice associated with phenotypes similar to those of the humans under study suggest a causal relationship. Furthermore, in most cases, mutant mice can be ordered immediately and can provide models of human diseases for laboratory studies. Most mutations cause loss of function (e.g., not gain of function or new functions), and since most mouse genes have human orthologs, i.e., homologs, many such cases may be resolved immediately. Thus, CE is a powerful resource for addressing the problem of “missing heritability” associated with immune abnormalities, and genes that control or mediate cellular metabolic processes can be prime candidates for consideration, as has been pointed out for 386 new genes.
[0056] The association between mutations and phenotypes, represented by genes having one or more mutant alleles, and flow cytometric parameters of peripheral blood leukocytes can be evaluated. By flow cytometric analysis, immune cell populations with specific functional correlation can be detected and measured, and insights into the developmental stages through which cells pass can be obtained. Abnormal flow cytometric patterns are often associated with immune dysfunction, and many immunodeficiency and autoimmune phenotypes can be initially detected by analysis of peripheral blood by flow cytometry rather than by functional screening itself. Human disease states showing similar or identical flow cytometric phenotypes have demonstrated many clinical relevance of flow cytometric abnormalities in mice. To date, approximately 5% genomic saturation has been achieved in the screening of 42 flow cytometric parameters, from which 1,004 genes with good / excellent phenotypic associations that have not been previously associated with immune function have been identified (from GO analysis of 1,279 genes, 275 were found to be associated with "immune system processes"). Therefore, even with a maximum false discovery rate of 20%, approximately 456 more immunologically important new genes will be discovered.
[0057] When broadly investigating all 1,279 genes with a relationship to at least one favorable / excellent phenotype, it can be observed that the proportion of genes with a relationship to one, two, or three favorable / excellent phenotypes (68.2%) is much higher compared to the proportion of genes with a relationship to a large number (≥20) of favorable / excellent phenotypes (2.3%). These findings suggest that most of the genes affecting immune cell populations in the blood are performing functions specific to cell types or phenotypes. The hypothesis that the same or similar combinations of phenotypes affected by two or more genes may indicate that those genes are functioning in a common molecular pathway is investigated. Despite performing uniform phenotype tests across all screens, favorable / excellent gene associations may not affect cell populations with equal frequency. For example, T cells may have a 4.8-fold higher likelihood of gene-related associations compared to conventional dendritic cells (DCs), 12.5-fold higher compared to plasmacytoid DCs, and 4.4-fold higher compared to neutrophils. A simple explanation is that significant phenotypic differences are detected less frequently in rarer blood cell populations, but another possibility reflecting cell biology is that, with respect to the number of these cells present at least in peripheral blood, T cells may be inherently less resistant to genetic mutations compared to conventional DCs, plasmacytoid DCs, or neutrophils. To gain insights into these issues, it is important to understand the functions of individual proteins and the pathways they regulate.
[0058] Most of the mice whose phenotypes were analyzed by flow cytometry have their phenotypes also analyzed in other screens, such as screens measuring response to immunization, innate immune response, body weight, blood pressure, heart rate, dextran sulfate sodium (DSS) sensitivity, circadian rhythm, and motor coordination. Data from the screening of skeletal phenotypes detected by DEXA scan are currently publicly accessible. In the future, data obtained from other screens may be made available to the general users of CE to interpret the wide range of phenotypic results arising from each mutation. All biomedically relevant phenotypic screens ultimately shed light on human phenotypic studies and help distinguish the mechanisms of phenotypes caused by specific alleles, as many mutations are scored in heterologous screens (e.g., immune function and body weight, or immune function and neurobehavioral function).
[0059] Further details regarding materials and methods that can be used to identify relevance are shown below. C57BL / 6 male mice at 8 to 10 weeks of age may be mutagenized by ENU. As described above, the mutagenized G0 males can be mated with C57BL / 6J females, and the resulting G1 males can be mated with C57BL / 6J females to produce G2 mice. When G2 females are backcrossed to their G1 males, G3 mice are born and their phenotypes are screened. Whole exome sequencing and mapping can be performed.
[0060] To generate mice with CRISPR / Cas9 target mutations, superovulation can be induced by injecting 6.5 units of pregnant mare serum gonadotropin (PMSG, Millipore) into female C57BL / 6J mice, followed by injection of 6.5 units of human chorionic gonadotropin (hCG, Sigma-Aldrich) 48 hours later. Superovulated mice were then mated overnight with C57BL / 6J male mice. The next day, fertilized eggs were collected from the oviducts, and in vitro-transcribed Cas9 mRNA (50 ng / μL) and small base-pairing guide RNA (50 ng / μL) were injected into the cytoplasm or pronuclei of the embryos. The injected embryos were cultured in M16 medium (Sigma-Aldrich) at 37 °C and 5% CO2. To generate mutant mice, two-cell stage embryos can be transferred to the ampulla of the oviducts of pseudopregnant Hsd:ICR (CD-1) (Harlan Laboratories) females (10 to 20 embryos per oviduct).
[0061] Peripheral blood can be collected from G3 mice over 6 weeks old by bleeding from the cheek. Red blood cells (RBCs) can be lysed using hypotonic buffer (eBioscience). The sample can be washed once using FACS staining buffer (phosphate buffered saline (PBS) supplemented with 1% [weight / volume] bulk separation analysis (BSA)), and then centrifuged at 500×g for 5 minutes. The RBC-depleted sample can be stained for 1 hour at 4°C in 100 μL of a 1:200 mixture of fluorescently labeled antibodies against 15 cell surface markers including major immune lineages B220 (BD, clone RA3-6B2), CD19 (BD, clone 1D3), IgM (BD, clone R6-60.2), IgD (BioLegend, clone 11-26c.2a), CD3ε (BD, clone 145-2C11), CD4 (BD, clone RM4-5), CD8α (BioLegend, clone 53-6.7), CD11b (BioLegend, clone M1 / 70), CD11c (BD, clone HL3), F4 / 80 (Tonbo, clone BM8.1), CD44 (BD, clone 1M7), CD62L (Tonbo, clone MEL-14), CD5 (BD, clone 53-7.3), CD43 (BD, clone S7), NK1.1 (BioLegend, clone OK136), and 1:200 Fc block (Tonbo, clone 2.4G2). Flow cytometry data can be collected using a cell analyzer (e.g., BD LSR Fortessa), and the proportion of immune cell populations in each G3 mouse can be analyzed using software for analyzing cytometry data. The resulting phenotypic data can be uploaded to a server (e.g., Mutagenetix) for automated mapping of the causative alleles.
[0062] The AMM can be performed as described herein. For example, the genotypes at all the mutant sites present in the exome of the G3 mouse can be determined prior to phenotypic screening. The DNA of the tail of the G1 male may be subjected to whole exome sequencing using a sequencing instrument (e.g., Illumina HiSEq.2500). Then, the genotypes of the G2 and G3 mice are determined at the identified mutant sites (e.g., using Ion PGM (Life Technologies)). After phenotypic screening, using a linkage analyzer program, linkage analysis can be performed for all the mutations within the pedigree using recessive, additive, and dominant genetic models. Also, using a linkage explorer program, scatter plots and Manhattan plots of the phenotypic data can be displayed. The p-value of the association between the genotype and the phenotype can be calculated by applying a likelihood ratio test and Bonferroni correction from a generalized linear model or a generalized linear mixed effect model.
[0063] In some embodiments, the CE prediction model can be constructed using a random forest algorithm (e.g., implemented in the R classification and regression training (caret) package). The linkage data obtained through screening can be released stepwise according to the phenotype.
[0064] The damage score is an ensemble score that uses a logistic regression model to integrate independent prediction scores (e.g., 38 scores). Thirty-seven of the prediction scores may be obtained from databases (e.g., human dbNSFP), including scores from the algorithms of SIFT, SIFT4G, Polyphen2-HDIV, Polyphen2-HVAR, LRT, MutationTaster2, MutationAssessor, FATHMM, MetaSVM, MetaLR, CADD, CADD_hg19, VEST4, PROVEAN, FATHMM-MKL coding, FATHMM-XF coding, fitCons (4 scores), LINSIGHT, DANN, GenoCanyon, Eigen, Eigen-PC, M-CAP, REVEL, MutPred, MVP, MPC, PrimateAI, GEOGEN2, BayesDel_addAF, BayesDel_noAF, ClinPred, LIST-S2, and ALoFT. The dataset may use the ranked scores of each algorithm converted by dbNSFP. The 38th prediction score is the probability of protein damage to the phenotypic variance caused by mouse mutations and is calculated as described herein. Among the 38 prediction scores, the score of MutPred is important together with the probability of protein damage to the phenotypic variance caused by mouse mutations and phastCons100way_vertebrate (conservation score). The damage score can be used as a quantitative prediction score to measure the likelihood that a mouse mutation is harmful.
[0065] If a mouse missense mutation is the same as a human mutation (e.g., changes in both nucleotides and amino acids), the effects of the mutation in humans and mice may be similar. Thus, human scores can be used to predict the likelihood of damage in mice. A set of mouse ENU mutations with cluster tags (e.g., known to be harmful or neutral) can be obtained from the Mutagenetix database. Known mutation cluster tags are obtained from the following four sources. 1) Physically isolated mutations that are within the range of essential genes and may be transmitted from heterozygous G2 females and their heterozygous G1 males to homozygous G3 mice at a rate that does not deviate significantly from Mendel's laws (linked to all other coding / splicing mutations within the family) are considered neutral. 2) Conversely, isolated mutations in important genes that are not transmitted to homozygotes are considered harmful to the extent that homozygotes are observed at a frequency significantly below the expected Mendelian ratio. 3) Mutations that cause a qualitative (usually visible) phenotype are considered harmful. 4) Mutations that have been verified to be significant in the phenotypic screening of CRISPR replacement alleles are also considered harmful.
[0066] Mutations tagged as harmful or neutral can be carried from the mouse genome into the human genome (translated into equivalent amino acids) and retained for mutations that result in the same nucleotide and amino acid changes in both genomes. Even using lift-over tools to convert genome coordinates and annotation files between assemblies, about 4% of mouse mutations may not map to corresponding human mutations. They may not be included in the final dataset for model training and testing. Point-biserial correlation can be used to estimate the relationship between mouse mutations tagged as harmful or neutral and the most important human mutation prediction scores. The correlation coefficient is 0.525, 95% CI: 0.50 to 0.55. Then, to obtain the scores of all available prediction methods, the corresponding human mutations can be searched in the dbNSFP database. The obtained scores are combined with the probability of phenotypically detectable damage caused by mouse mutations, integrated into the input dataset, and used to train and optimize a logistic regression model using the train function of the R caret package with 10-fold cross-validation. This process can be repeated multiple times (e.g., 3 times). Scaling of the data can be performed by a preprocessing function. Then, the constructed model (classifier) can be used to calculate the scores of a set of mutations whose class membership is unknown. The dataset used for prediction can be created in the same way as the dataset used for modeling. The scores predicted by the model represent the probability that a mutation falls into the harmful class. The higher the score, the higher the likelihood that the mutation is harmful.
[0067] The input dataset may contain mouse mutations (e.g., 3,334 mutations), some of which (e.g., 1,088) are harmful and some (e.g., 2,246) are neutral. To evaluate the performance of a model constructed when predicting membership in a new mutation category, the input dataset can be randomly split into two sets. One set consisting of mutations (e.g., 2,668 mutations, 80% of the original dataset, 871 harmful mutations, and 1,797 neutral mutations) may be used for training and validation of a logistic regression model, and a second set of the remaining 666 mutations may be used to test the performance of the established model. The 80 / 20 split for training and testing can be performed randomly multiple times (e.g., 10 times). The ROC curve may generate an AUC close to an average AUC value of 0.853 ± 0.014.
[0068] The E-score can be used to estimate the lethality probability of a mouse when a gene is knocked out. Essential and non-essential genes in mice can be distinguished by various independent features of the genes. The logistic regression method is used to fit the features of known essential and non-essential genes in mice to obtain a trained model for predicting the unknown essentiality of genes.
[0069] This model uses the following gene features: 1) From the Online Gene Essentiality (OGEE) database, gene conservation, connectivity in the protein - protein interaction network, expression stage during development, evolutionary age, GO terms, gene copy number, and gene product length. These features are associated with gene essentiality in many species including mice. 2) The essentiality of human orthologous genes, genes related to cell proliferation and viability in tested cell lines may be defined as important genes under specific conditions, and the frequency of importance in tested human cell lines may be used as a feature of the model. 3) The loss - of - function intolerance probability (pLI) score from the Exome Aggregation Consortium (ExAC), the closer the score is to 1, the higher the likelihood that the gene is essential for human survival. 4) The minimum p - value of ENU - targeted mouse genes obtained from lethal models by the linkage analyzer algorithm.
[0070] The phenotypic descriptions of 8,032 genes that may be knocked out in mice in Mouse Genome Informatics (MGI) are carefully considered, and the set of genes designated as "essential" or "non - essential" may be manually curated according to the following criteria. 1) If the homozygous knockout allele is explicitly described as causing fetal lethality, neonatal lethality, prenatal lethality, perinatal lethality, or pre - weaning lethality, the gene is considered necessary for survival before weaning and may be classified as an essential gene, and an E - score of 1 may be assigned to the gene. 2) If the homozygous knockout allele results in viability, normal growth, no obvious phenotype, or is compatible with some phenotype but has no obvious effect on survival rate, it may be classified as a non - essential gene, and an E - score of 0 may be assigned to the gene. Furthermore, an E - score of 1 may be assigned to genes verified to cause significant lethality before weaning in CRISPR knockout experiments, and an E - score of 0 may be assigned to genes verified to yield a normal Mendelian ratio in the mating of heterozygous mutants in CRISPR knockout experiments.
[0071] A set of 7,009 genes (2,587 of which are labeled as essential genes and 4,422 as non-essential genes) can be integrated with the described gene features. The resulting dataset can be used to train and optimize a logistic regression model using the train function of the R caret package with 10-fold cross-validation, and can be repeated multiple times (e.g., 3 times). Scaling of the data can be performed by a preprocessing function. The preProcess function estimates the parameters required at runtime. The constructed model can be used to predict the essentiality of the remaining mouse genes. The predicted scores are between 0 and 1. The closer the score is to 1, the higher the likelihood that the gene is essential.
[0072] To evaluate the performance of the model constructed when predicting the unknown essentiality of genes, the dataset used to construct the model can be randomly split into two sets. One set containing 5,608 genes (80% of the original dataset, 3,538 non-essential genes, and 2,070 essential genes) can be used for training and validation of the logistic regression model, and the remaining 1,401 genes can be used to test the performance of the model established in the training dataset. The 80 / 20 split of training and testing can be performed randomly multiple times (e.g., 10 times). The ROC curve generates an AUC close to an average AUC value of 0.891 ± 0.0087.
[0073] FIG. 11 is a flowchart showing an exemplary operation 1100 for mutagenesis according to a particular aspect of the concept of the present invention. Operation 1100 can be performed by a CE system such as, for example, processor 103 and memory 115.
[0074] In block 1102, the CE system may receive one or more input features including phenotypic data and mutation data. In block 1104, the CE system may generate a CE score indicating the probability of the relevance between a phenotype and a mutation based on one or more input features via a machine learning model.
[0075] In some embodiments, the mutation data may also include a damage score indicating the likelihood that the protein associated with the mutation is functionally impaired. For example, the CE system may generate a damage score via another machine learning model trained using known harmful and neutral mutations.
[0076] In some embodiments, one or more input features may also include an essentiality score indicating the likelihood of pre-weaning lethality in homozygous mice for a robust knockout allele of the gene associated with the mutation. The CE system may generate an essentiality score via another machine learning model trained using genes known to be non-essential for survival and genes known to be essential for survival. In some embodiments, one or more input features also include features associated with an algorithm score indicating the likelihood that the mutation is the cause.
[0077] In some embodiments, one or more input functions may also include linkage data generated using automated meiotic mapping (AMM) (e.g., as performed by a linkage analyzer algorithm or program). In one embodiment, if two or more mutations are co-segregated, the CE system may determine which of the two or more mutations is a more robust candidate cause of the phenotype by omitting instances of the shared homozygosity of the two or more mutations. In block 1104, a CE score may be generated based on that determination.
[0078] In some embodiments, one or more input features are the number of phenotypes having an algorithm score for a mutation that meets a threshold, where the algorithm score is the number of phenotypes indicating the likelihood of being caused by the mutation, the average number of automated meiotic mapping (AMM) operations that yield a p-value meeting the threshold for each allele of the gene associated with the mutation, the algorithm score of the mutation or phenotype, the number of AMM operations that yield a p-value meeting the threshold of the gene associated with the mutation, the damage score of the mutation, where the damage score indicates the likelihood that the protein associated with the mutation is functionally impaired, the number of families within the superfamily associated with the gene, and whether the p-value of the result of the AMM operation for the superfamily meets the threshold, the number of phenotypes having a p-value of the superfamily meeting the threshold, the number of families contributing to the p-value of the superfamily meeting the threshold, the number of families in the superfamily, the percentage of fluorescence-activated cell sorting (FACS) screens having a p-value meeting the threshold of the mutation, the minimum value of the p-value from the AMM operation, the percentage of mutant allele (VAR) mice whose screening results overlap with the screening results of B6 mice, whether the result of the AMM operation for the superfamily meets the threshold of the null allele and missense allele, whether the result of the AMM operation for the superfamily meets the threshold of the null allele, the percentage of VAR mice whose screening results overlap with the screening results of reference allele (REF) mice, the difference in the results of the AMM operation for heterozygous (HET) mice and VAR mice, the number of female REF mice used in the AMM operation, the percentage of body weight screens having a p-value meeting the threshold of the mutation, the number of female HET mice used in the AMM operation, or any one or any combination of the differences in the results of the AMM operation for REF mice and VAR mice may be included.
[0079] In block 1106, the processing system may output a CE score. For example, in some aspects, the processing system may generate a candidate status for the relevance between a phenotype and a mutation (e.g., an excellent candidate, a good candidate, a potential candidate, or a non - good candidate) based on the CE score and an algorithm score indicating the likelihood of being caused by a mutation.
[0080] These and various other configurations will be described in more detail herein. As will be understood by those of ordinary skill in the art upon reading the following disclosure, the various aspects described herein may be a method, a computer system, or a computer program product. Accordingly, these aspects can take the form of at least one implementation that is either a fully hardware implementation, a fully software implementation, or a combination of software and hardware aspects. Further, such aspects can take the form of a computer program product stored by one or more computer - readable storage media (e.g., non - transitory computer - readable media) having computer - readable program code or instructions contained within the storage media or on the storage media. Any suitable computer - readable storage media can be utilized, including hard disks, CD - ROMs, optical storage devices, magnetic storage devices, and / or any combination thereof. Additionally, the various signals representing the data or events described herein can be transferred between a source and a destination in the form of electromagnetic waves that travel through a signal - conducting medium such as metal wires, optical fibers, and / or wireless transmission media (e.g., air and / or space).
[0081] Implementations of the concepts of the present invention include the various steps described herein. These steps may be performed by hardware components or may be incorporated into machine-executable instructions, which may be used to cause a general-purpose or dedicated processor programmed with the instructions to perform the steps. Alternatively, these steps may be performed by a combination of hardware, software, and / or firmware.
[0082] Although specific implementations are described, it should be understood that this is for illustrative purposes only. Those skilled in the relevant art will recognize that other components and configurations may be used without departing from the spirit and scope of the concepts of the present invention. Accordingly, the following description and drawings are illustrative and should not be construed as limiting. Numerous specific details are set forth in order to provide a complete understanding of the concepts of the present invention. However, in some instances, well-known details or conventional details may not be described so as not to obscure the description. References to one implementation or an implementation in the concepts of the present invention may refer to the same implementation or any implementation, and such references mean at least one of the implementations.
[0083] References to "one implementation" or "an implementation" mean that the particular features, structures, or characteristics described in connection with the implementation are included in at least one implementation of the concepts of the present invention. When the phrase "in one implementation" is used in various places in this specification, it does not necessarily refer to the same implementation, nor does it refer to a separate or alternative implementation that is mutually exclusive with other implementations. Furthermore, various features are described that may be realized in some implementations and not in others.
[0084] The terms used in this specification generally have their ordinary meanings in the context of the concepts of the present invention and in the particular context in which each term is used. For any one or more of the terms described in this specification, alternative languages and synonyms may be used, and there is no special meaning whether the term is described in detail or explained in this specification. In some cases, synonyms for certain terms are provided. The listing of one or more synonyms does not exclude the use of other synonyms. Anywhere in this specification, including examples of any terms discussed herein, the use of examples is for illustrative purposes only and is not intended to further limit the scope and meaning of the concepts of the present invention or any of the illustrative terms. Similarly, the concepts of the present invention are not limited to the various implementations shown in this specification.
[0085] Although it is not intended to limit the scope of the concepts of the present invention, examples of instruments, devices, methods, and the results related thereto according to the implementations of the concepts of the present invention are shown below. Note that the title or subtitle may be used in the examples for the convenience of the reader, but it does not limit the scope of the concepts of the present invention. Unless otherwise defined, the technical and scientific terms used in this specification have the meanings generally understood by those skilled in the technical field related to the concepts of the present invention. In case of conflict, this specification including the definitions shall prevail.
[0086] Further features and advantages of the concepts of the present invention are described in the following description, some of which are apparent from the description, or can be learned by practicing the principles disclosed in this specification. The features and advantages of the concepts of the present invention can be realized and obtained by the means and combinations particularly pointed out in the appended claims. These and other features of the concepts of the present invention will become more fully apparent from the following description and the appended claims, or can be learned by practicing the principles described in this specification.
Claims
1. A method for mutation processing, comprising: receiving one or more input features including phenotypic data and mutation data; generating, via a machine learning model, a candidate explorer (CE) score indicating the probability of the relevance between a phenotype and a mutation based on the one or more input features; outputting an index of the relevance between the phenotype and the mutation based on the CE score; A method comprising the steps of.
2. The method according to claim 1, wherein the index of the relevance includes a candidate status of the relevance based on the CE score and an algorithm score indicating the possibility of being caused by the mutation.
3. The method according to claim 1, wherein the mutation data includes a damage score indicating the possibility that a protein associated with the mutation is functionally impaired.
4. The method according to claim 3, further comprising generating the damage score via another machine learning model trained using known harmful and neutral mutations.
5. The method according to claim 1, wherein the one or more input features further include an essentiality score indicating the possibility of pre-weaning lethality in homozygous mice for a robust knockout allele of a gene associated with the mutation.
6. The method according to claim 5, further comprising generating the essentiality score via another machine learning model trained using genes known to be non-essential for survival and genes known to be essential for survival.
7. The method according to claim 1, wherein the one or more input features further include features associated with an algorithm score indicating the possibility of being caused by the mutation.
8. The method according to claim 1, wherein the one or more input features further include linkage data generated using automated meiotic mapping (AMM).
9. When two or more mutations are co-segregated, determining which of the two or more mutations is a more robust candidate cause of the phenotype by omitting instances of the co-homozygosity of the two or more mutations, and generating the CE score based on the determination. The method according to claim 1, further comprising the steps of.
10. The one or more input features are The number of phenotypes having an algorithm score for the mutation that meets a threshold, where the algorithm score indicates the likelihood that the mutation is the cause. The average number of auto-meiotic mapping (AMM) operations that yield a p-value that meets the threshold for each allele of the gene associated with the mutation. The algorithm score for the mutation or phenotype. The number of AMM operations that yield a p-value that meets the threshold for the gene associated with the mutation. The damage score for the mutation, where the damage score indicates the likelihood that the protein associated with the mutation is functionally impaired. The number of families within a superfamily associated with the gene, and whether the p-value of the result of the AMM operation for the superfamily meets a threshold. The number of phenotypes having a p-value for the superfamily that meets a threshold. The number of families that contribute to the p-value of the superfamily that meets a threshold. The number of families in the superfamily. The proportion of fluorescence-activated cell sorting (FACS) screens having a p-value that meets the threshold for the mutation. The minimum value of the p-value from the AMM operations. The proportion of mutant allele (VAR) mice whose screening results overlap with the screening results of B6 mice. Whether the result of the AMM operation for the superfamily meets the thresholds for null alleles and missense alleles. Whether the result of the AMM operation for the superfamily meets the threshold for null alleles. The proportion of VAR mice whose screening results overlap with the screening results of reference allele (REF) mice. The difference in the results of AMM operations for heterozygous (HET) mice and VAR mice. The number of female REF mice used in the AMM operations. The proportion of body weight screens having a p-value that meets the threshold for the mutation. The number of female HET mice used in the AMM operations, and The difference in the results of AMM operations for REF mice and VAR mice. The method according to claim 1, comprising at least one of the above.
11. An apparatus for mutation processing, comprising a memory, and one or more processors coupled to the memory, which receive one or more input features including phenotype data and mutation data. Generating a candidate explorer (CE) score indicating the probability of an association between a phenotype and a mutation based on the one or more input features via a machine learning model, and Outputting an indicator of the association between the phenotype and the mutation based on the CE score, One or more processors configured to perform the above, and An apparatus comprising the above.
12. The apparatus according to claim 11, wherein the indicator of the association includes a candidate status of the association based on the CE score and an algorithm score indicating the likelihood that the mutation is the cause.
13. The apparatus according to claim 11, wherein the mutation data includes a damage score indicating the likelihood that a protein associated with the mutation is functionally impaired.
14. The apparatus according to claim 13, wherein the one or more processors are further configured to generate the damage score via another machine learning model trained using known harmful and neutral mutations.
15. The apparatus according to claim 11, wherein the one or more input features further include an essentiality score indicating the likelihood of pre-weaning lethality in homozygous mice for a robust knockout allele of a gene associated with the mutation.
16. The apparatus according to claim 15, wherein the one or more processors are further configured to generate the essentiality score via another machine learning model trained using genes known not to be essential for survival and genes known to be essential for survival.
17. The apparatus according to claim 11, wherein the one or more input features further include features associated with an algorithm score indicating the likelihood that the mutation is the cause.
18. The apparatus according to claim 11, wherein the one or more input features further include linkage data generated using automated meiotic mapping (AMM).
19. If two or more mutations are co-segregated, the one or more processors are further configured to determine which of the two or more mutations is a more robust candidate cause of the phenotype by omitting instances of homozygosity of the two or more mutations, and the one or more processors are configured to generate the CE score based on the determination. The apparatus according to claim 11.
20. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to receive one or more input features including phenotype data and mutation data; generate a candidate explorer (CE) score indicating the probability of an association between a phenotype and a mutation based on the one or more input features via a machine learning model; output an indicator of the association between the phenotype and the mutation based on the CE score; A non-transitory computer-readable medium that causes the above to be performed.