Detection of causal genes underlying variant associations

By integrating GWAS summary statistics with gene expression and interaction data, processor-implemented techniques effectively predict causal genes, overcoming limitations of existing methods and enhancing treatment selection and research targeting.

WO2025137320A1PCT designated stage expired Publication Date: 2025-06-26ILLUMINA INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/061085
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-22
Filing Date
2024-12-19
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

Current genome-wide association studies (GWAS) struggle to identify causal genes underlying variant associations, due to challenges such as linkage disequilibrium and incomplete regulatory element maps, which limits biological insight and effective treatment selection.

Method used

The development of processor-implemented techniques that integrate GWAS summary statistics with gene expression, biological pathway, and protein-protein interaction data to predict causal genes using similarity-based gene prioritization methods and polygenic priority scores.

Benefits of technology

These techniques improve the identification of causal genes by enhancing the accuracy of gene-level associations and providing improved polygenic priority scores, which can be used for treatment selection and research targeting, with a 10% improvement in precision-recall curve performance compared to conventional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024061085_26062025_PF_FP_ABST
    Figure US2024061085_26062025_PF_FP_ABST
Patent Text Reader

Abstract

Genome-wide association studies may allow for detection of variants that are statistically significantly associated with disease risk. However, inferring which are the genes underlying these variant associations may be difficult. The presently disclosed approaches utilize processor-implemented approaches for calculating polygenic priority scores (284) to facilitate causal gene identification in terms of both precision and recall compared to other techniques.
Need to check novelty before this filing date? Find Prior Art

Description

DETECTION OF CAUSAL GENES UNDERLYING VARIANT ASSOCIATIONSFIELD OF THE TECHNOLOGY DISCLOSED

[0001] The technology disclosed relates to the use of techniques that are implemented on computers and digital data processing systems for the purpose of detecting causal genes underlying variant associations from genome-wide association studies (GWAS).BACKGROUND

[0002] The subject matter discussed in this section should not be assumed to be prior art merely as a result of its mention in this section. Similarly, a problem mentioned in this section or associated with the subject matter provided as background should not be assumed to have been previously recognized in the prior art. The subject matter in this section merely represents different approaches, which in and of themselves can also correspond to implementations of the claimed technology.

[0003] Genetic variations can help explain many diseases, particularly complex or metabolic diseases that may present many and varied observable characteristics or symptoms, all of which may vary in severity between individuals. With respect to such complex diseases, every human has a unique genetic code and there are many genetic variants within a group of individuals, some of which may be linked to such complex diseases. Correspondingly, it may be difficult to identify which genes are likely to be of clinical interest in the context of a given genetic disease. By way of example, it may be difficult to detect and identify causal genes underlying the variant associations present within genome-wide association studies.

[0004] With respect to such studies, genome-wide association studies (GWAS) have been used to identify thousands of genetic loci associated with various traits, such as complex traits and genetic diseases. Such studies may be useful for understanding processes and pathways associated with complex traits, including complex (e.g., multigenic) genetic diseases. However, for the majority of loci identified using GWASthe identity of the causal gene or genes that underly the association remains unknown. Correspondingly, the biological insight that might be gained by such studies is limited.

[0005] In particular, there are various challenges to identifying a causal gene via such studies. For example, linkage disequilibrium (LD) between variants may obfuscate the identity of the causal variant. Further, most associated loci do not contain coding variants but, instead, the causal variant may act through gene regulatory mechanisms. Incomplete maps from a regulatory element to a respective gene may therefore hinder causal gene identification. With the preceding in mind, techniques for better identifying causal genes may be useful.BRIEF DESCRIPTION

[0006] The present approach relates the development and use of processor- implemented techniques to predict or identify causal genes from GW AS summary statistics. The techniques related herein may be employed to leam or identify similarities of expression, interaction, and pathway membership between genes to predict target genes, e.g., causal genes. By way of example, in certain embodiments of the techniques described herein a similarity-based gene prioritization method is described that may be used to derive a gene-level score (e.g., a polygenic priority score) that may be used in the identification of causal genes. In certain of the described embodiments causal genes may be identified by integrating GWAS summary statistics with gene expression, biological pathway, and predicted protein-protein interaction (PPI) data.

[0007] The techniques described herein may be employed to evaluate similarities of expression, interaction(s), and pathway membership between genes to predict target (e.g.. causal) genes. In particular, the techniques described herein are based on the assumption that causal genes share functional characteristics. For example, the present techniques are based on an assumption that genes whose physical locations on the genome are near associated SNPs and that share similar biological annotations are most likely to be causal.

[0008] With this in mind, and in contrast to prior approaches, the presently disclosed techniques employ a dependent variable that is an estimated probability that a respective gene is the correct target based upon distance (e.g., ranked distance) to an index variant (i.e., a leading significant GWAS variant). By way of example, such probability estimates may be based on distance to index single nucleotide polymorphisms (SNPs).

[0009] In certain embodiments a principal component analysis (PC A) or comparable statistical technique may be employed to reduce the dimensionality7of the feature space that is to be related to the estimated probabilities based upon distance. By way of example, PCA analysis may be employed in certain implementations to reduce a feature space of approximately 50,000 features to a feature space of approximately 4,000 to 9,000 features (e.g., an approximately 5-fold to 12-fold reduction).

[0010] With respect to associating or relating the estimated probabilities and this reduced-dimensionality feature space, in certain such embodiments the relationship between this reduced-dimensionality7feature space and the estimated probabilities may be established using a classification framework (e.g., a qualitative or binning framework) and an associated non-linear model, as opposed to linear modeling. The derived relationships, represented as polygenic priority7scores, may be further evaluated or processed to establish a per-gene interpretation of the contributions of features within the feature space (i.e., per-gene feature importance predictions), which may be useful in understanding the biological context or "‘story” being conveyed by the underlying data. In such implementations the disclosed processor-implemented approaches are useful in identifying causal genes, which may then be used in treatment or therapy selection or in identifying targets for future research, and providing improved performance compared to other methods conventionally employed in the field.

[0011] With the preceding in mind, and by way of summary, as part of the presently described techniques various gene-level associations may be computed from GWAS summary statistics (or other descriptive, comprehensive statistics) and may be used to learn joint polygenic enrichments of gene features derived from cell-type specific gene expression, biological pathways, and protein-protein interactions (PPI). A priorityscore may be assigned to every protein coding gene according to these enrichments and this score may be used to nominate causal genes for review or evaluation. Such scores and nominated causal genes may be presented to a reviewer, such as via a user interface of a processor-based system, and / or may be used as part of automatically selecting or recommending a diagnosis, treatment (e.g., pharmaceutical or biologic treatment), research target, prognosis, or recommendation for an individual. In some embodiments, such selections or recommendations may be presented as a ranked list or with associated probabilistic assessments for consideration by a reviewer.

[0012] With the preceding in mind, in one embodiment a processor-implemented method is provided for detecting causal genes. In accordance with this embodiment, an estimated probability that a respective gene is a causal gene is calculated. The estimated probability is based on a distance to an index variant. A feature space comprising gene and feature data is generated or accessed. The estimated probability and the feature space are processed using a classification framework to generate a fitted model characterizing the relationship between the estimated probability and one or more feature of the feature space. Predictions for per-gene interpretation of feature contributions are generated. Based on one or both of an output of the classification framework or per-gene interpretations of feature contributions, a determination is made that the respective gene is the causal gene of interest.

[0013] In a further embodiment, one or more tangible, machine-readable media storing processor-executable routines are provided. In accordance with this embodiment, the processor-executable routines, when executed by a processor, cause acts to be performed comprising: fitting a non-linear classification model configured to output vectors of joint polygenic enrichments of gene features that is used to generate a respective polygenic priority score, wherein the non-linear classification model fits: (1) a dependent variable comprising an estimated probability that a respective gene is a causal gene, wherein the estimated probability is based on a distance to an index variant; and (2) a plurality' of independent variables corresponding to a feature space comprising gene and feature data; and interpreting the polygenic priority scores or other outputs of the non-linear classification model to determine whether the respective gene is the causal gene.BRIEF DESCRIPTION OF DRAWINGS

[0014] These and other features, aspects, and advantages of the present invention will become better understood when the following detailed description is read with reference to the accompanying drawings in which like characters represent like parts throughout the drawings, wherein:

[0015] FIG. 1 depicts a schematic diagram of an implementation of a causal gene identification system in a networked or cloud computing environment, in accordance with aspects of the present technique;

[0016] FIG. 2 is a simplified block diagram of a processor-based system, in accordance with aspects of the present technique;

[0017] FIG. 3 illustrates aspects of derivation of an estimated probability that a respective gene is the correct target based upon distance to an index variant, in accordance with aspects of the present technique;

[0018] FIG. 4 illustrates aspects of derivation of a feature space, in accordance with aspects of the present technique;

[0019] FIG. 5 illustrates aspects of derivation of a priority score (e.g., a polygenic priority score) using distance-based probability estimates and a feature space, in accordance with aspects of the present technique; and

[0020] FIG. 6 illustrates au-PRC curves for techniques described herein and for other techniques, in accordance with aspects of the disclosure.DETAILED DESCRIPTION

[0021] The following discussion is presented to enable any person skilled in the art to make and use the technology disclosed, and is provided in the context of a particular application and its requirements. Various modifications to the disclosed implementations will be readily apparent to those skilled in the art, and the generalprinciples defined herein may be applied to other implementations and applications without departing from the spirit and scope of the technology disclosed. Thus, the technology disclosed is not intended to be limited to the implementations shown, but is to be accorded the widest scope consistent with the principles and features disclosed herein.

[0022] As discussed herein, genome-wide association studies (GWAS) allow variants to be detected that are statistically significantly associated with disease risk or with other forms of multigenic phenoty pe presentation. However, inferring which are the genes underlying these variant associations (i.e., the causal genes) is a challenging task. The techniques described herein may be employed on such GWAS data (or similar gene study data) to evaluate similarities of expression, interaction(s), and pathway membership betw een genes so as to predict target (e.g., causal) genes. In practice, these techniques substantially improve performance in terms of identifying such causal genes compared to other approaches. By way of example, the area under precision-recall curve (auPRC) for identify ing target genes may be improved from 0.64 to 0.70 (i.e., by approximately 10%), where precision is understood to be the ratio of true positives to declared positives (also referred to as the positive predictive value and corresponding to the complement of the false discovery’ rate) and recall is understood to be the number of true positive samples correctly classified by a model (also known as the true positive rate or sensitivity). Further, the ease with which predictions may be interpreted (e.g., understanding the biological context or “story '’) is also improved.

[0023] In particular, as discussed herein an estimated probability that a respective gene is the correct target gene is based upon distance (e.g., ranked distance) to an index variant (i.e., a leading significant GWAS variant). In practice, this distance may be based on distance to index single nucleotide polymorphisms (SNPs). In certain described embodiments this estimated probability is used as the dependent variable in the presently^ described assessment techniques. As further described, instead of a regression-based framework being used to model or fit the dependent variable and observed or known features, a classification-based framework may instead be employed to associate such estimated probabilities to a feature space that relates genes and features. As described herein, such features encompass underlying data describingexpression, interaction, and pathway membership between genes. In certain embodiments, a non-linear model (as opposed to a linear model) may be employed as part of the classification framework. Such a non-linear model may facilitate modeling or identification of non-linear features (e.g., interactions (such as for two-gene or multigene features) and / or re-scaling of features).

[0024] With respect to the feature space, to further improve the assessment process as well as the performance of the presently described techniques when implemented on a processor-based system, the dimensionality of the feature space may be reduced, such as using statistical techniques. By way of example, principal component analysis (PCA) may be employed to reduce the dimensionality of the feature space processed by the classification framework. Such dimensionality reduction approaches may be employed to reduce the feature space by a factor of between 5-fold and 12-fold, such as an approximately 7-fold reduction in features from -50,000 to -7,000.

[0025] The output of the classification framework may be characterized as a polygenic priority score and may be evaluated or processed to establish a per-gene interpretation of the contributions of features within the feature space (i.e., per-gene feature importance predictions). In practice these scores may be used to identify causal genes and provide improved performance compared to other methods conventionally employed in the field.

[0026] With the preceding in mind, and turning to the figures, FIG. 1 illustrates aspects of one embodiment of a processor-based (e.g., computational) pipeline for identifying target (i.e., causal) genes in genetic data (e.g., GWAS data). In particular, FIG. 1 depicts aspects of a cloud- or network-based approach by which GWAS data may be analyzed to derive one or more polygenic priority scores that may be used to predict or prioritize potential target genes. In such a context, all or part of the analysis may occur remotely or, alternatively, the analysis may be performed using both local and remote resources. For example, certain aspects of the processes described herein may be performed at the datacenter or remote server, while other aspects of the processed may be performed locally at the workstation or thin-client. Further, though a cloud- or network-based approach is described with respect to FIG. 1 so as to providea comprehensive example, in practice the processes and techniques described herein may be performed on a single processor-based device, either with or without a network connection. Thus, the example described with respect to FIG. 1 should be understood to not be limiting, but instead to provide context for one type of real-world implementation.

[0027] With this in mind, and turning to FIG. 1, a polygenic priority scoring framework 100 is depicted in accordance with embodiments of the present technique. More specifically, FIG. 1 illustrates an abstraction of a cloud platform infrastructure and local client interface to the cloud infrastructure, such as via a local network. In this example, a cloud-based platform 90 (such as may be instantiated at a datacenter or a remote server) is connected to a client device 98 via a network 92 to facilitate processing of GWAS data in response to a request 122 to generate one or more responses 124 (e.g., polygenic priority scores). Such a connection may be implemented via a web browser interface, a dedicated, standalone application, or other suitable program or data interfaces. In the depicted example, the client device 98 is itself part of or in communication with a local client netw ork 96 that is configured to communicate with the network 92 that allows communication outside the client network 96. As used herein, a server, workstation, or other processor-based device may be understood to be implemented as a virtual instance (e.g., a virtual server) or as a physical or hardware implementation, though it should be understood that virtual servers also have underlying physical memory' and processor aspects.

[0028] The implementation of the polygenic priority scoring framework 100 illustrated in FIG. 1 includes a score calculation engine 102 configured to implement the logic and processes described herein and one or more databases 106 (either w ithin a client instance, within the cloud-based platform 90 (e.g., within the datacenter or a related datacenter), or otherwise accessible by the instance and / or platform 90). The score calculation engine 102 may interact with a user of the client device 98 via requests 122 (e.g., requests to generate one or more polygenic priority' scores) and responses 124 (e.g., estimated polygenic priority scores).

[0029] For the embodiment illustrated in FIG. 1, the database 106 may be a standalone database, a database server instance, or a collection of database server instances. The illustrated database 106 may store or access GWAS data or panel results 108, one or more sources of expression data 110, one or more sources of interaction (e.g., protein-protein interaction (PPI) data 112, and / or one or more sources of pathwaymembership data 114. As discussed herein, the GWAS data and / or other data described may be used by the score calculation engine 102 to formulate a responsive reply 124, comprising a polygenic priority score or related metric or response, to the user of the client device 98.

[0030] With the preceding in mind, FIG. 2 depicts an example of a processor-based system 160 (e.g., a workstation, a server, a thin client, a computer system, and so forth) suitable for use as the client device(s) 98 or as part of the cloud-based platform 90 in accordance with the framework illustrated in FIG. 1. In this example system, a high- level hardware architecture is described for reference. Such hardware may be physically embodied as one or more computer systems (e.g., servers, workstations, and so forth). It should be appreciated that the present example may include components not found in all embodiments of such a system or may not illustrate all components that may be found in such a system. Further, in practice aspects of the present approach may be implemented in part or entirely in a virtual server or client environment or as part of a cloud platform. However, in such contexts the various virtual server or client instantiations will still be implemented on an underlying hardware platform as described with respect to FIG. 2. although certain functional aspects described may be implemented at the level of the virtual server or client.

[0031] With this in mind, FIG. 2 is a simplified block diagram of a processor-based system (e.g., a computer system) 160 that can be used to implement the technology disclosed. Such a computer system typically includes at least one processor (e.g.. microprocessor or CPU) 1 4 that communicates with a number of peripheral devices via bus subsystem 168. These peripheral devices can include a storage subsystem 172 including, for example, memory devices 176 (e.g., RAM 180 and ROM 184) and a file storage subsystem 188, user interface input devices 192, user interface output devices 196, and a network interface subsystem 198. The input and output devices allow userinteraction with computer system (e g., processing / storage system 160). Network interface subsystem 198 provides an interface to outside networks, including an interface to corresponding interface devices in other computer systems.

[0032] In the context of the depicted processor-based system 160, the user interface input devices 192 can include a keyboard; pointing devices such as a mouse, trackball, touchpad, or graphics tablet; a scanner; a touch screen incorporated into the display; audio input devices such as voice recognition systems and microphones; and other types of input devices. In general, use of the term “input device” may be construed as encompassing all possible types of devices and ways to input information into computer system.

[0033] User interface output devices 196 can include a display subsystem, a printer, a fax machine, or non-visual displays such as audio output devices. The display subsystem can include a cathode ray tube (CRT), a flat-panel device such as a liquid crystal display (LCD) or organic light emitting diode (OLED) display, a projection device, or some other mechanism for creating a visible image. The display subsystem can also provide a non-visual display such as audio output devices. In general, use of the term “output device” may be construed as encompassing all possible types of devices and ways to output information from the computer system to the user or to another machine or computer system.

[0034] Storage subsystem 172 stores programming and data constructs that provide the functionality of some or all of the modules and methods described herein, such as one or more of the score calculation engine 102, GWAS results 108, expression data 110, interaction data 112, pathway membership data 114, and so forth. Stored software modules are generally executed by a processor 164 alone or in combination with other processors 164. Data constructs or tables may be stored locally on the processor-based system 160 or accessed from a remote system on which they are stored in such a storage subsystem.

[0035] Memory 176 used in the storage subsystem 172 can include a number of memory structures or devices, such as a main random-access memory (RAM) 180 forstorage of instructions and data during program execution and a read only memory (ROM) 184 in which fixed instructions are stored. A file storage subsystem 188 can provide persistent storage for program and data files, and can include a hard disk drive, solid state data drives, a floppy disk drive along with associated removable media, a CD-ROM drive, an optical drive, or removable media cartridges. The modules implementing the functionality of certain implementations can be stored by file storage subsystem 188 in the storage subsystem 172, or in other machines accessible by the processor 164.

[0036] Bus subsystem 168 provides a mechanism for letting the various components and subsystems of computer system communicate with each other. Although bus subsystem 168 is shown schematically as a single bus, alternative implementations of the bus subsystem 168 can use multiple busses.

[0037] The processor-based system 160 itself can be of varying types including a personal computer, a portable computer, a w orkstation, a computer terminal, a netw ork computer, a thin client, a mainframe, a stand-alone server, a server farm, a widely- distributed set of loosely networked computers, or any other data processing system or user device. Due to the ever-changing nature of computers and networks, the description of a processor-based system 160 as depicted in FIG. 2 is intended only as an example for purposes of illustrating the functionality and types of components associated with the technology disclosed. Many other configurations of computer system are possible having more or less components or different components than the computer system depicted in FIG. 2.

[0038] With the preceding context in mind, the techniques discussed herein may be employed for processing genetic data (e.g., GWAS) data to facilitate the detection of causal genes. Such processing may allow identification or derivation of causal genes and / or of metabolic pathway memberships associated with a disease or condition and, correspondingly, of potential treatments or remediations that may be employed in treating a disease.

[0039] With this in mind, it may be noted that in certain existing approaches to assessing GWAS data for the detection of causal genes, metrics such as MAGMA (Multi-marker Analysis of GenoMic Annotation) scores have been derived from an input (e.g., a summary statistics linkage disequilibrium (LD) reference panel) and employed as a dependent variable (e.g., “y”) in the analysis. Such approaches may be characterized as employing similarity -based methods to prioritize likely causal genes for assessment by searching for global patterns in associated genes and nominating those for consideration that have similar gene expression, functions, biological or metabolic pathway membership(s), and / or protein-protein interaction (PPI) network connections.

[0040] Conversely, certain embodiments of the techniques disclosed herein instead employ an estimated probability' that a given gene is the correct target (i.e., a causal gene of interest) as the dependent variable of the assessment. In certain such implementations, this probability is based on distance (e.g., a ranked distance) to an index variant (e.g., an index single-nucleotide polymorphism (SNP)). In some embodiments, multiple imputations may be performed to model uncertainty. Derivation of such estimated probabilities, y, 220, are illustrated as a high-level schematic view in FIG. 3. In this example, the estimated probabilities are derived with reference to the summary statistics of a linkage disequilibrium (LD) reference panel 224 generated from GWAS data.

[0041] Such approaches based on distance to an index variant provide improved performance as distance-based metrics (e.g., distance to an index SNP) are better estimators of causality than measures based on similarity-based metrics. This observation is illustrated in Table 1, illustrating the results of three independent approaches to quantifying the distribution of ordinal rank for the causal gene from a lead GWAS SNP.TABLE 1With respect to Table 1, the "‘gene body"’ may be understood to be the region corresponding to the furthest upstream transcription start site (TSS) and furthest downstream transcription end site (TES) for that gene, with the distance being defined as distance to this TSS-TES region. The ABCmax corresponds to studies of Activity - by -Contact measures, pQTL corresponds to studies of protein quantitative trait loci studies, and Exons corresponds to measures derived from exome sequencing studies.

[0042] In further aspects, a feature space comprising gene feature data (e.g., a matrix or data structure associating or relating gene and feature data) may be used for comparison to the distance-based probability estimates as part of identifying causal genes. Such a feature space 250, illustrated as a high-level schematic view in FIG. 4, may encompass or represent gene expression data 254, protein-protein interaction (PPI) network data 258, and biological or metabolic pathway data 262 (e.g., presence or activity in a metabolic or biochemical pathway). In practice, such gene feature data may be derived from any suitable source, including but not limited to: (a) bulk and / or single-cell gene expression datasets, (b) curated biological pathways or pathway data, and / or (c) predicted protein-protein interaction (PPI) networks.

[0043] In practice, such a feature space 250 may encompass approximately 50.000 features or more, which may make computations and comparisons based on this feature space computationally intensive. In certain embodiments statistical grouping or simplification techniques may be employed to reduce the dimensionality of the feature space 250, taking into account factors such as redundancy of information, lack of discrimination or information added, by a gene or feature, inconsistency with other features within the feature space 250 and so forth.

[0044] By way of example, in certain embodiments analytic techniques may be employed that simplify the description or presentation of a set of interrelated variables(e.g., gene expression data, interaction data, and pathway membership data of a feature space). In one such example, principal component analysis (PCA) is one technique that may be employed to reduce the dimensionality of the feature space, such as by a factor of 5, 6, 7, 8, 9, 10, 11, 12 or greater. For instance, in one such example a feature space of approximately 50,000 features may be reduced to approximately 4,000 to 9000 features (e.g., -7.000 features) to be considered. Such a reduction of the dimensionality of the feature space improves computational efficiency in deriving priority scores and identifying likely causal genes in subsequent steps, as discussed herein.

[0045] With respect to the use of the distance-based estimated probability’ that a given gene is a causal gene of interest and the reduced dimensionality feature space 250, a classification framework is employed (as opposed to a regression framework employed in conventional approaches) in conjunction with a non-linear model to classify or relate the estimated probabilities and reduced-dimensionality feature space. As noted herein, such a non-linear model may facilitate modeling or identification of non-linear features (e.g., interactions (such as for two-gene or multi-gene features) and / or re-scaling of features).

[0046] With respect to the use of the classification framework, such an approach assumes the output variable has a discrete, qualitative value, such as a discrete state (e.g., [true / false], [causal / non-causal], [diseased / not diseased], [causal, likely causal, possibly causal, indeterminate, likely not causal, not causal]), such that the classification framework maps or sorts input variable(s) or value(s) to the appropriate discrete output variable or value. Such approaches are in contrast with regression-based approaches, which assume that the output value or variable may vary continuously across a range, as opposed to being discrete or binned. Correspondingly, in regression- based-approaches the focus is on identifying the best fit line describing the relationship between the input variable(s) and the output variable and on determining such continuous values based on the input variables. In classification-based approaches the focus is instead on identifying decision boundaries or thresholds that may be used to classify the dataset into two or more discrete categories.

[0047] In practice, such a non-linear, classification framework approach may be implemented using machine learning tools and approaches. In accordance with one example, a distributed gradient-boosted decision tree machine learning approach may be employed to implement a non-linear classification framework as described herein. By way of example, XGBoost is one such machine learning technique that may be employed to implement a suitable non-linear classification framework. Such approaches may be implemented as decision tree ensemble learning algorithms suitable for classification processing techniques, where the ensemble learning algorithms combine multiple machine learning algorithms to improve model performance. In such contexts gradient-boosting may be understood to refer to improving the performance of a single, weak model by combination with other weak models so at to generate a collective model that is stronger than its constituents. In a gradient-boosted decision tree technique, therefore, an ensemble of shallow decision trees may be iteratively trained such that over each iteration error residuals of the previous model are used to fit the next model. A weighted sum of all of the decision tree predictions may be provide as the final prediction. With this in mind, in one embodiment XGBoost in combination with gradient boosting may be employed as a non-linear classifier to process the estimated probability and the reduced dimensionality feature space 250. In the depicted example, the classification process, as an output, yields a vector of joint polygenic enrichments of gene features, / ?, that can be used to generate polygenic priority scores, y-

[0048] In one embodiment, the outputs of such a non-linear classification framework are used to generate predictions for per-gene interpretation of feature contributions. In certain embodiments, Tree-SHAP techniques may be employed to interpret the ensemble tree model outputs of a classification framework as described herein. Such interpretation may be helpful in furthering how much is understood about the nature of the respective features in question and may allow interpretation of the biological context conveyed by the model results. By way of example, such interpretation techniques may be useful in relating features to genes on a per-gene basis and / or in providing understanding of each features contribution to the phenotype(s) of a complex trait or disease.

[0049] In practice, such Tree-SHAP techniques may be used to estimate SHAP values for tree models, as well as ensembles of trees, under various different possible assumptions about feature dependence or relatedness. In this manner, such SHAP values (i.e., Shapley Additive exPlanations) may be employed to increase transparency and interpretability of such machine learning models. In particular, the SHAP value for a feature may be understood to be the average change in the output of a model by conditioning on the respective feature while introducing features (e.g., while introducing features one at a time) over some or all feature orderings. With this in mind, Tree-SHAP techniques may be employed to generate predictions, based on the outputs of the classification framework, that interpret or otherwise provide the per-gene feature contributions to the genetic disease or phenotype in question and helping to identify causal genes from among an otherwise prohibitively expansive solution set.

[0050] Turning to FIG. 5, both the estimated probability 220 that a given gene is the correct target (i. e. , a causal gene of interest), which is depicted as the dependent variable (y), and the reduced dimensionality feature space 250, which is depicted as the independent variable(s) (x), are illustrated as inputs to a non-linear classification framework 280, allowing the derivation of the coefficient / ?. corresponding to a vector of joint polygenic enrichments of gene features. Correspondingly, a prediction in the form of a polygenic priority7score 284 corresponding to y = X / 3 may be derived that can be used to determine whether a given gene is causal in the investigative context. In particular, polygenic priority scores are computed for each gene, g. on a chromosome, i, by multiplying its row vector of gene features, Xg, by ft-Chr i- For example, to compute priority scores for gene g on chromosome 1, yg— xg^_chr tis computed. In this example, ygis referred to as the polygenic priority score for gene g.

[0051] With the preceding in mind, and turning to FIG. 6. results of a study are graphically illustrated for the presently disclosed evaluation framework. In this study, the ten nearest genes in a GWAS locus were evaluated using different techniques and / or variations of these techniques. By way of example, the baseline technique, denoted “pops” employed a MAGMA score as the dependent variable, with no reduction in the dimensionality of the feature space and with a linear regression framework. The resultsof this approach are denoted by line 300. Conversely, the technique in accordance with some or all of the features disclosed herein (e.g.. a distance-based estimated probability as the dependent variable, reduction in dimensionality of the feature space, use of a non-linear classification framework, and so forth) is denoted “cgpops” in the depicted results and are shown by line 304. As shown in these results, the “cgpops” approach, relative to the baseline “pops” approach, improves the area under precision-recall curve (auPRC) for identifying the target gene from 0.64 to 0.70 (i.e., by approximately 10%), representing a substantial improvement in the identification of causal genes.

[0052] With the preceding in mind, and by way of summary, the presently described model provides improved performance at causal gene identification, allowing calculation of polygenic priority scores that may be useful in identifying causal genes of interest, such as from genome wide association studies (GW AS). In particular, the presently described techniques facilitate prioritization of causal genes underlying complex (e.g.. polygenic) traits and diseases by evaluating biologically relevant properties from multiple types of gene features.

[0053] While only certain features of the invention have been illustrated and described herein, many modifications and changes will occur to those skilled in the art. It is, therefore, to be understood that the appended claims are intended to cover all such modifications and changes as fall within the true spirit of the invention.

Claims

CLAIMS:1 . A processor-implemented method for detecting causal genes, comprising: calculating an estimated probability that a respective gene is a causal gene, wherein the estimated probability is based on a distance to an index variant; generating or accessing a feature space comprising gene and feature data; processing the estimated probability and the feature space using a classification framework to generate a fitted model characterizing the relationship between the estimated probability and one or more feature of the feature space; generating predictions for per-gene interpretation of feature contributions; and based on one or both of an output of the classification framework or per-gene interpretations of feature contributions, determining that the respective gene is the causal gene of interest.

2. The processor-implemented method of claim 1, wherein the distance is a ranked distance.

3. The processor-implemented method of claim 1, wherein the index variant comprises an index single nucleotide polymorphism (SNP).

4. The processor-implemented method of claim 1, wherein the estimated probability7is calculated based on summary statistics generated for genome-wide association studies (GWAS) data.

5. The processor-implemented method of claim 1 , wherein the feature space comprises one or more of gene expression data, protein-protein interaction (PPI) data, or biological or metabolic pathway data.

6. The processor-implemented method of claim 1, further comprising: reducing the dimensionality of the feature space prior to processing the estimated probability and the feature space using the classification framework.

7. The processor-implemented method of claim 6, wherein reducing the dimensionality of the feature space comprises performing principal component analysis (PCA) on the feature space.

8. The processor-implemented method of claim 6, wherein reducing the dimensionality of the feature space reduces the number of features by approximately 5- to 12-fold.

9. The processor-implemented method of claim 1, wherein the feature space is derived from one or more of bulk or single-cell gene expression datasets, curated biological pathways or pathway data, or predicted protein-protein interaction (PPI) networks.

10. The processor-implemented method of claim 1, wherein the classification framework incorporate anon-linear model.

11. The processor-implemented method of claim 1, wherein the classification framework is implemented using XGBoost.

12. The processor-implemented method of claim 1, wherein the classification framework outputs a vector of joint polygenic enrichments of gene features that is used to generate polygenic priority scores.

13. The processor-implemented method of claim 1, further comprising: generating a polygenic priority score for the respective gene based on an output of the classification framew ork.

14. The processor-implemented method of claim 1, wherein generating predictions for per-gene interpretation of feature contributions comprises: performing a Tree-SHAP (Shapley Additive exPlanations) operation to interpret outputs of the classification framework.

15. One or more tangible, machine-readable media storing processor-executable routines, wherein the processor-executable routines, when executed by a processor, cause acts to be performed comprising: fitting a non-linear classification model configured to output vectors of joint polygenic enrichments of gene features that is used to generate a respective polygenic priority score, wherein the non-linear classification model fits: a dependent variable comprising an estimated probability that a respective gene is a causal gene, wherein the estimated probability is based on a distance to an index variant; and a plurality of independent variables corresponding to a feature space comprising gene and feature data; and interpreting the polygenic priority' scores or other outputs of the non-linear classification model to determine whether the respective gene is the causal gene.

16. The one or more tangible, machine-readable media of claim 15, wherein the distance is a ranked distance.

17. The one or more tangible, machine-readable media of claim 15, wherein the index variant comprises an index single nucleotide polymorphism (SNP).

18. The one or more tangible, machine-readable media of claim 15, wherein the feature space comprises one or more of gene expression data, protein-protein interaction (PPI) data, or biological or metabolic pathway data.

19. The one or more tangible, machine-readable media of claim 15, wherein the processor-executable routines, when executed by a processor, cause further acts to be performed comprising: performing principal component analysis on the feature space to reduce the dimensionality of the feature space.

20. The one or more tangible, machine-readable media of claim 15, wherein the non-linear classification model is implemented using XGBoost.

21. The one or more tangible, machine-readable media of claim 15. wherein interpreting the polygenic priority scores or other outputs of the non-linear classification model comprises performing a Tree-SHAP (Shapley Additive exPlanations) operation.