Methods of predicting phenotypic traits from proteomic data and correlating the phenotypic traits to genomic data, analysis devices that perform the methods, and storage media that directs an analysis device to perform the methods
By applying a predictive model to proteomic data and identifying correlated genetic variants, phenotypic traits are predicted without direct observation, addressing limitations in existing methodologies and improving predictive accuracy.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SOMALOGIC OPERATING CO INC
- Filing Date
- 2025-09-23
- Publication Date
- 2026-05-21
AI Technical Summary
Existing methodologies for correlating organismal phenotypes to genetic variation require direct observation of phenotypic traits, limiting their applicability and necessitating improved methods for predicting phenotypic traits from proteomic data and correlating them to genomic data.
Applying a predictive model to proteomic data to predict phenotypic traits and identifying correlated genetic variants, allowing for the prediction of phenotypic traits without direct observation, using proteomic and genomic data analysis devices and storage media.
Enables the prediction of phenotypic traits prior to their expression, expanding knowledge of causal relationships between genetic variants and diseases, and enhancing prediction accuracy through higher-order protein abundance analysis.
Smart Images

Figure US2025047594_21052026_PF_FP_ABST
Abstract
Description
[0001] Koiitch Romano Dascenzo Gates LLC Attorney Docket No. SML315PCT
[0002] METHODS OF PREDICTING PHENOTYPIC TRAITS FROM PROTEOMIC DATA AND CORRELATING THE PHENOTYPIC TRAITS TO GENOMIC DATA, ANALYSIS DEVICES THAT PERFORM THE METHODS, AND STORAGE MEDIA THAT DIRECTS AN ANALYSIS DEVICE TO PERFORM THE METHODS
[0003] FIELD
[0004] This disclosure relates to methods of predicting phenotypic traits from proteomic data and correlating the phenotypic traits to genomic data, to analysis devices that perform the methods, and to storage media that directs an analysis device to perform the methods.
[0005] INTRODUCTION
[0006] Various methodologies, such as quantitative trail locus (QTL) analysis, may be utilized to correlate organismal phenotypes to potentially causal genetic variation. Such methodologies generally rely upon direct observation of the organismal phenotype within an individual to establish the correlation between the organismal phenotypes and the genetic variation. While effective in certain circumstances, reliance on direct observation of the organismal phenotype requires that an individual exhibit a given organismal phenotype to establish a corresponding correlation, severely limiting the overall applicability of these methodologies. Thus, there exists a need for improved methods of predicting phenotypic traits from proteomic data and correlating the phenotypic traits to genomic data, for analysis devices that perform the methods, and for storage media that directs analysis devices to perform the methods.
[0007] SUMMARY
[0008] The present disclosure provides analysis devices, storage media, and methods relating to prediction of phenotypic traits from proteomic data and correlating the phenotypic traits to genomic data. The proteomic data includes information regarding protein abundance for a plurality of distinct proteins for each individual of a plurality of individuals, and the genomic data includes information regarding presence of a plurality of distinct genetic variants for each individual. The methods include applying a predictive model to the proteomic data to obtain model results that predict at least one predicted Koiitch Romano Dascenzo Gates LLC Attorney Docket No. SML315PCT
[0009] phenotypic trait that is correlated to the proteomic data for at least one individual of the plurality of individuals. The applying includes correlating the protein abundance of at least two proteins of the plurality of distinct proteins to the predicted phenotypic trait for the at least one individual. The methods also include identifying, from the genomic data for the at least one individual, at least one genetic variant that is correlated to the at least one predicted phenotypic trait.
[0010] Features, functions, and advantages may be achieved independently in various embodiments of the present disclosure, or may be combined in yet other embodiments, further details of which can be seen with reference to the following description and drawings.
[0011] BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Fig. 1 is a flowchart depicting examples of methods of predicting phenotypic traits from proteomic data and correlating the phenotypic traits to genomic data, according to the present disclosure.
[0013] Fig. 2 is a table that summarizes examples of genomic data for a plurality of individuals.
[0014] Fig. 3 is a table that summarizes examples of proteomic data for a plurality of individuals.
[0015] Fig. 4 is a table that summarizes examples of coefficients for a linear model that may be utilized to predict a phenotypic trait from proteomic data.
[0016] Fig. 5 is a table that summarizes examples of scores from the linear model for the plurality of individuals.
[0017] Fig. 6 is a box plot that summarizes an example of a linear correlation between the scores from the linear model and a first single nucleotide polymorphism for the plurality of individuals.
[0018] Fig. 7 is a box plot that summarizes an example of a linear correlation between the scores from the linear model and a second single nucleotide polymorphism for the plurality of individuals.
[0019] Fig. 8 is a box plot that summarizes an example of a linear correlation between the scores from the linear model and a third single nucleotide polymorphism for the plurality of individuals. Koiitch Romano Dascenzo Gates LLC Attorney Docket No. SML315PCT
[0020] Fig. 9 is an example of a Manhattan plot generated utilizing methods, according to the present disclosure, and illustrating predicted probability of dementia for various genetic variants.
[0021] Fig. 10 is an example of a Manhattan plot generated utilizing methods, according to the present disclosure, and illustrating predicted probability of heart failure for various genetic variants.
[0022] DETAILED DESCRIPTION
[0023] Various aspects and examples of methods of predicting phenotypic traits from proteomic data and correlating the phenotypic traits to genomic data, analysis devices that perform the methods, and storage media that directs analysis devices to perform the methods are described below and illustrated in the associated drawings. Unless otherwise specified, a method in accordance with the present teachings, and / or its various components, may contain at least one of the structures, components, functionalities, steps, and / or variations described, illustrated, and / or incorporated herein. Furthermore, unless specifically excluded, the process steps, structures, components, functionalities, and / or variations described, illustrated, and / or incorporated herein in connection with the present teachings may be included in other similar devices and methods, including being interchangeable between disclosed embodiments. The following description of various examples is merely illustrative in nature and is in no way intended to limit the disclosure, its application, or uses. Additionally, the advantages provided by the examples and embodiments described below are illustrative in nature and not all examples and embodiments provide the same advantages or the same degree of advantages.
[0024] This Detailed Description includes the following sections, which follow immediately below: (1) Definitions; (2) Overview; (3) Examples, Components, and Alternatives; (4) Advantages, Features, and Benefits; and (5) Conclusion. The Examples, Components, and Alternatives section is further divided into subsections, each of which is labeled accordingly.
[0025] Definitions
[0026] The following definitions apply herein, unless otherwise indicated. Koiitch Romano Dascenzo Gates LLC Attorney Docket No. SML315PCT
[0027] “Comprising,” “including,” and “having” (and conjugations thereof) are used interchangeably to mean including but not necessarily limited to, and are open-ended terms not intended to exclude additional, unrecited elements or method steps.
[0028] Terms such as “first,” “second,” and “third” are used to distinguish or identify various members of a group, or the like, and are not intended to show serial or numerical limitation.
[0029] “Processing logic” describes any suitable device(s) or hardware configured to process data by performing one or more logical and / or arithmetic operations (e.g., executing coded instructions). For example, processing logic may include one or more processors (e.g., central processing units (CPUs) and / or graphics processing units (GPUs)), microprocessors, clusters of processing cores, FPGAs (field-programmable gate arrays), artificial intelligence (Al) accelerators, digital signal processors (DSPs), and / or any other suitable combination of logic hardware.
[0030] A “controller” or “electronic controller” includes processing logic programmed with instructions to carry out a controlling function with respect to a control element. For example, an electronic controller may be configured to receive an input signal, compare the input signal to a selected control value or setpoint value, and determine an output signal to a control element (e.g., a motor or actuator) to provide corrective action based on the comparison. In another example, an electronic controller may be configured to interface between a host device (e.g., a desktop computer, a mainframe, etc.) and a peripheral device (e.g., a memory device, an input / output device, etc.) to control and / or monitor input and output signals to and from the peripheral device.
[0031] A “phenotypic trait” is an observable characteristic of an organism, such as eye color or height. Phenotypic traits result from the interactions between the organism’s genes and environmental factors. Stated differently, phenotypic traits are physical manifestations of genetics and can be seen or measured. Additional examples of phenotypic traits that may be important to people and / or that may be predicted utilizing methods disclosed herein include diseases, such as dementia, cancer, heart failure, fatty liver disease, and cardiovascular disease. Koiitch Romano Dascenzo Gates LLC Attorney Docket No. SML315PCT
[0032] “Proteomic data” is data regarding protein presence and / or abundance for a plurality of distinct proteins. The plurality of distinct proteins is present within and / or generated by an individual, examples of which are disclosed herein.
[0033] “Genomic data” is data regarding the presence of one or more distinct genetic variants, or genes, within the individual.
[0034] An “individual” is an organism, such as a person or an animal, for which proteomic data and corresponding genomic data are available, for which the proteomic data and the corresponding genomic data may be generated, and / or for which it may be desirable to predict one or more phenotypic traits.
[0035] Overview
[0036] In general, methods in accordance with the present teachings predict phenotypic traits from proteomic data and correlate the phenotypic traits to genomic data. The proteomic data includes information regarding protein abundance for a plurality of distinct proteins for each individual of a plurality of individuals, and the genomic data includes information regarding presence of a plurality of distinct genetic variants for each individual. The methods include applying a predictive model to the proteomic data to obtain model results that predict at least one predicted phenotypic trait that is correlated to the proteomic data for at least one individual of the plurality of individuals. The applying includes correlating the protein abundance of at least two proteins of the plurality of distinct proteins to the predicted phenotypic trait for the at least one individual. The methods also include identifying, from the genomic data for the at least one individual, at least one genetic variant that is correlated to the at least one predicted phenotypic trait.
[0037] Technical solutions are disclosed herein for prediction of phenotypic traits from proteomic data and correlation of the phenotypic traits to genomic data. Specifically, the disclosed methods address a technical problem tied to genetics technology and arising in the realm of analysis of genomic data, namely the technical problem of predicting phenotypic traits from genomic data. The methods disclosed herein provide an improved solution to this technical problem by facilitating prediction of phenotypic traits prior to, or without, observation of the phenotypic traits.
[0038] The disclosed methods provide an integrated practical application of the principles discussed herein. Specifically, the disclosed systems and methods improve the Koiitch Romano Dascenzo Gates LLC Attorney Docket No. SML315PCT
[0039] functioning of a computer, improve prediction of phenotypic traits that are related to genetic variants, implement the solution with a particular machine that is integral to the invention, and describe a specific manner of analyzing proteomic data and genomic data that provides a specific improvement over prior systems and results in improved prediction of correlation between phenotypic traits and genomic data. Accordingly, the disclosed methods apply (or use) the relevant principles in a meaningfully limited way.
[0040] Aspects of the methods may be embodied as a computer method, analysis device, special-purpose analysis device, computer system, or computer program product. Accordingly, aspects of the methods may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, micro-code, and the like), or an embodiment combining software and hardware aspects, all of which may generally be referred to herein as a “circuit,” “module,” or “system.” Furthermore, aspects of the methods may take the form of a computer program product embodied in a computer-readable medium (or media) having computer-readable program code / instructions embodied thereon.
[0041] Any combination of computer-readable media may be utilized. Computer-readable media can be a computer-readable signal medium and / or a computer-readable storage medium. A computer-readable storage medium may include an electronic, magnetic, optical, electromagnetic, infrared, and / or semiconductor system, apparatus, or device, or any suitable combination of these. More specific examples of a computer-readable storage medium may include the following: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a readonly memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, and / or any suitable combination of these and / or the like. In the context of this disclosure, a computer-readable storage medium may include any suitable non-transitory, tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device.
[0042] A computer-readable signal medium may include a propagated data signal with computer-readable program code embodied therein, for example, in baseband or as part Koiitch Romano Dascenzo Gates LLC Attorney Docket No. SML315PCT
[0043] of a carrier wave. Such a propagated signal may take any of a variety of forms, including, but not limited to, electro-magnetic, optical, and / or any suitable combination thereof. A computer-readable signal medium may include any computer-readable medium that is not a computer-readable storage medium and that is capable of communicating, propagating, or transporting a program for use by or in connection with an instruction execution system, apparatus, or device.
[0044] Program code embodied on a computer-readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, and / or the like, and / or any suitable combination of these.
[0045] Computer program code for carrying out operations for aspects of the methods may be written in one or any combination of programming languages, including an object-oriented programming language (such as Java, C++), conventional procedural programming languages (such as C), and functional programming languages (such as Haskell). Mobile apps may be developed using any suitable language, including those previously mentioned, as well as Objective-C, Swift, C#, HTML5, and the like. The program code may execute entirely on a user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), and / or the connection may be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0046] Aspects of the methods may be described below with reference to flowchart illustrations and / or block diagrams. Each block and / or combination of blocks in a flowchart and / or block diagram may be implemented by computer program instructions. The computer program instructions may be programmed into or otherwise provided to processing logic (e.g., a processor of a general purpose computer, special purpose computer, field programmable gate array (FPGA), or other programmable data processing apparatus) to produce a machine, such that the (e.g., machine-readable) instructions, which execute via the processing logic, create means for implementing the functions / acts specified in the flowchart and / or block diagram block(s). Koiitch Romano Dascenzo Gates LLC Attorney Docket No. SML315PCT
[0047] Additionally or alternatively, these computer program instructions may be stored in a computer-readable medium that can direct processing logic and / or any other suitable device to function in a particular manner, such that the instructions stored in the computer-readable medium produce an article of manufacture including instructions which implement the function / act specified in the flowchart and / or block diagram block(s).
[0048] The computer program instructions can also be loaded onto processing logic and / or any other suitable device to cause a series of operational steps to be performed on the device to produce a computer-implemented process such that the executed instructions provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block(s).
[0049] Any flowchart and / or block diagram in the drawings is intended to illustrate the architecture, functionality, and / or operation of possible implementations according to aspects of the methods. In this regard, each block may represent a module, segment, or portion of code, which comprises one or more executable instructions for implementing the specified logical function(s). In some implementations, the functions noted in the block may occur out of the order noted in the drawings. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. Each block and / or combination of blocks may be implemented by special purpose hardware-based systems (or combinations of special purpose hardware and computer instructions) that perform the specified functions or acts.
[0050] Examples, Components, and Alternatives
[0051] The following sections describe selected aspects of illustrative methods as well as related analysis devices and / or storage media. The examples in these sections are intended for illustration and should not be interpreted as limiting the scope of the present disclosure. Each section may include one or more distinct embodiments or examples, and / or contextual or related information, function, and / or structure.
[0052] A. Illustrative Method
[0053] This section describes steps of an illustrative method 100 for predicting phenotypic traits from proteomic data and correlating the phenotypic traits to genomic data. Where Koiitch Romano Dascenzo Gates LLC Attorney Docket No. SML315PCT
[0054] appropriate, reference may be made to components and systems that may be used in carrying out each step. These references are for illustration and are not intended to limit the possible ways of carrying out any particular step of the method.
[0055] Fig. 1 is a flowchart illustrating steps performed in an illustrative method 100 and may not recite the complete process or all steps of the method. Although various steps of method 100 are described below and depicted in Fig. 1 , the steps need not necessarily all be performed, and in some cases may be performed simultaneously or in a different order than the order shown.
[0056] Methods 100 include methods of predicting phenotypic traits from proteomic data and correlating the phenotypic traits to genomic data. The proteomic data includes information regarding protein abundance for a plurality of distinct proteins for each individual of a plurality of individuals. The genomic data includes information regarding presence of a plurality of distinct genetic variants for each individual. Methods 100 may include collecting a biological sample at 110, obtaining genomic data at 120, obtaining proteomic data at 130, identifying data at 140, and / or filtering at 150. Methods 100 include applying a predictive model at 160, and methods 100 also may include transforming model results at 170. Methods 100 also include identifying at least one genetic variant at 180, and methods 100 also may include displaying results at 190.
[0057] Collecting the biological sample at 110 may include collecting a corresponding biological sample from each individual of the plurality of individuals. When methods 100 include the collecting at 110, the obtaining at 120 may include analyzing the corresponding biological sample from each individual to produce and / or generate the genomic data. Additionally or alternatively, and when methods 100 include the collecting at 110, the obtaining at 130 may include analyzing the corresponding biological sample from each individual to produce and / or generate the proteomic data. The biological sample may include and / or be any suitable biological sample, such as a biological sample that includes genomic components, proteomic components, and / or both genomic and proteomic components. Examples of the biological sample include a bodily fluid, blood, blood plasma, and / or blood serum.
[0058] Obtaining the genomic data at 120 may include obtaining the genomic data for the plurality of individuals, such as for each individual of the plurality of individuals. The Koiitch Romano Dascenzo Gates LLC Attorney Docket No. SML315PCT
[0059] obtaining at 1 0 may be performed in any suitable manner. As an example, the obtaining at 120 may include obtaining the genomic data from a database of genomic data. As additional examples, the obtaining the genomic data may include performing, or obtaining genomic data generated by, genome sequencing, whole genome sequencing, DNA microarray analysis, and / or genotype imputation. As a more specific example, and when methods 100 include the collecting at 110, the analyzing the corresponding biological sample from each individual to produce and / or generate the genomic data may include performing the genome sequencing, the whole genome sequencing, the DNA microarray analysis, and / or the genotype imputation.
[0060] Obtaining the proteomic data at 130 may include obtaining the proteomic data for the plurality of individuals, such as for each individual of the plurality of individuals. The obtaining at 130 may be performed in any suitable manner. As an example, the obtaining at 130 may include obtaining the proteomic data from a database of proteomic data. As additional examples, the obtaining the proteomic data may include performing, or obtaining proteomic data generated by, top-down proteomics, bottom-up proteomics, liquid chromatography, mass spectroscopy, affinity purification, bioinformatics, and / or proteomic microarray analysis. As a more specific example, and when methods 100 include the collecting at 110, the analyzing the corresponding biological sample from each individual to produce and / or generate the proteomic data may include performing the top-down proteomics, bottom-up proteomics, liquid chromatography, mass spectroscopy, affinity purification, bioinformatics, and / or proteomic microarray analysis.
[0061] Identifying the data at 140 may include identifying genomic data and proteomic data obtained from and / or associated with each individual. Additionally or alternatively, the identifying at 140 may include generating and / or utilizing a database of genomic data and proteomic data that identifies genomic data and proteomic data obtained from each individual. Stated differently, the identifying at 140 may include associating genomic data obtained from a given individual with proteomic data obtained from the same individual, such as to permit and / or facilitate the identifying at 180. The identifying at 140 may include anonymously identifying the data, such as via utilizing a corresponding anonymous identifier for each individual. Koiitch Romano Dascenzo Gates LLC Attorney Docket No. SML315PCT
[0062] Filtering at 150 may include filtering the genomic data. This may include filtering the genomic data to permit, to facilitate, and / or to approve an accuracy of the identifying at 180 and may be accomplished in any suitable manner. As examples, the filtering at 150 may include filtering the genomic data by minor allele frequency, by imputation quality, and / or via Hardy-Weinberg equilibrium. The filtering of the genomic data may be performed subsequent to the obtaining at 120 and / or prior to the identifying at 180.
[0063] Additionally or alternatively, the filtering at 150 may include filtering the proteomic data. This may include filtering the proteomic data to permit, to facilitate, and / or to improve an accuracy of the applying at 160 and may be accomplished in any suitable manner. The filtering the proteomic data may be performed subsequent to the obtaining at 130 and / or prior to the applying at 160.
[0064] Applying the predictive model at 160 may include applying the predictive model to the proteomic data. The applying at 160 may include applying the predictive model to predict, or to predict a probability of, at least one predicted phenotypic trait that is correlated to the proteomic data for at least one individual of the plurality of individuals. Additionally or alternatively, the applying at 160 may include correlating the protein abundance of at least two proteins of the plurality of distinct proteins to the predicted phenotypic trait for the at least one individual. The correlating the protein abundance additionally or alternatively may be referred to herein as associating the protein abundance with the predicted phenotypic trait and / or applying a formula to ascertain a probability of the predicted phenotypic trait based upon the protein abundance.
[0065] In some examples, the applying at 160 may include correlating the protein abundance of more than two distinct proteins of the plurality of proteins. As examples, the applying at 160 may include correlating the protein abundance of at least four, at least six, at least eight, at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 100, at least 1000, or at least 10,000 proteins of the plurality of distinct proteins to the predicted phenotypic trait for the at least one individual.
[0066] The applying at 160 may be performed in any suitable manner. As an example, the applying at 160 may include applying a predetermined predictive model that is configured to predict phenotypic traits based upon proteomic data. As another example, the applying at 160 may include applying a supervised machine learning model trained, Koiitch Romano Dascenzo Gates LLC Attorney Docket No. SML315PCT
[0067] or that previously was trained, with and / or on proteomic training data and corresponding phenotypic training data. In some such examples, the applying at 160 also may include training the supervised machine learning model.
[0068] In a specific example, the applying at 160 may include applying a linear model that includes corresponding weighting coefficients for the protein abundance of each protein of the plurality of distinct proteins. This may include applying the linear model to the protein abundance for each individual. Additionally or alternatively, this may include applying the linear model to obtain model results that predict, or that predict a likelihood of, the at least one predicted phenotypic trait.
[0069] Transforming the model results at 170 may include transforming the model results, obtained during the applying at 160, for each individual. This may include transforming the model results prior to the identifying at 180, to permit the identifying at 180, and / or to facilitate the identifying at 180. In some examples, the transforming at 170 may include transforming the model results, such as to a format that may be utilized with and / or during the identifying at 180. The transforming at 170 may be performed in any suitable manner. As examples, the transforming at 170 may include performing a rank inverse normal transformation of the model results, normalizing the model results, and / or utilizing, or only utilizing, a top quintile and a bottom quintile of the model results.
[0070] Identifying the at least one genetic variant at 180 may include identifying at least one genetic variant that is correlated to the at least one predicted phenotypic trait. Additionally or alternatively, the identifying at 180 may include identifying the at least one genetic variant from the genomic data for the at least one individual. The identifying at 180 may be performed in any suitable manner. As an example, the identifying at 180 may include utilizing a genetic association methodology to identify the at least one genetic variant that is correlated to the predicted phenotypic trait. Examples of the genetic association methodology include linear regression, logistic regression, linear regression with covariates, logistic regression with covariates, mixed-model generalized linear regression, mixed-model generalized logistic regression, quantitative trait locus analysis, and / or genome-wide association study analysis utilizing the at least one predicted phenotypic trait. Koiitch Romano Dascenzo Gates LLC Attorney Docket No. SML315PCT
[0071] In some examples, the identifying at 180 may include directly associating the proteomic data to the genomic data. In some examples, the identifying at 180 may include inferential^ associating the at least one predicted phenotypic trait to the genomic data, such as to permit and / or facilitate correlating the at least one predicted phenotypic trait to the at least one genetic variant for the at least one individual without, or without the need for, direct observation of the at least one predicted phenotypic trait in and / or by the at least one individual. Stated differently, methods 100 may permit and / or facilitate linking of organismal phenotypes to potentially causal genetic variation without, or prior to, expression and / or observation of the organismal phenotypes. Stated still differently, methods 100 may permit and / or facilitate performing the identifying at 180 without directly observing the at least one predicted phenotypic trait in the at least one individual and / or in any individual of the plurality of individuals.
[0072] Displaying the results at 190 may include displaying a correlation between the at least one genetic variant and the at least one predicted phenotypic trait and may be performed in any suitable manner. As an example, the displaying at 190 may include generating and / or displaying a Manhattan plot of the genomic data. In some such examples, the displaying at 190 also may include identifying, in and / or on the Manhattan plot, that at least one genetic variant that is correlated to the at least one predicted phenotypic trait.
[0073] Methods 100, or at least the applying at 160 and the identifying at 180, may be performed with, via, and / or utilizing an analysis device, which also may be referred to herein as a special-purpose analysis device. The analysis device may be programmed to perform any suitable step and / or steps of methods 100. Examples of the analysis device include a computing device, a special-purpose computing device, and / or an analysis device that includes the supervised machine learning model trained to perform the applying at 160. In some such examples, the supervised machine learning model may be trained with proteomic training data and corresponding phenotypic training data from a training population of individuals. In some such examples, the training population of individuals may differ from the plurality of individuals from which the genomic data and the proteomic data are obtained during the obtaining at 120 and the obtaining at 130, respectively. In some such examples, the corresponding phenotypic training data may be Koiitch Romano Dascenzo Gates LLC Attorney Docket No. SML315PCT
[0074] based, at least in part, on directly observed phenotypic traits within the training population of individuals.
[0075] B. Simplified Example
[0076] In a simplified example, the obtaining at 120 includes obtaining genomic data in the form of three single nucleotide polymorphisms (SNPs) for a plurality of individuals, such as is illustrated in Fig. 2. In Fig. 2, the diploid genotype for each SNP is indicated for each individual by two letters separated by a slash. For example, “A / T” indicates that the individual is heterozygous for this SNP and has both an “A” and a “T” allele at the given site on the genome.
[0077] In this example, the obtaining at 130 includes obtaining proteomic data in the form of protein abundance for 10 proteins for the plurality of individuals, a subset of which is illustrated in Fig. 3. Then, the applying at 160 includes applying a linear model with coefficients that are illustrated in Fig. 4 to arrive at the scores that are illustrated in Fig. 5 for each individual. These scores are indicative of, predictive of, or indicate a probability of, a specific and / or predetermined phenotypic trait for each individual.
[0078] For example, application of the linear model for Alice results in equation (1):
[0079] 131 (0.1 )+586(1 ,2)+346(1 ,3)+100(-1.9)+379(0.5)+984*(0.1 )+84(-0.8)+110(1.2) (1 )
[0080] Similarly, application of the linear model for Bob results in equation (2):
[0081] 128(0.1 )+570(1.2)+301 (1.3)+96(-1.9)+390(0.5)+975(0.1 )+28(-0.8)+192(1.2) (2)
[0082] Thus, the score from the linear model for Alice is 1328.8 and the score from the linear model for Bob is 1406.2. This is indicated in Fig. 5 together with the scores for the other individuals for which proteomic data and genomic data were obtained. It is noted that Prot9 and ProtIO were omitted from equations (1) and (2) (or have a coefficient of zero), as these two proteins are not included in the linear model. Stated differently, it is within the scope of the present disclosure that the obtaining at 130 may include obtaining protein abundance information for proteins that are not utilized during the applying at 160 and / or that are not predictive of the at least one predicted phenotypic trait. Koiitch Romano Dascenzo Gates LLC Attorney Docket No. SML315PCT
[0083] Subsequently, the identifying at 180 is performed. In this simplified example, a linear regression of phenotype predictor (i.e., the score for each individual as indicated in Fig. 5) is performed on genotype for each SNP. The results of this linear regression are illustrated in Figs. 6-8. As illustrated in Fig. 6, SNP1 lacks a strong association between genotype and score, with the linear regression providing a p-value for the linear regression of score on genotype of 0.75. As illustrated in Fig. 7, SNP2 shows a potentially weak association, with the linear regression providing a p-value of 0.06. However, Fig. 8 shows a strong association for SNP3, with the linear regression providing a p-value of 0.0009. Thus, this analysis suggests that SNP3 may be predictive of the predetermined phenotypic trait.
[0084] C. Historical Data Examples
[0085] Methods 100 also were applied to historical datasets of genomic data and proteomic data. The historical datasets included genetic variants known to be significant biomarkers for various diseases, and these examples were utilized to validate the ability of methods 100 to identify association between specific genetic variants and specific diseases. The results were displayed in Manhattan plots, and two examples are illustrated in Figs. 9-10. In particular, Fig. 9 plots the association for predicted dementia risk, while Fig. 10 plots the association for predicted risk of heart failure.
[0086] In these Figures, each point is one genetic variant, and the points are arranged along the X-axis according to the position of the genetic variant within the human genome. The Y-axis indicates the significance of association between the genetic variant and the predicted probability of the disease. Two dashed horizontal lines also are included, with the bottom dashed horizontal line indicating a generally accepted threshold for suggestive correlation (p-value of 10‘5) and the top dashed horizontal line indicating a generally accepted (and higher) threshold for significant correlation (p-value of 5x10-8).
[0087] As may be seen from the Figures, the analysis indicates that several genetic variants meet the threshold for suggestive correlation with each disease. Both Figures also include one peak that significantly exceeds the threshold for significant correlation. More specifically, and with reference to Fig. 9, the circled peak is centered around the ApoE gene, with the most significant SNP being rs429358, which is a known biomarker for dementia. In Fig. 10, the circled peak is centered around the IL1RL1 gene, with the Koiitch Romano Dascenzo Gates LLC Attorney Docket No. SML315PCT
[0088] most significant SNP being rs1420101 ; IL1RL1 is a known biomarker for heart failure. These results indicate that methods 100 do, in fact, permit identification of genetic associations with phenotypic traits, such as dementia risk and heart failure, utilizing proteomic data and genomic data and that the methods do not require direct observation of the phenotypic traits.
[0089] D. Illustrative Combinations and Additional Examples
[0090] This section describes additional aspects and features of the invented methods, presented without limitation as a series of paragraphs, some or all of which may be alphanumerically designated for clarity and efficiency. Each of these paragraphs can be combined with one or more other paragraphs, and / or with disclosure from elsewhere in this application in any suitable manner. Some of the paragraphs below expressly refer to and further limit other paragraphs, providing without limitation examples of some of the suitable combinations.
[0091] A1. A method of predicting phenotypic traits from proteomic data and correlating the phenotypic traits to genomic data, wherein the proteomic data includes information regarding protein abundance for a plurality of distinct proteins for each individual of a plurality of individuals, and further wherein the genomic data includes information regarding presence of a plurality of distinct genetic variants for each individual, the method comprising:
[0092] optionally obtaining genomic data for the plurality of individuals;
[0093] optionally obtaining proteomic data for the plurality of individuals;
[0094] applying a predictive model to the proteomic data to obtain model results that predict, or that predict a probability of, at least one predicted phenotypic trait that is correlated to the proteomic data for at least one individual of the plurality of individuals, wherein the applying the predictive model includes correlating the protein abundance of at least two proteins of the plurality of distinct proteins to the predicted phenotypic trait for the at least one individual; and
[0095] identifying, from the genomic data for the at least one individual, at least one genetic variant that is correlated to the at least one predicted phenotypic trait. Koiitch Romano Dascenzo Gates LLC Attorney Docket No. SML315PCT
[0096] A2. The method of paragraph A1 , wherein the method includes obtaining the genomic data and obtaining the proteomic data for each individual of the plurality of individuals.
[0097] A3. The method of any of paragraphs A1-A2, wherein the method includes identifying genomic data and proteomic data obtained from each individual.
[0098] A4. The method of any of paragraphs A1-A3, wherein the method includes generating a database of genomic data and proteomic data that identifies genomic data and proteomic data obtained from each individual.
[0099] A5. The method of any of paragraphs A1-A4, wherein the method further includes collecting a corresponding biological sample from each individual, wherein the obtaining the genomic data includes analyzing the corresponding biological sample from each individual to generate the genomic data, and further wherein the obtaining the proteomic data includes analyzing the corresponding biological sample from each individual to generate the proteomic data.
[0100] A6. The method of paragraph A5, wherein the corresponding biological sample includes at least one of a biological sample that includes both genomic and proteomic components, a bodily fluid, a blood plasma, and a blood serum.
[0101] A7. The method of any of paragraphs A1 -A6, wherein the obtaining the genomic data includes obtaining the genomic data from a database of genomic data.
[0102] A8. The method of any of paragraphs A1 -A7, wherein the obtaining the genomic data includes obtaining genomic data generated by at least one of genome sequencing, whole genome sequencing, DNA microarray analysis, and genotype imputation.
[0103] A9. The method of any of paragraphs A1 -A8, wherein the obtaining the genomic data includes performing at least one of genome sequencing, whole genome sequencing, DNA microarray analysis, and genotype imputation.
[0104] A10. The method of any of paragraphs A1-A9, wherein the obtaining the proteomic data includes obtaining the proteomic data from a database of proteomic data.
[0105] A11. The method of any of paragraphs A1-A10, wherein the obtaining the proteomic data includes obtaining proteomic data generated by at least one of top-down proteomics, bottom-up proteomics, liquid chromatography, mass spectroscopy, affinity purification, bioinformatics, and proteomic microarray analysis. Koiitch Romano Dascenzo Gates LLC Attorney Docket No. SML315PCT
[0106] A1 . The method of any of paragraphs A1-A11 , wherein the obtaining the proteomic data includes performing at least one of top-down proteomics, bottom-up proteomics, liquid chromatography, mass spectroscopy, affinity purification, bioinformatics, and proteomic microarray analysis.
[0107] A13. The method of any of paragraphs A1-A12, wherein, at least one of subsequent to the obtaining the genomic data and prior to the identifying, the method further includes filtering the genomic data.
[0108] A14. The method of paragraph A13, wherein the filtering the genomic data includes at least one of filtering by minor allele frequency, filtering by imputation quality, and filtering via Hardy-Weinberg equilibrium.
[0109] A15. The method of any of paragraphs A1-A14, wherein, at least one of subsequent to the obtaining the proteomic data and prior to the applying, the method further includes filtering the proteomic data.
[0110] A16. The method of any of paragraphs A1-A15, wherein the applying the predictive model includes applying a predetermined predictive model configured to predict phenotypic traits based upon the proteomic data.
[0111] A17. The method of any of paragraphs A1-A16, wherein the applying the predictive model includes applying a supervised machine learning model trained with proteomic training data and corresponding phenotypic training data.
[0112] A18. The method of paragraph A17, wherein the applying the predictive model further includes training the supervised machine learning model.
[0113] A19. The method of any of paragraphs A1-A18, wherein the applying the predictive model includes applying a linear model that includes corresponding weighting coefficients for the protein abundance of each protein of the plurality of distinct proteins, obtained from each individual, to obtain the model results.
[0114] A20. The method of any of paragraphs A1 -A19, wherein, prior to the identifying, the method further includes transforming the model results for each individual, optionally wherein the transforming includes at least one of a rank inverse normal transformation of the model results, normalization of the model results, and utilizing, or only utilizing, a top quintile and a bottom quintile of the model results. Koiitch Romano Dascenzo Gates LLC Attorney Docket No. SML315PCT
[0115] A21. The method of any of paragraphs A1-A20, wherein the identifying the at least one genetic variant includes utilizing a genetic association methodology to identify the at least one genetic variant that is correlated to the predicted phenotypic trait, optionally wherein the genetic association methodology includes at least one of linear regression, logistic regression, linear regression with covariates, logistic regression with covariates, mixed-model generalized linear regression, mixed-model generalized logistic regression, quantitative trait locus analysis, and genome-wide association study analysis utilizing the at least one predicted phenotypic trait.
[0116] A22. The method of any of paragraphs A1-A21, wherein the identifying the at least one genetic variant includes directly associating the proteomic data to the genomic data.
[0117] A23. The method of any of paragraphs A1-A22, wherein the identifying the at least one genetic variant includes inferentially associating the at least one predicted phenotypic trait to the genomic data.
[0118] A24. The method of any of paragraphs A1-A23, wherein the method further includes displaying a correlation between the at least one generic variant and the at least one predicted phenotypic trait.
[0119] A25. The method of paragraph A24, wherein the displaying includes displaying a Manhattan plot of the genomic data and identifying, in the Manhattan plot, the at least one genetic variant that is correlated to the at least one predicted phenotypic trait.
[0120] A26. The method of any of paragraphs A1-A25, wherein the method includes performing the identifying without directly observing the at least one predicted phenotypic trait at least one of in the at least one individual and in any individual of the plurality of individuals.
[0121] A27. The method of any of paragraphs A1-A26, wherein the method includes performing at least one of the applying the predictive model and the identifying the at least one genetic variant utilizing an analysis device, or a special-purpose analysis device, that is programmed to perform the method.
[0122] A28. The method of paragraph A27, wherein the analysis device includes, or is, a computing device. Koiitch Romano Dascenzo Gates LLC Attorney Docket No. SML315PCT
[0123] A29. The method of any of paragraphs A27-A28, wherein the analysis device includes a supervised machine learning model trained to perform the applying the predictive model, optionally wherein the supervised machine learning model is trained with proteomic training data and corresponding phenotypic training data from a training population of individuals, and further optionally wherein the corresponding phenotypic training data is based, at least in part, on directly observed phenotypic traits within the training population of individuals.
[0124] B1. An analysis device, or a special-purpose analysis device, programmed to perform any suitable step and / or steps of any of the methods of any of paragraphs A1-A29.
[0125] C1. Non-transitory computer-readable storage media including computerexecutable instructions that, when executed, direct an analysis device, or a specialpurpose analysis device, to perform any suitable step and / or steps of any of the methods of any of paragraphs A1 -A29.
[0126] Advantages, Features, and Benefits
[0127] The different embodiments and examples of the methods described herein provide several advantages over known solutions for associating phenotypic traits to genomic data. For example, illustrative embodiments and examples described herein permit identification of correlation and / or association between genomic data and phenotypic traits without the need to directly observe the phenotypic traits within the individual(s) from which the genomic data is obtained. As such, the methods described herein may be utilized to predict the phenotypic traits prior to expression of the phenotypic traits within the individual(s).
[0128] Additionally, and among other benefits, illustrative embodiments and examples described herein may be utilized to significantly expand the overall knowledge base with respect to potentially causal relationships between specific genetic variants and specific diseases by permitting these potentially causal relationships to be explored utilizing existing proteomic data and corresponding genomic data for which directly observed phenotypic traits may not be available.
[0129] Additionally, and among other benefits, illustrative embodiments and examples described herein utilize at least two proteins to predict the at least one predicted Koiitch Romano Dascenzo Gates LLC Attorney Docket No. SML315PCT
[0130] phenotypic trait. This analysis is sensitive to higher-order effects and / or interactions for protein abundance among the plurality of distinct proteins, thereby providing a high overall degree of prediction accuracy.
[0131] No known system or device can perform these functions. However, not all embodiments and examples described herein provide the same advantages or the same degree of advantage.
[0132] Conclusion
[0133] The disclosure set forth above may encompass multiple distinct examples with independent utility. Although each of these has been disclosed in its preferred form(s), the specific embodiments thereof as disclosed and illustrated herein are not to be considered in a limiting sense, because numerous variations are possible. To the extent that section headings are used within this disclosure, such headings are for organizational purposes only. The subject matter of the disclosure includes all novel and nonobvious combinations and subcombinations of the various elements, features, functions, and / or properties disclosed herein. The following claims particularly point out certain combinations and subcombinations regarded as novel and nonobvious. Other combinations and subcombinations of features, functions, elements, and / or properties may be claimed in applications claiming priority from this or a related application. Such claims, whether broader, narrower, equal, or different in scope to the original claims, also are regarded as included within the subject matter of the present disclosure.
Claims
Koiitch Romano Dascenzo Gates LLC Attorney Docket No. SML315PCTCLAIMS1. A method of predicting phenotypic traits from proteomic data and correlating the phenotypic traits to genomic data, wherein the proteomic data includes information regarding protein abundance for a plurality of distinct proteins for each individual of a plurality of individuals, and further wherein the genomic data includes information regarding presence of a plurality of distinct genetic variants for each individual, the method comprising:applying a predictive model to the proteomic data to obtain model results that predict at least one predicted phenotypic trait that is correlated to the proteomic data for at least one individual of the plurality of individuals, wherein the applying the predictive model includes correlating the protein abundance of at least two proteins of the plurality of distinct proteins to the predicted phenotypic trait for the at least one individual; and identifying, from the genomic data for the at least one individual, at least one genetic variant that is correlated to the at least one predicted phenotypic trait.
2. The method of claim 1 , wherein the applying the predictive model includes at least one of:(i) applying a predetermined predictive model configured to predict phenotypic traits based upon the proteomic data; and(ii) applying a supervised machine learning model trained with proteomic training data and corresponding phenotypic training data, optionally wherein the applying the predictive model further includes training the supervised machine learning model.
3. The method of any of claims 1 -2, wherein the applying the predictive model includes applying a linear model that includes corresponding weighting coefficients for the protein abundance of each protein of the plurality of distinct proteins, obtained from each individual, to obtain the model results.
4. The method of any of claims 1-3, wherein, prior to the identifying, the method further includes transforming the model results for each individual, optionally wherein the transforming includes at least one of a rank inverse normal transformation ofKoiitch Romano Dascenzo Gates LLC Attorney Docket No. SML315PCTthe model results, normalization of the model results, and only utilizing a top quintile and a bottom quintile of the model results.
5. The method of any of claims 1-4, wherein the identifying the at least one genetic variant includes utilizing a genetic association methodology to identify the at least one genetic variant that is correlated to the predicted phenotypic trait.
6. The method of claim 5, wherein the genetic association methodology includes at least one of linear regression, logistic regression, linear regression with covariates, logistic regression with covariates, mixed-model generalized linear regression, mixed-model generalized logistic regression, quantitative trait locus analysis, and genome-wide association study analysis utilizing the at least one predicted phenotypic trait.
7. The method of any of claims 1-6, wherein the identifying the at least one genetic variant includes directly associating the proteomic data to the genomic data, and further wherein the identifying the at least one genetic variant includes inferential^ associating the at least one predicted phenotypic trait to the genomic data.
8. The method of any of claims 1-7, wherein the method further includes displaying a correlation between the at least one generic variant and the at least one predicted phenotypic trait.
9. The method of claim 8, wherein the displaying includes displaying a Manhattan plot of the genomic data and identifying, in the Manhattan plot, the at least one genetic variant that is correlated to the at least one predicted phenotypic trait.
10. The method of any of claims 1-9, wherein the method includes performing the identifying without directly observing the at least one predicted phenotypic trait at least one of in the at least one individual and in any individual of the plurality of individuals.Koiitch Romano Dascenzo Gates LLC Attorney Docket No. SML315PCT11. The method of any of claims 1-10, wherein the method includes performing at least one of the applying the predictive model and the identifying the at least one genetic variant utilizing a special-purpose analysis device that is programmed to perform the method.
12. The method of claim 11 , wherein the special-purpose analysis device includes a supervised machine learning model trained to perform the applying the predictive model, wherein the supervised machine learning model is trained with proteomic training data and corresponding phenotypic training data from a training population of individuals, and further wherein the corresponding phenotypic training data is based, at least in part, on directly observed phenotypic traits within the training population of individuals.
13. The method of any of claims 1-12, wherein the method includes obtaining the genomic data and obtaining the proteomic data for each individual of the plurality of individuals.
14. The method of claim 13, wherein the method further includes collecting a corresponding biological sample from each individual, wherein the obtaining the genomic data includes analyzing the corresponding biological sample from each individual to generate the genomic data, and further wherein the obtaining the proteomic data includes analyzing the corresponding biological sample from each individual to generate the proteomic data.
15. The method of claim 13, wherein the obtaining the genomic data includes obtaining the genomic data from a database of genomic data, and further wherein the obtaining the proteomic data includes obtaining the proteomic data from a database of proteomic data.