Systems and methods for identifying biomarkers using continuous gene expression profiles

By converting non-continuous gene expression data into a continuous profile through projection and decomposition, the method addresses inconsistencies in existing technologies, enabling accurate biomarker identification and classification for patient conditions.

WO2025207529A1PCT designated stage Publication Date: 2025-10-02CEPHEID INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/021170
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-25
Filing Date
2025-03-24
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

Existing solutions for predicting patient conditions using gene expression data fail to account for fixed effects, stochastic nature, and random noise, and require strong distribution assumptions, leading to inconsistent and inaccurate results.

Method used

Convert non-continuous gene expression data into a continuous gene expression profile by projecting onto a reference group, decomposing into intrinsic mode functions, and selecting biomarkers based on these profiles to represent patient conditions.

Benefits of technology

Generates a continuous gene expression profile that filters out noise and bias, allowing for accurate identification of biomarkers indicative of patient conditions, enabling effective classification and development of medical assays.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025021170_02102025_PF_FP_ABST
    Figure US2025021170_02102025_PF_FP_ABST
Patent Text Reader

Abstract

Systems and methods for identifying biomarkers using continuous gene expression profiles are provided. In example aspects, non-continuous gene expression data for a patient is converted into a continuous gene expression profile. Genes indicative of a patient condition are selected based on the continuous gene expression profile. In an example, a selection model selects a plurality of combinations of genes based on the continuous gene expression profile, and a plurality of classification models are trained based on the selected genes. The results of the classification models are used to determine biomarkers.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEMS AND METHODS FOR IDENTIFYING BIOMARKERS USING CONTINUOUS GENE EXPRESSION PROFILES CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 569,518, filed March 25, 2024, the disclosure of which is hereby incorporated by reference in its entirety. BACKGROUND

[0002] Gene expression data has many uses, such as identifying patient conditions. However, gene expression data is not maintained in a uniform manner, as different platforms and operators may treat gene expression data differently. For example, gene expression count data obtained from RNA sequencing may be stochastic in nature such that it should not be treated as an absolute number. Some gene expression datasets may include normalized gene expression counts to account for differences in the sample read depth and the length of the gene whereas other datasets include unnormalized gene expression counts. Depending on how the gene expression counts are normalized, the data may display different patterns.

[0003] Additionally, existing solutions for predicting patient conditions using gene expression data suffer from several issues. For example, solutions that use normalized gene expression counts do not account for fixed effects (i.e., unobserved heterogeneity that is constant within an individual but different across groups), the stochastic nature of gene sequence reads, or random noise. In another example, solutions that use unnormalized gene expression counts do not take fixed effects into account and assume that prior means and overall trends of dispersion-mean dependence across different samples are comparable. Other statistical approaches to predicting patient conditions using gene expression data are parametric in nature, requiring strong assumptions about gene expression distributions. SUMMARY

[0004] In general terms, this disclosure is directed to systems and methods for identifying biomarkers using continuous gene expression profiles. In example aspects,non-continuous gene expression data is processed into a continuous gene expression profile. Biomarkers may be selected based on continuous gene expression profiles.

[0005] In a first aspect, a method for identifying biomarkers based on patient conditions is provided. Non-continuous gene expression data is projected onto an average gene expression data from a reference group. The projected gene expression data is converted into a function with local minima and maxima. The function with local minima and local maxima is decomposed into one or more intrinsic mode functions. A continuous gene expression profile is generated based on at least one of the one or more intrinsic mode functions. The continuous gene expression profile represents the non-continuous gene expression data as a continuous function. One or more biomarkers are selected based on the continuous gene expression profile. The one or more biomarkers are indicative of a patient condition.

[0006] In a second aspect, a system for identifying biomarkers based on patient conditions is provided. The system includes one or more processors and one or more computer readable storage devices storing data instructions. The data instructions, when executed by the one or more processors, cause the system to project non- continuous gene expression data onto an average gene expression data from a reference group, convert the projected gene expression data into a function with local minima and local maxima, decompose the function with local minima and local maxima into one or more intrinsic mode functions, generate a continuous gene expression profile based on at least one of the one or more intrinsic mode functions, and select one or more biomarkers based on the continuous gene expression profile. The continuous gene expression profile represents the non-continuous gene expression data as a continuous function. The one or more selected biomarkers are indicative of a patient condition.

[0007] In a third aspect, a non-transitory computer-readable medium is provided. The computer-readable medium stores data instructions that, when executed by one or more processors, cause the one or more processors to process non-continuous gene expression data into a continuous gene expression profile, and select one or more biomarkers based on the continuous gene expression profile. The continuous gene expression profile represents the non-continuous gene expression data as a continuous function. The one or more biomarkers are indicative of a patient condition.

[0008] This summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Fig.1 illustrates an embodiment of a gene analysis system.

[0010] Fig.2 illustrates an example of selecting biomarkers based on continuous gene expression profiles.

[0011] Fig.3 illustrates a flowchart of an example method for converting gene expression data into a continuous gene expression profile.

[0012] Fig.4 illustrates a flowchart of an example method for selecting biomarkers based on continuous gene expression profiles.

[0013] Fig.5 illustrates a graph of example gene expression data.

[0014] Fig.6 illustrates a graph of example gene expression data transformed into a function with local minima and maxima for ensemble empirical mode decomposition.

[0015] Fig.7 illustrates graphs of example intrinsic mode functions and a residual derived from transformed gene expression data.

[0016] Fig.8 illustrates a graph of energies calculated from example intrinsic mode functions.

[0017] Fig.9 illustrates a graph of an example continuous gene expression profile.

[0018] Fig.10 illustrates a graph of example averaged continuous gene expression profiles of a plurality of patient conditions.

[0019] Fig.11 illustrates an example embodiment of a computing device on which aspects of the present disclosure can be implemented. DETAILED DESCRIPTION

[0020] Various embodiments will be described in detail with reference to the drawings, wherein like reference numerals represent like parts and assemblies throughout the several views. Reference to various embodiments does not limit the scope of the claims attached hereto. Additionally, any examples set forth in thisspecification are not intended to be limiting and merely set forth some of the many possible embodiments for the appended claims.

[0021] As used herein, the term “including” as used herein should be read to mean “including, without limitation,” “including but not limited to,” or the like.

[0022] As briefly described above, embodiments of the present disclosure are directed to systems and methods for identifying biomarkers using continuous gene expression profiles. In example aspects, non-continuous gene expression data for a patient is converted into a continuous gene expression profile after filtering out fixed effects and random noise in the data. In further aspects, genes with expression differences between one or more patient conditions are selected based on continuous gene expression profiles. For example, selection models and classification models may be used to select the genes that are indicative of the patient condition based on the continuous gene expression profiles.

[0023] Turning now to FIG.1, a gene analysis system 100 is shown. In examples, the gene analysis system 100 converts gene expression data 122 into continuous gene expression profiles (CGEPs) 124. In further examples, the gene analysis system 100 selects biomarkers—i.e., genes that differentiate between patient conditions—based on the continuous gene expression profiles 124. In the illustrated embodiment, the gene analysis system 100 includes a gene analysis server 110 connected to a computing device 104 and a gene analysis database 120. In an example embodiment, the computing device 104, the gene analysis server 110, and the gene analysis database 120 are connected over one or more networks. For example, the computing device 104 may be connected to the gene analysis server 110 over the Internet, and the gene analysis server 110 may be connected to the gene analysis database 120 over a local area network. In examples, the gene analysis server 110 is a cloud server, and the gene analysis database 120 is a cloud database.

[0024] In the illustrated example, the gene analysis server 110 includes a gene data processor 112 and a gene selector 114. The gene selector 114 includes a selection model 116 and a classification model 118. The gene analysis database 120 stores data used by the gene analysis server 110. In the illustrated embodiment, the gene analysis database 120 includes gene expression data 122 and gene expression profiles 124.

[0025] The gene data processor 112 processes gene expression data 122 to generate one or more continuous gene expression profiles 124. The gene expression data 122 includes gene expression counts, which may be representative of patient conditions, such as if a patient is healthy, infected with a virus, infected with bacteria, or symptomatic but not infected.

[0026] In an example, the gene expression data 122 includes data for a plurality of patients. While examples described herein describe processing gene expression data 122 for a single patient to generate a continuous gene expression profile 124 for the patient, the gene data processor 112 may process gene expression data 122 for a plurality of patients to generate a plurality of continuous gene expression profiles 124, and the gene data processor 112 may use the plurality of continuous gene expression profiles to generate an average continuous gene expression profile for a patient condition.

[0027] In an embodiment, the gene data processor 112 converts the gene expression data 122 associated with a patient into a function with local minima and maxima and uses the function with local minima and maxima to generate a continuous gene expression profile 124 for the patient. In an example, the gene data processor 112 projects the patient’s gene expression data 122 onto a reference group. For example, the reference group may be healthy gene expression data 122. In an embodiment, the patient’s gene expression data is projected onto an average gene expression data 122 of one or more healthy patients—e.g., average gene expression counts for healthy patients. In an alternative embodiment, the patient’s gene expression data 122 is projected onto gene expression data 122 of an individual healthy patient. In such examples, the user 102 may select the healthy gene expression data 122 onto which the patient’s gene expression data 122 is projected.

[0028] In an example, an orthogonal portion (vN) of the patient’s gene expression data 122 (v) are calculated based on the projection onto the reference gene expression data 122 (a) according to the following equation, where the patient’s gene expression data 122 (v), the reference gene expression data 122 (a), and the orthogonal portion (vN) are vectors: 1) ^^ = ^^ ௩∙^ே − ^∙^ ∙ ^^

[0029] In this example, the function (V(x)) with local minima and maxima is produced by summing the components of the orthogonal portion (vN) for each gene index (x) in the patient’s gene122 according to the following equation: 2) ^^^^^^ = ∑௫^ୀ^ ^^ே[^^]

[0030] The function with local minima and maxima is then decomposed into intrinsic mode functions using ensemble empirical mode decomposition. In an example, the function with local minima and maxima is a continuous function. Examples of ensemble empirical mode decomposition are described in “Denoising traffic collision data using Ensemble Empirical Mode Decomposition (EEMD) and its application for constructing Continuous Risk Profile (CRP),” by Nam-Seog Kim et al., and published on May 28, 2014, by Elsevier Ltd., which is hereby incorporated by reference in its entirety.

[0031] In alternative embodiments, the gene expression data 122 is prepared for decomposition by cumulatively adding the gene expression data to create a non- decreasing function (C(d)), and using the non-decreasing function to create a transformed function (B(d)) that includes extrema based on the number of genes indexed (D) according to the following equation: 3) ^^^^^^ = ^^^^^^ − ^^^^^ × ^^

[0032] noise (wj(d)) is added to the transformed function (B(d)) to create a number (j) of artificial observations (Bj(d)), according to the following equation: 4) ^^^^^^^ = ^^^^^^ + ^^^(^^)

[0033] The noise requires two input parameters: the amplitude of the noise and the number of ensembles. In embodiments, the number of artificial observations depends on the number of ensembles. For example, the number of artificial observations may be greater than zero and less than the number of ensembles.

[0034] In an example, the number of ensembles is large enough to eventually cancel out the added noise and small enough to make the ensemble empirical mode decomposition computationally efficient. Similarly, in an example, the amplitude is large enough that the ensemble empirical mode decomposition can effectively separate the different frequencies embedded in the transformed function and small enough that it does not alter the original signal.

[0035] In an embodiment, the amplitude of the noise (An) and the number of ensembles (Nesb) are determined by comparing the input parameters with a value of a final standard deviation of error (^n) (i.e., a difference between an input signal and corresponding intrinsic mode functions), according to the following equation: 5) ^^^^^^ ^^^ + ଶ ^^^^^^^^^ = 0

[0036] Because the determination of the two input parameters based on the above equation requires a comparison with the standard deviation of error based on the intrinsic mode functions, multiple rounds of ensemble empirical mode decomposition may be performed to optimize the input parameters. During an initial round of ensemble empirical mode decomposition for which there is not a standard deviation of error to which the input parameters can be compared, the user 102 may set the initial input parameters. In an example, the amplitude of the noise is 0.2 standard deviations of the original signal and the number of ensembles is 500; however, these values may vary depending on the gene expression data 122 being analyzed.

[0037] In an embodiment, a sifting process is used to decompose the function with local minima and maxima into intrinsic mode functions. The process described herein describes decomposing the function with local minima and maxima into intrinsic mode functions; however, a similar process can be used to decompose artificial observations. In embodiments in which adding noise creates multiple artificial observations, the process can be repeated for each artificial observation.

[0038] The sifting process is initiated by constructing upper and lower boundaries for the function with local minima and maxima. For example, cubic spline interpolation may be used to determine the upper and lower boundaries. The mean of the upper and lower boundaries (m0(x)) is compared against the function with local minima and maxima (V(x)) to determine a first intrinsic mode function (c1(x)). The difference (h0(x)) between the function with local minima and maxima and the mean of the upper and lower boundaries becomes the first intrinsic mode function if two conditions are met: (1) a difference between a number of extrema and a zero-crossing equals zero or one, and (2) the mean value of the upper and lower boundaries is zero at any data point.

[0039] If these conditions are not met, upper and lower boundaries are determined for the difference (h0(x)). A new mean (m1(x)) is calculated based on the new upper andlower boundaries, and a new difference (h1(x)) is calculated as a difference between the first difference (h0(x)) and the new mean (m1(x)).

[0040] This process repeats until properties of a difference (hk(x)) satisfy the two conditions described above. Once the conditions are satisfied, the difference (hk(x)) becomes the first intrinsic mode function (c1(x)). The first intrinsic mode function is subtracted from the function with local minima and maxima to create a residual (r1(x)).

[0041] The decomposition process is repeated, determining further intrinsic mode functions (ci(x)) from residuals (ri(x)), until a final residual (R(x)) becomes or only one extremum remains. In an example, eight intrinsic mode functions are determined. The function with local minima and maxima (V(x)) is a function of the intrinsic mode functions (ci(x)) and the final residual (R(x)), according to the following equation: 6) ^^(^^) = ∑^^ୀ^ ^^^(^^) + ^^(^^)

[0042] are multiple artificial observations, theprocess artificial observation, and the intrinsic mode functions from each of the artificial observations are averaged respectively—e.g., the first intrinsic mode functions from each artificial observation are averaged, and so on.

[0043] The continuous gene expression profile 124 is generated based on the intrinsic mode functions. In an embodiment, energies are calculated for the intrinsic mode functions, and intrinsic mode functions are selected based on the energies to filter out noise and bias from the gene expression that is unique to the patient. For example, intrinsic mode functions that have low energies are treated as noise, and the residual is treated as bias. The following equation may be used to calculate a mean energy (Ei) for an intrinsic mode function (ci(k)) with a number (m) of data points: 7) ^^ = ^ ∑^ [^ ଶ^ ^ ^ୀ^ ^^(^^)]

[0044] the intrinsic mode functions are ranked based on the associated energies, and the intrinsic mode functions with associated energies that are greater than a threshold are selected to be used to generate the continuous gene expression profile 124. In an example, the threshold is 4% of the largest calculated energy. In another example, the threshold is set by the user 102. In alternative embodiments, the energies for the intrinsic mode functions are presented to the user102 on the computing device 104, and the user 102 manually selects which intrinsic mode functions on which to base the continuous gene expression profile 124.

[0045] The continuous gene expression profile 124 is generated based on the selected intrinsic mode functions. In an example, the continuous gene expression profile 124 is a sum of the selected intrinsic mode functions. Because the low energy intrinsic mode functions and the residual are filtered out, the continuous gene expression profile 124 includes the gene expression data of the patient substantially without noise or bias. Accordingly, continuous gene expression profiles 124 can be used to classify patients based on patient conditions and identify characteristic biomarkers that display consistent patterns among different patient conditions.

[0046] The gene selector 114 uses continuous gene expression profiles 124 to select biomarkers. In embodiments, continuous gene expression profiles 124 are compared to select genes that are indicative of patient conditions. For example, genes that are indicative of patient conditions may be identified by local minima and maxima in the continuous gene expression profiles 124 that are unique to the patient conditions. In some embodiments, averages of continuous gene expression profiles 124 for different patient conditions may be used to select the biomarkers.

[0047] In examples, the biomarkers selected by the gene selector 114 may be used to develop medical assays. Accordingly, the gene selector 114 may be limited in the number of biomarkers that can be selected. While selecting a higher number of biomarkers may increase the accuracy of the medical assay, selecting too many biomarkers for the medical may increase technical complexity, decrease assay performance and robustness, or increase cost. In example embodiments, the gene selector 114 may select between six and twenty biomarkers. In an example, the gene selector 114 selects nine biomarkers. In other examples, the gene selector 114 selects eighteen biomarkers.

[0048] In an example, the biomarkers chosen by the gene selector 114 represent a variety of characteristics in the continuous gene expression profile 124 that are associated with different biological pathways to minimize redundancy. If the biomarkers selected by the gene selector 114 are highly correlated with one another— for example, participating in the same biological pathway—the biomarkers may represent the same characteristics, and may therefore be redundant. Such genes may besubject to the same or similar regulation. For the reasons stated above, the number of biomarkers that are selected by the gene selector 114 may be limited. Accordingly, in some examples, the gene selector 114 may select biomarkers that are not highly correlated with one another. A subset of biomarkers (e.g., between six and nine biomarkers) may also be selected from a larger set of candidates (e.g., fifty candidates) to optimize performance for the technology to be used for the assay. For example, biomarkers may be selected to optimize for stability of the transcript, expression level above a threshold, transcript length or ease of primer and probe design.

[0049] In an embodiment, the gene selector 114 uses an iterative process to select the biomarkers. In an example, the iterative process includes three steps: gene selection, classification, and evaluation. In this process, the gene selection and classification steps may be repeated for a number of iterations, and after all iterations are complete, the results may be evaluated during the evaluation step.

[0050] In the illustrated embodiment, the gene selector 114 uses a selection model 116 to select genes in the gene selection step. In an embodiment, the selection model 116 includes a random selection model. For example, the random selection model may randomly select genes from the continuous gene expression profiles 124. In an example, the random selection model randomly selects nine genes. In alternative examples, the random selection model randomly selects between six and twenty genes. For example, in an embodiment, the random selection model may select eighteen genes. In further examples, any number of genes may be selected. The number of genes selected by the random selection model may be determined by the user 102.

[0051] In an alternative embodiment, the selection model 116 includes a clustering model. The clustering model clusters genes from the continuous gene expression profiles 124 to maximize differences between clusters such that similar values in the continuous gene expression profiles 124 are clustered together and different values fall into different clusters.

[0052] A representative gene may be selected from each cluster. In an example, the representative gene is selected based on correlations between the genes in the cluster and patient conditions. In an example, the representative gene has the highest correlation with patient conditions. In an alternative example, a neighborhood of the ten highest correlation values in the cluster is considered, and the representative gene isselected randomly from the neighborhood. In further embodiments, the representative gene is randomly selected from among all genes in the cluster. By randomly selecting the representative gene from the cluster, different solutions can be tested across multiple iterations of the biomarker selection process. Alternatively, the user 102 may select the representative gene.

[0053] The number of clusters used by the clustering model may be determined by the user 102. In an example, the number of clusters is determined based on the number of genes that the user 102 wants to be selected. In an example, the clustering model uses nine clusters. In some examples, the clustering model uses eighteen clusters. In alternative examples, the clustering model uses between six and twenty clusters. In further examples, any number of clusters may be used.

[0054] In an example, the clustering model is an unsupervised learning clustering model. Examples of unsupervised learning clustering models include k-means clustering models, hierarchical clustering models, and DBSCAN.

[0055] In an alternative embodiment, the selection model 116 includes a partial least squares discriminant analysis model. In an example, the partial least squares discriminant analysis model performs dimension reduction from a high-dimensional to a low-dimensional space of characteristic components that are representative of patient conditions and are orthogonal to each other. Genes that contribute to the orthogonal components represent a variety of different characteristics that may be associated with different patient conditions.

[0056] In an example, the genes are selected based on the contributions—e.g., loadings and explained variance. In examples, the nine genes with the highest contributions are selected. In some examples, the eighteen genes with the highest contributions are selected. In alternative examples, between six and twenty genes are selected. In further examples, any number of genes may be selected.

[0057] In further embodiments, the selection model 116 may include other models or algorithms. Additionally, in some embodiments, the selection model 116 may include multiple models. For example, the selection model 116 may include both a random selection model and a clustering model. Because the gene selection step may be performed multiple times during the iterative process, different models may be used during different iterations.

[0058] In some embodiments, the genes may be selected by a user, rather than the selection model 116. For example, a user may compare continuous gene expression profiles 124 associated with different patient conditions to determine which local minima and maxima in the continuous gene expression profiles 124 are unique to each patient condition.

[0059] After genes are selected by the selection model 116, the classification model 118 performs the classification step. In embodiments, the classification model 118 is trained to detect patient conditions based on the selected genes.

[0060] Performance of the classification model 118 can then be determined by classifying gene expression data 122 to calculate one or more evaluation statistics, such as accuracy, for the classification model 118 when using the selected genes to detect patient conditions. For examples, the classifications output by the classification model 118 may be compared against ground truth values to determine an accuracy for the classification model 118. In alternative embodiments, the classification model receives continuous gene expression profiles 124 as input and determines patient conditions based on the continuous gene expression profiles 124. In an embodiment, k-fold cross validation is used to evaluate the classification model 118.

[0061] In an embodiment, the classification model 118 includes a multinomial logistic regression model. In alternative embodiments, the classification model 118 includes a gaussian process model or a decision tree, such as an extreme gradient boosting decision tree. In further embodiments, the classification model 118 may include any machine learning model or statistical model. In some examples, the classification model 118 includes a plurality of models.

[0062] In some examples, the gene expression data 122 is compiled from multiple datasets, and the datasets may represent the gene expression data 122 differently. For example, a first dataset may include normalized gene expression data 122 while a second dataset may include gene expression data 122 that is not normalized. When the classification model 118 is evaluated using gene expression data 122, if the gene expression data 122 includes data from multiple datasets, the classification model 118 may include parameters that are fit on each dataset differently to enable adjustment to each specific dataset; however, the classification model 118 may still include the samemodel architecture and be trained on the same selected genes regardless of the dataset from which the gene expression data 122 originated.

[0063] As described above, the gene selector 114 may iterate through the gene selection step and the classification step multiple times so that multiple different combinations of potential biomarkers are tested. In an example, 10,000 iterations are performed by the gene selector 114. After the iterations are complete, the gene selector 114 evaluates the results of the iterations to select genes as biomarkers.

[0064] For example, the results of the iterations may be evaluated based on a calculated accuracy. In an embodiment, the genes that were used during the iteration in which the accuracy was the highest are selected as biomarkers. If multiple iterations are tied for the highest accuracy, the gene selector 114 may use tiebreakers to select the biomarkers. For example, the gene selector 114 may perform more testing rounds with the classification model 118 trained on the genes from the tied iterations until the iterations are no longer tied. In another example, the gene selector 114 may select any genes common among the tied iterations as biomarkers and use random selection from among the remaining genes in the tied iterations to select the rest of the biomarkers. In further examples, the genes from the tied iterations are ranked based on the frequency at which they appear in the tied iterations and selected based on the ranking. In other examples, biomarkers are randomly selected from the tied iterations, or one of the tied iterations is randomly selected and the genes from the selected iteration are used as biomarkers.

[0065] In an alternative embodiment, the gene selector 114 considers the genes selected during iterations in which the accuracy was above a threshold—e.g., 85% accuracy. In an example, the user 102 determine the threshold. The gene selector 114 then selects the most frequently used genes (e.g., the nine or eighteen most selected genes) from among the iterations in which the accuracy was above the threshold as the biomarkers.

[0066] If multiple genes are tied in frequency of use, the gene selector 114 may use tiebreakers to select the biomarkers. For example, the gene selector 114 may select the genes that are associated with higher accuracy iterations from among the tied genes. In another example, the gene selector 114 determines an average accuracy for the tied genes, as described below, and selects from among the tied genes based on the averageaccuracies. In further examples, the gene selector randomly selects from among the tied genes.

[0067] In a further embodiment, the gene selector 114 determines an average accuracy for each gene based on the accuracy of the iterations in which the gene was selected. For example, if a gene was selected in two iterations with accuracies of 90% and 80%, the average accuracy for the gene would be 85%. The gene selector 114 can then rank the genes based on the associated average accuracies and select biomarkers based on the ranking (e.g., the nine or eighteen genes with the highest average accuracies).

[0068] If multiple genes are tied for average accuracies, the gene selector 114 may use tiebreakers to select the biomarkers. In an example, the gene selector 114 selects the genes that were select most frequently from among the tied genes. In an alternative example, the gene selector 114 selects the genes that are associated with the highest accuracy iterations. In further examples, the gene selector 114 randomly selects from among the tied genes.

[0069] After biomarkers are selected by the gene selector 114, the biomarkers may be used to train models to determine patient conditions based on gene expression data 122. Similarly, in another example, medical assays can be developed based on the selected biomarkers. The biomarkers can be used in a nucleic acid amplification test to identify a patient as having a particular condition, such as a viral or bacterial infection.

[0070] In some embodiments, in addition to or alternative to using the continuous gene expression profiles 124 to select biomarkers, the continuous gene expression profiles 124 may be used to determine patient conditions. In an embodiment, the continuous gene expression profiles 124 are developed based on gene expression data 122 as described above. Once the continuous gene expression profiles 124 are developed, a classification model 118 is trained using the continuous gene expression profiles 124, rather than a selected set of biomarkers, to determine patient conditions. In an embodiment, the classification model 118 includes a random forest model. For example, the classification model 118 may include a random forest model with 1000 trees. In an example, the classification model 118 is trained using a dataset of gene expression profiles 124 that has been bootstrapped 25 times. In alternative embodiments, the classification model 118 includes any statistical model or machinelearning model. In further embodiments, the classification model 118 may include multiple models.

[0071] By using the continuous gene expression profiles 124 to classify patient conditions rather than the raw gene expression data 122, the classification model 118 does not need to include dataset specific parameters based on the dataset from which the gene expression data 122 is compiled. As described above, gene expression data 122 may be treated differently in different data sets (e.g., some datasets may normalize the gene expression data 122 while other datasets may not), so classification models 118 that receive gene expression data 122 may need specific parameters to handle different types of gene expression data 122. Because the gene analysis server 110 converts the gene expression data 122 into continuous gene expression profiles 124 that are the same regardless of how the initial gene expression data 122 is handled by the datasets from which the gene expression data 122 is compiled, the classification model 118 does not need to include different parameters for different datasets.

[0072] Fig.2 illustrates an example biomarker selection through multiple iterations of gene selection, classification, and evaluation. As described above, in each iteration, a selection model 116 uses the continuous gene expression profiles 124 to select genes 202 on which a classification model 118 is trained. In embodiments, the biomarkers that are selected are the selected genes 202 from the iteration in which the classification model 118 performed the best (i.e., the iteration in which the classification model 118 had the highest accuracy.

[0073] In the illustrated embodiment, three iterations are shown for ease of illustration; however, different numbers of iterations may be used in alternative embodiments (e.g., 10,000 iterations). In each iteration the selection model 116 selects genes 202 based on the continuous gene expression profiles 124. For example, the selection model 116 may select genes 202 based on similarity of gene counts and whether the gene is indicative of a patient condition. In an example, the selection model 116 includes a k-means clustering model. In the illustrated example, for ease of illustration, the selection model 116 selects three genes during each iteration; however, in alternative examples, different numbers of genes may be selected by the selection model 116 (e.g., nine or eighteen genes). In the first iteration, the selected genes 202a include genes “xyz,” “abc,” and “ytg.” In the second iteration, the selected genes 202binclude genes “abh,” “khs,” and “uys.” In the third iteration, the selected genes 202c include genes “yhy,” “jos,” and “uys.”

[0074] In each iteration, the selected genes 202 are used to train the classification model. In an example, the classification model 118 includes a logistic regression model. After the classification model 118 is trained on the selected genes 202, performance of the classification model 118 is determined. For example, an accuracy of the classification model 118 may be determined.

[0075] As described above, biomarkers may be selected based on the performance of the classification model 118 during different iterations. For example, the selected genes 202 that were used to train the classification model 118 during the iteration in which the accuracy of the classification model 118 is highest may be selected as biomarkers. In the illustrated embodiment, the classification model 118 had accuracies of 88%, 70%, and 95%, respectively, during the three iterations. Accordingly, in an example, the selected genes 202c that were used to train the classification model 118 during the third iteration (i.e., genes “yhy,” “jos,” and “uys”) may be selected as biomarkers, as the classification model 118 performed best when trained on those genes.

[0076] Turning to Fig.3, a flowchart of an example method 300 for processing gene expression data into a continuous gene expression profile is shown. In the illustrated embodiment, the method 300 includes operations 302, 304, 306, 308. In an embodiment, the method 300 is performed by a gene data processor.

[0077] The operation 302 includes projecting gene expression data onto average gene expression data for a reference condition. In an embodiment, the average gene expression data for the reference condition may include average gene expression counts for patients with the reference condition—e.g., healthy patients. As described above orthogonal portions of the gene expression data are determined by projecting the gene expression data onto the reference gene expression data. In some embodiments the average gene expression data is from healthy patients, but different reference groups may be used.

[0078] The operation 304 includes converting the projected gene expression data into a function with local minima and maxima. In an example, the orthogonal portions of the projected gene expression data are cumulatively added to create the function withlocal minima and maxima. In an example, the function with local minima and maxima is a continuous function.

[0079] The operation 306 includes decomposing the function with local minima and maxima into intrinsic mode functions and a residual. As described above, in examples, ensemble empirical mode decomposition is used to decompose the function.

[0080] The operation 308 includes generating a continuous gene expression profile based on the intrinsic mode functions. In an example, energies are calculated for each of the intrinsic mode functions. The energies of the intrinsic mode functions are compared to a threshold, and the intrinsic mode functions with energies greater than the threshold are selected. Alternatively, a user may evaluate the energies and select the intrinsic mode functions based on the energies. In an embodiment, the gene expression profile is a sum of the selected intrinsic mode functions.

[0081] Fig.4 illustrates a flowchart of an example method 400 for selecting biomarkers using gene expression profiles. In the illustrated embodiment, the method 400 includes operations 402, 404, 406, 408, 410. In an embodiment, the method 400 is performed by a gene selector with a selection model and a classification model.

[0082] The operations 402, 404, 406, 408 may repeat iteratively to determine accuracies for using different genes to determine patient conditions. The operation 402 includes selecting genes for a classifier. In an embodiment, a clustering algorithm is used to generate clusters of genes from continuous gene expression profiles, and a representative gene is selected from each cluster, as described above. In alternative embodiments, random selection or partial least squares discriminant analysis may additionally or alternatively be used. In an example, nine genes are selected. In other examples, other numbers of genes may be selected. For example, eighteen genes may be selected.

[0083] The operation 404 includes building a classifier based on the selected genes. In an example, a multinomial logistic regression model is trained based on the selected genes. Alternative examples of classifiers that can be used include gaussian process models, decision trees, and neural networks. In some embodiments, different types of classifiers are used during different iterations of the operation 404.

[0084] The operation 406 includes determining the performance of the classifier trained during the operation 504. In an embodiment, continuous gene expressionprofiles are input into the classifier, and the classifier predicts patient conditions based on the input continuous gene expression profiles. An accuracy of the classifier is determined based on the predictions—e.g., by comparing the predictions to ground truth values. In alternative examples, the classifier receives gene expression data as an input and predicts patient conditions based on the input gene expression data.

[0085] The operation 408 includes determining if more iterations are to be performed. In an example, 10,000 iterations are performed of the operations 402, 404, 406. In an embodiment, a user determines the number of iterations. If more iterations are to be performed, the method 400 returns to the operation 402, and the method 400 continues to loop. If all the iterations have been performed, the method 400 proceeds to the operation 410.

[0086] The operation 410 includes evaluating the performance of the classifiers to select biomarkers. In an example, the biomarkers are selected as the genes that were used in the best performing classifier—i.e., the classifier with the highest accuracy. In another example, genes that were used by the classifiers that performed better than a threshold (e.g., 85% accuracy) are considered, and the genes that were used most frequently among those classifiers are selected.

[0087] The selected biomarkers can then be used, for example, to develop medical assays. The identified biomarkers can be used to develop assays to measure an expression value for each of the selected biomarkers. The measurements may be normalized against one or more control targets that may be present in the sample or added during the assay. The assays may be nucleic acid amplification tests such as real- time PCR assays and may utilize a cartridge system such as developed by Cepheid, with its principal place of business in Sunnyvale, California. Depending on the format of the test (e.g., the system used for the test), it may interrogate 2 or more targets, such as between 2 and 30 targets, 2 to 25 targets, 2 to 20 targets, 2 to 15 targets, 2 to 9 targets, 5 to 30 targets, 5 to 20 targets, 5 to 8 targets, 6 to 20 targets, 6 to 15 targets, or 6 to 8 targets. The test may interrogate RNA or DNA targets. Larger numbers of markers may be used for array-based detection systems which can interrogate more than 20, more than 30, more than 50 or more than 100 targets in a single reaction volume. Individual target probes are spotted or attached to the array in known locations in an addressable array. The patient sample may be prepared for PCR amplification byutilization of a cartridges. The proposed approach narrows the search down from tens of thousands of genes in a microarray or RNASeq dataset to a few representative biomarkers that can distinguish between different disease categories (bacterial vs. viral vs. symptomatic but not infected vs. healthy). The outcome of the assay may be used to guide treatment decisions. For example, if the patient is determined to have a bacterial infection, antibiotics may be administered. If the patient is determined to have a viral infection and no bacterial infection the treatment selection may be antivirals or no treatment, resulting in reduced use of antibiotics for viral infections where they have no benefit and additional risk from antimicrobial resistance. The biomarker discovery approach can be easily extended to other types of biomarker types like proteins or single cell RNA-Seq. The methods may also be used to identify biomarkers for measuring drug efficacy from a dataset derived from patients treated with the investigational drug or the placebo and the outcome data for whether or not the drug was efficacious. The biomarkers can also be used to monitor the treatment of infectious diseases (e.g., targeted antibiotic use based on host response for tuberculosis infection or treatments of autoimmune or inflammatory diseases).

[0088] Turning now to Figs.5-10, an example of gene expression data being converted into a continuous gene expression profile is shown. Fig.5 illustrates a graph 500 showing an example of gene expression data (e.g., normalized gene count data). In the illustrated example, the gene expression data is non-continuous data—e.g., data for which the independent variable (e.g., the gene index in the illustrated example) is defined at a countable number of distinct values. The graph 500 shows expression counts for genes that were read. The counts of the genes in the gene expression data for the patient may vary based on conditions of the patient. For example, the gene expression counts may vary based on whether the patient is healthy, infected with a virus, infected with bacteria, or symptomatic but not infected. In examples, the expression counts are normalized based on a read depth and a length of the gene that was read. In alternative examples, the expression counts are not normalized.

[0089] Fig.6 illustrates a graph 600 showing the gene expression data from Fig.5 after the data is transformed into a function with local minima and maxima for ensemble empirical mode decomposition. In an example, as described above, to prepare the gene expression data for ensemble empirical mode decomposition, the geneexpression data is projected onto data from a reference group, such as healthy patients, and cumulatively summed to create a function with local minima and maxima. In alternative embodiments, the gene expression data is transformed into a non-decreasing function, and noise is added to create extrema.

[0090] Fig.7 illustrates a plurality of graphs 700 showing intrinsic mode functions and a residual decomposed from transformed gene expression data, such as the data shown in Fig.6. In an embodiment, the intrinsic mode functions and the residual are decomposed from transformed gene expression data using ensemble empirical mode decomposition, as explained above.

[0091] Fig.8 illustrates a graph 800 of the calculated energies for the intrinsic mode functions shown in Fig.7. As described above, intrinsic mode functions may be selected based on the calculated energies to generate a continuous gene expression profile. For example, the intrinsic mode functions with energies greater than a threshold may be selected to generate the continuous gene expression profile. In the illustrated example, intrinsic mode functions 5-8 may be selected as the associated energies are greater than a threshold (e.g., 4% of the maximum calculated energy).

[0092] Fig.9 illustrates an example graph 900 showing a continuous gene expression profile. The continuous gene expression profile represents the gene expression data as a continuous function. As described above, regardless of how the gene expression data is compiled (e.g., normalized or non-normalized), the continuous gene expression profile will be the same. This allows gene expression data from multiple datasets in which the gene expression data is compiled differently to be compared. Additionally, as described above, the continuous gene expression profile may be used to select biomarkers and classify patient conditions.

[0093] Fig.10 illustrates an example graph 1000 showing average continuous gene expression profiles from patients of various patient conditions. As can be seen in the graph 1000, the continuous gene expression profiles may have different extrema depending on the patient condition. These extrema, or other values in the continuous gene expression profiles, may be used to identify biomarkers and classify patient conditions. The peaks and valleys that vary between the conditions can be used to identify genes in the variable regions that can be used as biomarkers.

[0094] Fig.11 illustrates an example computing device 1100 on which aspects of the present disclosure may be implemented. The computing device 1100 can be used, for example, to implement computing devices such as the computing device 104, the gene analysis server 110, or any other computing device useable as described above in connection with Fig.1.

[0095] In the example of Fig.11, the computing device 1100 includes a memory 1102, a processing system 1104, a secondary storage device 1106, a network interface card 1108, a video interface 1110, a display unit 1113, an external component interface 1114, and a communication medium 1116. The memory 1102 includes one or more computer storage media capable of storing data and / or instructions. In different embodiments, the memory 1102 is implemented in different ways. For example, the memory 1102 can be implemented using various types of computer storage media, and generally includes at least some tangible media. In some embodiments, the memory 1102 is implemented using entirely non-transitory media.

[0096] The processing system 1104 includes one or more processing units, or programmable circuits. A processing unit is a physical device or article of manufacture comprising one or more integrated circuits that selectively execute software instructions. In various embodiments, the processing system 1104 is implemented in various ways. For example, the processing system 1104 can be implemented as one or more physical or logical processing cores. In another example, the processing system 1104 can include one or more separate microprocessors. In yet another example embodiment, the processing system 1104 can include an application-specific integrated circuit (ASIC) that provides specific functionality. In yet another example, the processing system 1104 provides specific functionality by using an ASIC and by executing computer-executable instructions.

[0097] The secondary storage device 1106 includes one or more computer storage media. The secondary storage device 1106 stores data and software instructions not directly accessible by the processing system 1104. In other words, the processing system 1104 performs an I / O operation to retrieve data and / or software instructions from the secondary storage device 1106. In various embodiments, the secondary storage device 1106 includes various types of computer storage media. For example, the secondary storage device 1106 can include one or more magnetic disks, magnetictape drives, optical discs, solid-state memory devices, and / or other types of tangible computer storage media.

[0098] The network interface card 1108 enables the computing device 1100 to send data to and receive data from a communication network. In different embodiments, the network interface card 1108 is implemented in different ways. For example, the network interface card 1108 can be implemented as an Ethernet interface, a fiber optic network interface, a wireless network interface (e.g., WiFi, WiMax, Bluetooth, etc.), or another type of network interface.

[0099] In optional embodiments where included in the computing device 1100, the video interface 1110 enables the computing device 1100 to output video information to the display unit 1113. The display unit 1113 can be various types of devices for displaying video information, such as an LCD display panel, a plasma screen display panel, a touch-sensitive display panel, an LED or OLED screen, a cathode-ray tube display, or a projector. The video interface 1110 can communicate with the display unit 1113 in various ways, such as via a Universal Serial Bus (USB) connector, a VGA connector, a digital visual interface (DVI) connector, an S-Video connector, a High- Definition Multimedia Interface (HDMI) interface, or a DisplayPort connector.

[0100] The external component interface 1114 enables the computing device 1100 to communicate with external devices. For example, the external component interface 1114 can be a USB interface and / or another type of interface that enables the computing device 1100 to communicate with external devices or peripheral devices integrated within the same housing (e.g., in the case of mobile devices). In various embodiments, the external component interface 1114 enables the computing device 1100 to communicate with various external components, such as external storage devices, input devices, speakers, modems, media player docks, other computing devices, scanners, digital cameras, and fingerprint readers.

[0101] The communication medium 1116 facilitates communication among the hardware components of the computing device 1100. The communication medium 1116 facilitates communication among the memory 1102, the processing system 1104, the secondary storage device 1106, the network interface card 1108, the video interface 1110, and the external component interface 1114. The communication medium 1116 can be implemented in various ways. For example, the communication medium 1116can include a PCI bus, a PCI Express bus, an accelerated graphics port (AGP) bus, a serial Advanced Technology Attachment (ATA) interconnect, a parallel ATA interconnect, a Fiber Channel interconnect, a USB bus, a Small Computing system Interface (SCSI) interface, or another type of communications medium.

[0102] The memory 1102 stores various types of data and / or software instructions. The memory 1102 stores a Basic Input / Output System (BIOS) 1118 and an operating system 1120. The BIOS 1118 includes a set of computer-executable instructions that, when executed by the processing system 1104, cause the computing device 1100 to boot up. The operating system 1120 includes a set of computer-executable instructions that, when executed by the processing system 1104, cause the computing device 1100 to provide an operating system that coordinates the activities and sharing of resources of the computing device 1100. Furthermore, the memory 1102 stores application software 1122. The application software 1122 includes computer-executable instructions, that when executed by the processing system 1104, cause the computing device 1100 to provide one or more applications. The memory 1102 also stores program data 1124. The program data 1124 is data used by programs that execute on the computing device 1100.

[0103] Although particular features are discussed herein as included within an electronic computing device 1100, it is recognized that in certain embodiments not all such components or features may be included within a computing device executing according to the methods and systems of the present disclosure. Furthermore, different types of hardware and / or software systems could be incorporated into such an electronic computing device.

[0104] In accordance with the present disclosure, the term computer readable media as used herein may include computer storage media and communication media. As used in this document, a computer storage medium is a device or article of manufacture that stores data and / or computer-executable instructions. Computer storage media may include volatile and nonvolatile, removable and non-removable devices or articles of manufacture implemented in any method or technology for storage of information, such as computer readable instructions, data structures, program modules, or other data. By way of example, and not limitation, computer storage media may include various types of dynamic random access memory (DRAM), solid state memory, read-only memory(ROM), electrically-erasable programmable ROM, magnetic disks (e.g., hard disks, floppy disks, etc.), and other types of devices and / or articles of manufacture that store data. Communication media may be embodied by computer readable instructions, data structures, program modules, or other data in a modulated data signal, such as a carrier wave or other transport mechanism, and includes any information delivery media. The term “modulated data signal” may describe a signal that has one or more characteristics set or changed in such a manner as to encode information in the signal. By way of example, and not limitation, communication media may include wired media such as a wired network or direct-wired connection, and wireless media such as acoustic, radio frequency (RF), infrared, and other wireless media.

[0105] It is noted that, in some embodiments of the computing device 1100 of Fig. 11, the computer-readable instructions are stored on devices that include non-transitory media. In particular embodiments, the computer-readable instructions are stored on entirely non-transitory media.

[0106] The various embodiments described above are provided by way of illustration only and should not be construed to limit the claims attached hereto. Those skilled in the art will readily recognize various modifications and changes that may be made without following the example embodiments and applications illustrated and described herein, and without departing from the full scope of the following claims.

Claims

WHAT IS CLAIMED IS:

1. A method for identifying biomarkers based on patient conditions, the method comprising: projecting non-continuous gene expression data onto an average gene expression data from a reference group; converting the projected gene expression data into a function with local minima and local maxima; decomposing the function with local minima and local maxima into one or more intrinsic mode functions; generating a continuous gene expression profile based on at least one of the one or more intrinsic mode functions, wherein the continuous gene expression profile represents the non- continuous gene expression data as a continuous function; and selecting one or more biomarkers based on the continuous gene expression profile, wherein the one or more biomarkers are indicative of a patient condition.

2. The method of claim 1, wherein converting the projected gene expression data into a function with local minima and local maxima includes cumulatively adding orthogonal portions of the projected gene expression data.

3. The method of claim 1, wherein generating the continuous gene expression profile based on at least one of the one or more intrinsic mode functions includes: calculating an energy for each of the one or more intrinsic mode functions; comparing the energies of the one or more intrinsic mode functions to a threshold; and generating the continuous gene expression profile based on the one or more intrinsic mode functions for which the energy is greater than the threshold.

4. The method of claim 1, wherein decomposing the function with local minima and local maxima into one or more intrinsic mode functions includes performing ensemble empirical mode decomposition on the function with local minima and local maxima.

5. The method of claim 1, wherein selecting one or more biomarkers based on the continuous gene expression profile includes:clustering genes in the continuous gene expression profile; selecting a representative gene from each cluster, wherein the selected representative gene is a biomarker.

6. The method of claim 1, wherein selecting one or more biomarkers based on the continuous gene expression profile includes performing partial least squares discriminant analysis of the continuous gene expression profile.

7. The method of claim 1, wherein selecting one or more biomarkers based on the gene expression profile includes comparing continuous gene expression profile to a second gene expression profile to identify local minima and local maxima that are variable between patient conditions.

8. The method of claim 1, wherein selecting one or more biomarkers based on the continuous gene expression profile includes: performing a predetermined number of evaluation cycles, each cycle including: selecting one or more genes from the continuous gene expression profile; building a classifier based on the selected one or more genes, wherein the classifier predicts a patient condition based on the selected one or more genes; and determining an accuracy of the classifier; and selecting the one or more genes on which the classifier with a highest accuracy was built as biomarkers.

9. The method of claim 8 wherein the classifier is one of a multinomial logistic regression model, a gaussian process model, or a decision tree.

10. The method of claim 8, wherein selecting one or more genes from the continuous gene expression profile includes: clustering genes in the continuous gene expression profile; and selecting a representative gene from each cluster.

11. The method of claim 8, wherein selecting one or more genes from the continuous gene expression profile includes applying dimension reduction to the continuous gene expression profile.

12. A system for identifying biomarkers based on patient conditions, the system comprising: one or more processors; and one or more computer-readable storage devices storing data instructions that, when executed by the one or more processors, cause the system to: project non-continuous gene expression data onto an average gene expression data from a reference group; convert the projected gene expression data into a function with local minima and local maxima; decompose the function with local minima and local maxima into one or more intrinsic mode functions; generate a continuous gene expression profile based on at least one of the one or more intrinsic mode functions, wherein the continuous gene expression profile represents the gene expression data as a continuous function; and select one or more biomarkers based on the continuous gene expression profile, wherein the one or more biomarkers are indicative of a patient condition.

13. The system of claim 12, wherein to select one or more biomarkers based on the continuous gene expression profile includes to: cluster genes in the continuous gene expression profile; and select a representative gene from each cluster, wherein the selected representative gene is a biomarker.

14. The system of claim 12, wherein to select one or more biomarkers based on the continuous gene expression profile includes to perform partial least squares discriminant analysis of the continuous gene expression profile.

15. A non-transitory computer-readable medium having stored thereon data instructions that, when executed by one or more processors, cause the one or more processors to: process non-continuous gene expression data into a continuous gene expression profile, wherein the continuous gene expression profile represents the non-continuous gene expression data as a continuous function; and select one or more biomarkers based on the continuous gene expression profile, wherein the one or more biomarkers are indicative of a patient condition.

16. The computer-readable medium of claim 15, wherein to process non-continuous gene expression data into a continuous gene expression profile includes to: at least partially remove bias and noise from the non-continuous gene expression data.

17. The computer-readable medium of claim 15, wherein to process non-continuous gene expression data into a gene expression profile include to: project the non-continuous gene expression data onto an average gene expression data; convert the projected gene expression data into a function with local minima and local maxima; decompose the function with local minima and local maxima into one or more intrinsic mode functions; and generate the continuous gene expression profile based on at least one of the one or more intrinsic mode functions.

18. The computer-readable medium of claim 17, wherein to convert the projected gene expression data into a function with local minima and local maxima includes to cumulatively add orthogonal portions of the projected gene expression data.

19. The computer-readable medium of claim 17, wherein to generate the continuous gene expression profile based on at least one of the one or more intrinsic mode functions includes to: calculate an energy for each of the one or more intrinsic mode functions; compare the energies of the one or more intrinsic mode functions to a threshold; andgenerate the continuous gene expression profile based on the one or more intrinsic mode functions for which the energy is greater than the threshold.

20. The computer-readable medium of claim 17, wherein to decompose the function with local minima and local maxima into one or more intrinsic mode functions includes to perform ensemble empirical mode decomposition on the function with local minima and maxima.

Citation Information

Patent Citations

  • US202463569518P