Systems and methods for identifying prognosis indicating isoforms
The system identifies prognosis indicating isoforms through multivariant analysis to enhance processing efficiency and provide accurate risk stratification for personalized cancer treatment recommendations.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- THE RGT UNIV OF MICHIGAN
- Filing Date
- 2025-10-27
- Publication Date
- 2026-05-07
AI Technical Summary
Existing gene-based prognostic scoring systems for cancer are ineffective in risk stratifying patients across various cancer types, particularly muscle invasive bladder cancer, and do not consider the full extent of gene mutations and expression changes associated with patient outcomes, leading to inefficient processing and limited utility in advanced cancers.
A system that identifies prognosis indicating isoforms by generating expression and correlation indicators for gene and isoform sequences, using multivariant analysis to refine a set of isoforms associated with patient survival, and generating prognostic scores for accurate risk stratification and personalized treatment recommendations.
The system provides improved processing efficiency and accurate risk stratification of patient survival, enabling personalized treatment approaches by identifying relevant isoforms rather than genes, reducing processing resources, and improving clinical outcomes.
Smart Images

Figure US2025052584_07052026_PF_FP_ABST
Abstract
Description
Docket No. 30275 / 70688 / PCSYSTEMS AND METHODS FOR IDENTIFYING PROGNOSIS INDICATING ISOFORMSCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of U.S. Application No. 63 / 713,649, filed October 30, 2024, and entitled “Systems and Methods for Identifying Prognosis Indicating Isoforms”, which is incorporated herein by reference in its entirety.FIELD
[0002] The present disclosure generally relates to systems for analyzing gene sequence data to identify features that are associated with subject survival of a target condition, and, more particularly, to systems and methods for identifying prognosis indicating isoforms of the gene sequence data that are associated with subject survival of a target condition.STATEMENT OF GOVERNMENT INTEREST
[0003] This invention was made with government support under CA259763, CA201335, and CA273138 awarded by the National Institutes of Health. The government has certain rights in the invention.BACKGROUND
[0004] Existing systems and methods for predicting disease progression utilize sequencing based prognostic algorithms that focus on expression of a subset of prognostic genes. For example, gene based prognostic scoring systems that help risk stratify and guide treatments of prostate, breast and other cancer types are known. However, these gene based prognostic scoring and classification systems suffer from several problems. First, these systems and methods have not proven effective at identifying statistically meaningful stratification of patient outcomes in every type of cancer or disease. In particular, gene based prognostic scoring has proven ineffective at risk stratifying patients with muscle invasive bladder cancer. Likewise, gene based prognostic scoring has limited utility in advanced or metastatic breast, prostate, and / or colon cancers. Second, the gene based prognostic scoring and classification systems typically identify a large number of input features that need to be examined and quantified for every patient, which can be burdensome on computing systems operating these models. Third, the gene based prognostic scoring and classification systems do not consider the full extent of revealed gene mutations and expression changes from molecular profiling of exome, transcriptome, and other molecular modalities that are associated with patient outcomes. For example, the gene based systems do not consider the effect of tumorDocket No. 30275 / 70688 / PC heterogeneity with respect to expression of alternative splicing events on tumor biology and patient outcomes.SUMMARY
[0005] In some aspects, the techniques described herein relate to a computer system including: one or more processors; and one or more non-transitory, computer- readable media storing instructions that, when executed by the one or more processors, cause the computer system to: obtain gene sequence data, the gene sequence data associated with a plurality of research subjects diagnosed with a target condition; generate a respective expression indicator for identified isoforms in the gene sequence data; generate gene correlation indicators for gene sequences in the gene sequence data, the gene correlation indicators indicating presence or absence of a gene level correlation with survival of the target condition for an associated subject of the plurality of research subjects; generate isoform correlation indicators for the identified isoforms, the isoform correlation indicators indicating presence or absence of an isoform level correlation with survival of the target condition for the associated subject; identify an initial set of prognosis indicating isoforms based on (1) the respective expression indicator for each identified isoform, (2) the isoform correlation indicators, and (3) the gene correlation indicators; identify a refined set of prognosis indicating isoforms based on a multivariant analysis of the initial set of prognosis indicating isoforms relative to survival of the target condition by the plurality of research subjects; generate prognostic scores for the plurality of research subjects based on the refined set of prognosis indicating isoforms; identify different survival risk groupings for the target condition that are linked to respective ranges of the prognostic scores; and save the survival risk groupings in a data store for use in selecting future patient treatment programs for the target condition.
[0006] In some aspects, the techniques described herein relate to a computer system including: one or more processors; and one or more non-transitory, computer- readable media storing instructions that, when executed by the one or more processors, cause the computer system to: receive patient gene sequence data taken from a patient diagnosed with a target condition; identify, in a data store, a refined set of prognosis indicating isoforms that are associated with the target condition; generate a respective expression indicator for the refined set of prognosis indicating isoforms that are present in the patient gene sequence data; generate a patient prognostic score using the respective expression indicator for the refined set of prognosis indicating isoforms present in the patient gene sequence data and isoform correlation indicators for the refined set of prognosis indicating isoforms that are stored in the data store; select one of a plurality of survival risk groupings for the target condition based onDocket No. 30275 / 70688 / PC the patient prognostic score; and present, on a display device, a recommended treatment program for the patient based on the selected one of the plurality of survival risk groupings.
[0007] In some aspects, the techniques described herein relate to a computer implemented method including: obtaining gene sequence data, the gene sequence data associated with a plurality of research subjects diagnosed with a target condition; generating a respective expression indicator for identified isoforms in the gene sequence data; generating gene correlation indicators for gene sequences in the gene sequence data, the gene correlation indicators indicating presence or absence of a gene level correlation with survival of the target condition for an associated subject of the plurality of research subjects; generating isoform correlation indicators for the identified isoforms, the isoform correlation indicators indicating presence or absence of an isoform level correlation with survival of the target condition for the associated subject; identifying an initial set of prognosis indicating isoforms based on (1) the respective expression indicator for each identified isoform, (2) the isoform correlation indicators, and (3) the gene correlation indicators; identifying a refined set of prognosis indicating isoforms based on a multivariant analysis of the initial set of prognosis indicating isoforms relative to survival of the target condition by the plurality of research subjects; generating prognostic scores for the plurality of research subjects based on the refined set of prognosis indicating isoforms; identifying different survival risk groupings for the target condition that are linked to respective ranges of the prognostic scores; and saving the survival risk groupings in a data store for use in selecting future patient treatment programs for the target condition.BRIEF DESCRIPTION OF THE DRAWINGS
[0008] The Figures described below depict various aspects of the system and methods disclosed therein. It should be understood that each Figure depicts an embodiment of a particular aspect of the disclosed system and methods, and that each of the Figures is intended to accord with a possible embodiment thereof. Further, wherever possible, the following description refers to the reference numerals included in the following Figures, in which features depicted in multiple Figures are designated with consistent reference numerals.
[0009] There are shown in the drawings arrangements which are presently discussed, it being understood, however, that the present embodiments are not limited to the precise arrangements and instrumentalities shown, wherein:
[0010] FIG. 1 illustrates a block diagram of a system for identifying prognosis indicating isoforms, in accordance with various embodiments disclosed herein.Docket No. 30275 / 70688 / PC
[0011] FIG. 2 illustrates a chart showing gene correlation indicators generate by the system of FIG. 1.
[0012] FIG. 3 illustrates a chart showing isoform correlation indicators generate by the system of FIG. 1.
[0013] FIG. 4 is a flow diagram of a method for identifying prognosis indicating isoforms, in accordance with various embodiments disclosed herein.
[0014] FIG. 5 is a flow diagram of a method for selecting a treatment program using prognosis indicating isoforms identified in accordance with various embodiments disclosed herein.
[0015] The Figures depict preferred embodiments for purposes of illustration only. Alternative embodiments of the systems and methods illustrated herein may be employed without departing from the principles of the invention described herein.DETAILED DESCRIPTION
[0016] The systems and methods described herein relate to bioinformatic tools to comprehensively describe the isoform expression within a subject cohort diagnosed with a particular target condition such as muscle invasive bladder cancer (MIBC). The systems and methods described herein then utilize this isoform expression to identify unique association with study subject outcomes where the cognate genes do not show any similar association. The systems and methods described herein also relate to a risk prediction tool which accurately stratifies patients into high and low-risk groups across different datasets using the identified isoforms. This accurate stratification of patient survival risk enables the systems and methods to recommend divergent treatments and patient monitoring programs as a function of a particular patient’s risk level instead of a one size fits all standard treatment.
[0017] For example, in the case of MIBC, all patients typically receive aggressive neoadjuvant therapy which provides a relatively modest clinical benefit in trials. However, the risk prediction tools described herein may identify patients eligible for a more conservative treatment approach that reduces the harshest side effects of treatment in patients with low risk of recurrent disease progression. Conversely, the risk prediction tools may also identify higher risk patients in need of additional monitoring and a more aggressive treatment approaches from the outset to improve overall survival.
[0018] Furthermore, the systems and methods described herein provide improved processing efficiency in identifying prognosis indicating isoforms and in computing prognostic scores for a patient using the identified prognosis indicating isoforms as compared to gene based prognostic scoring and classification systems. For example, the systemsDocket No. 30275 / 70688 / PC described herein when identifying prognosis indicating isoforms for bladder cancer identified 22 relevant isoforms as compared to 161 possibly relevant genes identified using similar processes. Because the prognosis score calculation includes terms for each relevant gene or isoform reducing the relevant terms from 161 genes to 22 isoforms provides improvements in processing efficiency while also producing a more consistently significant risk stratification for patients. Additionally, the systems and methods described herein may reduce processing resources by filtering potentially relevant isoforms out of the analysis at different stages of the process. For example, in some embodiments, only isoforms with an expression value above a preconfigured threshold and that occur on a gene with additional identified isoforms are further analyzed for correlation with subject survival data.
[0019] FIG. 1 shows a block diagram of a system 100 for identifying prognosis indicating isoforms and related subject survival risk groupings with respect to a target condition. The system 100 includes a computing system 102 such as a local server, remote cloud server, computer, tablet, etc. The computing system 102 may include a processing unit 104 and a memory unit 106.
[0020] Processing unit 104 includes one or more processors, each of which may be a programmable microprocessor or the like that executes software or other computing instructions stored in memory unit 106 to execute some or all of the functions of the system 100 as described herein. Processing unit 104 may include one or more graphics processing units (GPUs) and / or one or more central processing units (CPUs), for example. Alternatively, or in addition, one or more processors in processing unit 104 may be other types of processors (e.g., application- specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), etc.), and some of the functionality of the system 100 as described herein may instead be implemented in hardware.
[0021] Memory unit 106 may include one or more volatile and / or non-volatile memories. Any suitable memory type or types may be included in memory unit 106, such as read-only memory (ROM) and / or random access memory (RAM), flash memory, a solid-state drive (SSD), a hard disk drive (HDD), and so on. Collectively, memory unit 106 may store one or more software applications, the data received / used by those applications, and the data output / generated by those applications. In particular, the memory unit 106 may store instructions for specific modules of the computing system 102 as described herein.
[0022] As shown in FIG. 1, the processing unit 104 may obtain gene sequence data 108 from a data store data store 110 for processing and analysis by the computing system 102. In some embodiments, the processing unit 104 may obtain the gene sequence data 108 directlyDocket No. 30275 / 70688 / PC from a gene sequencing device or similar system that converts patient or subject samples into computer readable gene sequence data. In general, the gene sequence data 108 may be associated with a plurality of research subjects diagnosed with a target condition. For example, the gene sequence data 108 may include gene sequence data such as mRNA data, transcriptome data, phenotype data, clinical metadata files, Binary Alignment Map (BAM) files, etc for subjects previously diagnosed with the target condition. In some embodiments, the processing unit 104 may receive the gene sequence data 108 as an unaligned output file from a sequencing machine (e.g., a FASTQ file); process, sort, and align the unaligned file to the humane genome to generate a SAM file; and convert the SAM file into a binary format BAM file.
[0023] Furthermore, the target condition may include Bladder cancer (BLCA) and the gene sequence data 108 may include data as profiled by the Cancer Genome Atlas (TCGA) and retrieved from the database of Genotypes and Phenotypes (dbGaP) by the processing unit 104 for immediate use by the computing system 102 and / or for storage in the data store 110 for later processing and analysis. It should be appreciated that other target conditions besides BLCA may be analyzed using the computing system 102 in the manner described herein. These target conditions may include tumor types having a large number of samples with RNA-sequencing information including, but not limited to, the 33 tumor types in The Cancer Genome Atlas. It should also be appreciated that non-cancer related target conditions may be analyzed using the computing system 102 in the manner described herein.
[0024] The data store 110 may be implemented as a database, data lake, memory, or other digital storage medium known in the art. Accordingly, the data store 110 may be file system data store, an object-based data store, or other type of data store utilized in the art. Depending on the embodiment, the data store 110 may be implemented locally at the computing system 102, externally at an external data storage service, or a combination thereof. The computing system 102, via the processing unit 104, may be in wired or wireless communication with the external data storage service.
[0025] Once the processing unit 104 obtains the gene sequence data 108, the processing unit 104 may employ an expression quantification module 112 to generate gene expression indicators 113 for genes in the gene sequence data 108 and isoform expression indicators 114 for identified isoforms in the gene sequence data 108. In some embodiments, the expression quantification module 112 may execute a StringTie quantification algorithm to generate the gene expression indicators 113 and the isoform expression indicators 114. However, the expression quantification module 112 may also utilize other quantification tools or algorithmsDocket No. 30275 / 70688 / PC known in the art such as Salmon with Bias Correction vO.9.1, Cufflinks, etc. to generate the gene expression indicators 113 and the isoform expression indicators 114.
[0026] In some embodiments the expression quantification module 112 may generate the gene expression indicators 113 and the isoform expression indicators 114 with respect to reference sequence definitions for genes and unique isoforms. For example, the reference sequence definitions may include 69,272 unique isoforms and 25,522 gene definitions.
[0027] The gene expression indicators 113 may include a numerical quantification or other indicator of the degrees to which each gene in the gene sequence data 108 is expressed. Similarly, the isoform expression indicators 114 may include a numerical quantification or other indicator showing degrees to which each associated isoform identified in the gene sequence data 108 is expressed within an associated gene sequence of the gene sequence data 108. In some embodiments, the gene expression indicators 113 and the isoform expression indicators 114 can be normalized by subtracting the mean of the log2 of the transcript per million (TPM) value for all genes or isoforms in a given subject from the log2(TPM) value for a given gene or isoform in the same subject and dividing by the standard deviation.
[0028] In some embodiments, the processing unit 104 may convert the gene sequence data 108 into appropriate files types such as converting FASTQ files into BAM files as described herein. The processing unit 104 may also align the gene sequence data 108 to a human reference genome such as by aligning sequencing reads embodied in FASTQ files to the GRCh38 human reference genome to procedure SAM files as described herein.
[0029] Once the expression quantification module 112 generates the gene expression indicators 113 and the isoform expression indicators 114, the processing unit 104 may employ a univariant analysis module 116 to analyze the gene expression indicators 113 and the isoform expression indicators 114 relative to subject survival data 118. The subject survival data 118 includes data indicating the course of the target condition for each of the research subjects associated with the gene sequence data 108. For example, the subject survival data 118 may indicate whether a research subject died or survived the target condition, a severity of the condition at the time of diagnosis, a time line from diagnosis until death of the subject, subject specific factors (ethnicity, gender, age, etc.), etc.
[0030] As shown in FIG. 1, the univariant analysis module 116 may generate and output gene correlation indicators 120 and isoform correlation indicators 122. The gene correlation indicators 120 may indicate presence or absence of a gene level correlation between particular genes in the gene sequence data 108 and the subject survival data 118 for an associated subject of the plurality of research subjects. The isoform correlation indicators 122Docket No. 30275 / 70688 / PC may indicate presence or absence of an isoform level correlation between isoforms identified in the gene sequence data 108 and the subject survival data 118 for an associated subject of the plurality of research subjects.
[0031] In some embodiments, the univariant analysis module 116 may perform univariant cox regression analyses and log rank tests on the gene expression indicators 113 and the isoform expression indicators 114 and associated subject survival data 118 relative to control data to generate the gene correlation indicators 120 and the isoform correlation indicators 122. In some embodiments, the log rank tests utilize quantile expression thresholds of 25, 50 and 75. In these embodiments, the gene correlation indicators 120 and isoform correlation indicators 122 may include hazard ratio results of the univariant cox regression analysis, the p-value results of the univariant cox regression analysis, and / or the p- value results of the log rank test for a particular gene or isoform in the gene sequence data 108.
[0032] In some embodiments, the isoform expression indicators 114 may indicate that the isoform level correlation with survival of the target condition is present when (1) results of the univariant cox regression for the respective isoform indicate a hazard ratio greater than one and a p-value less than 0.05; and (2) results of any of the log rank tests for the respective isoform indicate a p-value less than 0.2 (e.g., any of the log-rank tests run at each quantile separation of the respective isoform). Similar thresholds for univariant cox regression and the log rank test may be used to determine that the gene expression indicators 113 indicate that the gene level correlation with survival of the target condition is present.
[0033] Example gene expression indicators 113 for the BLCA target condition are shown in FIG. 2 and example isoform expression indicators 114 for the BLCA target condition are shown in FIG. 3. In particular, FIG. 2 shows a chart 200 of the hazard ratio results 202 of the univariant cox regression and the associated cox p-value 204 and log rank p-value 206 for different genes in the gene sequence data 108. Similarly, FIG. 3 shows a chart 300 of the hazard ratio results 302 of the univariant cox regression and the associated cox p-value 304 and log rank p-value 306 for different isoforms in the gene sequence data 108.
[0034] With reference again to FIG. 1, the processing unit 104 may use an isoform selection module 124 to select an initial set of prognosis indicating isoforms 126 based on the gene correlation indicators 120, the isoform correlation indicators 122, and the isoform expression indicators 114. In some embodiments, the isoform selection module 124 may identify the initial set of prognosis indicating isoforms 126 by first selecting a group of the identified isoforms in the gene sequence data 108 based on the isoform expression indicators 114. Then, the isoform selection module 124 may identify candidate isoforms as isoforms inDocket No. 30275 / 70688 / PC the selected group of the identified isoforms for which the associated ones of the isoform correlation indicators 122 indicate that the isoform level correlation with the prognosis of the target condition is present (e.g., the hazard ratio results 302 is greater than 1, the cox p-value 304 is less than 0.05, and the log rank p-value 306 is less than 0.2).
[0035] Next, the isoform selection module 124 may select the initial set of prognosis indicating isoforms 126 as the candidate isoforms for which the gene correlation indicators 120 of the respective gene sequence on which the candidate isoforms are present indicates that the gene level correlation with the prognosis of the target condition is absent (e.g., at least one of the hazard ratio results 202 is less than 1, the cox p-value 204 is greater than 0.05, or the log rank p-value 206 is greater than 0.2). By excluding isoforms present on genes that demonstrate the gene level correlation with the subject survival data 118 from the initial set of prognosis indicating isoforms 126, the computing system 102 can provide a high confidence that any correlation to survival is a result of the presence of the isoforms in the initial set of prognosis indicating isoforms 126. It should be appreciated that isoform selection module 124 may perform these steps in different orders than described above to produce the initial set of prognosis indicating isoforms 126.
[0036] In some embodiments, the isoform selection module 124 may select the group of the identified isoforms based on the isoform expression indicators 114 by identifying a filtered set of the isoforms identified in the gene sequence data 108. This filtered set may be those isoforms that are associated with a gene sequence of the gene sequence data 108 for which two or more of the identified isoforms are present. The isoform selection module 124 may then select the isoforms in the filtered set of the identified isoforms that have a respective expression indicator 114 above a preconfigured threshold as the group of the identified isoforms. For example, in some embodiments, only isoforms with an expression value greater than 0.1*log2(TPM) may be included in the group of the identified isoforms.
[0037] In some embodiments, the group of isoforms may be identified before processing by the univariant analysis module 116 so that the isoform correlation indicators 122 are generated only for those isoforms that belong to a gene with two or more recognized isoforms, and those that had an expression of greater than 0.1 log2(TPM). Identifying the group of isoforms before processing by the univariant analysis module 116 may result in significant saving of processing time and resources by limiting the number of univariant cox regressions and log rank tests that need to be performed. For example, in analysis of the BLCA target condition using TCGA data as described herein, 69,272 individual isoforms may be initially identified and processed by the expression quantification module 112 toDocket No. 30275 / 70688 / PC generate the isoform expression indicators 114. These 69,272 individual isoforms may then be filtered down to 33,263 isoforms (e.g. approximately 48% of the original total) that belong to a gene with two or more recognized isoforms, and those that had an expression of greater than 0.1 log2(TPM). As such filtering the isoforms before executing the univariant analysis module 116 may utilize approximately 62% of the processing resources needed to calculate the isoform correlation indicators 122 for all 69,272 individual isoforms identified in the gene sequence data 108.
[0038] Once the isoform selection module 124 selects the initial set of prognosis indicating isoforms 126, the processing unit 104 may utilize a regularized multivariant analysis module128 to generate multivariant isoform correlation indicators 130. The multivariant analysis may further filter the initial set of prognosis indicating isoforms 126 to improve accuracy of the system and identify isoforms that have a true association to the subject survival data 118. In particular, the regularized multivariant analysis prevents overfitting of the data, avoids colinearity within the model (e.g., cases where two isoforms are effecting survivability via the same biological mechanism and, therefore, are linearly correlated or not independent), and prevents false-positive isoform identification by controlling for the effects of all included isoforms to identify those of the initial set of prognosis indicating isoforms 126 that are the most statistically relevant to the subject survival data 118.
[0039] For example, in some embodiments, the multivariant analysis module 128 may perform a LASSO regression to analyze effect of the isoform expression indicators 114 for each isoform in the initial set of prognosis indicating isoforms 126 relative to the subject survival data 118. The LASSO regression may be optimized using overall survival of the subjects as the dependent variable, and the normalized expression data of the isoform expression indicators 114 for each isoform in the initial set of prognosis indicating isoforms129 as the independent variables. The LASSO regression utilizes regularization to remove input features that are not important in modeling subject survival. In some embodiments, the LASSO regression may use a lambda parameter that is optimized for the input data using a grid search. In some embodiments, the optimized lambda parameter may include 0.01321941.
[0040] Once the multivariant analysis module 128 generates the multivariant isoform correlation indicators 130, the processing unit 104 may utilize the isoform selection module 124 to select a refined set of prognosis indicating isoforms 132 based on the multivariant isoform correlation indicators 130. For example, the isoform selection module 124 may select isoforms in the initial set of prognosis indicating isoforms 129 with a negative betaDocket No. 30275 / 70688 / PC coefficient in the results of the LASSO regression (e.g., the multivariant isoform correlation indicators 130) as the refined set of prognosis indicating isoforms 132.
[0041] In analysis of the BLCA target condition using TCGA data as described herein, the initial set of prognosis indicating isoforms 126 may include 34 isoforms and the refined set of prognosis indicating isoforms 132 may include 22 isoforms of those originally identified 34 isoforms. In particular, the refined set of prognosis indicating isoforms 132 for the BLCA target condition may include NM_001017425, NM_001149, NM_001164319, NM_001166215, NM_001171089, NM_001243773, NM_001276379, NM_001288570, NM_001317856_2, NM_001322027, NM_001323969, NM_001323970, NM_001328685, NM_001364608, NM_001365412, NM_001699, NM_006665, NM_032053, NM_057164, NR_045028, NR_109973, and NR_134658.
[0042] As shown in FIG. 1, the processing unit 104 may generate prognostic scores 134 for each of the plurality of research subjects based on the isoform expression indicators 114 and the isoform correlation indicators 122 for the refined set of prognosis indicating isoforms 132 using a prognostic scoring module 136. Each of the prognostic scores 134 may generally quantify a combined influence of the refined set of prognosis indicating isoforms 132 on an individual subject. In some embodiments, the prognostic scoring module 136 may calculate the prognostic scores 134 for each subject by summing together the products of at least one of the isoform correlation indicators 122 and the respective expression indicator 114 for each isoform in the refined set of prognosis indicating isoforms 132 present in an associated subject. In particular, the prognostic scores 134 may be determined using equation 1 below and defined as the summation of the beta coefficients ( ?,) from the univariate cox regression for a given isoform multiplied by a subjects normalized expression value (nevi for the refined set of prognosis indicating isoforms 132.Equation 1: prognostic score = Yn=22t * nevi
[0043] Once the prognostic scoring module 136 generates the prognostic scores 134, the processing unit 104 may identify survival risk groupings 138 for the target condition based on the prognostic scores 134 and the subject survival data 118 using a risk group generation module 140. The survival risk groupings 138 may be linked to respective ranges of the prognostic scores 134 that define different statistically relevant survival groupings for the plurality of subjects.Docket No. 30275 / 70688 / PC
[0044] In particular, the risk group generation module 140 may be configured to normalize all the prognostic scores 134 so that they fall between 0 and 1 according to equation 2 below, where minPscore is the minimum value in the determined prognostic scores 134 and maxPscore is the maximum value in the determined prognostic scores 134. After normalizing the prognostic scores 134, the risk group generation module 140 may test multiple different divisions of the normalized prognostic scores 134 (e.g., 2 groups, 3 groups, 4 groups, etc.) to identify a set of groups that is most significantly associated with different subject survival outcomes as indicated by the subject survival data 118. These groupings may include equal or unequal divisions of the normalized 0-1 range of the prognostic scores. For example, in the example case of the BLCA target condition as described herein, the risk group generation module 140 may identify 3 distinct survival risk groupings 138. These risk groupings may include a low risk grouping for normalized prognostic scores 134 less than 0.326, a medium risk grouping for normalized prognostic scores 134 between 0.326 and 0.716, and a high risk grouping for normalized prognostic scores greater than 0.716. In this embodiment, the medium risk grouping may be associated with a median survival of 24.11 months and the high risk grouping may be associated with a median survival of 8.33 months.„ > _T7. , „Equation 2: Normalized Score; =
[0045] The processing unit 104 may save the survival risk groupings 138 in the data store 110 for use in selecting future patient treatment programs for the target condition. The processing unit 104 may also link together the different survival risk groupings 138 with particular treatment regiments suitable for the different survival risk of the target condition. For example, in the case of the BLCA target condition as described herein, the low risk one of the survival risk groupings 138 can be associated with a recommended treatment program that includes a tumor removal surgery without accompanying radiation or chemotherapy. Additionally, patients assigned to the high-risk grouping may be selected for more intensive treatment with additional therapies including immunotherapy, radiation or increased surveillance. Patients with low-risk groupings may be recommended for lower intensity alternative therapy to surgery such as radiation or endoscopic resection. Finally, in other tumor types, high versus low risk groupings may be used to assign treatments of higher or lower intensity including chemotherapy, surgery or radiation.
[0046] The processing unit 104 may also present the survival risk groupings 138 on a display device 142 of the system 100. The display device 142 may include a computerDocket No. 30275 / 70688 / PC monitor or similar graphical display system known in the art that is operably connected to the computing system 102 via wired or wireless means. In some embodiments, the display device 142 may be part of a client or user device (e.g., a personal computer, mobile phone, tablet, etc.). The client device may be operatively coupled to the computing system 102 via wired or wireless means known in the art and may include a user interface. For example, the client device may execute a dedicated application, web browser, etc. as known in the art configured to interface with the computing system 102. In response to user interactions with the computing system 102, the processing unit 104 may provide the survival risk groupings 138 and / or other data associated with the computing system 102 to the client device for presentation on the display device 142. Furthermore, in some embodiments, user interaction with the client device may generate the association between the survival risk groupings 138 and the different patient treatment programs for the target condition.
[0047] In some embodiments, the computing system 102 may also be used to generate a patient prognostic score for a patient diagnosed with the target condition and present a recommended treatment program for the patient based on the generated prognostic score. In these embodiments, the processing unit 104 may receive patient gene sequence data taken from the patient. The patient gene sequence data may be received from the data store 110 or another device or system that converts a patient sample into computer readable gene sequence data. The processing unit 104 may then use the expression quantification module 112 to generate a respective expression indicator for the refined set of prognosis indicating isoforms 132 that are present in the patient gene sequence data.
[0048] The processing unit 104 may use the prognostic scoring module 136 to generate a patient prognostic score using the respective expression indicator for the refined set of prognosis indicating isoforms 129 present in the patient gene sequence data and at least one of the isoform correlation indicators 122 for the refined set of prognosis indicating isoforms 132. In particular, the prognostic scoring module 136 may employ equation 1 described above to generate the patient prognostic score.
[0049] Once the prognostic scoring module 136 has generated the prognostic score for the patient, the processing unit 104 may normalize the generated patient prognostic score using equation 2 as described above and the minimum and maximum values for the previously generated prognostic scores 134. The processing unit 104 may then use this normalized patient prognostic score to determine which of the survival risk groupings 138 the patient falls into and present, on the display device 142, the recommended treatment program for the patient that is associated with the identified survival risk grouping 138.Docket No. 30275 / 70688 / PC
[0050] FIG. 4 shows a method 400 for generating the refined set of prognosis indicating isoforms 132 and the survival risk groupings 138 using the computing system 102. The method 400 may be performed by the processing unit 104 executing instructions stored on the memory unit 106.
[0051] At block 410, the method 400 includes obtaining gene sequence data, the gene sequence data associated with a plurality of research subjects diagnosed with a target condition.
[0052] At block 420, the method 400 includes generating a respective expression indicator for identified isoforms in the gene sequence data.
[0053] At block 430, the method 400 includes generating gene correlation indicators for gene sequences in the gene sequence data. The gene correlation indicators indicating presence or absence of a gene level correlation with survival of the target condition for an associated subject of the plurality of research subjects.
[0054] At block 440, the method 400 includes generating isoform correlation indicators for the identified isoforms, the isoform correlation indicators indicating presence or absence of an isoform level correlation with survival of the target condition for the associated subject.
[0055] At block 450, the method 400 includes identifying an initial set of prognosis indicating isoforms based on (1) the respective expression indicator for each identified isoform, (2) the isoform correlation indicators, and (3) the gene correlation indicators.
[0056] At block 460, the method 400 includes identifying a refined set of prognosis indicating isoforms based on a multivariant analysis of the initial set of prognosis indicating isoforms relative to survival of the target condition by the plurality of research subjects.
[0057] At block 470, the method 400 includes generating prognostic scores for the plurality of research subjects based on the refined set of prognosis indicating isoforms.
[0058] At block 480, the method 400 includes identifying different survival risk groupings for the target condition that are linked to respective ranges of the prognostic scores.
[0059] At block 490, the method 400 includes saving the survival risk groupings in a data store for use in selecting future patient treatment programs for the target condition.
[0060] FIG. 5 shows a method 500 for selecting a treatment program using the refined set of prognosis indicating isoforms 132 and the survival risk groupings 138 using the computing system 102. The method 500 may be performed by the processing unit 104 executing instructions stored on the memory unit 106.
[0061] At block 510, the method 500 includes receiving patient gene sequence data taken from a patient diagnosed with the target condition.Docket No. 30275 / 70688 / PC
[0062] At block 520, the method 500 includes generating a respective expression indicator for the refined set of prognosis indicating isoforms 132 that are present in the patient gene sequence data.
[0063] At block 530, the method 500 includes generating a patient prognostic score using the respective expression indicator for the refined set of prognosis indicating isoforms 132 present in the patient gene sequence data and at least one of the isoform correlation indicators 122 for the refined set of prognosis indicating isoforms.
[0064] At block 540, the method 500 includes selecting one of the survival risk groupings 138 based on the patient prognostic score.
[0065] At block 550, the method 500 includes presenting, on the display device 142, a recommended treatment program for the patient based on the selected one of the survival risk groupings 138.
[0066] It is understood that the blocks of the methods 400 and 500 need not occur strictly in the order shown.
[0067] EXPERIMENTS AND VALIDATION
[0068] The configuration of the computing system 102 described herein and the clinical advantages of isoform based prognosis scoring and classification were tested with respect to different BLCA specimen datasets.
[0069] Initially, BLCA RNA Sequence data from 408 subjects was acquired from the TCGA dataset. This sequence data was then formatted for analysis by the computing system 102. The BLCA TCGA sequence data was processed by the expression quantification module 112 to generate gene expression indicators 113 and isoform expression indicators 114 for the BLCA TCGA sequence data using reference sequence data on 69,272 unique isoforms and 25,522 gene definitions. The StringTie v2.1.1 method was used and found to provide quantifications that had the strongest concordance of splice junction coverage and quantification (e.g., a median odds ratio greater than 25) as compared with other methods such as Salmon with Bias Correction vO.9.1, Cufflinks, etc.
[0070] The resulting isoform expression indicators 114 were then used to filter the identified unique isoforms in the BLCA TCGA sequence data down to 33,263 unique isoforms by removing isoforms where with an expression of less than 0.1 mean log2(TPM)) and those with only a single known isoform where cancer state specific alternative splicing events were not possible.
[0071] The gene expression indicators 113 for the BLCA TCGA sequence data and the isoform expression indicators 114 for the filtered group of 33,263 unique isoforms were thenDocket No. 30275 / 70688 / PC processed by the univariant analysis module 116 to generate gene correlation indicators 120 and isoform correlation indicators 122 for the BLCA TCGA sequence data. In particular, the univariant analysis module 116 used univariate Cox regression and log rank models as described herein.
[0072] The isoform selection module 124 then used the gene correlation indicators 120, the isoform correlation indicators 122 for the BLCA TCGA sequence data, and subject survival data for the 408 subjects to select an initial set of 34 prognosis indicating isoforms 126. The isoform selection module 124 selected the 34 isoforms as those for which the cox regression hazard ratio was greater than 1, the cox p-value was less than 0.05, the log rank statistic evaluated at any of the 3 quantile thresholds (25, 50 and 75) had a p-value less than 0.2, and gene-level expression of the parent gene could not be significantly associated with worse prognosis.
[0073] After quantifying isoform expression for the 69,272 isoforms in the BLCA TCGA sequence data using the expression quantification module 112, the isoform expression indicators 114 were validated by looking for corresponding read coverage across the unique splice junctions of previously validated isoforms. In particular, sequencing data from the highest expressing subject was aligned to a reference file containing every exon-exon junction among the gene’s isoforms. The results were then compared to isoform quantifications. From this comparison, the most differentiated splice junction for the 34 isoforms of interest were identified and then read coverage across the junctions for APLP2, FLNB, and TIAL1 was visualized to look for qualitative coverage to support isoform identification.
[0074] To further validate the 34 isoforms identified as the initial set of prognosis indicating isoforms 126, the APLP2, TIAL1, and FLNB isoforms were analyzed in more detail via lab experimentation and RT-PCR quantification. APLP2 and TIAL2 each showed two isoforms associated with prognosis; APLP2: NM_1328684 & NM_1328685 and TIAL1: NM_00 1323969 & NM_001323970. The lab experimentation confirmed that each of the APLP2 isoforms corresponded to splice variants lacking exon 14. The lab experimentation confirmed that the FLNB isoform (NM_001164319) corresponded to a unique exon 30 skipping event. The lab experimentation additionally confirmed that the TIAL1 prognostic isoform carried an alternative exon 2 which resulted in an alternative coding sequence start leading to loss of a carboxy terminal domain. This quantification and additional visual inspection confirmed the presence of the prognostic isoforms of APLP2, FLNB and TIAL1 isoforms in TCGA patient samples.Docket No. 30275 / 70688 / PC
[0075] After initial validation the 34 isoforms identified as the initial set of prognosis indicating isoforms 126 were further refined into a refined set of 22 isoforms (e.g., the refined set of isoforms 132) using the multivariant analysis module 128 and isoform selection module 124 as described herein. In particular, this process identified 22 isoforms (NM_001017425, NM_001149, NM_001164319, NM_001166215, NM_001171089, NM_001243773, NM_001276379, NM_001288570, NM_001317856_2, NM_001322027, NM_001323969, NM_001323970, NM_001328685, NM_001364608, NM_001365412, NM_001699, NM_006665, NM_032053, NM_057164, NR_045028, NR_109973, NR_134658) as negatively correlated with survival.
[0076] The prognostic scoring module 136 and risk group generation module 140 were then utilized to generate prognostic scores 134 and survival risk groupings 138 for the 408 subjects associated with the BLCA TCGA sequence data. The resulting survival risk groupings 138 split the subject into three groups: patients with a prognostic score less than 0.326 (low risk), patients with a prognostic score between 0.326 and 0.716 (medium risk), and samples with a prognostic score greater than 0.716 (high risk). Overall survival was then compared across these groups via Kaplan-Meier analysis. These 3 cohorts had markedly different median overall survival with a median survival of 24.11 months, and 8.33 months in the medium and high risk groups, respectively. Median survival could not be calculated for the low risk group, as survival probability remained at -75% past 160 months (p-value = 2.07e-25). To confirm that this effect was not unique to TCGA BLCA cohort or due to overfitting, the same scoring method was applied to RNA seq data from the IMVigor 210 cohort as a confirmation set. This analysis again showed significant differences in survival across the same three risk stratified groups with a median survival of 14.75, 9.00, and 3.17 months in the low, medium, and high risk groups, respectively (p-value = 1.17e-03).
[0077] The prognostic scores 134 and survival risk groupings 138 were additionally analyzed to determine independence from possible confounding variables, such as molecular subtyping, disease stage, age, sex, ethnicity, and race. This analysis found that prognostic scores 134 and survival risk groupings 138 retained prognostic importance irrespective of these variables. Overall, these results confirmed that the identified prognostic isoforms had independent prognostic value from gene level classifiers.
[0078] Finally, to further validate effectiveness the isoform approach was compared to a gene-based approach. In particular, a similar protocol to that described herein for the computing system 102 was utilized to develop a gene-based system. This system identified 600 genes as associated with worse prognosis in the BLCA TCGA data using similar metricsDocket No. 30275 / 70688 / PC as those applied in our isoform analysis described herein. All 600 genes were modeled via LASSO to select important features, and 161 ultimately retained a negative beta coefficient and were therefore included in the gene-based prognostic score. The TCGA cohort was then split into three groups based on the gene-based prognostic score: patients with a prognostic score less than 0.288 (low risk), patients with a prognostic score between 0.288 and 0.506 (medium risk), and samples with a prognostic score greater than 0.506 (high risk). Stratifying patients by a gene-based prognostic score was moderately effective, with a significant difference in overall survival and a median survival of 104.57 months in the low risk group, 64.76 months in the medium risk group, and 17.97 months in the high risk group (p-value = 1.085376e-13) (Sup Fig 7A). However, when applied to the IMVIGOR 210 dataset, the difference in survival does not reach significance, with a median overall survival of 8.08 months in the low risk group, 7.71 months in the medium risk group, and 13.27 months in the high risk group (p-value = 0.1369256) (Sup Fig 7B). Taken together, these results show that our isoform-based prognostic scoring method is more effective at stratifying patients by overall risk than traditional methods, while also using significantly fewer input features.
[0079] ADDITIONAL CONSIDERATIONS
[0080] Throughout this specification, plural instances may implement components, operations, or structures described as a single instance. Although individual operations of one or more methods are illustrated and described as separate operations, one or more of the individual operations may be performed concurrently, and nothing requires that the operations be performed in the order illustrated. Structures and functionality presented as separate components in example configurations may be implemented as a combined structure or component. Similarly, structures and functionality presented as a single component may be implemented as separate components. These and other variations, modifications, additions, and improvements fall within the scope of the subject matter herein.
[0081] The systems and methods described herein are directed to an improvement to computer functionality, and improve the functioning of conventional computers. Additionally, certain embodiments are described herein as including logic or a number of routines, subroutines, applications, or instructions. These may constitute either software (e.g., code embodied on a non-transitory, machine-readable medium) or hardware. In hardware, the routines, etc., are tangible units capable of performing certain operations and may be configured or arranged in a certain manner. In example embodiments, one or more computer systems (e.g., a standalone, client or server computer system) or one or more hardware modules of a computer system (e.g., a processor or a group of processors) may be configuredDocket No. 30275 / 70688 / PC by software (e.g., an application or application portion) as a hardware module that operates to perform certain operations as described herein.
[0082] In various embodiments, a hardware module may be implemented mechanically or electronically. For example, a hardware module may comprise dedicated circuitry or logic that is permanently configured (e.g., as a special-purpose processor, such as a field programmable gate array (FPGA) or an application- specific integrated circuit (ASIC)) to perform certain operations. A hardware module may also comprise programmable logic or circuitry (e.g., as encompassed within a general-purpose processor or other programmable processor) that is temporarily configured by software to perform certain operations. It will be appreciated that the decision to implement a hardware module mechanically, in dedicated and permanently configured circuitry, or in temporarily configured circuitry (e.g., configured by software) may be driven by cost and time considerations.
[0083] Accordingly, the term “hardware module” should be understood to encompass a tangible entity, be that an entity that is physically constructed, permanently configured (e.g., hardwired), or temporarily configured (e.g., programmed) to operate in a certain manner or to perform certain operations described herein. Considering embodiments in which hardware modules are temporarily configured (e.g., programmed), each of the hardware modules need not be configured or instantiated at any one instance in time. For example, where the hardware modules include a general-purpose processor configured using software, the general-purpose processor may be configured as respective different hardware modules at different times. Software may accordingly configure a processor, for example, to constitute a particular hardware module at one instance of time and to constitute a different hardware module at a different instance of time.
[0084] Hardware modules can provide information to, and receive information from, other hardware modules. Accordingly, the described hardware modules may be regarded as being communicatively coupled. Where multiple of such hardware modules exist contemporaneously, communications may be achieved through signal transmission (e.g., over appropriate circuits and buses) that connect the hardware modules. In embodiments in which multiple hardware modules are configured or instantiated at different times, communications between such hardware modules may be achieved, for example, through the storage and retrieval of information in memory structures to which the multiple hardware modules have access. For example, one hardware module may perform an operation and store the output of that operation in a memory device to which it is communicatively coupled. A further hardware module may then, at a later time, access the memory device to retrieve and processDocket No. 30275 / 70688 / PC the stored output. Hardware modules may also initiate communications with input or output devices, and can operate on a resource (e.g., a collection of information).
[0085] The various operations of example methods described herein may be performed, at least partially, by one or more processors that are temporarily configured (e.g., by software) or permanently configured to perform the relevant operations. Whether temporarily or permanently configured, such processors may constitute processor-implemented modules that operate to perform one or more operations or functions. The modules referred to herein may, in some example embodiments, comprise processor-implemented modules.
[0086] Similarly, the methods or routines described herein may be at least partially processor-implemented. For example, at least some of the operations of a method may be performed by one or more processors or processor-implemented hardware modules. The performance of certain of the operations may be distributed among the one or more processors, not only residing within a single machine, but deployed across a number of machines. In some example embodiments, the processor or processors may be located in a single location (e.g., within a home environment, an office environment or as a server farm), while in other embodiments the processors may be distributed across a number of locations.
[0087] The performance of certain of the operations may be distributed among the one or more processors, not only residing within a single machine, but deployed across a number of machines. In some example embodiments, the one or more processors or processor- implemented modules may be located in a single geographic location (e.g., within a home environment, an office environment, or a server farm). In other example embodiments, the one or more processors or processor-implemented modules may be distributed across a number of geographic locations.
[0088] It should also be understood that, unless a term is expressly defined in this patent using the sentence “As used herein, the term ‘ ’ is hereby defined to mean...” or a similar sentence, there is no intent to limit the meaning of that term, either expressly or by implication, beyond its plain or ordinary meaning, and such term should not be interpreted to be limited in scope based upon any statement made in any section of this patent (other than the language of the claims). To the extent that any term recited in the claims at the end of this disclosure is referred to in this disclosure in a manner consistent with a single meaning, that is done for sake of clarity only so as to not confuse the reader, and it is not intended that such claim term be limited, by implication or otherwise, to that single meaning.
[0089] Unless specifically stated otherwise, discussions herein using words such as “processing,” “computing,” “calculating,” “determining,” “presenting,” “displaying,” or theDocket No. 30275 / 70688 / PC like may refer to actions or processes of a machine (e.g., a computer) that manipulates or transforms data represented as physical (e.g., electronic, magnetic, or optical) quantities within one or more memories (e.g., volatile memory, non-volatile memory, or a combination thereof), registers, or other machine components that receive, store, transmit, or display information.
[0090] As used herein any reference to “one embodiment” or “an embodiment” means that a particular element, feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment. The appearances of the phrase “in one embodiment” in various places in the specification are not necessarily all referring to the same embodiment.
[0091] As used herein, the terms “comprises,” “comprising,” “includes,” “including,” “has,” “having” or any other variation thereof, are intended to cover a non-exclusive inclusion. For example, a process, method, article, or apparatus that comprises a list of elements is not necessarily limited to only those elements but may include other elements not expressly listed or inherent to such process, method, article, or apparatus. Further, unless expressly stated to the contrary, “or” refers to an inclusive or and not to an exclusive or. For example, a condition A or B is satisfied by any one of the following: A is true (or present) and B is false (or not present), A is false (or not present) and B is true (or present), and both A and B are true (or present).
[0092] In addition, use of the “a” or “an” are employed to describe elements and components of the embodiments herein. This is done merely for convenience and to give a general sense of the description. This description, and the claims that follow, should be read to include one or at least one and the singular also may include the plural unless it is obvious that it is meant otherwise.
[0093] Upon reading this disclosure, those of skill in the art will appreciate still additional alternative structural and functional designs through the principles disclosed herein. Therefore, while particular embodiments and applications have been illustrated and described, it is to be understood that the disclosed embodiments are not limited to the precise construction and components disclosed herein. Various modifications, changes and variations, which will be apparent to those skilled in the art, may be made in the arrangement, operation and details of the method and apparatus disclosed herein without departing from the spirit and scope defined in the appended claims.
[0094] The patent claims at the end of this patent application are not intended to be construed under 35 U.S.C. § 112(f) unless traditional means-plus-function language isDocket No. 30275 / 70688 / PC expressly recited, such as “means for” or “step for” language being explicitly recited in the claim(s).
Claims
Docket No. 30275 / 70688 / PCWhat is claimed is:
1. A computer system comprising: one or more processors; and one or more non-transitory, computer-readable media storing instructions that, when executed by the one or more processors, cause the computer system to: obtain gene sequence data, the gene sequence data associated with a plurality of research subjects diagnosed with a target condition; generate a respective expression indicator for identified isoforms in the gene sequence data; generate gene correlation indicators for gene sequences in the gene sequence data, the gene correlation indicators indicating presence or absence of a gene level correlation with survival of the target condition for an associated subject of the plurality of research subjects; generate isoform correlation indicators for the identified isoforms, the isoform correlation indicators indicating presence or absence of an isoform level correlation with survival of the target condition for the associated subject; identify an initial set of prognosis indicating isoforms based on (1) the respective expression indicator for each identified isoform, (2) the isoform correlation indicators, and (3) the gene correlation indicators; identify a refined set of prognosis indicating isoforms based on a multivariant analysis of the initial set of prognosis indicating isoforms relative to survival of the target condition by the plurality of research subjects; generate prognostic scores for the plurality of research subjects based on the refined set of prognosis indicating isoforms; identify different survival risk groupings for the target condition that are linked to respective ranges of the prognostic scores; and save the survival risk groupings in a data store for use in selecting future patient treatment programs for the target condition.
2. The computer system of claim 1 wherein to identify the initial set of prognosis indicating isoforms, the instructions, when executed by the one or more processors, cause the computer system to: select a group of the identified isoforms based on the respective expression indicator for each identified isoform in the subject gene sequence data;Docket No. 30275 / 70688 / PC identify candidate isoforms as isoforms in the selected group of the identified isoforms for which the associated isoform correlation indicator indicates that the isoform level correlation with the prognosis of the target condition is present; and select the initial set of prognosis indicating isoforms as the candidate isoforms for which the generated gene correlation indicator of the respective gene sequence on which the candidate isoforms is present indicates that the gene level correlation with the prognosis of the target condition is absent.
3. The computer system of claim 2 wherein to select the group of the identified isoforms based on the respective expression indicator for each identified isoform, the instructions, when executed by the one or more processors, cause the computer system to: identify a filtered set of the identified isoforms that are associated with a gene sequence of the gene sequence data for which two or more of the identified isoforms are present; and select isoforms in the filtered set of the identified isoforms with a respective expression indicator above a preconfigured threshold as the group of the identified isoforms.
4. The computer system of any one of claims 1-3 wherein the respective expression indicator includes a normalized quantification of a degree to which each of the identified isoforms is expressed within an associated gene sequence of the gene sequence data and wherein to generate the respective expression indicator, the instructions, when executed by the one or more processors, cause the computer system to: align the gene sequence data to a human reference genome; and process the aligned gene sequence data with a StringTie quantification algorithm to generate the respective expression indicator.Docket No. 30275 / 70688 / PC5. The computer system of any one of claims 1-4 by the one or more processors, cause the computer system to: perform a univariant cox regression for each of the identified isoforms that analyzes survival data for the plurality of research subjects for the target condition relative to control data; and perform a log rank test using quantile thresholds of 25, 50 and 75 for each of the identified isoforms that analyzes the survival data for the plurality of research subjects for the target condition relative to the control data.
6. The computer system of claim 5 wherein the isoform correlation indicators associated with a respective isoform of the identified isoforms indicate that the isoform level correlation with survival of the target condition is present when (1) results of the univariant cox regression for the respective isoform indicate a hazard ratio greater than one and a p-value less than 0.05; and (2) results of the log rank test for the respective isoform indicate a p-value less than 0.2.
7. The computer system of any one of claims 1-6 wherein to identify the refined set of prognosis indicating isoforms based on the multivariant analysis of the initial set of prognosis indicating isoforms, the instructions, when executed by the one or more processors, cause the computer system to: perform a LASSO regression to analyze effect of the respective expression indicator for each isoform in the initial set of prognosis indicating isoforms relative to survival of the target condition; and select isoforms with a negative beta coefficient in results of the LASSO regression as the refined set of prognosis indicating isoforms.
8. The computer system of any one of claims 1-7 wherein to generate prognostic scores for the plurality of research subjects based on the refined set of prognosis indicating isoforms, the instructions, when executed by the one or more processors, cause the computer system to: sum together, for each of the plurality of research subjects, products of at least one of the isoform correlation indicators and the respective expression indicator for each isoform in the refined set of prognosis indicating isoforms present in an associated subject.Docket No. 30275 / 70688 / PC9. The computer system of any one of claims 1-8 wherein the instructions, when executed by the one or more processors, cause the computer system to: receive patient gene sequence data taken from a patient diagnosed with the target condition; generate a respective expression indicator for the refined set of prognosis indicating isoforms that are present in the patient gene sequence data; generate a patient prognostic score using the respective expression indicator for the refined set of prognosis indicating isoforms present in the patient gene sequence data and at least one of the isoform correlation indicators for the refined set of prognosis indicating isoforms; select one of the survival risk groupings based on the patient prognostic score; and present, on a display device, a recommended treatment program for the patient based on the selected one of the survival risk groupings.
10. The computer system of claim 9 wherein: the target condition is bladder cancer; and the recommended treatment program includes a tumor removal surgery without accompanying radiation or chemotherapy when the select one of the survival risk groupings is a lowest risk of the survival risk groupings.
11. The computer system of any one of claims 1-10 wherein: the target condition is bladder cancer; the gene sequence data is taken from bladder cancer tumor cells of the research subjects; and the refined set of prognosis indicating isoforms includes NM_001017425, NM_001149, NM_001164319, NM_001166215, NM_001171089, NM_001243773, NM_001276379, NM_001288570, NM_001317856_2, NM_001322027, NM_001323969, NM_001323970, NM_001328685, NM_001364608, NM_001365412, NM_001699, NM_006665, NM_032053, NM_057164, NR_045028, NR_109973, NR_134658.
12. A computer system comprising: one or more processors; and one or more non-transitory, computer-readable media storing instructions that, when executed by the one or more processors, cause the computer system to:Docket No. 30275 / 70688 / PC receive patient gene sequence data taken from a patient diagnosed with a target condition; identify, in a data store, a refined set of prognosis indicating isoforms that are associated with the target condition; generate a respective expression indicator for the refined set of prognosis indicating isoforms that are present in the patient gene sequence data; generate a patient prognostic score using the respective expression indicator for the refined set of prognosis indicating isoforms present in the patient gene sequence data and isoform correlation indicators for the refined set of prognosis indicating isoforms that are stored in the data store; select one of a plurality of survival risk groupings for the target condition based on the patient prognostic score; and present, on a display device, a recommended treatment program for the patient based on the selected one of the plurality of survival risk groupings.
13. The computer system of claim 12 wherein: the target condition is bladder cancer; and the recommended treatment program includes a tumor removal surgery without accompanying radiation or chemotherapy when the select one of the survival risk groupings is a lowest risk of the survival risk groupings.
14. The computer system of either claim 12-13 wherein to generate patient the prognostic score, the instructions, when executed by the one or more processors, cause the computer system to: sum together products of the isoform correlation indicators and the respective expression indicator for each isoform in the refined set of prognosis indicating isoforms present in the patient.
15. The computer system of any one of claims 12-14 wherein to identify the refined set of prognosis indicating isoforms, the instructions, when executed by the one or more processors, cause the computer system to: identify an initial set of prognosis indicating isoforms from subject gene sequence data based on (1) a respective expression indicator for identified isoforms in the subject gene sequence data, (2) gene correlation indicators for gene sequences in the subject geneDocket No. 30275 / 70688 / PC sequence data, and (3) isoform correlation indicators for the identified isoforms in the subject gene sequence data, wherein: the subject sequence gene data is associated with a plurality of research subjects diagnosed with the target condition, the gene correlation indicators indicating presence or absence of a gene level correlation with survival of the target condition for an associated subject of the plurality of research subjects, and the isoform correlation indicators indicating presence or absence of an isoform level correlation with survival of the target condition for the associated subject; and select the refined set of prognosis indicating isoforms from the initial set of prognosis indicating isoforms based on results of a multivariant analysis of the initial set of prognosis indicating isoforms.
16. The computer system of claim 15 wherein to identify the initial set of prognosis indicating isoforms, the instructions, when executed by the one or more processors, cause the computer system to: select a group of the identified isoforms based on the respective expression indicator for each identified isoform in the subject gene sequence data; identify candidate isoforms as isoforms in the selected group of the identified isoforms for which the associated isoform correlation indicator indicates that the isoform level correlation with the prognosis of the target condition is present; and select the initial set of prognosis indicating isoforms as the candidate isoforms for which the generated gene correlation indicator of the respective gene sequence on which the candidate isoforms is present indicates that the gene level correlation with the prognosis of the target condition is absent.
17. The computer system of claim 16 wherein to select the group of the identified isoforms based on the respective expression indicator for each identified isoform, the instructions, when executed by the one or more processors, cause the computer system to: identify a filtered set of the identified isoforms that are associated with a gene sequence of the gene sequence data for which two or more of the identified isoforms are present; andDocket No. 30275 / 70688 / PC select isoforms in the filtered set of the identified isoforms with a respective expression indicator above a preconfigured threshold as the group of the identified isoforms.
18. A computer implemented method comprising: obtaining gene sequence data, the gene sequence data associated with a plurality of research subjects diagnosed with a target condition; generating a respective expression indicator for identified isoforms in the gene sequence data; generating gene correlation indicators for gene sequences in the gene sequence data, the gene correlation indicators indicating presence or absence of a gene level correlation with survival of the target condition for an associated subject of the plurality of research subjects; generating isoform correlation indicators for the identified isoforms, the isoform correlation indicators indicating presence or absence of an isoform level correlation with survival of the target condition for the associated subject; identifying an initial set of prognosis indicating isoforms based on (1) the respective expression indicator for each identified isoform, (2) the isoform correlation indicators, and (3) the gene correlation indicators; identifying a refined set of prognosis indicating isoforms based on a multivariant analysis of the initial set of prognosis indicating isoforms relative to survival of the target condition by the plurality of research subjects; generating prognostic scores for the plurality of research subjects based on the refined set of prognosis indicating isoforms; identifying different survival risk groupings for the target condition that are linked to respective ranges of the prognostic scores; and saving the survival risk groupings in a data store for use in selecting future patient treatment programs for the target condition.
19. The computer implemented method of claim 18 wherein identifying the initial set of prognosis indicating isoforms includes: selecting a group of the identified isoforms based on the respective expression indicator for each identified isoform in the subject gene sequence data; identifying candidate isoforms as isoforms in the selected group of the identified isoforms for which the associated isoform correlation indicator indicates that the isoform level correlation with the prognosis of the target condition is present; andDocket No. 30275 / 70688 / PC selecting the initial set of prognosis indicating isoforms as the candidate isoforms for which the generated gene correlation indicator of the respective gene sequence on which the candidate isoforms is present indicates that the gene level correlation with the prognosis of the target condition is absent.
20. The computer implemented method of either claim 18 or claim 19 further comprising: receiving patient gene sequence data taken from a patient diagnosed with the target condition; generating a respective expression indicator for the refined set of prognosis indicating isoforms that are present in the patient gene sequence data; generating a patient prognostic score using the respective expression indicator for the refined set of prognosis indicating isoforms present in the patient gene sequence data and at least one of the isoform correlation indicators for the refined set of prognosis indicating isoforms; selecting one of the survival risk groupings based on the patient prognostic score; and presenting, on a display device, a recommended treatment program for the patient based on the selected one of the survival risk groupings.
Citation Information
Patent Citations
Spring control and shock absorber
CA201335A
Electric cigar lighter
CA259763A
Fire extinguisher
CA273138A
Ovarian cancer prognostic subgrouping
WO2016066797A2
Systems for and methods of treatment selection
WO2022081923A2