Methods For Normalizing And Correcting Rna Expression Data

AU2026205141A1Pending Publication Date: 2026-07-16TEMPUS AI INC

Patent Information

Authority / Receiving Office
AU · AU
Patent Type
Applications
Current Assignee / Owner
TEMPUS AI INC
Filing Date
2026-06-30
Publication Date
2026-07-16

Smart Images

  • Figure 00000032_0000
    Figure 00000032_0000
  • Figure 00000033_0000
    Figure 00000033_0000
  • Figure 00000034_0000
    Figure 00000034_0000
Patent Text Reader

Abstract

20 26 20 51 41 30 J un 2 02 6 A B S T R A C T 2 0 2 6 2 0 5 1 4 1 3 0 J u n 2 0 2 6
Need to check novelty before this filing date? Find Prior Art

Claims

2026205141   30 Jun 2026What is Claimed:

1. A computer-implemented method comprising:generating, from a comparison of a normalized RNA sequence dataset against a standard RNA sequence dataset, at least one conversion factor for applying to a next RNA sequence dataset; andcorrecting RNA sequence data of the next RNA sequence dataset using the at least one conversion factor.

2. The computer-implemented method of claim 1, further comprising:including the corrected RNA sequence data of the next RNA sequence dataset into the standard RNA sequence dataset.

3. The computer-implemented method of claim 1, further comprising:obtaining a gene expression dataset comprising the RNA sequence data for one or more genes, the RNA sequence data including gene length data, guanine-cytosine (GC) content data, and depth of sequencing data; andnormalizing the RNA sequence data against the standard RNA sequence data by comparing the RNA sequence data for the one or more genes to sequence data in the standard RNA sequence dataset.

4. The computer-implemented method of claim 3, wherein normalizing the RNA sequence data comprises:normalizing the gene length data for the one or more genes to reduce systematic bias;normalizing the GC content data for the one or more genes to reduce systematic bias; and normalizing the depth of sequencing data for the RNA sequence data.

5. The computer-implemented method of claim 4, wherein normalizing the gene length data comprises using a quantile normalization procedure.

6. The computer-implemented method of claim 4, wherein normalizing the GC content data comprises using a quantile normalization procedure.

7. The computer-implemented method of claim 4, wherein normalizing the depth ofsequencing data comprises:determining a ratio of expression data to reference geometric mean expression data obtained from the standard RNA sequence dataset;2026205141   30 Jun 2026determining ratios of expression data to reference geometric mean expression data for a plurality of additional RNA sequence data corresponding to the at least one gene, to develop a set of ratios for the gene expression dataset; anddetermining a size factor as a median of the set of ratios.

8. The computer-implemented method of claim 3, wherein normalizing the RNA sequence data comprises applying a Reads Per Kilobase Million (RPKM) normalization, a Fragments Per Kilobase Million (FPKM) normalization, or a Transcripts Per Kilobase Million (TPM) normalization.

9. The computer-implemented method of claim 1, wherein the RNA sequence dataset is a raw dataset.

10. The computer-implemented method of claim 1, wherein the RNA sequence dataset is a Cancer Genome Atlas (TCGA) dataset.

11. The computer-implemented method of claim 1, wherein the RNA sequence dataset is a Genotype-Tissue Expression (GTEx) dataset.

12. The computer-implemented method of claim 1, wherein generating the at least one conversion factor comprises:for a sample gene, obtaining sample data from a normalized RNA sequence dataset and obtaining sample data from the standard RNA sequence dataset;determining a statistical mapping between the sample data of the normalized RNA sequence dataset and the sample data of the standard RNA sequence dataset; anddetermining the at least one conversion factor using the statistical mapping.

13. The computer-implemented method of claim 12, wherein determining the statistical mapping comprises determining a linear mapping model between the sample data of the normalized RNA sequence dataset and the sample data of the standard RNA sequence dataset, the method further comprising:determining an intercept and a beta value for the linear mapping model; anddetermining the at least one conversion factor using the statistical mapping from the intercept and the beta value.

14. The computer-implemented method of claim 12, wherein generating the at least one conversion factor comprises:2026205141   30 Jun 2026(i) for a sample gene, obtaining sample data from normalized RNA sequence dataset and obtaining sample data from the standard RNA sequence dataset;(ii) determining a linear mapping model between the sample data of the normalized RNA sequence dataset and the sample data of the standard RNA sequence dataset;(iii) determining an intercept and a beta value for the linear mapping model;(iv) performing (i) - (iv) a plurality of times for the sample gene; and(iv) determining a gene specific conversion factor from a mean intercept and a mean beta value for the plurality of times.

15. A computing device comprising one or more memories and one or more processors configured to:generate, from a normalization of an RNA sequence data against a standard RNA sequence dataset, at least one conversion factor for applying to a next RNA sequence dataset; andcorrect RNA sequence data of the next RNA sequence dataset using the at least one conversion factor.

16. The computing device of claim 15, wherein the one or more processors are configured to:include the corrected RNA sequence data of the next RNA sequence dataset into the standard RNA sequence dataset.

17. The computing device of claim 15, wherein the one or more processors are configured to:obtain a gene expression dataset comprising the RNA sequence data for one or more genes, the RNA sequence data including gene length data, guanine-cytosine (GC) content data, and depth of sequencing data; andcorrect the RNA sequence data against the standard RNA sequence data by comparing the RNA sequence data for the one or more genes to sequence data in the standard RNA sequence dataset.

18. The computing device of claim 17, wherein the one or more processors are configured to normalize the RNA sequence data by being configured to:normalize the gene length data for the one or more genes to reduce systematic bias;normalize the GC content data for the one or more genes to reduce systematic bias; andnormalize the depth of sequencing data for the RNA sequence data.2026205141   30 Jun 202619. The computing device of claim 15, wherein the one or more processors are configured to generate the at least one conversion factor by being configured to:for a sample gene, obtain sample data from a normalized RNA sequence dataset and obtaining sample data from the standard RNA sequence dataset;determine a statistical mapping between the sample data of the normalized RNA sequence dataset and the sample data of the standard RNA sequence dataset; anddetermine the at least one conversion factor using the statistical mapping.

20. The computing device of claim 19, wherein the one or more processors are configured to determine the statistical mapping by being configured to determine a linear mapping model between the sample data of the normalized RNA sequence dataset and the sample data of the standard RNA sequence dataset, the one or more processors being further configured to:determine an intercept and a beta value for the linear mapping model; anddetermine the at least one conversion factor using the statistical mapping from the intercept and the beta value.

21. A computer-implemented method comprising:receiving, at one or more processors, a gene expression dataset;identifying within the gene expression dataset, using a regression technique implemented by the one or more processors, gene expression data having multiple modal expression peaks;for the gene expression data, normalizing, using the one or more processors, a spacing between each of the multiple model expression peaks to form a normalized gene expression data; andstoring the normalized gene expression data in a normalized gene expression dataset.

22. The computer-implemented method of claims 21, wherein the gene expression dataset is an RNA sequence dataset.

23. The computer-implemented method of claims 21, the method further comprising normalizing, using the one or more processors, a reference baseline expression based on the multiple model expression peaks, wherein the reference baseline expression identifies overexpressed gene expression data and under-expressed gene expression data.

24. The computer-implemented method of claims 21, wherein the gene expression data has a bimodal distribution, and the multiple model expression peaks consist of two expression peaks.2026205141   30 Jun 202625. The computer-implemented method of claims 24, wherein normalizing the spacing between the two expression peaks comprises setting the spacing to 1.

26. The computer-implemented method of claims 25, the method further comprising normalizing, using the one or more processors, a reference baseline expression for the two expression peaks by setting a zero expression value between the two expression peaks.

27. The computer-implemented method of claims 24, wherein the two expression peaks of the bimodal distribution of the gene expression data correspond to a tumor specific expression peak and a tissue specific expression peak.

28. The computer-implemented method of claims 24, wherein the regression technique is a two-leaf decision tree regressor.

29. The computer-implemented method of claims 21, wherein the regression technique is a multiple-leaf decision tree regressor.

30. The computer-implemented method of claims 21, wherein the gene expression dataset is an RNA sequence dataset comprising a plurality of gene expression data each corresponding to different gene and each having a bimodal distribution.

31. A computer-implemented method comprising:receiving, at one or more processors, a RNA sequence dataset;identifying within the gene expression dataset, using a regression technique implemented by the one or more processors, a plurality of RNA expression data each having a bimodal distribution comprising two expression peaks;for each of the plurality of RNA expression data, normalizing, using the one or more processors, a spacing between the two expression peaks such that each of the plurality of RNA expression data has the same spacing between the two expression peaks; andstoring the normalized RNA expression data in a normalized RNA sequence dataset.

32. The computer-implemented method of claims 31, the method further comprising shifting each of the plurality of RNA expression data to have a zero expression value between the two expression peaks.