K-value and miniquant to index and reduce isoform quantification error

WO2026206656A1PCT designated stage Publication Date: 2026-10-01THE RGT UNIV OF MICHIGAN
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2026/019294
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-27
Filing Date
2026-03-16
Publication Date
2026-10-01

Smart Images

  • Figure US2026019294_01102026_PF_FP_ABST
    Figure US2026019294_01102026_PF_FP_ABST
Patent Text Reader

Abstract

A method includes sequencing a sample for genes likely to have low gene isoform and / or repetitive element quantification errors by identifying gene isoforms and regions corresponding to a gene, generating read-isoform alignment probabilities for nucleic acid sequencing reads that originated from each gene isoform, determining a generalized condition number for the gene based on the read-isoform alignment probabilities, and sequencing a subset of genes based on the generalized condition number for each of the genes. Another method quantifies isoforms and repetitive elements of a sample using sequencing data by mapping long and short reads to gene communities, estimating isoform abundance within each gene community, and determining isoform expression levels of gene isoforms in the sample. A further method quantifies isoforms using long reads by generating a long-read likelihood function and estimating isoform abundance based on the long-read likelihood function.
Need to check novelty before this filing date? Find Prior Art

Description

PATENT APPLICATION Attorney Docket No.: 30275 / 70852 / PC K- VALUE AND MINIQUANT TO INDEX AND REDUCE ISOFORM QUANTIFICATION ERROR CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of U.S. Application No. 63 / 778,909, filed March 27, 2026, and entitled " K-Value and MiniQuant to Index and Reduce Isoform Quantification Error", which is incorporated herein by reference in its entirety.STATEMENT OF GOVERNMENTAL INTEREST

[0002] This invention was made with government support under HG011469 awarded by the National Institute of Health. The Government has certain rights in the invention.FIELD OF THE INVENTION

[0003] The present disclosure generally relates to methods for identifying and / or ranking complex genes using a generalized condition number, and more particularly, to techniques for quantifying gene isoforms in a sample using long reads and / or short reads.BACKGROUND

[0004] The background description provided herein is for the purpose of generally presenting the context of the disclosure. Work of the presently named inventors, to the extent it is described in this background section, as well as aspects of the description that may not otherwise qualify as prior art at the time of filing, are neither expressly nor impliedly admitted as prior art against the present disclosure.

[0005] Recent advancements in RNA sequencing (RNA-seq) technologies, particularly long-read RNA-seq offered by platforms such as Pacific Biosciences (PacBio) and Oxford Nanopore Technologies (ONT), have shown significant promise in identifying gene isoforms more effectively than short-read RNA-seq. This is mainly due to their ability to span entire transcript lengths, thereby reducing the ambiguity in read alignments that is often encountered with short reads. Despite these technological leaps, challenges persist in accurately quantifying gene isoforms using long-read RNA-seq. Short-read RNA-seq, while suffering from limitations in identifying complex gene isoforms due to read alignment ambiguity, has been extensively used for gene isoform quantification due to its higher throughput and lower cost.PATENT APPLICATION Attorney Docket No.: 30275 / 70852 / PC

[0006] Quantifying gene isoforms accurately remains a critical yet unresolved question in the field, especially for complex genes where isoforms share exonic sequences. The reliability of quantification directly impacts the understanding of gene function and regulation, and inaccuracies can lead to misinterpretation of gene isoform dynamics in biological processes. Moreover, integrating the complementary strengths of long and short reads to enhance quantification accuracy has emerged as a potential avenue, yet determining the optimal combination of these technologies for specific genes and datasets poses another layer of complexity. These challenges underscore the need for innovative approaches that can leverage the advantages of both long and short reads, taking into consideration factors such as read length, gene structure, and data specificity.

[0007] Given the limitations and challenges associated with current methodologies for gene isoform quantification, there are significant opportunities for the development of improved platforms and technologies. These advancements could offer more robust frameworks for addressing the complexities inherent in accurately quantifying gene isoforms, particularly those associated with complex genes. By enhancing the reliability of gene isoform quantification, such innovations hold the potential to significantly advance our understanding of gene function and regulation across diverse biological contexts.SUMMARY

[0008] Despite the advancements in long-read RNA-seq, which have significantly improved the identification of gene isoforms, quantifying these isoforms accurately remains a complex task due to several unresolved issues. To tackle these challenges, the disclosure introduces a novel approach that utilizes a generalized condition number (also referred to herein as a “K- value”) to measure quantification errors arising from the ambiguity of read alignments.

[0009] This approach is embodied in a software tool named miniQuant, designed to rank genes by quantification errors and to quantify gene isoforms using either long reads alone or in combination with short reads. The effectiveness of miniQuant is demonstrated through rigorous mathematical proofs, extensive simulation data, experimental validations, and the analysis of over 17,000 public datasets from major consortia such as GTEx, TCGA, and ENCODE. The application of miniQuant has revealed critical insights into isoform switches during thePATENT APPLICATION Attorney Docket No.: 30275 / 70852 / PC differentiation of human embryonic stem cells, showcasing its utility in uncovering complex gene expression dynamics.

[0010] By integrating long and short reads for gene isoform quantification, the miniQuant tool not only improves the accuracy of quantification by reducing the deconvolution and sampling errors but also enables a more comprehensive understanding of gene expression patterns. This approach leverages the strengths of both sequencing methods. Long reads offer about a more complete view of gene isoform structure, while short reads provide depth and coverage. By mapping these reads to gene communities and estimating isoform abundance as described in more detail below, the techniques achieve a more accurate quantification of gene isoforms. This is further enhanced by the use of machine learning models to determine community-specific weights for long and short reads, optimizing the integration based on the characteristics of the genes and the sequencing data. This approach allows for the optimal utilization of available sequencing data, thereby enhancing the reliability of gene isoform quantification and facilitating more accurate biological interpretations.

[0011] Moreover, the disclosure highlights the importance of addressing sampling errors, particularly in the context of long read-based quantification. Despite the advantages of long reads in reducing deconvolution uncertainty, their relatively low throughput poses challenges in accurately quantifying lowly-expressed gene isoforms. The disclosure provides strategies to mitigate these limitations, emphasizing the critical role of data integration in achieving accurate and reliable gene isoform quantification.

[0012] Additionally, the disclosure provides a comprehensive solution to the challenges associated with gene isoform quantification in long-read RNA-seq. By integrating long and short reads in a gene- and data-specific manner, the present embodiments significantly improve the accuracy and reliability of quantification. This advancement has profound implications for our understanding of gene expression dynamics and their role in biological processes, offering new avenues for research and potential therapeutic interventions.

[0013] In one aspect, a method for sequencing a sample includes: (1) for a plurality of genes: (i) identifying a set of gene isoforms corresponding to the gene; (ii) identifying a set of regions within the gene; (iii) generating, by one or more processors, probabilities that a nucleic acid sequencing read that originated from each gene isoform in the set of gene isoforms correspondsPATENT APPLICATION Attorney Docket No.: 30275 / 70852 / PC to each region in the set of regions; and (iv) determining, by the one or more processors, a generalized condition number for the gene based on the probabilities; (2) identifying, by the one or more processors, a subset of the plurality of genes based on the generalized condition number; and (3) sequencing a sample for the identified subset of genes.

[0014] In another aspect, a method of quantifying isoforms of one or more genes in a sample using nucleic acid sequencing data includes: (1) obtaining, by one or more processors, nucleic acid sequencing data comprising short reads and long reads for a sample; (2) mapping, by the one or more processors, the short reads and the long reads to one of a plurality of gene communities, wherein each gene community includes one or more genes and gene isoforms thereof; (3) for each of the plurality of gene communities, estimating, by the one or more processors, the isoform abundance of the gene isoforms in the gene community based on the short reads and the long reads mapped to the gene community; and (4) for each of a plurality of gene isoforms in each of the plurality of gene communities, determining, by the one or more processors, an isoform expression level of the isoform within the gene community in the sample based on a combination of a relative abundance of the gene community to other gene communities and a relative isoform abundance of the gene isoform to other gene isoforms within the gene community to quantify abundance of the plurality of gene isoforms in the sample.

[0015] In yet another aspect, a method of quantifying isoforms of one or more genes in a sample using nucleic acid sequencing data includes: (1) obtaining, by one or more processors, nucleic acid sequencing data comprising long reads for a sample; (2) generating, by the one or more processors, a long-read likelihood function indicating likelihoods of isoform abundances of gene isoforms given the long reads from the sample; and (3) estimating, by the one or more processors, the isoform abundance of the gene isoforms based on the long-read likelihood function.

[0016] Advantages will become more apparent to those of ordinary skill in the art from the following description of the preferred embodiments which have been shown and described by way of illustration. As will be realized, the present embodiments may be capable of other and different embodiments, and their details are capable of modification in various respects.Accordingly, the drawings and description are to be regarded as illustrative in nature and not as restrictive.PATENT APPLICATION Attorney Docket No.: 30275 / 70852 / PC BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The figures described below depict various aspects of the system and methods disclosed herein. It should be understood that each figure depicts an embodiment of a particular aspect of the disclosed system and methods, and that each of the figures is intended to accord with a possible embodiment thereof.

[0018] There are shown in the drawings arrangements which are presently discussed, it being understood, however, that the present embodiments are not limited to the precise arrangements and instrumentalities shown, wherein:

[0019] Fig. 1 depicts a block diagram of an example computing environment for generating generalized condition numbers for genes and quantifying gene isoforms in a sample, according to some aspects.

[0020] Fig. 2 depicts example exon-intron structures of a gene, gene isoforms corresponding to the gene, and short and long reads mapped to the gene, according to some aspects.

[0021] Fig. 3 depicts an example process for using a read-isoform alignment conditional probability matrix to determine a generalized condition number (K- value) for a gene, according to some aspects.

[0022] Fig. 4A depicts an example chart of the number of genes and gene isoforms thereof with various K-values based on a GENCODE annotation of the human genome, according to some aspects.

[0023] Fig. 4B depicts an example comparison of the accuracy of 5 different short-read software tools for genes within different K-value groups across 9 sequencing depths using an mean absolute relative difference (MARD) between estimated isoform abundances and true isoform abundances for each of the short-read software tools, according to some aspects.

[0024] Fig. 4C depicts an example chart of the MARD for 10 different genes with low and high K-values across 9 different sequencing depths quantified by a short-read software tool, according to some aspects.

[0025] Fig. 4D depicts an example comparison of the performance of differentially expressed gene isoforms with high and low K-values, according to some aspects.PATENT APPLICATION Attorney Docket No.: 30275 / 70852 / PC

[0026] Fig. 4E depicts an example comparison of the performance of differentially expressed gene isoforms across 9 different sequencing depths within the different K-value groups of Fig.4B for 5 different short-read software tools, according to some aspects.

[0027] Fig. 4F depicts an example comparison of the accuracy and irreproducibility of shortread software tools for genes within the different K-value groups of Fig. 4B for 3 different gene annotation consortia, according to some aspects.

[0028] Fig. 5 depicts an example process for quantifying isoform abundances of gene isoforms in a sample, which may be performed by the miniQuant software tool, according to some aspects.

[0029] Fig. 6 depicts an example combined block and logic diagram that depicts the generation of a community-specific weight for long reads, ac, using a machine learning model, according to some aspects.

[0030] Fig. 7A depicts an example comparison of the accuracy of a long-read alone mode of the miniQuant (miniQuant-L) software tool compared to other long-read software tools using an MARD between the estimated isoform abundances and the true isoform abundances, according to some aspects.

[0031] Fig. 7B depicts an example comparison of the accuracy of a hybrid mode of the miniQuant (miniQuant-H) software tool compared to other long-read and short-read software tools using the MARD between the estimated isoform abundances and the true isoform abundances, according to some aspects.

[0032] Fig. 8 is a flow diagram of an example method for sequencing a sample for genes likely to have certain gene isoform and / or repetitive element quantification errors, which may be implemented at least in part by a server device, according to some aspects.

[0033] Fig. 9 is a flow diagram of an example method for quantifying isoforms and / or repetitive elements of one or more genes in a sample using nucleic acid sequencing data, which may be implemented by a server device, and more specifically, a miniQuant software tool, according to some aspects.PATENT APPLICATION Attorney Docket No.: 30275 / 70852 / PC

[0034] The Figures depict preferred implementations for purposes of illustration only.Alternative implementations of the systems and methods illustrated herein may be employed without departing from the principles of the invention described herein.DETAILED DESCRIPTION

[0035] Although the following text discloses a detailed description of implementations of methods, apparatuses and / or articles of manufacture, it should be understood that the legal scope of the property right is defined by the words of the claims set forth at the end of this document. Accordingly, the following detailed description is to be construed as examples only and does not describe every possible implementation, as describing every possible implementation would be impractical, if not impossible. Numerous alternative implementations could be implemented, using either current technology or technology developed after the filing date of this patent. It is envisioned that such alternative implementations would still fall within the scope of the claims.

[0036] Given the complexity and diversity of gene isoforms, accurately quantifying their abundance presents a significant challenge in genomics and transcriptomics. The advent of long-read RNA sequencing (RNA-seq) technologies has ushered in a new era of gene isoform identification and quantification. These technologies offer the potential to overcome the limitations of short-read RNA-seq, particularly in accurately quantifying gene isoforms that arise from alternative splicing — a process pivotal to increasing the diversity of RNAs, proteins, and phenotypic traits in humans. Despite the qualitative advantages of long-read RNA-seq in identifying gene isoforms, quantification accuracy remains hindered by several unresolved challenges. These include the ambiguity of read alignments due to shared exonic sequences among isoforms, leading to significant quantification errors, especially for complex genes.

[0037] To address these challenges, a novel approach, encapsulated in the software tool miniQuant, has been developed. MiniQuant leverages the complementary strengths of long and short reads through a sophisticated integration mechanism, significantly enhancing the accuracy of gene isoform quantification. This approach is grounded in the concept of a generalized condition number (K-value), which serves as a proxy for quantification error attributable to the ambiguity of read alignments. By ranking genes according to their K-values, miniQuant enablesPATENT APPLICATION Attorney Docket No.: 30275 / 70852 / PC researchers to identify and focus on genes that can be quantified reliably, thereby improving the overall reliability of gene isoform studies.

[0038] The utility of miniQuant extends beyond mere quantification. MiniQuant facilitates the exploration of isoform switches during critical biological processes, such as the differentiation of human embryonic stem cells into specific cell types. This capability is demonstrated through the application of miniQuant in uncovering isoform switches during the differentiation of human embryonic stem cells to pharyngeal endoderm and primordial germ celllike cells. Such insights are invaluable for understanding the nuanced roles of gene isoforms in development and disease.

[0039] One of the significant improvements introduced by these techniques is the reduction of quantification errors caused by the ambiguity of read alignments, a common issue in gene isoform quantification. By utilizing a generalized condition number as a proxy for quantification error, the techniques enable the identification of genes that are likely to be quantified with higher accuracy. This is particularly beneficial for genes with complex exon-isoform structures, where traditional short-read sequencing methods may fall short due to alignment ambiguity.

[0040] Moreover, the integration of long-read and short-read sequencing data plays a crucial role in overcoming the limitations of each sequencing technology. Long-read sequencing data, known for its ability to span entire gene isoforms, reduces the uncertainty of data deconvolution and improves gene isoform quantification accuracy. However, the relatively low throughput and high cost of long-read sequencing pose challenges, particularly for lowly expressed genes and isoforms. The present techniques address these challenges by integrating short-read sequencing data, which offers higher throughput at a lower cost, thereby reducing sampling error and enhancing the overall accuracy of gene isoform quantification.

[0041] The techniques also introduce a machine learning approach to adaptively find the gene-and data-specific weighting scheme between long-read and short-read data, further optimizing the quantification process. This gene- and data-specific integration not only improves quantification accuracy but also enables the identification of isoform switches during cellular differentiation processes, revealing subtle but meaningful alterations in isoform usage that are key to understanding isoform-specific functions.PATENT APPLICATION Attorney Docket No.: 30275 / 70852 / PC

[0042] The present techniques represent a significant advancement in the field of genomics, offering a comprehensive and accurate approach to gene isoform and repetitive element quantification. By addressing the challenges associated with read alignment ambiguity and integrating the strengths of long-read and short- read sequencing technologies, these techniques pave the way for more reliable and insightful genomic analyses.EXEMPLARY COMPUTING ENVIRONMENT

[0043] Fig. 1 is a block diagram of an example computing environment 100 for generating generalized condition numbers for genes and quantifying gene isoforms in a sample, as described herein. The example computing environment 100 may include a sample 32, a short-read sequencing device 20, a long-read sequencing device 30, a server device 60, a client device 10, and a network 18.

[0044] The sample 32 may be a biological sample which may be a tissue sample, blood sample, saliva sample, skin sample, hair sample, cheek sample, or any other type of biological material from which genetic material can be extracted. Target nucleic acids (e.g., polyA mRNA molecules) may be captured for the sample 32 which may be then reverse transcribed into DNA. This is followed by library preparation, which typically involves amplification (e.g., PCR). The sequencing library is then subjected to sequencing. As used herein, the term “sequencing” includes sequencing and / or analysis of the sequencing data to annotate nucleotide sequences.

[0045] More specifically, one portion of the library may be provided to a short-read sequencing device 20 for short-read sequencing to generate a set of short reads for the sample 32 and another portion of the library may be provided to a long-read sequencing device 30 for long-read sequencing to generate a set of long reads for the sample 32. Optionally, a set of short reads or long reads for a sample from the same biological context or from the same cell type(s) as the sample 32 may be provided.

[0046] As used herein, the term “long reads” or “long-read sequence” may refer to a fragment of a genomic sequence that is longer than “short reads” or a “short-read sequence” which is typically between 75 and 300 bp. A long-read sequence may be 1,000 bp, 2,000 bp, 5,000 bp, 10,000 bp, 100,000 bp, 1 MB, etc.PATENT APPLICATION Attorney Docket No.: 30275 / 70852 / PC

[0047] In any event, the short reads and long reads for the sample 32 may be provided from the short-read sequencing device 20 and the long-read sequencing device 30, respectively, to the server device 60. In some versions, the short-read sequencing device 20 may utilize sequencing-by-synthesis (SBS) or sequencing-by-binding (SBB) to sequence nucleotide fragments into nucleotide reads and determine nucleobase calls for the nucleotide reads. The short-read sequencing device 20 may further store the nucleobase calls as part of base-call data that is formatted as a binary base call (BCL) / binary alignment map (BAM) file and send the BCL / BAM file to the server device 60. The long-read sequencing device 30 may utilize nanopore sequencing or Single Molecule, Real-Time (SMRT) sequencing to sequence nucleotide fragments into nucleotide reads and determine nucleobase calls for the nucleotide reads. The long-read sequencing device 30 may further store the nucleobase calls as part of base-call data that is formatted as a Hierarchical Data Format (HDF5) / POD5 / BAM file and send the HDF5 / POD5 / B AM file to the server device 60.

[0048] The short-read sequencing device 20 and the long-read sequencing device 30 may communicate the BCL / HDF5 / POD5 / BAM file and / or other data to the server device 60 via one or more network(s) 18 which can be any suitable local or wide area network(s) including a WiFi network, a Bluetooth network, a cellular network such as 3G, 4G, Long-Term Evolution (LTE), 5G, the Internet, etc., or directly (e.g., bypassing the one or more network(s) 18). In other implementations, the short-read sequencing device 20 and the long-read sequencing device 30 may communicate the BCL / HDF5 / POD5 / BAM file and / or other data to the server device 60 directly. For example, the short-read sequencing device 20 and the long-read sequencing device 30 may communicate the BCL / HDF5 / POD5 / BAM file and / or other data to the server device 60 via a wired connection (e.g., via a Universal Serial Bus (USB)).

[0049] The server device 60 receives the nucleic acid sequencing data including short reads and long reads from the short-read sequencing device 20 and long-read sequencing device 30, respectively. As mentioned above, the short and long reads may include a target nucleic acid (e.g., cDNA, DNA, RNA, or mRNA). The server device 60 may include a memory 64, one or more processors (CPUs) 142, a network interface 144, and an EG module 148 which may be a keyboard or a touchscreen, for example. The server device 60 may also be communicatively coupled to a gene isoform database 80 which may store a list of gene isoforms for a set of genes,PATENT APPLICATION Attorney Docket No.: 30275 / 70852 / PC probabilities that reads annotated to a particular gene originated from each of the gene isoforms, a set of gene regions for each gene, etc. The memory 64 can be a non-transitory memory and can include one or several suitable memory modules, such as random access memory (RAM), read-only memory (ROM), flash memory, other types of persistent memory, etc.

[0050] The memory 64 may store an operating system (OS) 152, which can be any type of suitable mobile or general-purpose operating system. The memory 64 also stores a generalized condition number generator 146 that generates generalized condition numbers (K- values) for genes based on their gene isoforms and / or exonic / intronic regions, for example using read-isoform alignment conditional probabilities. These K-values help identify genes that are difficult to quantify accurately, serving as a measure of quantification difficulty. Genes with higher K-values may be more complex and have higher quantification errors than genes with lower K-values.

[0051] Additionally, the server device 60 runs the miniQuant software tool 154, which ranks genes based on their K-values and quantifies isoforms of genes in a sample. Mini Quant 154 can operate in two modes: miniQuant-L, which uses long reads only, and miniQuant-H, which combines both long and short reads for enhanced quantification accuracy.

[0052] The server device 60 may then transmit a ranked list of genes according to their respective K-values to the client device 10. The client device 10 may include a memory and one or more processors (CPUs). The memory can be a non-transitory memory and can include one or several suitable memory modules, such as random access memory (RAM), read-only memory (ROM), flash memory, other types of persistent memory, etc.

[0053] The memory may store an operating system (OS), which can be any type of suitable mobile or general-purpose operating system. The memory also stores a transcriptomics application which may present the ranked list of genes according to their respective K-values. In some implementations, a subset of the genes ranked above a threshold ranking or having a K-value above or below a threshold K- value may be identified and selected for sequencing. Then the short-read sequencing device 20 or the long-read sequencing device 30 may sequence a sample 32 for the identified subset of the genes.PATENT APPLICATION Attorney Docket No.: 30275 / 70852 / PC

[0054] In other implementations, the server device 60 may transmit expression levels of gene isoforms in a sample 32 determined using mini Quant 154 to the client device 10 to quantify the abundance of gene isoforms in the sample 32. The client device 10 may display the expression levels of gene isoforms in the sample 32 to a user, allowing researchers to interpret the data and make informed decisions about their studies. Additionally or alternatively, the client device 10 may present warnings for genes with high K-values (above a threshold value) indicating potential quantification errors, results of a differential expression analysis using the K- value, etc.

[0055] While the generalized condition number generator 146 and the miniQuant software tool 154 are included as part of the server device 60 in Fig. 1, this is merely on example for ease of illustration only. The generalized condition number generator 146 may be a part of the server device 60, the client device 10, or a combination of the server device 60 and the client device 10. Similarly, the miniQuant software tool 154 may be a part of the server device 60, the client device 10, or a combination of the server device 60 and the client device 10.

[0056] Fig. 2 illustrates example exon-intron structures of a gene, gene isoforms corresponding to the gene, and short and long reads mapped to the gene. A gene is composed of exons and introns. Exons are sequences within a gene that are coded for proteins, while introns are non-coding sequences that are removed during the RNA splicing process. A single gene can give rise to multiple gene isoforms through a process known as alternative splicing, where different combinations of exons are joined together, excluding certain introns or exons in the process.

[0057] As shown in Fig. 2, a particular gene includes four exons 202-208, referred to as exon 1 (ref. no. 202), exon 2 (ref. no. 204), exon 3 (ref. no. 206), and exon 4 (ref. no. 208). This gene can produce three different isoforms based on which exons are included in the final mRNA transcript. Isoform A includes exons 1 (ref. no. 202a), 2 (ref. no. 204a), and 3 (ref. no. 206a); isoform B is made up of exons 2 (ref. no. 204b), 3 (ref. no. 206b), and 4 (ref. no. 208b); and isoform C comprises exons 1 (ref. no. 202c), 2 (ref. no. 204c), and 4 (ref. no. 208c). These isoforms represent the diversity of protein products that can be generated from a single gene, depending on which exons are spliced together during mRNA processing.

[0058] To further elucidate the concept, Fig. 2 also showcases how sequencing reads, both long and short, map to these gene structures. Long reads provide more extensive coveragePATENT APPLICATION Attorney Docket No.: 30275 / 70852 / PC across the gene, allowing for a clearer identification of which exons are present in a particular isoform. In the figure, two long reads 212, 214 map unambiguously to isoform A, as they cover exons 1, 2, and 3 completely or at least partially. Similarly, one long read 216 maps unambiguously to isoform B, covering exons 2, 3, and 4, and two long reads 218, 220 map unambiguously to isoform C, spanning portions of exons 1, 2, and 4. These long reads are crucial for identifying the presence of specific isoforms within a sample.

[0059] However, not all reads provide clear-cut information about isoform identity. Fig. 2 illustrates this with a first long read 222 that is ambiguous because it includes exons 1 and 3 but is missing portions of exon 2, making it unclear which isoform it represents. A second long read 224 is also ambiguous, covering only exons 2 and 3, again leaving isoform identity uncertain. Additionally, several short reads are depicted, which only include portions of exons 1, 2, 3, and 4. Due to their limited length, these short reads are ambiguous, as they do not span enough of the gene to unambiguously determine their origin from specific isoforms.

[0060] Fig. 2 conveys the intricacies of gene structure, the generation of isoforms, and the challenges and opportunities presented by sequencing technologies in mapping reads to specific gene isoforms. This illustration underscores the complexity of gene expression analysis and the importance of understanding both gene structure and the capabilities of sequencing technologies in interpreting genetic information.

[0061] To quantify the complexity of genes according to their isoforms, regions (e.g., exonintron structures), and reads, the generalized condition number generator 146 generates generalized condition numbers (K- values) for the genes as a proxy for gene isoform quantification error. Fig. 3 depicts an example process 300 for using a read-isoform alignment conditional probability matrix to determine a generalized condition number (K- value) for a gene.

[0062] The process 300 includes generating a read-isoform alignment conditional probability matrix 302. Each column of the conditional probability matrix 302 may represent a different isoform of a gene and each row may represent a different region of the gene. These regions can include various parts of the gene, such as exonic regions (coding regions of the gene).

[0063] Based on the transcript annotation, a gene is partitioned into fragments with segmentations at exon-exon junction points and start / end / junction points of all of the isoforms ofPATENT APPLICATION Attorney Docket No.: 30275 / 70852 / PC the gene. Given the read length for a read, the server device 60 identifies the fragments to which the read is aligned for each possible read start site or possible position of read alignment. A single read may be aligned to more than one fragment, and read start sites aligned to the same combination of fragments form a region.

[0064] Then each element of the conditional probability matrix 302 indicates the probability that a read ( / ?n) corresponds to a particular region s given the read originated from a particular isoform t, Ast— P RnG s| / ?ne t). The matrix 302 is represented as A G RSX Twhich is a region-by-isoform matrix with regions defined by read-isoform alignment.

[0065] The generalized condition number generator 146 then decomposes the conditional probability matrix 302, Ast, into three distinct matrices: an s x s unitary matrix (U), an s x t singular value matrix (E) (ref. no. 304), and a t x t unitary matrix (V), such that Ast= U EV*. The singular value matrix (E) (ref. no. 304) includes a set of non-negative real numbers on its diagonal (oi, 02, ... or), and zeroes in the other elements of the matrix 304.

[0066] Then the generalized condition number generator 146 determines the K-value (k(A)) 306 for the conditional probability matrix 302 by first identifying the largest singular value (omax(A)) and smallest positive singular value (or(A)) within the singular value matrix 304 (E). The largest singular value represents the highest degree of variability or spread in the data, while the smallest singular value indicates the least variability. The generalized condition number generator 146 then divides largest singular value by the smallest positive singular value to obtain the K-value, k(A) =<rr(A)

[0067] This K-value serves as an indicator of the stability and reliability of the gene isoform readings. A higher K-value suggests that the gene isoform readings are more susceptible to variations and may be less reliable, whereas a lower K-value indicates more stable and reliable readings.

[0068] Fig. 4A depicts an example chart 400 of the number of genes and gene isoforms thereof with various K-values based on a GENCODE annotation of the human genome. The chart 400 illustrates a distribution of the number genes and a corresponding number of gene isoforms having particular K-values. These K-values represent the complexity of gene expression. A K-value of 1 indicates a simple scenario, with approximately one gene isoformPATENT APPLICATION Attorney Docket No.: 30275 / 70852 / PC per gene. Conversely, a K-value larger than 25 suggests a highly complex scenario, with about 20 different gene isoforms per gene on average.

[0069] More specifically, the chart 400 indicates that 29,121 genes and 29,152 gene isoforms have a K-value of 1; 3,931 genes and 9,538 gene isoforms have a K-value ranges from 1 to 2; 2,712 genes and 9,785 gene isoforms have a K-value ranging from 2 to 3; 3,438 genes and 17,825 gene isoforms have a K-value ranging from 3 to 5; 3,111 genes and 21,676 gene isoforms have a K-value ranging from 5 to 8; 3,382 genes and 29,836 gene isoforms have a K-value ranging from 8 to 13; 3,770 genes and 43,867 gene isoforms have a K-value ranging from 13 to 25; and 3,433 genes and 69,412 gene isoforms have a K-value larger than 25. This chart effectively illustrates the spectrum of gene complexity, from simple to highly complex, based on the number of isoforms associated with each gene.

[0070] Fig. 4B depicts an example boxplot 410 showing the median MARD of genes within 7 different K-value groups across 9 sequencing depths for 5 different short-read software tools: kallisto, Salmon, RSEM, Cufflinks, and StringTie. The median MARD is determined by comparing the estimated isoform abundances by the short-read software tools to their true isoform abundances, which are known from a simulated dataset. The lower the MARD value, the closer the estimated abundance is to the true abundance, indicating higher accuracy in isoform quantification.

[0071] The MARD values for each of the competing software tools are plotted against multiple genes, each characterized by varying levels of complexity, denoted as K-values. Across each of the software tools, the MARD values increase as the K-values increase. This indicates that the quantification error is directly correlated to the K-value. For example, in each of the software tools, the MARD values are the lowest for genes with K-values between 1 and 2 and the highest for genes with K-values greater than 25.

[0072] Fig. 4C depicts an example chart 420 of the MARD for 10 different genes with low and high K-values across 9 different sequencing depths quantified by a short-read software tool, kallisto. The barplots for genes with low K-values (e.g., less than 5) are highlighted in red and genes with high K-values e.g., greater than 5) are highlighted in blue. Moreover, the K-value for each gene is included next to the gene in parenthesis. More specifically, STAT3 has a K-value of 756.59, DDR1 has a K-value of 111.49, DNM1 has a K-value of 107.95, MYND8 has aPATENT APPLICATION Attorney Docket No.: 30275 / 70852 / PC K-value of 106.12, F0XP1 has a K-value of 91.43, ID1 has a K-value of 3.31, MEDIO has a K-value of 2.02, Z7C3 has a K-value of 1.65, F0X.H1 has a K-value of 1.25, and LEFTY 1 has a K-value of 1.18.

[0073] The overall MARD (transcriptome-wide) was calculated based on the median MARD the genes. As shown in Fig. 4C, the MARD values for genes with low K- values (e. ., less than 5) are significantly lower than the MARD values for genes with high K-values (e.g., greater than 5).

[0074] Fig. 4D depicts an example chart 430 illustrating a comparison of the performance of differentially expressed gene isoforms with high and low K-values. The chart 430 compares performance metrics of gene isoforms for differential expression analysis, specifically comparing those with high and low K-values. In this context, performance is measured in terms of quantification accuracy, precision, recall, and F-l scores. Gene isoforms with low K-values, which represent simpler gene expressions are quantified with higher levels of accuracy, precision, and F-l scores. This suggests that the simpler the gene expression, the more accurately and precisely it can be quantified. On the other hand, gene isoforms with high K-values, which indicate more complex gene expressions, are associated with lower performance metrics.

[0075] Fig. 4E depicts an example chart 440 comparing the performance of differentially expressed gene isoforms across 9 different sequencing depths within the different K-value groups of Fig. 4B for 5 different short-read software tools: kallisto, Salmon, RSEM, Cufflinks, and StringTie. The performance is measured in terms of quantification accuracy, precision, recall, and F-l scores. Across each of the 5 different short-read software tools, gene isoforms with low K-values, which represent simpler gene expressions are quantified with higher levels of accuracy, precision, recall and F-l scores. For example, in each of the software tools, the precision values are the highest for genes with K-values between 1 and 2 and the lowest for genes with K-values greater than 25. Similarly, in each of the software tools, the quantification accuracy levels are the highest for genes with K-values between 1 and 2 and the lowest for genes with K-values greater than 25. Still further, in each of the software tools, the recall values are the highest for genes with K-values between 1 and 2 and the lowest for genes with K-valuesPATENT APPLICATION Attorney Docket No.: 30275 / 70852 / PC greater than 25. Finally, in each of the software tools, the F-l scores are the highest for genes with K-values between 1 and 2 and the lowest for genes with K-values greater than 25.

[0076] Fig. 4F depicts an example chart 450 comparing the accuracy and irreproducibility of short-read software tools for genes within the different K-value groups of Fig. 4B for 3 different gene annotation consortia: GTEx, TCGA, ENCODE. The accuracy is determined using MARD values. RSEM was employed to estimate isoform abundances of data from GTEx and TCGA, while kallisto was employed to estimate isoform abundance of data from ENCODE. Across each of the gene annotation consortia, the MARD values increase as the K-values increase. This indicates that the quantification error is directly correlated to the K-value. For example, in each of the software tools, the MARD values are the lowest for genes with K-values between 1 and 2 and the highest for genes with K-values greater than 25.

[0077] Similarly, across each of the gene annotation consortia, the irreproducibility values increase as the K-values increase. For example, in each of the software tools, the irreproducibility values are the lowest for genes with K-values between 1 and 2 and the highest for genes with K-values greater than 25.

[0078] In some implementations, the miniQuant software tool 154 may rank genes based on their respective K-values and provide a ranked list of genes for display. The ranked list may then be used to select a subset of the genes (e.g., genes ranked above a threshold ranking or having K-values above or below a threshold value) to sequence in a sample, for example using the shortread sequencing device 20 and / or the long-read sequencing device 30.

[0079] The miniQuant software tool 154 may also provide warning messages or highlight genes having K-values above or below a threshold value indicating a likelihood of high quantification error for those genes. Still further, the K-value may be used for data collection design to guide sequencing coverage and technology. In yet other implementations, the K-value may be used for other downstream analysis, such as a differential expression analysis. For example, the mini Quant software tool 154 may perform a differential expression analysis of gene isoforms for an identified subset of genes having K-values above or below a threshold value. Also as mentioned above, the K-value may be used for quantifying isoforms in a sample, for example in the miniQuant-H mode by selecting a weight for combining long-read likelihoods and short-read likelihoods based on the K-value using machine learning techniques. The K-PATENT APPLICATION Attorney Docket No.: 30275 / 70852 / PC value may be applicable to any transcripts that have similar nucleotide sequences including gene isoforms, repetitive elements (e.g., transposable elements), etc.

[0080] Fig. 5 depicts an example process 500 for quantifying isoform abundances of gene isoforms in a sample 32, which may be performed by the mini Quant software tool 154. The miniQuant software tool 154 obtains several inputs 502 including isoform annotations, which provide detailed information about the various gene isoforms present within the genome. The isoform annotation can be GENCODE annotation or any other suitable annotation (e.g., samplespecific annotation). Additionally, the inputs 502 include M long reads and N short reads from the sample 32, which may be nucleic acid sequences (e.g., cDNA, DNA, RNA, or mRNA) used to identify and quantify the presence of specific gene isoforms in the sample.

[0081] Then the mini Quant software tool 154 maps each of the M long reads and N short reads to a gene community, c, of C gene communities. Each gene community includes one or more genes and gene isoforms thereof and / or repetitive elements (e.g., transposable elements) or other transcripts having similar nucleotide sequences that share read alignments so that the ambiguity of the reads only exists within a gene community. Each read is assigned to at most one gene community and there is no overlap between the gene communities, such that there are no shared reads across the gene communities. Within a gene community, a read can be mapped to multiple gene isoforms or genes.

[0082] For a particular gene community, c, there are Ncshort reads and Mclong reads that can be mapped to T" isoforms having abundances 0C= (0^, ... , 0 ) such that Xt=i= 1- Accordingly, the isoform abundance of an isoform in a gene community is a fractional representation of the abundance of the isoform relative to the abundances of the other isoforms in the gene community.

[0083] To determine the abundance of isoforms in a gene community, the miniQuant software tool 154 generates a long-read likelihood function which indicates the likelihoods of the isoform abundances 0Cof the gene isoforms in the gene community given the Mclong reads from the sample mapped to the particular gene community, c. The long-read likelihood function may be represented as LLR QC\R), where R includes the sequences of the reads in the community (e.g., the sequences of the Mclong reads). In the miniQuant-L mode, the long-read likelihood function is the only likelihood function used to determine the abundance of isoforms in the genePATENT APPLICATION Attorney Docket No.: 30275 / 70852 / PC community. The miniQuant software tool 154 may determine probabilities that each of the Mclong reads corresponds to each of the T gene isoforms in the particular gene community c. Then the miniQuant software tool 154 may generate the long-read likelihood function based on the probabilities that each of the Mclong reads corresponds to each of the 'T gene isoforms.

[0084] In some implementations, such as when the miniQuant software tool 154 is in the miniQuant-L mode, the miniQuant software tool 154 does not map reads to gene communities. Instead, the miniQuant software tool generates the long-read likelihood function for gene isoforms across all genes associated with the long-reads rather than for gene isoforms in a particular gene community.

[0085] In the miniQuant-H mode, the mini Quant software tool 154 generates a short- read likelihood function which indicates the likelihoods of the isoform abundances 9Cof the gene isoforms in the gene community given the Ncshort reads from the sample mapped to the particular gene community, c. The short-read likelihood function may be represented as LSR QC\R), where R includes the sequences of the reads in the community (e.g., the sequences of the Ncshort reads). The mini Quant software tool 154 may determine probabilities that each of the Ncshort reads corresponds to each of the T gene isoforms in the particular gene community c. Then the miniQuant software tool 154 may generate the short- read likelihood function based on the probabilities that each of the Ncshort reads corresponds to each of the T gene isoforms.

[0086] The miniQuant software tool 154 applies a community-specific weight, ac, to the long-read likelihood function, applies one minus the community-specific weight, 1- ac, to the shortread likelihood function and combines the weighted long-read and short-read likelihood functions to generate a hybrid likelihood function 504 to estimate the abundance of isoforms in the gene community. More specifically, the mini Quant software tool 154 generates the hybrid likelihood function 504 using the following equation:L(ec|7?) = [LLR(ec| / ?)]«c[Weq / ?)]1-^

[0087] In the miniQuant-L mode, acis 1 such that L(QC\R) = LLR(QC\R). In the miniQuant-H mode, the community-specific weight, etc, may be determined using a machine learning model, as described in more detail below with reference to Fig. 6.PATENT APPLICATION Attorney Docket No.: 30275 / 70852 / PC

[0088] In any event, the miniQuant software tool 154 determines a maximum likelihood estimate of the isoform abundances, 0C, based on the likelihood function using an Expectation Maximization algorithm. The estimated abundances for a particular community-specific weight, de, are denoted as §c’ac.

[0089] The miniQuant software tool 154 uses the relative abundances of gene isoforms in the gene community to determine abundances of gene isoforms in the sample 506. More specifically, miniQuant software tool 154 may determine a relative abundance pcof the gene community c to the other gene communities in the set of gene communities, C. The relative abundances of the C gene communities may be denoted as P = (Pi, ... Pc), where ^^=1 Pc = 1 • In gene community c, the estimated relative abundance of the Iegene isoforms are 6C=... , 0 ) such that = 1. Each gene isoform in the gene community has an effective length, Z(' .

[0090] In the miniQuant-L mode, the mini Quant software tool 154 may estimate the relative abundance pcof the gene community c based on the number of long reads in the community, Mc,Mcrelative to the total number of long reads in the sample, where (T = — ?. In the miniQuant-Hmode, the miniQuant software tool 154 may estimate the relative abundance / ?cof the gene community c based on the number of short reads in the community, Nc, and the effective lengths of the short reads in the community, lct, relative to the total number of short reads in the sample,VTC Nt ~ Lt=i ^c and the effective lengths of the short reads in each of the communities I , where (P = -t—,VCVTC Wi 2<c=l^i=i]Cwhere AZf = Nc*

[0091] The mini Quant software tool 154 combines the estimated relative abundance / ?cof the gene community c with the estimated isoform abundance of a gene isoform, 0 , in the gene community to estimate the abundance of the gene isoform t in the sample. The miniQuant software tool 154 may determine the transcript per million abundance of a gene isoform t, TPMt,as TPMt = pc * * 106- The mini Quant software tool 154 may repeat this process for each of the gene isoforms to determine the abundances of each of the gene isoforms in the sample.PATENT APPLICATION Attorney Docket No.: 30275 / 70852 / PC

[0092] As mentioned above, the mini Quant software tool 154 may determine the communityspecific weight, etc, using a machine learning model. Fig. 6 is a combined block and logic diagram that depicts the generation of a community-specific weight for long reads, etc, using a machine learning model 620. The machine learning model 620 may be generated by a machine learning engine 602 included in the miniQuant software tool 154. Some of the blocks in Fig. 6 represent hardware and / or software components (e.g., block 602), other blocks represent data structures or memory storing these data structures, registers, or state variables (e.g., blocks 604, 620), and other blocks represent output data (e.g., block 606). Input signals are represented by arrows labeled with corresponding signal names.

[0093] To generate the machine learning model 620, the machine learning engine 602 receives training data including input vectors (xc) for multiple gene communities (community 1 -community n). The input vectors for each community include read and gene structure features for the communities. These features may include the generalized condition number (K-value) for each gene within the community, the length, start position, and orientation of reads, and / or any other suitable read and gene structure features for the community. The training data also includes known true relative abundances (9C) of the gene isoforms for each input vector (xc) according to simulation data from various short-read and long-read RNA-seq datasets generated with GENCODE annotation and ground truths. The miniQuant software tool 154 estimates isoform abundances of gene isoforms 0 in the gene community for the input vector (xc), and adjusts the estimated isoform abundances using a community- specific weight acwith acE Ac= {0, 0,05, ... , 0.95, 1}. Then the miniQuant software tool 154 calculates the actual Mean Absolute Relative Difference (MARD) yc= (yc,ac) based on the difference between the true relative abundances (0C) of the gene isoforms for input vector (xc) and the estimated isoform abundancesof gene isoforms 0 ' in the gene community for the input vector (xc). The mini Quant software tool 154 identifies the community-specific weight o.cfrom the set of candidate communityspecific weights (Ac— {0, 0,05, ... , 0.95, 1}) with the minimum actual MARD (yc>ac).

[0094] For example, the training data may include read and gene structure features 622 of community 1 (xi) and a known true relative abundance (01) and actual MARDfor community 1. The training data may also include read and gene structure features 624 of community 2 (X2.) and a known true relative abundance (02) and actual MARD (y2,a2) forPATENT APPLICATION Attorney Docket No.: 30275 / 70852 / PC community 2. Furthermore, the training data includes read and gene structure features 626 of community 3 (x?) and a known true relative abundance (03) and actual MARD (y3>a3) for community 3. Still further, the training data includes read and gene structure features 628 of community n (xn.) and a known true relative abundance (0n) and actual MARD (yn,an) for community n.

[0095] While the example training data includes input vectors, known true relative abundances, and actual MARDs for four gene communities 622-628, this is merely an example for ease of illustration only. The training data may include input vectors, known true relative abundances, and actual MARDs for any suitable number of communities.

[0096] The machine learning engine 602 then analyzes the training data to generate a machine learning model 620 for generating a predicted MARD (yc) and community-specific weight acfor a community c given an input vector (xc) for the community c. While the machine learning model 620 is illustrated as a linear regression model, the machine learning model may be another type of regression model such as a logistic regression model, a decision tree, several decision trees, a neural network, a hyperplane, a Hidden Markov Model, or any other suitable machine learning model. The machine learning engine 602 may train the machine learning model using any suitable machine learning techniques such as linear regression, polynomial regression, logistic regression, random forests, boosting such as adaptive boosting, gradient boosting, and extreme gradient boosting, nearest neighbors, Bayesian networks, neural networks, support vector machines, or any other suitable machine learning technique.

[0097] Then for a community c having unknown true relative abundances of gene isoforms, the miniQuant software tool 154 obtains an input vector (xe) 604 including read and gene structure features of community c. These read and gene structure features may include the generalized condition number (K-value) for each gene within community c, the length, start position, and orientation of reads, and / or any other suitable read and gene structure features for community c.

[0098] The machine learning engine 602 may then apply the input vector (xc) to the trained machine learning model 620 to generate a predicted MARD yc= yc a) . For example, ac6 Acthe machine learning engine 602 may generate predicted MARDs ycfor each of the weights inPATENT APPLICATION Attorney Docket No.: 30275 / 70852 / PC the set (Ac= {0, 0,05, ... , 0.95, 1}). Then the machine learning engine 602 identifies the community-specific weight ac606 as the optimal community-specific weighta°ptimalbased on the minimum predicted MARD of the predicted MARDs. In this manner, the optimalcommunity-specific weight ac‘ = ar gmmc(cyc a.

[0099] The miniQuant software tool 154 may then apply the community-specific weight ac606 for gene community c when generating the hybrid likelihood function 504 as shown in Fig. 5 to estimate the relative abundances of gene isoforms in gene community c.

[0100] Fig. 7A illustrates a comparative analysis of the accuracy of isoform quantification among various long-read software tools compared to the mini Quant software tool 154 in the miniQuant-L mode. The metric used for this comparison is the MARD, which serves as a quantitative measure of the deviation between the estimated isoform abundances produced by each software tool and the true isoform abundances, which are known a priori.

[0101] The charts compare the mini Quant software tool 154 in the miniQuant-L mode against the LIQA, StringTie2, TALON, FLAMES, FLAIR, Bambu, and IsoQuant long-read based tools. The metric used for this comparison is the mean absolute relative difference (MARD), which serves as a quantitative measure of the deviation between the estimated isoform abundances produced by each software tool and the true isoform abundances, which are known a priori in this comparative study.

[0102] More specifically, each chart compares the MARD for one of the long-read based tools (LIQA, StringTie2, TALON, FLAMES, FLAIR, Bambu, and IsoQuant) to the MARD for the miniQuant software tool 154 in the miniQuant-L mode.

[0103] The charts within Fig. 7A are organized to show the performance of each software tool across a range of sequencing depths, specifically at 0.5 million, 1 million, 2 million, 3 million, 4 million, 5 million, 10 million, 20 million, and 30 million long reads.

[0104] Across all sequencing depths, the miniQuant-L software tool 154 exhibits lower MARD values in comparison to the other long-read software tools, since the AM ARD is positive in all instances indicating that the MARDs for miniQuant-L 154 were lower than the MARD for the other long read-based tools. The lower MARD values associated with miniQuant-L 154PATENT APPLICATION Attorney Docket No.: 30275 / 70852 / PC indicate a closer alignment with the true isoform abundances, underscoring its superior accuracy in isoform quantification.

[0105] The consistent outperformance of miniQuant-L 154 across various sequencing depths, as evidenced by its lower MARD values when compared to other long read-based tools, establishes its potential as a more reliable tool for isoform quantification in genomic research. This finding is particularly relevant for researchers and practitioners in the field of genomics, where accurate isoform quantification is critical for understanding complex genetic information and its implications for health and disease.

[0106] Fig. 7B illustrates a comparative analysis of the accuracy the miniQuant software tool 154 in the miniQuant-H mode in estimating isoform abundances against a range of established short read-based and long-read based tools. This comparison is quantitatively measured using the MARD between the estimated isoform abundances provided by each software tool and the true isoform abundances, which are known from a controlled dataset. The lower the MARD value, the closer the estimated abundance is to the true abundance, indicating higher accuracy in isoform quantification.

[0107] Fig. 7B divided into two charts 750, 760 for clarity and comprehensive analysis. The first chart 750 compares the miniQuant-H software tool 154 with various short- read software tools, namely kallisto, Salmon, RSEM, Cufflinks, and StringTie.

[0108] The second chart 760 presents a comparison between the miniQuant-H software tool 154 and several long-read software tools, including miniQuant-L (Long-read mode of the same tool), IsoQuant, Bambu, FLAIR, FLAMES, TALON, StringTie2, andLIQA.

[0109] The short read-based tools quantified isoforms for 40 million short reads and the long-read-based tools quantified isoforms for 5 million long reads. The miniQuant-H software tool 154 quantified isoforms for 5 million long reads and 40 million short reads. In both charts, the MARD values for each of the competing software tools are plotted against multiple genes, each characterized by varying levels of complexity, denoted as K- values.

[0110] Across all tools and gene complexity levels, the miniQuant-H software tool consistently shows lower MARD values compared to both the short-read and long-read software tools it is compared against. This indicates that miniQuant-H is more accurate in estimatingPATENT APPLICATION Attorney Docket No.: 30275 / 70852 / PC isoform abundances, as it deviates less from the true isoform abundances than its counterparts. The consistent performance of miniQuant-H across different complexity levels and against both categories of software tools underscores its versatility and superior accuracy in isoform quantification, making it a valuable tool for genomic research and analysis.

[0111] Additionally, across several of the tools, the miniQuant-H software tool 154 performed better with larger differences in MARD as the gene complexity levels increased. For example, when compared to the StringTie short read-based tool, the MARD for miniQuant-H was lower by about 0.25 for genes having a K- value of 1. However, when compared to the same StringTie short read-based tool, the MARD for miniQuant-H was lower by about 0.5 for genes having a K-value greater than 25.

[0112] Fig. 8 is a flow diagram of an example method 800 for sequencing a sample for genes likely to have certain gene isoform and / or repetitive element quantification errors, which may be implemented at least in part by a server device 60 and more specifically, a generalized condition number generator 146.

[0113] At block 802, the generalized condition number generator 146 identifies a set of gene isoforms corresponding to a gene. The generalized condition number generator 146 also identifies a set of regions within the gene (block 804). The regions may include various parts of the gene, such as exonic regions (coding regions of the gene).

[0114] Then at block 806, the generalized condition number generator 146 generates probabilities that a nucleic acid sequencing read that originated from each gene isoform in the set of gene isoforms correspond to each region in the set of regions. For example, the generalized condition number generator 146 may generate a read-isoform alignment conditional probability matrix, where each column represents a different isoform of a gene and each row represents a different region of the gene. Each element of the conditional probability matrix indicates the probability that a read (Rn) corresponds to a particular region s given the read originates from a particular isoform t, Ast= P RnEs\RnE t).

[0115] At block 808, the generalized condition number generator 146 determines a generalized condition number (K-value) for the gene based on the probabilities. More specifically, the generalized condition number generator 146 may decompose the conditionalPATENT APPLICATION Attorney Docket No.: 30275 / 70852 / PC probability matrix, Ast, into three distinct matrices: an s x s unitary matrix (U), an s x t singular value matrix (S), and a t x t unitary matrix (V), such that Ast= U SV*. The singular value matrix (S) includes a set of non-negative real numbers on its diagonal (01, 02, ... OT), and zeroes in the other elements of the matrix.

[0116] Then the generalized condition number generator 146 determines the K-value (k(A)) 306 for the conditional probability matrix 302 by first identifying the largest singular value (omax(A)) and smallest positive singular value (or(A)) within the singular value matrix (S). The generalized condition number generator 146 then divides largest singular value by the smallest positive singular value to obtain the K-value, k(A) =•

[0117] At block 810, the generalized condition number generator 146 determines whether a K-value has been determined for each gene in a set of genes and proceeds to determine K- values until a K-value has been determined for each gene in the set.

[0118] Once K-values have been determined for each gene in the set, the generalized condition number generator 146 may provide a list of the genes and their respective K-values for display on a client device 10. In some implementations, the generalized condition number generator 146 may rank the list of genes in accordance with their respective K-values and provide the ranked list for display. The generalized condition number generator 146 may also label the list, for example with warnings for genes having high K-values above a threshold value that there are likely to be large quantification errors when quantifying isoforms for such genes.

[0119] Still further, the generalized condition number generator 146 may identify a subset of genes based on their K-values (block 812). For example, the generalized condition number generator 146 may identify genes having K-values below a threshold value which are unlikely to have large quantification errors when quantifying isoforms for the identified subset of genes.

[0120] Then at block 814, a sample 32 may be sequenced for the identified subset of genes. For example, the short-read sequencing device 20 or the long-read sequencing device 30 may sequence the sample 32 for the identified subset of genes.

[0121] Fig. 9 is a flow diagram of an example method 900 for quantifying isoforms of one or more genes in a sample using nucleic acid sequencing data, which may be implemented by a server device 60, and more specifically, a miniQuant software tool 154 executed by the serverPATENT APPLICATION Attorney Docket No.: 30275 / 70852 / PC device 60. While the method 900 is for quantifying isoforms, the method 900 may be applicable to any transcripts that have similar nucleotide sequences including gene isoforms, repetitive elements (e.g., transposable elements), etc. Accordingly, the miniQuant software tool 154 in the miniQuant-L or miniQuant-H modes may quantify gene isoforms as well as repetitive elements or any transcripts that have similar nucleotide sequences using the method 900.

[0122] The method 900 begins with obtaining nucleic acid sequencing data comprising both short reads and long reads for a sample 32 (block 902). Next, the method 900 involves mapping the short reads and the long reads to one of a plurality of gene communities (block 904). Each gene community includes one or more genes and gene isoforms thereof.

[0123] For each of the plurality of gene communities, the method 900 estimates the isoform abundance of the gene isoforms in the gene community based on the short reads and the long reads mapped to the gene community (block 906). This estimation is achieved through a combination of a long-read likelihood function and / or a short-read likelihood function, which indicate the likelihoods of the isoform abundances of the gene isoforms given the mapped reads. In the miniQuant-L mode, the miniQuant software tool 154 only uses the long-read likelihood function to estimate the isoform abundance of the gene isoforms in the gene community.

[0124] In the miniQuant-H mode, the mini Quant software tool 154 combines the long-read and short-read likelihood functions to estimate the isoform abundance of the gene isoforms in the gene community. The mini Quant software tool 154 also determines a community-specific weight oic and applies the community- specific weight acto the long-read likelihood function and 1 — etc to the short-read likelihood function to generate a hybrid likelihood function to estimate the isoform abundance of the gene isoforms in the gene community. For example, the miniQuant software tool 154 may determine the community- specific weight acusing a trained machine learning model by providing gene and read structure features (e.g., a K-value) for the community to the trained machine learning model.

[0125] In any event, the miniQuant software tool 154 may perform a maximum likelihood estimate of the long-read likelihood function or the hybrid likelihood function using an Expectation Maximization algorithm to estimate the isoform abundance of the gene isoforms in the gene community.PATENT APPLICATION Attorney Docket No.: 30275 / 70852 / PC

[0126] At block 908, the mini Quant software tool 154 determines whether the isoform abundances have been estimated for each of the C gene communities mapped to the long and short reads in the sample 32. The miniQuant software tool 154 then proceeds to block 906 to estimate isoform abundances for the next gene communities until the isoform abundances for each of the C gene communities have been estimated.

[0127] One the isoform abundances for each of the C gene communities have been estimated, the method 900 determines, for each of a plurality of gene isoforms in each of the plurality of gene communities, an isoform expression level of the isoform within the gene community in the sample (block 910). This determination is based on a combination of the relative abundance / ?cof the gene community to other gene communities and the relative isoform abundance of the gene isoform Qctto other gene isoforms within the gene community.ADDITIONAL CONSIDERATIONS

[0128] Although the disclosure herein sets forth a detailed description of numerous different implementations, it should be understood that the legal scope of the description is defined by the words of the claims set forth at the end of this patent and equivalents. The detailed description is to be construed as exemplary only and does not describe every possible implementation since describing every possible implementation would be impractical. Numerous alternative implementations may be implemented, using either current technology or technology developed after the filing date of this patent, which would still fall within the scope of the claims.

[0129] The following additional considerations apply to the foregoing discussion. Throughout this specification, plural instances may implement components, operations, or structures described as a single instance. Although individual operations of one or more methods are illustrated and described as separate operations, one or more of the individual operations may be performed concurrently, and nothing requires that the operations be performed in the order illustrated. Structures and functionality presented as separate components in example configurations may be implemented as a combined structure or component. Similarly, structures and functionality presented as a single component may be implemented as separate components. These and other variations, modifications, additions, and improvements fall within the scope of the subject matter herein.PATENT APPLICATION Attorney Docket No.: 30275 / 70852 / PC

[0130] Additionally, certain implementations are described herein as including logic or a number of routines, subroutines, applications, or instructions. These may constitute either software (e.g., code embodied on a machine-readable medium or in a transmission signal) or hardware. In hardware, the routines, etc., are tangible units capable of performing certain operations and may be configured or arranged in a certain manner. In example implementations, one or more computer systems (e.g., a standalone, client or server computer system) or one or more hardware modules of a computer system (e.g., a processor or a group of processors) may be configured by software (e.g., an application or application portion) as a hardware module that operates to perform certain operations as described herein.

[0131] The various operations of example methods described herein may be performed, at least partially, by one or more processors that are temporarily configured (e.g., by software) or permanently configured to perform the relevant operations. Whether temporarily or permanently configured, such processors may constitute processor-implemented modules that operate to perform one or more operations or functions. The modules referred to herein may, in some example implementations, comprise processor-implemented modules.

[0132] Similarly, the methods or routines described herein may be at least partially processor implemented. For example, at least some of the operations of a method may be performed by one or more processors or processor-implemented hardware modules. The performance of certain of the operations may be distributed among the one or more processors, not only residing within a single machine, but deployed across a number of machines. In some example implementations, the processor or processors may be located in a single location, while in other implementations the processors may be distributed across a number of locations.

[0133] The performance of certain of the operations may be distributed among the one or more processors, not only residing within a single machine, but deployed across a number of machines. In some example implementations, the one or more processors or processor-implemented modules may be located in a single geographic location (e.g., within a home environment, an office environment, or a server farm). In other implementations, the one or more processors or processor-implemented modules may be distributed across a number of geographic locations.PATENT APPLICATION Attorney Docket No.: 30275 / 70852 / PC

[0134] This detailed description is to be construed as exemplary only and does not describe every possible implementation, as describing every possible implementation would be impractical, if not impossible. A person of ordinary skill in the art may implement numerous alternate implementations, using either current technology or technology developed after the filing date of this application.

[0135] Those of ordinary skill in the art will recognize that a wide variety of modifications, alterations, and combinations may be made with respect to the above described implementations without departing from the scope of the invention, and that such modifications, alterations, and combinations are to be viewed as being within the ambit of the inventive concept.

[0136] The patent claims at the end of this patent application are not intended to be construed under 35 U.S.C. § 112(f) unless traditional means-plus-function language is expressly recited, such as “means for” or “step for” language being explicitly recited in the claim(s). The systems and methods described herein are directed to an improvement to computer functionality and improve the functioning of conventional computers.

Claims

PATENT APPLICATION Attorney Docket No.: 30275 / 70852 / PC CLAIMSWhat is claimed is:

1. A method for sequencing a sample, the method comprising:for a plurality of genes:identifying a set of gene isoforms corresponding to the gene;identifying a set of regions within the gene;generating, by one or more processors, probabilities that a nucleic acid sequencing read that originated from each gene isoform in the set of gene isoforms corresponds to each region in the set of regions; anddetermining, by the one or more processors, a generalized condition number for the gene based on the probabilities;identifying, by the one or more processors, a subset of the plurality of genes based on the generalized condition number; andsequencing a sample for the identified subset of genes.

2. The method of claim 1 , wherein sequencing the sample for the identified subset of genes includes performing a differential expression analysis of gene isoforms of the identified subset of genes.

3. The method of either claim 1 or claim 2, further comprising:generating a probability matrix of the probabilities that the nucleic acid sequencing read that originated from each gene isoform in the set of gene isoforms corresponds to each region in the set of regions;performing a singular value decomposition of the probability matrix to generate a set of singular values; anddetermining a ratio of a largest of the set of singular values, omax, to a smallest of the set of singular values, or, as the generalized condition number for the gene.PATENT APPLICATION Attorney Docket No.: 30275 / 70852 / PC 4. The method of any one of claims 1-3, wherein the set of regions of the gene are exonic regions.

5. The method of any one of claims 1-4, further comprising:ranking the plurality of genes in accordance with the generalized condition number for each of the plurality of claims; andproviding a ranked list of the plurality of genes for display.

6. The method of any one of claims 1-4, wherein identifying a subset of the plurality of genes includes identifying the subset having a generalized condition number below a threshold value to sequence the sample for genes likely to have low gene isoform and / or repetitive element quantification errors.

7. A method of quantifying isoforms of one or more genes in a sample using nucleic acid sequencing data, the method comprising:obtaining, by one or more processors, nucleic acid sequencing data comprising short reads and long reads for a sample;mapping, by the one or more processors, the short reads and the long reads to one of a plurality of gene communities, wherein each gene community includes one or more genes and gene isoforms thereof; andfor each of the plurality of gene communities, estimating, by the one or more processors, the isoform abundance of the gene isoforms in the gene community based on the short reads and the long reads mapped to the gene community; andfor each of a plurality of gene isoforms in each of the plurality of gene communities, determining, by the one or more processors, an isoform expression level of the isoform within the gene community in the sample based on a combination of a relative abundance of the gene community to other gene communities and a relative isoform abundance of the gene isoform to other gene isoforms within the gene community to quantify abundance of the plurality of gene isoforms in the sample.PATENT APPLICATION Attorney Docket No.: 30275 / 70852 / PC 8. The method of claim 7, wherein the gene community, C, includes A6short reads and Mclong reads from the sample mapped to the gene community, the gene community includes 7 gene isoforms, and each of the ' gene isoforms has an abundance, 0CT, represented as a fraction of a total isoform abundance of the T*7gene isoforms of the gene community, 9C.

9. The method of claim 8, wherein estimating the isoform abundance of gene isoforms in the gene community includes:generating, by the one or more processors, a long-read likelihood function indicating likelihoods of the isoform abundances of the gene isoforms in the gene community given the A / *7long reads from the sample mapped to the gene community;generating, by the one or more processors, a short-read likelihood function indicating likelihoods of the isoform abundances of the gene isoforms in the gene community given the Nrshort reads from the sample mapped to the gene community; andcombining, by the one or more processors, the long-read likelihood function and the short-read likelihood function to estimate the isoform abundance of the gene isoforms in the gene community.

10. The method of claim 9, further comprising:determining, by the one or more processors, a community-specific weight for long reads, ac;applying, by the one or more processors, the community-specific weight for long reads to the long-read likelihood function to generate a weighted long-read likelihood function for the gene community;applying, by the one or more processors, one minus the community-specific weight for long reads, 1 - ac, to the short-read likelihood function to generate a weighted short-read likelihood function for the gene community; andcombining, by the one or more processors, the weighted long-read likelihood function and the weighted short-read likelihood function to estimate the isoform abundance of the gene isoforms in the gene community.PATENT APPLICATION Attorney Docket No.: 30275 / 70852 / PC 11. The method of claim 10, wherein determining the community-specific weight for long reads, ac, includes:obtaining, by the one or more processors, features of the one or more genes and the short reads and long reads in the gene community; andapplying, by the one or more processors, the features of the one or more genes and the short reads and long reads to a trained machine learning model to generate the communityspecific weight for long reads, ac.

12. The method of claim 11, further comprising:training, by the one or more processors, the machine learning model using one or more features of each of a plurality of genes and the short reads and long reads and known true relative abundances for gene isoforms of each of the plurality of genes.

13. The method of claim 11, wherein the features of the one or more genes and the short reads and long reads include a generalized condition number for the one or more genes.

14. The method of claim 13, further comprising:determining, by the one or more processors, the generalized condition number for one of the one or more genes by:identifying a set of gene isoforms corresponding to the gene;identifying a set of regions within the gene;generating a probability matrix of probabilities that a nucleic acid sequencing read that originated from each gene isoform in the set of gene isoforms corresponds to each region in the set of regions;performing a singular value decomposition of the probability matrix to generate a set of singular values; anddetermining a ratio of a largest of the set of singular values to a smallest of the set of singular values as the generalized condition number for the gene.

15. The method of any one of claims 9-14, wherein estimating the isoform abundance of gene isoforms in the gene community includes:PATENT APPLICATION Attorney Docket No.: 30275 / 70852 / PC performing a maximum likelihood estimate of the combined long-read likelihood and short-read likelihood functions using an Expectation Maximization algorithm to estimate the isoform abundance of the gene isoforms in the gene community, 0C.

16. The method of any one of claims 9-15, further comprising:for each of the A6short reads from the sample mapped to the gene community, determining probabilities that the short read corresponds to each of the / gene isoforms, wherein the short-read likelihood function is generated based on the probabilities that each of the Acshort reads corresponds to each of the 7<gene isoforms; andfor each of the Mclong reads from the sample mapped to the gene community, determining probabilities that the long read corresponds to each of the T6gene isoforms, wherein the long-read likelihood function is generated based on the probabilities that each of the Mclong reads corresponds to each of the 7’cgene isoforms.

17. The method of any one of claims 7-16, wherein determining the relative abundance of the gene community to other gene communities includes:determining the relative abundance of the gene community, Pc, based on a number of short reads in the gene community, A0, and respective lengths of the short reads in the gene community, Zc, relative to a total number of short reads in each of the plurality of gene communities, A0, and respective lengths of the short reads in each of the plurality of gene communities, lc.

18. The method of claim 17, wherein determining the relative abundance of the gene community to other gene communities further includes:determining a number of short reads in the gene community corresponding to a particular gene isoform, Atc, based on the number of short reads in the gene community, A0, and a relative abundance of the particular gene isoform in the gene community, 0tc, combined with an effective length of the particular gene isoform in the gene community, ltc, determined by a length of the particular gene isoform and respective lengths of the short reads, relative to a summation of relative abundances of each of the gene isoforms in the gene community, 0C, combined with respective effective lengths of each of the gene isoforms in the gene community, lr.PATENT APPLICATION Attorney Docket No.: 30275 / 70852 / PC19. A method of quantifying isoforms of one or more genes in a sample using nucleic acid sequencing data, the method comprising:obtaining, by one or more processors, nucleic acid sequencing data comprising long reads for a sample;generating, by the one or more processors, a long-read likelihood function indicating likelihoods of isoform abundances of gene isoforms given the long reads from the sample; and estimating, by the one or more processors, the isoform abundance of the gene isoforms based on the long-read likelihood function.

20. The method of claim 19, wherein estimating the isoform abundance of the gene isoforms includes:performing a maximum likelihood estimate of the long-read likelihood function using an Expectation Maximization algorithm to estimate the isoform abundance of the gene isoforms.

21. The method of either claim 19 or claim 20, further comprising:for each of the long reads from the sample, determining probabilities that the long read corresponds to each of the gene isoforms,wherein the long-read likelihood function is generated based on the probabilities that each of the long reads corresponds to each of the gene isoforms.